IP Library › Granted Patent US 9,710,698
Granted Patent B2
US 9,710,698 · App. 14/408,729 · Granted Jul 18, 2017

Method, apparatus and computer program product for human-face features extraction

Inventors: Yong Ma (Beijing, CN); Yan Ming Zou (Beijing, CN); Kong Qiao Wang (Beijing, CN)
Assignee: Nokia Technologies Oy
G06K9/00281G06K9/00261G06K9/00288
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,710,698
App. No.
14/408,729
Granted
Jul 18, 2017
Kind
B2
Abstract

The present invention provides a method for obtaining a human-face feature vector from a video image sequence, comprising: detecting a same human-face image in a plurality of image frames of the video sequence; dividing the detected human-face image into a plurality of local patches with a predetermined size, wherein each local patch is around or near a human-face feature point; determining a correspondence relationship between respective local patches of the same human-face image in the plurality of image frames of the video sequence; and using human-face local feature vector components extracted from respective local patches having a mutual correspondence relationship to form human-face local feature vectors representing facial points corresponding to the local patches. Besides, the present invention further provides an apparatus for obtaining a human-face feature vector from a video image sequence and a corresponding computer program product.

Claims (56)

1. A method comprising:

detecting a same human-face image in a plurality of image frames of a video sequence;

dividing the detected human-face image into a plurality of local patches with a predetermined size, wherein each local patch is around or near a human-face feature point;

determining a correspondence relationship between respective local patches of the same human-face image in the plurality of image frames of the video sequence;

determining a first local feature from a first frame of the plurality of image frames of the same human-face image;

determining a second local feature, different than the first local feature, from a second frame of the plurality of image frames of the same human-face image;

using human-face local feature vector components extracted from respective local patches having a mutual correspondence relationship based on the first local feature and the second local feature of the first and the second frames of the same human-face image to form human-face local feature vectors representing the facial points corresponding to the local patches, wherein formation of the human-face local feature vectors comprises resizing the plurality of local patches to obtain the human-face local feature vectors representing the facial points corresponding to respective resized local patches;

and

combining the resulting human-face local feature vectors representing the facial points corresponding to respective local patches to form human-face global feature vectors.

2. The method of claim 1 , wherein using human-face local feature vector components extracted from respective local patches having a mutual correspondence relationship to form the human-face local feature vectors representing the facial points corresponding to the local patches comprises:

identifying different poses of the same human-face image in a plurality of image frames of a video sequence;

extracting human-face local feature vector components only from respective un-occluded local patches having a mutual correspondence relationship based on the identified different poses of the same human face in the plurality of image frames; and

combining the human-face local feature vector components extracted from respective un-occluded local patches having a mutual correspondence relationship to form a human-face local feature vectors representing the facial points corresponding to the local patches.

3. The method of claim 1 , wherein using human-face local feature vector components extracted from respective local patches having a mutual correspondence relationship to form human-face local feature vectors representing the facial points corresponding to the local patches comprises:

identifying different poses of the same human-face image in a plurality of image frames of a video sequence; and

weight combining the human-face local feature vector components extracted from respective local patches having a mutual correspondence relationship based on the identifed different poses of the same human face in different image frames to form human-face local feature vectors representing the facial points corresponding to the local patches.

4. The method of claim 1 , wherein dividing the detected human-face image into a plurality of local patches with a predetermined size is implemented by adding a grid on the detected human-face image.

5. The method of claim 1 , further comprising combining the resulting human-face local feature vectors representing the facial points corresponding to respective resized local patches to form human-face global feature vectors.

6. The method of claim 5 , further comprising combining human-face global feature vectors obtained under different local patch sizes to form a human-face global feature vector set.

7. The method of claim 1 , further comprising performing human-face recognition using the resulting human-face global feature vectors.

8. An apparatus, comprising:

at least one processor and at least one memory including a computer program code, wherein the at least one memory including the computer program code are configured, with the at least one processor, to cause the apparatus to:

detect a same human-face image in a plurality of image frames of a video sequence;

divide the detected human-face image into a plurality of local patches with a predetermined size, wherein each local patch is around or near a human-face feature point;

determine a correspondence relationship between respective local patches of the same human-face image in the plurality of image frames of the video sequence;

determine a first local feature from a first frame of the plurality of image frames of the same human-face image;

determine a second local feature, different than the first local feature, from a second frame of the plurality of image frames of the same human-face image;

use human-face local feature vector components extracted from respective local patches having a mutual correspondence relationship based on the first local feature and the second local feature of the first and the second frames of the same human-face image to form human-face local feature vectors representing facial points corresponding to the local patch, wherein formation of the human-face local feature vectors comprises resizing the plurality of local patches to obtain the human-face local feature vectors representing the facial points corresponding to respective resized local patches;

and

combine the resulting human-face local feature vectors representing the facial points corresponding to respective local patches to form human-face global feature vectors.

9. The apparatus of claim 8 , wherein dividing the detected human-face into a plurality of local patches with a predetermined size is implemented by adding a grid on the detected human-face image.

10. The apparatus of claim 8 , wherein the at least one memory including the computer program code are configured, with the at least one processor, to further cause the apparatus to combine the resulting human-face local feature vectors representing the facial points corresponding to respective resized local patches to form a human-face global feature vector.

11. The apparatus of claim 10 , wherein the at least one memory including the computer program code are configured, with the at least one processor, to further cause the apparatus to combine the resulting human-face global feature vectors under different local patch sizes to form a human-face global feature vector set.

12. The apparatus of claim 8 , wherein the at least one memory including the computer program code are configured, with the at least one processor, to further cause the apparatus to use the resulting human-face global feature vectors to perform human-face recognition.

13. The apparatus of claim 11 , wherein the at least one memory including the computer program code are configured, with the at least one processor, to further cause the apparatus to use the resulting human-face global feature vector set to perform human-face recognition.

14. An apparatus comprising:

at least one processor and at least one memory including a computer program code, wherein the at least one memory including the computer program code are configured, with the at least one processor, to cause the apparatus to:

detect a same human-face image in a plurality of image frames of a video sequence;

divide the detected human-face image into a plurality of local patches with a predetermined size, wherein each local patch is around or near a human-face feature point;

determine a correspondence relationship between respective local patches of the same human-face image in the plurality of image frames of the video sequence;

determine a first local feature from a first frame of the plurality of image frames of the same human-face image;

determine a second local feature, different than the first local feature, from a second frame of the plurality of image frames of the same human-face image; and

use human-face local feature vector components extracted from respective local patches having a mutual correspondence relationship based on the first local feature and the second local feature of the first and the second frames of the same human-face image to form human-face local feature vectors representing facial points corresponding to the local patch, wherein the apparatus is further caused to:

identify different poses of the same human-face image in a plurality of image frames of a video sequence;

extract the human-face local feature vector components only from respective un-occluded local patches having a mutual correspondence relationship based on the identified different poses of the same human face in the plurality of image frames; and

combine the human-face local feature vector components extracted from respective un-occluded local patches having a mutual correspondence relationship to form the human-face local feature vectors representing facial points corresponding to the local patches.

15. An apparatus

at least one processor and at least one memory including a computer program code, wherein the at least one memory including the computer program code are configured, with the at least one processor, to cause the apparatus to:

detect a same human-face image in a plurality of image frames of a video sequence;

divide the detected human-face image into a plurality of local patches with a predetermined size, wherein each local patch is around or near a human-face feature point;

determine a correspondence relationship between respective local patches of the same human-face image in the plurality of image frames of the video sequence;

determine a first local feature from a first frame of the plurality of image frames of the same human-face image;

determine a second local feature, different than the first local feature, from a second frame of the plurality of image frames of the same human-face image; and

use human-face local feature vector components extracted from respective local patches having a mutual correspondence relationship based on the first local feature and the second local feature of the first and the second frames of the same human-face image to form human-face local feature vectors representing facial points corresponding to the local patch, wherein the apparatus is further caused to:

identify different poses of the same human-face image in a plurality of image frames of a video sequence; and

weight combine the human-face local feature vector components extracted from respective local patches having a mutual correspondence relationship based on the identified different poses of the same human face in different image frames to form human-face local feature vectors representing the facial points corresponding to the local patches.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2016
From: NOKIA CORPORATION
To: NOKIA TECHNOLOGIES OY
Reel/Frame 038306/0086 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2015
From: MA, YONG; ZOU, YAN MING; WANG, KONG QIAO
To: NOKIA CORPORATION
Reel/Frame 035073/0513 →
Priority Claims (1)
CN 2012 1 0223706 · Jun 25, 2012 · national
Continuity (1)
Related Publication 20150205997A1 · Jul 23, 2015