IP Library Granted Patent US 11,430,259
Granted Patent B2
US 11,430,259 · App. 16/584,769 · Granted Aug 30, 2022

Object detection based on joint feature extraction

Inventors: Dong Chen (Beijing, CN); Fang Wen (Beijing, CN); Gang Hua (Beijing, CN)
Assignee: Microsoft Technology Licensing, LLC
G06V40/171G06K9/6271G06V10/454G06V40/161G06V40/165G06V40/168
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,430,259
App. No.
16/584,769
Granted
Aug 30, 2022
Kind
B2
Abstract

In implementations of the subject matter described herein, a solution for object detection is proposed. First, a feature(s) is extracted from an image and used to identify a candidate object region in the image. Then another feature(s) is extracted from the identified candidate object region. Based on the features extracted in these two stages, a target object region in the image and a confidence for the target object region are determined. In this way, the features that characterize the image from the whole scale and a local scale are both taken into consideration in object recognition, thereby improving accuracy of the object detection.

Claims (71)

1. A computer-implemented method comprising: performing, by at least one processor, a first extraction stage comprising:

extracting a first feature from an image of a whole scale, wherein the first feature characterizes information about context of the image with numeric values, and

identifying one or more candidate face regions of the image based on the extracted first feature;

performing a second extraction stage comprising extracting a second feature from the identified one or more candidate face regions, wherein the second feature characterizes information in a local scale, which is at the same scale as the whole scale, of the image with numeric values, comprises normalized face poses, and provides details of a rotation or scale variation inside the identified one or more candidate face regions;

performing a joint feature stage comprising:

obtaining the first feature, obtaining the second feature, and

identifying at least one face region of the image based on the obtained first feature and the obtained second feature.

2. The method of claim 1 , wherein the joint feature stage further comprises:

determining a confidence of the identified at least one face region based on the obtained first feature and the obtained second feature.

3. The method of claim 1 , wherein the identifying the one or more candidate face regions identifies a plurality of candidate face regions of the image based on the extracted first feature.

4. The method of claim 1 , further comprising:

transforming positions of facial landmarks to canonical positions, wherein the canonical positions are used to normalize face poses.

5. The method of claim 1 , wherein the second feature is extracted from a plurality of candidate face regions.

6. The method of claim 1 , wherein the performing a first extraction stage further comprises:

extracting another first feature from the image of the whole scale, wherein the another first feature characterizes information of the image with numeric values, and

identifying one or more candidate face regions of the image based on the extracted another first feature.

7. The method of claim 6 , wherein the performing the joint feature stage further comprises:

obtaining the another first feature, and

identifying at least one face region of the image based on the obtained first feature, the obtained another first feature, and the obtained second feature.

8. The method of claim 1 , wherein extracting the first feature comprises:

identifying a set of patches with a predefined size in the image;

constructing a first mask by binarizing values of pixels in the image based on a first patch in the set of patches, wherein the first patch is identified as a candidate object;

masking the image with the first mask; and

extracting the first feature from the masked image.

9. The method of claim 8 , wherein constructing the first mask further comprises:

generating an enlarged patch by increasing the predefined size of the first patch of the set of patch; and

constructing the first mask by binarizing values of pixels in the image based on the enlarged patch.

10. The method of claim 8 , wherein extracting the first feature from the masked image further comprises:

downsampling the image of the whole scale with a predefined sampling rate;

identifying a new set of patches with a defined size in the downsampled image;

constructing a second mask by binarizing values of pixels in the downsampled image based on a second patch in the new set of patches that is identified as a candidate object in the downsampled image;

masking the downsampled image with the second mask; and

extracting the first feature from the masked downsampled image.

11. The method of claim 8 , wherein the first feature is extracted by a first process and the at least one second feature is extracted by a second process, and the first and second processes are different.

12. A device comprising:

a processing unit;

a memory coupled to the processing unit and storing instructions thereon, the instructions, when executed by the processing unit, causing the device to:

perform a first extraction stage comprising:

extracting a first feature from an image of a whole scale, wherein the first feature characterizes information about context of the image with numeric values, and

identifying one or more candidate face regions of the image based on the extracted first feature;

perform a second extraction stage comprising extracting a second feature from the identified one or more candidate face regions using a second objective function, wherein the second feature characterizes information in a local scale, which is at the same scale as the whole scale, of the image with numeric values, comprises normalized face poses, and provides details of a rotation or scale variation inside the identified one or more candidate face regions;

perform a joint feature stage comprising:

obtaining the first feature,

obtaining the second feature, and

identifying at least one face region of the image based on the obtained first feature and the obtained second feature.

13. The device of claim 12 , wherein the joint feature stage further comprises:

determining a confidence of the identified at least one face region based on the obtained first feature and the obtained second feature.

14. The device of claim 12 , wherein the identifying the one or more candidate face regions identifies a plurality of candidate face regions of the image based on the extracted first feature.

15. The device of claim 12 , wherein a totality of the one or more candidate face regions comprises less than the whole scale of the image.

16. The device of claim 12 , wherein the second feature is extracted from a plurality of candidate face regions.

17. The device of claim 12 , wherein the performing a first extraction stage further comprises:

extracting another first feature from the image of the whole scale, wherein the another first feature characterizes information of the image with numeric values, and

identifying one or more candidate face regions of the image based on the extracted another first feature.

18. The device of claim 17 , wherein the performing the joint feature stage further comprises:

obtaining the another first feature, and

identifying at least one face region of the image based on the obtained first feature, the obtained another first feature, and the obtained second feature.

19. The device of claim 12 , wherein extracting the first feature comprises:

downsampling the image with a predefined sampling rate;

identifying patches with a defined size in the downsampled image;

constructing a second mask by binarizing values of pixels in the downsampled image based on one of the patches that is identified as a candidate object;

masking the downsampled image with the second mask; and

extracting the first feature from the masked downsampled image.

20. A computer program product being tangibly stored on a machine-readable storage medium and comprising machine-executable instructions, the instructions, when executed on at least one processor of a device, causing the device to:

perform a first extraction stage comprising:

extracting a first feature from an image of a whole scale, wherein the first feature characterizes information about context of the image with numeric values, and

identifying candidate face regions of the image based on the extracted first feature, a totality of the identified candidate face regions comprising less than the whole scale using a first objective function;

perform a second extraction stage comprising extracting a second feature from a plurality of the identified candidate face regions using a second objective function, wherein the second feature characterizes information in a local scale, which is at the same scale as the whole scale, of the image with numeric values, comprises normalized face poses, and provides details of a rotation or scale variation inside the identified one or more candidate face regions;

perform a joint feature stage comprising:

obtaining the first feature,

obtaining the second feature, and

identifying at least one face region of the image based on the obtained first feature and the obtained second feature.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2019
From: CHEN, DONG; WEN, FANG; HUA, GANG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 050508/0319 →
Continuity (2)
Continuation 15261761 · Sep 9, 2016
Related Publication 20200026907A1 · Jan 23, 2020