IP Library › Granted Patent US 11,182,592
Granted Patent B2
US 11,182,592 · App. 16/734,336 · Granted Nov 23, 2021

Target object recognition method and apparatus, storage medium, and electronic device

Inventors: Qixing Li (Beijing, CN); Fengwei Yu (Beijing, CN); Junjie Yan (Beijing, CN)
Assignee: BEIJING SENSETIME TECHNOLOGY DEVELOPMENT CO., LTD.
G06K9/00248G06K9/4604G06K9/629G06K9/6256G06T7/73G06T2207/10016G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,182,592
App. No.
16/734,336
Granted
Nov 23, 2021
Kind
B2
Abstract

A target object recognition method includes: performing target object detection on an object of an image to be detected to obtain target object prediction information of the object, where the target object prediction information is confidence information that the detected object is the target object; performing key point detection on the object of the image to be detected to obtain key point prediction information of the object, where the key point prediction information is confidence information that a key point of the detected object is a key point of the target object; fusing the target object prediction information with the key point prediction information to obtain comprehensive prediction information of the object; and recognizing the target object according to the comprehensive prediction information.

Claims (92)

1. A method for target object recognition, comprising:

obtaining an image area corresponding to an object of an image to be detected;

performing target object detection on the image area corresponding to the object of the image to be detected to obtain target object prediction information of the object, wherein the target object prediction information is confidence information that the detected object is a target object;

performing key point detection on the image area corresponding to the object of the image to be detected to obtain key point prediction information of the object, wherein the key point prediction information is confidence information that a key point of the detected object is a key point of the target object;

fusing the target object prediction information with the key point prediction information to obtain comprehensive prediction information of the object; and

recognizing the target object according to the comprehensive prediction information,

wherein after obtaining the image area corresponding to the object of the image to be detected, and before fusing the target object prediction information with the key point prediction information to obtain the comprehensive prediction information of the object, the method further comprises:

detecting deflection angle information of the object from the image area; and

the fusing the target object prediction information with the key point prediction information to obtain the comprehensive prediction information of the object comprises:

fusing the target object prediction information and the key point prediction information with the deflection angle information to obtain comprehensive prediction information of the object.

2. The method according to claim 1 , wherein the fusing the target object prediction information with the key point prediction information to obtain the comprehensive prediction information of the object comprises:

multiplying the target object prediction information and the key point prediction information to obtain the comprehensive prediction information of the object.

3. The method according to claim 1 , wherein the performing key point detection on the object of the image to be detected to obtain the key point prediction information of the object comprises:

performing, through a neural network model for positioning a key point, key point detection on the object of the image to be detected to obtain the key point prediction information of the object.

4. The method according to claim 1 , wherein the detecting deflection angle information of the object from the image area comprises:

detecting the deflection angle information of the object from the image area by using a neural network model for object classification.

5. The method according to claim 1 , wherein

the target object is a face;

before fusing the target object prediction information with the key point prediction information to obtain the comprehensive prediction information of the object, the method further comprises:

detecting at least one of a pitch angle of the face or a yaw angle of the face from the image area;

the fusing the target object prediction information with the key point prediction information to obtain the comprehensive prediction information of the object comprises:

normalizing at least one of the pitch angle of the face or the yaw angle of the face according to an applicable exponential function; and

multiplying the target object prediction information, the key point prediction information, and the normalized pitch angle of the face to obtain the comprehensive prediction information of the object;

or,

multiplying the target object prediction information, the key point prediction information, and the normalized yaw angle of the face to obtain the comprehensive prediction information of the object;

or,

multiplying the target object prediction information, the key point prediction information, the normalized pitch angle of the face, and the normalized yaw angle of the face to obtain the comprehensive prediction information of the object.

6. The method according to claim 1 , wherein the image to be detected is a video frame image;

after recognizing the target object according to the comprehensive prediction information, the method further comprising:

tracking the target object according to a result of recognizing the target object from multiple video frame images;

or, select a video frame image having the highest comprehensive prediction quality from the multiple video frame images as a captured image according to comprehensive prediction information separately obtained for the multiple video frame images;

or,

selecting a predetermined number of video frame images from the multiple video frame images according to the comprehensive prediction information separately obtained for the multiple video frame images, and performing feature fusion on the selected video frame images.

7. An apparatus for target object recognition, comprising:

a processor; and

a memory for storing instructions executable by the processor;

wherein the processor is configured to:

obtain an image area corresponding to an object of an image to be detected;

perform target object detection on the image area corresponding to the object of the image to be detected to obtain target object prediction information of the object, wherein the target object prediction information is confidence information that the detected object is a target object;

perform key point detection on the image area corresponding to the object of the image to be detected to obtain key point prediction information of the object, wherein the key point prediction information is confidence information that a key point of the detected object is a key point of the target object;

fuse the target object prediction information with the key point prediction information to obtain comprehensive prediction information of the object; and

recognize the target object according to the comprehensive prediction information,

wherein after obtaining the image area corresponding to the object of the image to be detected, and before fusing the target object prediction information with the key point prediction information to obtain the comprehensive prediction information of the object, the processor is further configured to:

detect deflection angle information of the object from the image area; and

the processor is specifically configured to: fuse the target object prediction information and the key point prediction information with the deflection angle information to obtain the comprehensive prediction information of the object.

8. The apparatus according to claim 7 , wherein the processor is configured to multiply the target object prediction information and the key point prediction information to obtain the comprehensive prediction information of the object.

9. The apparatus according to claim 7 , wherein the processor is configured to perform, using a neural network model for positioning the key point, key point detection on the object of the image to be detected to obtain the key point prediction information of the object.

10. The apparatus according to claim 7 , wherein the processor is configured to detect the deflection angle information of the object from the image area by using the neural network model for object classification.

11. The apparatus according to claim 7 , wherein

the target object is a face;

before the fusing the target object prediction information with the key point prediction information to obtain comprehensive prediction information of the object, the processor is further configured to:

detect at least one of a pitch angle of the face or a yaw angle of the face from the image area; and

normalize the at least one of the pitch angle of the face or the yaw angle of the face according to the applicable exponential function; and

multiply the target object prediction information, the key point prediction information, and the normalized pitch angle of the face to obtain the comprehensive prediction information of the object;

or,

multiply the target object prediction information, the key point prediction information, and the normalized yaw angle to obtain the comprehensive prediction information of the object;

or,

multiply the target object prediction information, the key point prediction information, the normalized pitch angle of the face, and the normalized yaw angle of the face to obtain the comprehensive prediction information of the object.

12. The apparatus according to claim 7 , wherein the image to be detected is a video frame image;

after the recognizing the target object according to the comprehensive prediction information, the processor is further configured to:

track the target object according to a result of recognizing the target object from the multiple video frame images;

or,

select a video frame image having the highest comprehensive prediction quality from the multiple video frame images as a captured image according to the comprehensive prediction information separately obtained for the multiple video frame images;

or,

select the predetermined number of video frame images from the multiple video frame images according to the comprehensive prediction information separately obtained for the multiple video frame images, and perform feature fusion on the selected video frame images.

13. A non-transitory computer-readable storage medium, having computer program instructions stored thereon, wherein the program instructions, when being executed by a processor, implement a method for target object recognition, the method comprising:

obtaining an image area corresponding to an object of an image to be detected;

performing target object detection on the image area corresponding to the object of the image to be detected to obtain target object prediction information of the object, wherein the target object prediction information is confidence information that the detected object is a target object;

performing key point detection on the image area corresponding to the object of the image to be detected to obtain key point prediction information of the object, wherein the key point prediction information is confidence information that a key point of the detected object is a key point of the target object;

fusing the target object prediction information with the key point prediction information to obtain comprehensive prediction information of the object; and

recognizing the target object according to the comprehensive prediction information,

wherein after obtaining the image area corresponding to the object of the image to be detected, and before fusing the target object prediction information with the key point prediction information to obtain the comprehensive prediction information of the object, the method further comprises:

detecting deflection angle information of the object from the image area; and

the fusing the target object prediction information with the key point prediction information to obtain the comprehensive prediction information of the object comprises:

fusing the target object prediction information and the key point prediction information with the deflection angle information to obtain comprehensive prediction information of the object.

14. The non-transitory computer-readable storage medium according to claim 13 , wherein the fusing the target object prediction information with the key point prediction information to obtain the comprehensive prediction information of the object comprises:

multiplying the target object prediction information and the key point prediction information to obtain the comprehensive prediction information of the object.

15. The non-transitory computer-readable storage medium according to claim 13 , wherein the performing key point detection on the object of the image to be detected to obtain the key point prediction information of the object comprises:

performing, using a neural network model for positioning a key point, key point detection on the object of the image to be detected to obtain the key point prediction information of the object.

16. The non-transitory computer-readable storage medium according to claim 13 , wherein the detecting deflection angle information of the object from the image area comprises:

detecting the deflection angle information of the object from the image area by using a neural network model for object classification.

17. The non-transitory computer-readable storage medium according to claim 13 , wherein

the target object is a face;

before fusing the target object prediction information with the key point prediction information to obtain the comprehensive prediction information of the object, the method further comprises:

detecting at least one of a pitch angle of the face or a yaw angle of the face from the image area;

the fusing the target object prediction information with the key point prediction information to obtain the comprehensive prediction information of the object comprises:

normalizing at least one of the pitch angle of the face or the yaw angle of the face according to an applicable exponential function; and

multiplying the target object prediction information, the key point prediction information, and the normalized pitch angle of the face to obtain the comprehensive prediction information of the object;

or,

multiplying the target object prediction information, the key point prediction information, and the normalized yaw angle of the face to obtain the comprehensive prediction information of the object;

or,

multiplying the target object prediction information, the key point prediction information, the normalized pitch angle of the face, and the normalized yaw angle of the face to obtain the comprehensive prediction information of the object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2020
From: LI, QIXING; YU, FENGWEI; YAN, JUNJIE
To: BEIJING SENSETIME TECHNOLOGY DEVELOPMENT CO., LTD.
Reel/Frame 053095/0709 →
Priority Claims (1)
CN 201711181299.5 · Nov 23, 2017 · national
Continuity (2)
Continuation PCTCN2018111513 · Oct 23, 2018
Related Publication 20200143146A1 · May 7, 2020
Cited By (1)
US 12,399,266