IP Library Granted Patent US 10,755,079
Granted Patent B2
US 10,755,079 · App. 15/856,983 · Granted Aug 25, 2020

Method and apparatus for acquiring facial information

Inventors: Leilei Gao (Beijing, CN); Shuhan Luan (Beijing, CN); Zhike Zhang (Beijing, CN); Fei Wang (Beijing, CN); Jing Li (Beijing, CN); Xiangtao Jiang (Beijing, CN); Yue Liu (Beijing, CN)
Assignee: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
G06K9/00228G06F16/583G06K9/00288G10L15/1815G10L15/22G10L2015/088G10L2015/223G10L2015/226
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,755,079
App. No.
15/856,983
Granted
Aug 25, 2020
Kind
B2
Abstract

Embodiments of the present disclosure disclose a method and apparatus for acquiring facial information. A specific embodiment of the method comprises: acquiring to-be-processed voice information and a to-be-processed image, and performing voice recognition on the to-be-processed voice information to acquire query information, the to-be-processed image comprising a plurality of facial images, and the query information being used to instruct querying facial information of a specified facial image in the to-be-processed image; recognizing semantically the query information to acquire a keyword set for querying the facial information; processing the to-be-processed image to acquire facial information in the to-be-processed image; and acquiring facial information corresponding to the to-be-processed voice information from the facial information by using the keyword set. The present embodiment realizes recognition on a plurality of facial images contained in a to-be-processed image and determines facial information by using a keyword, thereby improving the accuracy of acquiring a facial image by means of voice.

Claims (43)

1. A method for acquiring facial information, the method comprising:

acquiring to-be-processed voice information and a to-be-processed image, and performing voice recognition on the to-be-processed voice information to acquire query information, the to-be-processed image comprising a plurality of facial images, and the query information acquired from the to-be-processed voice information being used to instruct querying facial information of a specified facial image in the to-be-processed image;

recognizing semantically the query information to acquire a keyword set for specifying the specified facial image in the to-be-processed image;

performing facial recognition on the to-be-processed image, and determining each facial image in the to-be-processed image;

scanning the to-be-processed image to acquire coordinate information of the each facial image in the to-be-processed image;

setting position information of each facial image in the to-be-processed image according to the coordinate information of the each facial image, comprising:

dividing the to-be-processed image into grids of m rows and n columns according to the coordinate information of each facial image, a number of the facial image in each grid being not more than one; and

setting a row number and/or a column number of a grid where the each facial image is located as position information of the facial image;

determining target position information of a facial image corresponding to the query information in the to-be-processed image through a keyword in the keyword set; and

assigning the facial image corresponding to the target position information as a target facial image corresponding to the to-be-processed voice image, and acquiring facial information of the target facial image.

2. The method according to claim 1 , wherein the recognizing semantically the query information to acquire a keyword set for querying the facial information comprises:

semantically recognizing the query information to acquire semantically recognized information, and dividing the semantically recognized information into phrases to acquire a phrase set comprising at least one phrase; and

selecting at least one keyword from the phrase set for determining a facial image in the to-be-processed image, and combining the at least one keyword into a keyword set, the keyword including a position keyword.

3. The method according to claim 1 , wherein the method further comprises:

querying personal information corresponding to the each facial image, and assigning the personal information as facial information corresponding to the facial image, the personal information comprising a name and a gender.

4. The method according to claim 1 , further comprising: displaying the target facial image, the displaying the target facial image comprising:

displaying, in the each facial image in the to-be-processed image position, information corresponding to the each facial image; and

highlighting, in response to selection information of a user selecting position information of the target facial image, the target facial image corresponding to the position information as indicated by the selection information in the to-be-processed image.

5. The method according to claim 4 , further comprising:

displaying and/or playing, in response to determination information of the user determining the target facial image, facial information of the target facial image as indicated by the determination information.

6. An apparatus for acquiring facial information, the apparatus comprising:

at least one processor; and

a memory storing instructions, which when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising:

acquiring to-be-processed voice information and a to-be-processed image, and performing voice recognition on the to-be-processed voice information to acquire query information, the to-be-processed image comprising a plurality of facial images, and the query information acquired from the to-be-processed voice information being used to instruct querying facial information of a specified facial image in the to-be-processed image;

recognizing semantically the query information to acquire a keyword set for specifying the specified facial image in the to-be-processed image;

performing facial recognition on the to-be-processed image, and determining each facial image in the to-be-processed image;

scanning the to-be-processed image to acquire facial coordinate information of the each facial image in the to-be-processed image;

setting position information of each facial image in the to-be-processed image according to the coordinate information of the each facial image, comprising:

dividing the to-be-processed image into grids of m rows and n columns according to the coordinate information of each facial image, a number of the facial image in each grid being not more than one; and

setting a row number and/or a column number of a grid where the each facial image is located as position information of the facial image;

determining target position information of a facial image corresponding to the query information in the to-be-processed image through a keyword in the keyword set; and

assigning the facial image corresponding to the target position information as a target facial image.

7. The apparatus according to claim 6 , wherein the recognizing semantically the query information to acquire a keyword set for querying the facial information comprises:

semantically recognizing the query information to acquire semantically recognized information, and dividing the semantically recognized information into phrases to acquire a phrase set comprising at least one phrase; and

selecting at least one keyword from the phrase set for determining a facial image in the to-be-processed image, and combining the at least one keyword into a keyword set, the keyword including a position keyword.

8. The apparatus according to claim 6 , wherein the operations further comprise:

querying personal information corresponding to the each facial image, and assigning the personal information as facial information corresponding to the facial image, the personal information comprising a name and a gender.

9. The apparatus according to claim 6 , further comprising: displaying the target facial image, the displaying the target facial image comprising:

displaying, in the each facial image in the to-be-processed image position, information corresponding to the each facial image; and

highlighting, in response to selection information of a user selecting position information of the target facial image, the target facial image corresponding to the position information as indicated by the selection information in the to-be-processed image.

10. The apparatus according to claim 9 , further comprising:

displaying and/or playing, in response to determination information of the user determining the target facial image, facial information of the target facial image as indicated by the determination information.

11. A computer readable storage medium storing a computer program, wherein the program, when executed by a processor, causes the processor to perform the method according to claim 1 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2021
From: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.; SHANGHAI XIAODU TECHNOLOGY CO. LTD.
Reel/Frame 056811/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2017
From: GAO, LEILEI; LUAN, SHUHAN; ZHANG, ZHIKE; WANG, FEI; LI, JING; JIANG, XIANGTAO; LIU, YUE
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 044995/0553 →
Priority Claims (1)
CN 2017 1 1137585 · Nov 16, 2017 · national
Continuity (1)
Related Publication 20190147222A1 · May 16, 2019