IP Library › Granted Patent US 12,131,584
Granted Patent B2
US 12,131,584 · App. 17/642,781 · Granted Oct 29, 2024

Expression recognition method and apparatus, electronic device, and storage medium

Inventors: Yanhong Wu (Beijing, CN); Guannan Chen (Beijing, CN); Pablo Navarrete Michelini (Beijing, CN); Lijie Zhang (Beijing, CN)
Assignee: BOE TECHNOLOGY GROUP CO., LTD
G06V40/174G06V40/166G06V40/172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,131,584
App. No.
17/642,781
Granted
Oct 29, 2024
Kind
B2
Abstract

An expression recognition method is described that includes acquiring a face image to be recognized, and inputting the face image into N different recognition models arranged in sequence for expression recognition and outputting an actual expression recognition result, the N different recognition models being configured to recognize different target expression types, wherein N is an integer greater than 1.

Claims (60)

1. An expression recognition method, comprising:

acquiring a face image to be recognized; and

inputting the face image into N different recognition models arranged in sequence for expression recognition and outputting an actual expression recognition result, the N different recognition models being configured to recognize different target expression types, wherein N is an integer greater than 1;

wherein the inputting the face image into the N different recognition models arranged in sequence for expression recognition and outputting the actual expression recognition result comprises: inputting the face image into ith recognition model for expression recognition and outputting a first recognition result, wherein i is an integer ranging from 1 to N−1, and an initial value of i is 1;

determining whether the first recognition result and the target expression type corresponding to the ith recognition model are same, outputting the first recognition result as the actual expression recognition result when the first recognition result is same as the target expression type corresponding to the ith recognition model, and inputting the face image into (i+1)th recognition model for expression recognition when the first recognition result is different from the target expression type corresponding to the ith recognition model; and

in response to inputting the face image into Nth recognition model, outputting a second recognition result as the actual expression recognition result by performing expression recognition on the face image through the Nth recognition model, wherein the Nth recognition model is configured to recognize a plurality of target expression types, and the second recognition result is one of the plurality of target expression types.

2. The method according to claim 1 , wherein, in any two adjacent recognition models, a recognition accuracy of a former recognition model is greater than a recognition accuracy of a latter recognition model.

3. The method according to claim 1 , wherein each of previous N−1 recognition models is configured to recognize one target expression type.

4. The method according to claim 1 , wherein outputting the second recognition result by performing the expression recognition on the face image through the Nth recognition model comprises:

processing the face image through the Nth recognition model to obtain a plurality of target expression types and a plurality of probability values corresponding thereto; and

obtaining a maximum probability value by comparing the plurality of probability values, and outputting a target expression type corresponding to the maximum probability value as the second recognition result.

5. The method according to claim 1 , wherein each of the recognition models comprises a Gabor filter.

6. The method according to claim 1 , wherein each of the recognition models comprises: 16 convolutional layers, 1 global average pooling layer and 1 fully connected layer, and the convolutional layer comprises 3×3 convolution kernels.

7. The method according to claim 1 , wherein, before inputting the face image into N different recognition models arranged in sequence for expression recognition, the method further comprises:

acquiring a facial expression training data set, wherein the facial expression training data set comprises: a plurality of face images and target expression types corresponding to each of the plurality of face images;

determining a division order of each target expression type based on a proportion of the each target expression type in the facial expression training data set; and

sequentially generating the N recognition models by performing training based on the facial expression training data set and the division order.

8. The method according to claim 7 , wherein determining the division order of the each target expression type based on the proportion of the each target expression type in the facial expression training data set comprises:

obtaining a proportion order by sorting the proportions of the each target expression type in the facial expression training data set in descending order; and

determining an order of the each target expression type corresponding to the proportion order as the division order of the each target expression type.

9. The method according to claim 7 , wherein the sequentially generating the N recognition models by performing training based on the facial expression training data set and the division order comprises:

using the facial expression training data set as a current training data set;

dividing the current training data set according to jth expression type in the division order, and obtaining a first subset of expression type corresponding to the jth target expression type, and a second subset corresponding to other target expression types other than the jth target expression type, wherein an initial value of j is 1;

training a jth original recognition model by using the first subset and the second subset as training sets to obtain jth recognition model, wherein a target expression type corresponding to the jth recognition model is the jth target expression type;

adding 1 to a value of j, using the second subset as the current training data set as updated, and returning the step of dividing the current training data set according to the j-th target expression type in the division order, until the (N−1)th recognition model being determined; and

train an Nth original recognition model by using the current training set updated for N−1 times to obtain the Nth recognition model.

10. The method according to claim 7 , wherein the determining the division order of the each target expression type based on the proportion of each target expression type in the facial expression training data set comprises:

in the proportion of the each target expression type in the facial expression training data set, when a maximum value is greater than a proportion threshold, arranging a target expression type corresponding to the maximum value in a first place, and randomly arranging other target expression types to obtain a plurality of division orders;

performing binary classification division of the facial expression training data set according to each of the plurality of division orders to obtain a plurality of subsets, and determining impurity of divided data set according to the plurality of subsets; and

in the obtained impurities of the divided data set corresponding to the plurality of division orders, determining a division order corresponding to a minimum value of the obtained impurities as the division order of the each target expression type.

11. The method according to claim 10 , wherein after the determining the impurity of the divided data set, the method further comprises:

sorting the impurities corresponding to the plurality of division orders in ascending order, and determining division orders corresponding to previous L impurities as L target division orders, wherein L is an integer greater than 1;

performing training, according to the facial expression training data set and each of the L target division orders, to generate a plurality of target models corresponding to the each of the L target division orders; and

testing the plurality of target models corresponding to the each of the L target division orders through a test set, and determining the plurality of target models with a highest accuracy as the N recognition models, wherein a number of the plurality of target models with the highest accuracy rate is N.

12. The method according to claim 1 , wherein N is 5, and the target expression types to be recognized by previous four recognition models in 5 recognition models as sequentially arranged are: happy, surprised, neutral, and sad; and the target expression types to be recognized by a fifth recognition model are: angry, disgusted, and fearful.

13. An electronic device comprising:

at least one hardware processor; and

a memory configured to store executable instructions for the at least one hardware processor that, when executed, directs the at least one hardware processor to:

acquire a face image to be recognized; and

input the face image into N different recognition models arranged in sequence for expression recognition and outputting an actual expression recognition result, the N different recognition models being configured to recognize different target expression types, wherein N is an integer greater than 1;

input the face image into ith recognition model for expression recognition and outputting a first recognition result, wherein i is an integer ranging from 1 to N−1, and an initial value of i is 1;

determine whether the first recognition result and the target expression type corresponding to the ith recognition model are same, output the first recognition result as the actual expression recognition result when the first recognition result is same as the target expression type corresponding to the ith recognition model, and input the face image into (i+1)th recognition model for expression recognition when the first recognition result is different from the target expression type corresponding to the ith recognition model; and

in response to the face image into Nth recognition model being inputted, output a second recognition result as the actual expression recognition result by performing expression recognition on the face image through the Nth recognition model, wherein the Nth recognition model is configured to recognize a plurality of target expression types, and the second recognition result is one of the plurality of target expression types.

14. The device according to claim 13 , wherein, in any two adjacent recognition models, a recognition accuracy of a former recognition model is greater than a recognition accuracy of a latter recognition model.

15. The device according to claim 13 , wherein each of previous N−1 recognition models is configured to recognize one target expression type.

16. The device according to claim 13 , wherein the at least one hardware processor is further directed to:

process the face image through the Nth recognition model to obtain a plurality of target expression types and a plurality of probability values corresponding thereto; and

obtain a maximum probability value by comparing the plurality of probability values, and output a target expression type corresponding to the maximum probability value as the second recognition result.

17. The device according to claim 13 , wherein each of the recognition models comprises a Gabor filter.

18. The device according to claim 13 , wherein each of the recognition models comprises: 16 convolutional layers, 1 global average pooling layer and 1 fully connected layer, and the convolutional layer comprises 3×3 convolution kernels.

19. The device according to claim 13 , wherein the at least one hardware processor is further directed to:

acquire a facial expression training data set, wherein the facial expression training data set comprises: a plurality of face images and target expression types corresponding to each of the plurality of face images;

determine a division order of each target expression type based on a proportion of the each target expression type in the facial expression training data set; and

sequentially generate the N recognition models by performing training based on the facial expression training data set and the division order.

20. A non-transitory computer-readable storage medium on which a computer program is stored, wherein the computer program, when being executed by at least one hardware processor, is used for performing an expression recognition method, comprising:

acquiring a face image to be recognized; and

inputting the face image into N different recognition models arranged in sequence for expression recognition and outputting an actual expression recognition result, the N different recognition models being configured to recognize different target expression types, wherein Nis an integer greater than 1;

wherein the inputting the face image into the N different recognition models arranged in sequence for expression recognition and outputting the actual expression recognition result comprises: inputting the face image into ith recognition model for expression recognition and outputting a first recognition result, wherein i is an integer ranging from 1 to N−1, and an initial value of i is 1;

determining whether the first recognition result and the target expression type corresponding to the ith recognition model are same, outputting the first recognition result as the actual expression recognition result when the first recognition result is same as the target expression type corresponding to the ith recognition model, and inputting the face image into (i+1)th recognition model for expression recognition when the first recognition result is different from the target expression type corresponding to the ith recognition model; and

in response to inputting the face image into Nth recognition model, outputting a second recognition result as the actual expression recognition result by performing expression recognition on the face image through the Nth recognition model, wherein the Nth recognition model is configured to recognize a plurality of target expression types, and the second recognition result is one of the plurality of target expression types.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2024
From: WU, YANHONG; CHEN, GUANNAN; NAVARRETE MICHELINI, PABLO; ZHANG, LIJIE
To: BOE TECHNOLOGY GROUP CO., LTD.
Reel/Frame 068052/0081 →
Priority Claims (1)
CN 202010364481.X · Apr 30, 2020 · national
Continuity (1)
Related Publication 20220319233A1 · Oct 6, 2022