IP Library › Granted Patent US 12,236,712
Granted Patent B2
US 12,236,712 · App. 17/473,887 · Granted Feb 25, 2025

Facial expression recognition method and apparatus, electronic device and storage medium

Inventors: Yanbo Fan (Guangdong, CN); Yong Zhang (Guangdong, CN); Le Li (Guangdong, CN); Baoyuan Wu (Guangdong, CN); Zhifeng Li (Guangdong, CN); Wei Liu (Guangdong, CN)
Assignee: Tencent Technology (Shenzhen) Company Limited
G06V40/174G06F18/214G06F18/2193G06F18/253G06N3/045G06V10/255G06V10/44G06V10/56G06V40/171
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,712
App. No.
17/473,887
Granted
Feb 25, 2025
Kind
B2
Abstract

A facial expression recognition method includes extracting a first feature from color information of pixels in a first image, and extracting a second feature of facial key points from the first image. The method further includes combining the first feature and the second feature, to obtain a fused feature, and determining, by processing circuitry of an electronic device, a first expression.

Claims (81)

1. A facial expression recognition method comprising:

extracting a first feature of a facial texture from color information of pixels in a first image;

extracting a second feature of facial key points from the first image;

combining the first feature and the second feature, to obtain a fused feature; and

determining, by processing circuitry of an electronic device, a first expression type of a face in the first image that corresponds to the fused feature, the first expression type being determined from a plurality of facial expression types according to a correlation between the facial texture of the first feature and the second feature of the facial key points indicated by the fused feature, each of the plurality of facial expression types being indicated by a different correlation between the facial texture of the first feature and the second feature of the facial key points.

2. The method according to claim 1 , wherein

the extracting the first feature comprises:

extracting, through a convolutional neural network (CNN) of a trained neural network model, the first feature representing the facial texture in the first image from the color information of the pixels in the first image;

the extracting the second feature comprises:

extracting, through a graph neural network (GNN) of the trained neural network model, the second feature representing correlations between the facial key points, the facial key points representing components and/or a facial contour of the face;

the combining the first feature and the second feature comprises:

performing, through a fusion layer of the trained neural network model, a feature fusion operation on the first feature and the second feature, to obtain the fused feature; and

the determining comprises:

recognizing, through a classification network of the trained neural network model, from the plurality of facial expression types, the first expression type corresponding to the fused feature.

3. The method according to claim 2 , wherein the extracting the first feature representing the facial texture in the first image comprises:

using color encoding of the pixels in the first image as an input of the CNN, the CNN performing a convolution operation on the color encoding of the pixels in the first image, to obtain the first feature; and

obtaining the first feature outputted by the CNN.

4. The method according to claim 3 , wherein the using the color encoding of the pixels in the first image as an input of the CNN comprises:

performing, in response to a determination that a position of a reference point in the first image is different from a position of a reference point in an image template, a cropping operation and/or a scaling operation on the first image to obtain a second image, such that a position of the reference point in the second image is the same as the position of the reference point in the image template; and

using the color encoding of the pixels in the second image as the input of the CNN.

5. The method according to claim 2 , wherein the extracting the second feature representing the correlations between the facial key points comprises:

adding positions of the facial key points as nodes in the first image to obtain a face image, the face image comprising the nodes representing the facial key points, edges located between the nodes and representing correlation relations between the facial key points, and correlation weights of the edges; and

performing feature extraction on the face image, to obtain the second feature.

6. The method according to claim 5 , wherein, the obtaining the face image comprises:

determining the facial key points, the correlation relations between the facial key points, and correlation weights between the facial key points according to a plurality of third images, the third images being images identified with expression types; and

taking the facial key points as the nodes, connecting the edges located between the nodes and representing the correlation relations between the facial key points, and using the correlation weights between the facial key points as the correlation weights of the edges, to obtain the face image.

7. The method according to claim 1 , wherein the combining the first feature and the second feature comprises:

performing a weighted summation of the first feature and the second feature based on weights of the first feature and the second feature, and using a weighted summation result as the fused feature; or

concatenating the first feature and the second feature, to obtain the fused feature.

8. The method according to claim 1 , further comprising:

obtaining a training set, training images in the training set being identified with expression types and having the same color encoding type as the first image;

using the training images in the training set as an input of a neural network model, and training the neural network model to obtain an initial neural network model, the initial neural network model being obtained after initializing weights of network layers in the neural network model by using the training images in the training set as the input and using the expression types identified in the training images as an expected output;

obtaining, by using test images in a test set as an input of the initial neural network model, second expression types outputted by the initial neural network model, the test images in the test set being identified with expression types and having the same color encoding type as the first image;

using, in response to a determination that a matching accuracy rate between the second expression types outputted by the initial neural network model and the expression types identified in the test images in the test set is equal to or greater than a target threshold, the initial neural network model as a trained neural network model configured to perform the extracting of the first feature, the extracting of the second feature, the combining the first feature and the second feature, and the determining the first expression type; and

using, in response to a determination that the matching accuracy rate between the second expression types outputted by the initial neural network model and the expression types identified in the test images in the test set is less than the target threshold, the training images in the training set as the input of the initial neural network model, and continuing to train the initial neural network model until the matching accuracy rate between the second expression types outputted by the initial neural network model and the expression types identified in the test images in the test set is equal to or greater than the target threshold.

9. The method according to claim 2 , further comprising:

returning the determined first expression type to a terminal;

obtaining feedback information from the terminal, the feedback information indicating whether the determined first expression type is correct; and

further training, in response to a determination that the feedback information indicates that the determined first expression type is incorrect, the trained neural network model using an image with the same facial expression type or the same background type as the first image.

10. A facial expression recognition apparatus, comprising:

processing circuitry configured to:

extract a first feature of a facial texture from color information of pixels in a first image;

extract a second feature of facial key points from the first image;

combine the first feature and the second feature, to obtain a fused feature; and

determine a first expression type of a face in the first image that corresponds to the fused feature, the first expression type being determined from a plurality of facial expression types according to a correlation between the facial texture of the first feature and the second feature of the facial key points indicated by the fused feature, each of the plurality of facial expression types being indicated by a different correlation between the facial texture of the first feature and the second feature of the facial key points.

11. The apparatus according to claim 10 , wherein

the processing circuitry comprises a trained neural network model is configured to:

extract, through a convolutional neural network (CNN) of the trained neural network model, the first feature representing the facial texture in the first image from the color information of the pixels in the first image;

extract, through a graph neural network (GNN) of the trained neural network model, the second feature representing correlations between the facial key points, the facial key points representing components and/or a facial contour of the face;

perform, through a fusion layer of the trained neural network model, a feature fusion operation on the first feature and the second feature, to obtain the fused feature; and

recognize, through a classification network of the trained neural network model, from the plurality of facial expression types, the first expression type corresponding to the fused feature.

12. The apparatus according to claim 11 , wherein color encoding of the pixels in the first image is used as an input of the CNN, the CNN performing a convolution operation on the color encoding of the pixels in the first image, to obtain the first feature.

13. The apparatus according to claim 12 , wherein the processing circuitry is configured to:

perform, in response to a determination that a position of a reference point in the first image is different from a position of a reference point in an image template, a cropping operation and/or a scaling operation on the first image to obtain a second image, such that a position of the reference point in the second image is the same as the position of the reference point in the image template; and

use the color encoding of the pixels in the second image as the input of the CNN.

14. The apparatus according to claim 11 , wherein the trained neural network model is further configured to:

add positions of the facial key points as nodes in the first image to obtain a face image, the face image comprising the nodes representing the facial key points, edges located between the nodes and representing correlation relations between the facial key points, and correlation weights of the edges; and

perform feature extraction on the face image, to obtain the second feature.

15. The apparatus according to claim 14 , wherein, the trained neural network model is configured to obtain the face image by:

determining the facial key points, the correlation relations between the facial key points, and correlation weights between the facial key points according to a plurality of third images, the third images being images identified with expression types; and

taking the facial key points as the nodes, connecting the edges located between the nodes and representing the correlation relations between the facial key points, and using the correlation

weights between the facial key points as the correlation weights of the edges, to obtain the face image.

16. The apparatus according to claim 10 , wherein the processing circuitry is configured to combine the first feature and the second feature by:

performing a weighted summation of the first feature and the second feature based on weights of the first feature and the second feature, and using a weighted summation result as the fused feature; or

concatenating the first feature and the second feature, to obtain the fused feature.

17. The apparatus according to claim 10 , wherein the processing circuitry is further configured to:

obtain a training set, training images in the training set being identified with expression types and having the same color encoding type as the first image;

use the training images in the training set as an input of a neural network model, and train the neural network model to obtain an initial neural network model, the initial neural network model being obtained after initializing weights of network layers in the neural network model by using the training images in the training set as the input and using the expression types identified in the training images as an expected output;

obtain, by using test images in a test set as an input of the initial neural network model, second expression types outputted by the initial neural network model, the test images in the test set being identified with expression types and having the same color encoding type as the first image;

use, in response to a determination that a matching accuracy rate between the second expression types outputted by the initial neural network model and the expression types identified in the test images in the test set is equal to or greater than a target threshold, the initial neural network model as a trained neural network model configured to perform the extracting of the first feature, the extracting of the second feature, the combining the first feature and the second feature, and the determining the first expression type; and

use, in response to a determination that the matching accuracy rate between the second expression types outputted by the initial neural network model and the expression types identified in the test images in the test set is less than the target threshold, the training images in the training set as the input of the initial neural network model, and continue to train the initial neural network model until the matching accuracy rate between the second expression types outputted by the initial neural network model and the expression types identified in the test images in the test set is equal to or greater than the target threshold.

18. The apparatus according to claim 11 , wherein the processing circuitry is further configured to:

return the determined first expression type to a terminal;

obtain feedback information from the terminal, the feedback information indicating whether the determined first expression type is correct; and

further train, in response to a determination that the feedback information indicates that the determined first expression type is incorrect, the trained neural network model using an image with the same facial expression type or the same background type as the first image.

19. A non-transitory computer-readable storage medium, storing computer-readable instructions thereon, which, when executed by a processor, cause the processor to perform a facial expression recognition method comprising:

extracting a first feature of a facial texture from color information of pixels in a first image;

extracting a second feature of facial key points from the first image;

combining the first feature and the second feature, to obtain a fused feature; and

determining a first expression type of a face in the first image that corresponds to the fused feature, the first expression type being determined from a plurality of facial expression types according to a correlation between the facial texture of the first feature and the second feature of the facial key points indicated by the fused feature, each of the plurality of facial expression types being indicated by a different correlation between the facial texture of the first feature and the second feature of the facial key points.

20. The method according to claim 1 , wherein the first feature is of a facial expression and the second feature is of at least one of a facial component or facial contour.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2021
From: FAN, YANBO; ZHANG, YONG; LI, LE; WU, BAOYUAN; LI, ZHIFENG; LIU, WEI
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 057919/0675 →
Priority Claims (1)
CN 201910478195.3 · Jun 3, 2019 · national
Continuity (2)
Continuation PCTCN2020092593 · May 27, 2020
Related Publication 20210406525A1 · Dec 30, 2021
References Cited (37)
US 20080260212A1 · Moskal · 2008 [cited by examiner]
US 20170116467A1 · Li et al. · 2017 [cited by applicant]
US 20170262472A1 · Goldenberg · 2017 [cited by examiner]
US 20180144185A1 · Yoo · 2018 [cited by examiner]
US 20180211102A1 · Alsmadi · 2018 [cited by examiner]
CN 101620669A · 2010 [cited by applicant]
CN 106256469A · 2016 [cited by applicant]
CN 103824059B · 2017 [cited by examiner]
CN 107169413A · 2017 [cited by applicant]
CN 107358169A · 2017 [cited by applicant]
CN 107657204A · 2018 [cited by applicant]
CN 108038456A · 2018 [cited by examiner]
CN 108090460A · 2018 [cited by applicant]
CN 108256469A · 2018 [cited by examiner]
CN 109117795A · 2019 [cited by examiner]
CN 109446980A · 2019 [cited by applicant]
CN 109684911A · 2019 [cited by applicant]
CN 109711378A · 2019 [cited by applicant]
CN 109786840A · 2019 [cited by applicant]
CN 109815924A · 2019 [cited by examiner]
CN 110245573A · 2019 [cited by examiner]
CN 110263681A · 2019 [cited by applicant]
CN 108268838B · 2020 [cited by examiner]
WO WO2017045157A1 · 2017 [cited by applicant]
Chinese Office Action issued Nov. 16, 2020 in Chinese Application No. 201910478195.3, with English translation, 19 pgs. [cited by applicant]
Chinese Office Action issued Feb. 22, 2021 in Chinese Application No. 201910478195.3, with English translation, 15 pgs. [cited by applicant]
International Search Report Issued Sep. 3, 2020 in International No. PCT/CN2020/092593, with English translation, 11 pgs. [cited by applicant]
Zhang et al., Facial Landmark Detection by Deep Multi-task Learning, ECCV 2014, Part VI, LNCS 8694, Dept. of Information Engineering, The Chinese University of Hong Kong, Hong Kong, China pp. 94-108, 2014, (18 pgs). [cited by applicant]
Yon et al., Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition, The Thirty-Second AAAI Conference on Artificial Intelligence (AAAI-18), Association for the Advancement of Artificial Inte… [cited by applicant]
Zhang, et al., Multimodal learning for facial expression recognition, Pattern Recognition 48 (2015 )3191-3202.2015 Elsevier Ltd. journalhomepage: www.elsevier.com/locate/pr http://dk.doi.org/10.1016/j.pateog.2015.04.012… [cited by applicant]
Jung, et al., Joint Fine-Tuning in Deep Neural Network for Facial Expression Recognition, CVF, This ICCV paper is the Open Access version, provided by the Computer Vision Foundation, Except for this watermark, it is ide… [cited by applicant]
He. et al. Deep Residual Learning for Image Recognition, CVF, This ICCV paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the version available o… [cited by applicant]
Khan, Facial Expression Recognition using Facial Landmark Detection and Feature Extraction via Neural Networks, arXIV-1812.04510v3 [cs.CV] Jul. 15, 2020. [cited by applicant]
Zhang, et al., Facial Expression Recognition Based on Deep Evolutional Spatial-Temporal Networks, IEEE Transactions on Image Processing, vol. 26, No. 9, Sep. 2017, (11 pgs). [cited by applicant]
Day, Exploiting Facial Landmarks for Emotion Recognition in the Wild, Department of Electronics, University of York, UK, This paper was originally accepted to the ACM International Conference on Multimodal Interaction (… [cited by applicant]
Li et al., Deep Facial Expression Recognition: A Survey, arXiv:1804.08848v2 [ cs.CV] Oct. 22, 2018, The authors are with the Pattern Recognition and Intelligent System Laboratory, School of Information and Communication… [cited by applicant]
Mollahosseini, et al. , AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild, IEEE Transactions on Affective Computing, airXiv:1708.03985v4 [cs.CV], Oct. 9, 2017, Authors are with the … [cited by applicant]