IP Library › Granted Patent US 12,354,405
Granted Patent B1
US 12,354,405 · App. 19/061,635 · Granted Jul 8, 2025

Expression recognition method and system based on multi-scale features and spatial attention

Inventors: Zhaowei Liu (Yantai, CN); Haonan Wen (Yantai, CN); Yongchao Song (Yantai, CN); Wenhan Hou (Yantai, CN); Xinxin Zhao (Yantai, CN); Tengjiang Wang (Yantai, CN); Diantong Liu (Yantai, CN); Weiqing Yan (Yantai, CN); Peng Song (Yantai, CN); Anzuo Jiang (Yantai, CN); Hang Su (Yantai, CN)
Assignee: YANTAI UNIVERSITY
G06V40/175G06V10/82G06V40/171
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,405
App. No.
19/061,635
Granted
Jul 8, 2025
Kind
B1
Abstract

The present invention relates to the technical field of expression recognition, and in particular, to an expression recognition method and system based on multi-scale features and spatial attention. The method includes: performing feature extraction on acquired facial image data by using an HNFER neural network model to obtain an original input feature map; performing pooling and concatenation on extracted features based on a CoordAtt attention mechanism to obtain a feature map; performing deep convolution processing on the feature map to obtain an attention map, and then performing element-by-element multiplication to obtain a final feature map; and performing feature transformation and normalization on the final feature map to obtain an expression category probability and output the expression category probability. In the present invention, by integrating scale perception and spatial attention technologies, the model can recognize and classify different emotional states more accurately and maintain high performance even under complex environmental conditions.

Claims (22)

1. An expression recognition method based on multi-scale features and spatial attention, comprising:

acquiring facial image data;

constructing an HNFER neural network model;

performing feature extraction on acquired facial image data by using the HNFER neural network model to obtain an original input feature map;

performing pooling and concatenation on extracted features based on a CoordAtt attention mechanism to obtain a feature map;

performing deep convolution processing on the feature map to obtain an attention map, and then performing element-by-element multiplication to obtain a final feature map; and

performing feature transformation and normalization on the final feature map to obtain an expression category probability and output the expression category probability.

2. The expression recognition method based on multi-scale features and spatial attention according to claim 1 , wherein the performing feature extraction on acquired facial image data by using the HNFER neural network model comprises: performing three stages of feature extraction on the facial image data to transform a facial image with a size of H×W×3 to a facial image with a size of H/4×W/4×256, wherein H is a height, and W is a width.

3. The expression recognition method based on multi-scale features and spatial attention according to claim 2 , wherein the performing pooling and concatenation on extracted features based on a CoordAtt attention mechanism to obtain a feature map comprises: performing horizontal pooling and vertical pooling on the extracted features by the CoordAtt attention mechanism of the HNFER neural network model, respectively, and performing concatenation on pooled features to generate attention weights that act on a height and a width of the original input feature map, respectively.

4. The expression recognition method based on multi-scale features and spatial attention according to claim 3 , wherein the performing deep convolution processing on the feature map to obtain an attention map comprises: generating 512 feature maps from a feature map processed by CoordAtt through a convolutional layer, wherein a size of each feature map is transformed to H/4×W/4×512; and processing the feature maps through spatial weights at different scales based on an SAFM mechanism of the model to obtain convolutionally processed feature maps.

5. The expression recognition method based on multi-scale features and spatial attention according to claim 4 , wherein the performing element-by-element multiplication to obtain a final feature map comprises: restoring the convolutionally processed feature maps to an original resolution through upsampling, and concatenating with a feature map at a first scale, and after concatenating, generating an attention map through feature fusion, and performing element-by-element multiplication on the attention map and the original input feature map to enhance spatial features at multiple scales and obtain the final feature map.

6. The expression recognition method based on multi-scale features and spatial attention according to claim 5 , wherein the performing feature transformation and normalization on the final feature map to obtain an expression category probability comprises: generating 1024 feature maps from the final feature map through a convolutional layer, wherein a size of each feature map is transformed to H/4×W/4×1024; then, sequentially performing feature transformation on the feature maps through two fully connected layers, and outputting 6 nodes through a third fully connected layer to correspond to 6 expressions.

7. The expression recognition method based on multi-scale features and spatial attention according to claim 6 , wherein the performing feature transformation and normalization on the final feature map to obtain an expression category probability further comprises: using a softmax function to convert values of the 6 nodes that are output into probabilities, representing probabilities that images belong to each expression category.

8. An expression recognition system based on multi-scale features and spatial attention, comprising:

a data acquisition module, configured to acquire facial image data;

a modeling module, configured to construct an HNFER neural network model;

a feature extraction module, configured to perform feature extraction on acquired facial image data by using the HNFER neural network model to obtain an original input feature map;

a feature concatenation module, configured to perform pooling and concatenation on extracted features based on a CoordAtt attention mechanism to obtain a feature map;

a convolution module, configured to perform deep convolution processing on the feature map to obtain an attention map, and then perform element-by-element multiplication to obtain a final feature map; and

a calculation module, configured to perform feature transformation and normalization on the final feature map to obtain an expression category probability and output the expression category probability.

9. A non-transitory computer-readable storage medium, having a plurality of instructions stored therein, wherein the instructions are adapted to be loaded by a processor of a terminal device and to execute an expression recognition method based on multi-scale features and spatial attention according to claim 1 .

10. A terminal device, comprising a processor and a computer-readable storage medium, the processor being used for implementing various instructions, and the computer-readable storage medium being used for storing a plurality of instructions, wherein the instructions are adapted to be loaded by the processor and to execute an expression recognition method based on multi-scale features and spatial attention according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 24, 2025
From: LIU, ZHAOWEI; WEN, HAONAN; SONG, YONGCHAO; HOU, WENHAN; ZHAO, XINXIN; WANG, TENGJIANG; LIU, DIANTONG; YAN, WEIQING; SONG, PENG; JIANG, ANZUO; SU, HANG
To: YANTAI UNIVERSITY
Reel/Frame 070310/0246 →
Priority Claims (1)
CN 202410710860.8 · Jun 4, 2024 · national
Continuity (1)
Continuation PCTCN2024135203 · Nov 28, 2024
References Cited (6)
US 20230290134A1 · Hu · 2023 [cited by examiner]
US 20240338974A1 · Lee · 2024 [cited by examiner]
US 20250111696A1 · Jiang · 2025 [cited by examiner]
CN 113781385A · 2021 [cited by applicant]
CN 117058734A · 2023 [cited by applicant]
CN 117275074A · 2023 [cited by applicant]
Cited By (1)
US 12,731,285