IP Library › Granted Patent US 12,731,285
Granted Patent B2
US 12,731,285 · App. 18/834,647 · Granted Sep 8, 2026

Face pose estimation method and apparatus, electronic device, and storage medium

Inventors: Zhanbo Yang (Shenzhen, CN); Zeyuan Huang (Shenzhen, CN); Xiaoting Qi (Shenzhen, CN); Zhao Jiang (Shenzhen, CN)
Assignee: SHENZHEN XUMI YUNTU SPACE TECHNOLOGY CO., LTD.
G06T7/73G06T7/10G06T2207/20081G06T2207/20132G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,285
App. No.
18/834,647
Granted
Sep 8, 2026
Kind
B2
Abstract

The provided are a face pose estimation method and apparatus, an electronic device, and a storage medium. The method includes: obtaining a target image containing face information, and inputting the target image into a pre-constructed pose estimation model; performing feature extraction on the target image by using a shallow densely connected layer to obtain a plurality of first feature maps containing shallow feature information; performing an information fusion operation on the plurality of first feature maps by using a deep feature multiplexing layer to obtain a second feature map, such that deep feature information is fused into the shallow feature information; and extracting face pose information in the second feature map by using an attention layer to obtain a third feature map containing the face pose information, performing prediction on the third feature map by using a classifier, and determining a face pose in the target image.

Claims (22)

1 . A face pose estimation method, comprising:

obtaining a target image containing face information, and inputting the target image into a pre-constructed pose estimation model;

in the pre-constructed pose estimation model, performing feature extraction on the target image by using a shallow densely connected layer to obtain a plurality of first feature maps containing shallow feature information;

taking the plurality of first feature maps as input of a deep feature multiplexing layer, and performing an information fusion operation on the plurality of first feature maps by using the deep feature multiplexing layer to obtain a second feature map, wherein deep feature information is fused into the shallow feature information; and

extracting face pose information in the second feature map by using an attention layer to obtain a third feature map containing the face pose information, performing prediction on the third feature map by using a classifier to obtain a face pose prediction result corresponding to the third feature map, and determining a face pose in the target image according to the face pose prediction result.

2 . The face pose estimation method according to claim 1 , wherein the pose estimation model is constructed by:

acquiring an original image containing face information, detecting the original image by using a face detection model to obtain a face image and a face box corresponding to the original image, acquiring face pose information in the original image, and generating a first data set by using the face image, position coordinates of the face box and the face pose information;

cutting the original image in a preset cutting mode based on the original image and the position coordinates of the face box to obtain a cut face image, and generating a second data set by using the cut face image, the position coordinates of the face box and the face pose information; and

combining the first data set and the second data set to obtain a training set, and training the pose estimation model by using the training set to obtain a trained pose estimation model.

3 . The face pose estimation method according to claim 2 , wherein the training set comprises a face image and annotation information, the annotation information is configured as a label during model training, and the annotation information comprises a plurality of annotation points corresponding to the face box and a plurality of pose angles;

the plurality of annotation points of the face box comprise coordinates of any corner point corresponding to the face box as well as width and height of the face box, and the plurality of pose angles comprise a pitch angle, a yaw angle and a roll angle.

4 . The face pose estimation method according to claim 1 , wherein a step of, in the pre-constructed pose estimation model, performing the feature extraction on the target image by using the shallow densely connected layer to obtain the plurality of first feature maps containing the shallow feature information comprises:

the shallow densely connected layer comprising a plurality of convolution modules which are connected in sequence, sequentially performing a convolution operation on feature maps input to the plurality of convolution modules using the plurality of convolution modules, taking output of each of the plurality of convolution modules as input of a next convolution module, the input of each of the plurality of convolution modules further comprising the output of a previous convolution module, and taking the output of the last plurality of convolution modules in the shallow densely connected layer as the plurality of first feature maps.

5 . The face pose estimation method according to claim 1 , wherein a step of performing the information fusion operation on the plurality of first feature maps by using the deep feature multiplexing layer to obtain the second feature map, wherein the deep feature information is fused into the shallow feature information comprises:

the deep feature multiplexing layer comprising convolution modules with a number corresponding to the number of the plurality of first feature maps, performing convolution transformation on the plurality of first feature maps to obtain the second feature map by the convolution modules of the deep feature multiplexing layer, wherein the deep feature information is fused into the second feature map containing the shallow feature information; and performing global average pooling on the second feature map to obtain the corresponding second feature map after the global average pooling.

6 . The face pose estimation method according to claim 1 , wherein the attention layer comprises a squeeze-and-excitation (SE) attention module and a feature transformation module, and a step of extracting the face pose information in the second feature map by using the attention layer to obtain the third feature map containing the face pose information comprises:

performing weight calculation on a feature channel in the second feature map by the SE attention module, and weighting the feature channel according to a channel weight to obtain a weighted second feature map; and

performing feature extraction on the weighted second feature map by using the feature transformation module to obtain the third feature map containing effective feature information, the effective feature information containing the face pose information.

7 . The face pose estimation method according to claim 1 , wherein a step of performing the prediction on the third feature map by using the classifier to obtain the face pose prediction result corresponding to the third feature map, and determining the face pose in the target image according to the face pose prediction result comprises:

each face pose corresponding to a plurality of third feature maps, each of the plurality of third feature maps corresponding to a plurality of classifiers, and each of the plurality of classifiers being configured to predict a plurality of angle values according to the plurality of third feature maps, calculating a pose angle predicted by each of the plurality of classifiers according to the plurality of angle values, summing the pose angles predicted by all the plurality of classifiers to obtain the pose angle corresponding to each face pose, and taking the pose angles corresponding to three face poses as an estimation result of the face pose in the target image.

8 . An electronic device, comprising a memory, a processor and a computer program stored on the memory and runnable on the processor, wherein the processor, when executing the computer program, implements the face pose estimation method according to claim 1 .

9 . A non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the face pose estimation method according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2024
From: YANG, ZHANBO; HUANG, ZEYUAN; QI, XIAOTING; JIANG, ZHAO
To: SHENZHEN XUMI YUNTU SPACE TECHNOLOGY CO., LTD.
Reel/Frame 068131/0398 →
Priority Claims (1)
CN 202210129785.7 · Feb 11, 2022 · national
Continuity (1)
Related Publication 20250037305A1 · Jan 30, 2025
References Cited (23)
US 6959109B2 · Moustafa · 2005 [cited by examiner]
US 10467458B2 · Wang · 2019 [cited by examiner]
US 12205317B2 · Fu · 2025 [cited by examiner]
US 12354405B1 · Liu · 2025 [cited by examiner]
US 12361584B2 · Chen · 2025 [cited by examiner]
US 20080152218A1 · Okada · 2008 [cited by examiner]
US 20150347822A1 · Zhou · 2015 [cited by examiner]
US 20180150681A1 · Wang · 2018 [cited by examiner]
US 20190026538A1 · Wang · 2019 [cited by examiner]
US 20240338974A1 · Lee · 2024 [cited by examiner]
CN 108256454A · 2018 [cited by applicant]
CN 110837773A · 2020 [cited by applicant]
CN 108256454B · 2020 [cited by applicant]
CN 112766186A · 2021 [cited by applicant]
CN 113901884A · 2022 [cited by applicant]
CN 114519881A · 2022 [cited by applicant]
JP 2020113000A · 2020 [cited by applicant]
Fang et al., MR-CapsNet: A Deep Learning Algorithm for Image-Based Head Pose Estimation on CapsNet, IEEE Access, Digital Object identifier 10.1109/ACCESS.2011.3119615, pp. 141245-141257 . (Year: 2021). [cited by examiner]
Shi et al. Human Pose Estimation Based on the Multistage Learning and Dense Connection, 2020 IEEE, ISBN (Electronic): 978-0-7381-0545-1, 2020 13th International Congress on Image and Signal Processing, BioMedical Engine… [cited by examiner]
Chaoqun Hong, et al., Multimodal Face-Pose Estimation with Multitask Manifold Deep Learning, IEEE Transactions on Industrial Informatics, 2019, pp. 3952-3961, vol. 15, No. 7. [cited by applicant]
Zhansheng Xiong, et al., Deep Spatial-Temporal Field for Human Head Orientation Estimation, (Eds.): ICONIP 2019, LNCS 11955, 2019, pp. 499-509. [cited by applicant]
Heng Song, et al., An multi-task head pose estimation algorithm, 2021 5th Asian Conference on Artificial Intelligence Technology (ACAIT), 2021, pp. 174-181. [cited by applicant]
Bharindra Kamanditya, et al., Convolution Neural Network for Pose Estimation of Noisy Three-Dimensional Face Images, 2018 IEEE 5th International Conference on Engineering Technologies & Applied Sciences, 2018. [cited by applicant]