IP Library › Granted Patent US 12,525,005
Granted Patent B2
US 12,525,005 · App. 18/019,450 · Granted Jan 13, 2026

Method and system of multiple facial attributes recognition using highly efficient neural networks

Inventors: Ping Hu (Beijing, CN); Anbang Yao (Beijing, CN); Xiaolong Liu (Beijing, CN); Yurong Chen (Beijing, CN); Dongqi Cai (Beijing, CN)
Assignee: Intel Corporation
G06V10/82G06N3/0464G06V10/7715G06V40/171
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,525,005
App. No.
18/019,450
Granted
Jan 13, 2026
Kind
B2
Abstract

A method and system of multiple facial attributes recognition using highly efficient neural networks.

Claims (29)

1 . At least one non-transitory machine-readable medium comprising a plurality of instructions that in response to being executed on a computing device, cause the computing device to operate by:

obtaining at least one image with at least one facial region; and

detecting multiple facial attributes on the at least one facial region using a neural network with at least two blocks each having at least one network layer,

wherein one or more individual blocks have at least one individual layer with multiple kernels with varying sizes,

wherein one or more of the individual blocks perform at least one per-block fractional attention operation, and

wherein the detecting includes having at least one block of the neural network perform both channel expansion and then contraction transformation and spatial contraction and then expansion transformation while using at least one of the individual blocks.

2 . The medium of claim 1 , wherein the individual blocks are bottleneck blocks.

3 . The medium of claim 1 , wherein at least one of the kernels is dilated to fit a larger size than an initial size of the kernel.

4 . The medium of claim 1 , wherein the detecting comprises grouping channels into groups and providing at least two different kernels among the groups.

5 . The medium of claim 4 , wherein each group has a kernel of a different size.

6 . The medium of claim 4 , wherein the channels are grouped into 3 or 4 groups each with a different kernel.

7 . The medium of claim 4 , wherein the kernels comprise a 3×3 kernel, 5×5 kernel, and 3×3 kernel dilated to a 7×7 area by a dilation rate of 3.

8 . The medium of claim 4 , wherein the fractional attention operation comprises channel attention, spatial attention, or both.

9 . A computer-implemented neural network comprising:

a plurality of blocks operated by at least one processor and comprising at least one bottleneck block receiving block input features of image data and having at least one convolutional layer generating block output features that represent multiple attributes, wherein the at least one convolutional layer having multiple kernels with varying sizes applied to the input features, at least one block of the plurality of blocks to detect multiple facial attributes by performing both channel expansion and then contraction transformation and spatial contraction and then expansion transformation using at least one of individual blocks; and

at least one per-block fractional attention operation using a version of the block input features to generate weights to be applied to the block output features.

10 . The network of claim 9 wherein the individual blocks are bottleneck blocks.

11 . The network of claim 9 , wherein at least one of the kernels is dilated to fit a larger size than an initial size of the kernel.

12 . The network of claim 9 , wherein the fractional attention comprises channel attention, spatial attention, or both.

13 . A method of image processing comprising:

obtaining at least one image with at least one facial region; and

detecting multiple facial attributes on the at least one facial region using a neural network with at least two blocks each having at least one network layer,

wherein one or more of individual blocks have at least one individual layer with multiple kernels with varying sizes,

wherein one or more of the individual blocks perform at least one per-block fractional attention operation, and

wherein the detecting includes having at least one block of the neural network perform both channel expansion and then contraction transformation and spatial contraction and then expansion transformation while using at least one of the individual blocks.

14 . The method of claim 13 , wherein the detecting comprises grouping channels into groups and providing a different kernel for each group.

15 . The method of claim 13 , wherein the fractional attention comprises channel attention, spatial attention, or both.

16 . The method of claim 14 , wherein results of each group are concatenated together to form input channels of a next layer.

17 . The method of claim 13 , wherein a block of the at least two blocks is repeated at least four times.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 2, 2023
From: HU, PING; YAO, ANBANG; LIU, XIAOLONG; CHEN, YURONG; CAI, DONGQI
To: INTEL CORPORATION
Reel/Frame 062577/0075 →
Continuity (1)
Related Publication 20230290134A1 · Sep 14, 2023
References Cited (38)
US 20220277558A1 · Li · 2022 [cited by applicant]
CN 103824054 · 2014 [cited by applicant]
CN 106203395 · 2016 [cited by applicant]
CN 106529402 · 2017 [cited by applicant]
CN 109947960 · 2019 [cited by applicant]
CN 110678873 · 2020 [cited by applicant]
CN 111339818 · 2020 [cited by applicant]
CN 111339818A · 2020 [cited by examiner]
WO 2019183758 · 2019 [cited by applicant]
WO 2021120028A1 · 2021 [cited by applicant]
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, Andrew Rabinovich; Proceedings of the IEEE Conference on Computer Vision and Pattern Recognitio… [cited by examiner]
Jie Hu, Li Shen, Gang Sun; Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7132-7141 (Year: 2018). [cited by examiner]
Chim S, Lee JG, Park HH. Dilated Skip Convolution for Facial Landmark Detection. Sensors (Basel). Dec. 4, 2019;19(24):5350. doi: 10.3390/s19245350. PMID: 31817213; PMCID: PMC6960628. (Year: 2019). [cited by examiner]
International Search Report and Written Opinion for PCT Application No. PCT/CN2020/117788, dated Jun. 23, 2021. [cited by applicant]
Dai, J.F., et al. , “Deformable Convolutional Networks” , arXiv preprint arXiv:1703.06211; 2017. [cited by applicant]
Gao, H., et al., “Deformable kernels: Adapting effective receptive fields for object deformation”, arXiv preprint arXiv:1910.02940; 2019. [cited by applicant]
Gunther, M., et al. , “AFFACT—alignment free facial attribute classification technique”, arXiv preprint arXiv: 1611.06158; 2016. [cited by applicant]
Han, H., et al., “Heterogeneous face attribute estimation: A deep multi-task learning approach”, IEEE TPAMI, 2017. [cited by applicant]
Hand, E.M., et al., “Attributes for improved attributes: A multi-task network utilizing implicit and explicit relationships for facial attribute classification”, AAAI, 2017. [cited by applicant]
He, K.M., et al., “Deep residual learning for image recognition”, arXiv:1512.03385, 2015. [cited by applicant]
Howard, A.G., et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications”, https://arxiv.org/abs/1704.04861, Apr. 17, 2017. [cited by applicant]
Hu, J., et al., “Squeeze-and-Excitation Networks”, arXiv preprint arXiv:1709.01507; 2017. [cited by applicant]
Hu, P., et al., “Learning supervised scoring ensemble for emotion recognition in the wild”, ICMI, 2017. [cited by applicant]
Huang, G., et al., “Densely Connected Convolutional Networks”, arXiv:1608.06993; 2016. [cited by applicant]
Kalayeh, M.M., et al. , “Improving facial attribute prediction using semantic segmentation”, CVPR, 2017. [cited by applicant]
Kang, S., et al., “Face attribute classification using attribute-aware correlation map and gated convolutional neural networks”, ICIP, 2015. [cited by applicant]
Krizhevsky, A., et al., “ImageNet Classification with Deep Convolutional Neural Networks”, In Advances in Neural Information Processing systems (NIPS); pp. 1-9; 2012. [cited by applicant]
Lee, C.Y., et al., “Deeply-supervised nets”, arXiv:1409.5185, 2014. [cited by applicant]
Li, D., et al., “HBONet: Harmonious Bottleneck on Two Orthogonal Dimensions”, ICCV, 2019. [cited by applicant]
Liu, Z., et al., “Deep learning face attributes in the wild”, ICCV, 2015. [cited by applicant]
Ma, N., et al., “Shufflenet v2: Practical guidelines for efficient cnn architecture design”, ECCV, 2018. [cited by applicant]
Rudd, E.M., et al., “Moon: A mixed objective optimization network for the recognition of facial attributes”, ECCV, 2016. [cited by applicant]
Sandler, M., et al., “Mobilenetv2: Inverted residuals and linear bottlenecks”, CVPR (2018). [cited by applicant]
Xiao, T.J., et al., “The Application of Two-Level Attention Models in Deep Convolutional Neural Network for Fine-grained Image Classification”, arXiv preprint arXiv:1411.6447; 2014. [cited by applicant]
Zhang, X., et al., “Shufflenet: An extremely efficient convolutional neural network for mobile devices”, CVPR, 2018. [cited by applicant]
Zhong, Y., et al., “Face attribute prediction using off-the-shelf cnn features”, Proceedings of the IEEE International Conference on Biometrics (ICB), pp. 1-7; IEEE (2016). [cited by applicant]
Zhong, Y., et al., “Leveraging mid-level deep representations for predicting face attributes in the wild”, arXiv:1602.01827, 2016. [cited by applicant]
International Preliminary Report on Patentability for PCT Patent Application No. PCT/CN2020/117788, dated Apr. 6, 2023. [cited by applicant]