IP Library › Granted Patent US 12,205,317
Granted Patent B2
US 12,205,317 · App. 17/759,939 · Granted Jan 21, 2025

Light-weight pose estimation network with multi-scale heatmap fusion

Inventors: Yun Fu (Newton, MA); Songyao Jiang (Everett, MA); Bin Sun (Everett, MA)
Assignee: Northeastern University
G06T7/70G06T7/74G06V10/454G06V10/7715G06V10/82G06V20/64G06V40/10G06V40/103G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,317
App. No.
17/759,939
Granted
Jan 21, 2025
Kind
B2
Abstract

Embodiments identify joints of a multi-limb body in an image. One such embodiment unifies depth of a plurality of multi-scale feature maps generated from an image of a multi-limb body to create a plurality of feature maps each having a same depth. In turn, for each of the plurality of feature maps having the same depth, an initial indication of one or more joints in the image is generated. The one or more joints are located at an interconnection of a limb to the multi-limb body or at an interconnection of a limb to another limb. To continue, a final indication of the one or more joints in the image is generated using each generated initial indication of the one or more joints.

Claims (50)

1. A computer-implemented method of identifying joints of a multi-limb body in an image, the method comprising:

creating a plurality of feature maps each having a same depth by unifying depth of a plurality of feature maps generated from an image of a multi-limb body, wherein each feature map of the plurality of feature maps generated from the image has a respective scale;

for each of the plurality of feature maps having the same depth, generating an initial indication of one or more joints in the image, the one or more joints being located at an interconnection of a limb to the multi-limb body or at an interconnection of a limb to another limb; and

generating a final indication of the one or more joints in the image using each generated initial indication of the one or more joints.

2. The computer-implemented method of claim 1 further comprising:

from the generated final indication of the one or more joints in the image, generating an indication of one or more limbs in the image.

3. The computer-implemented method of claim 2 further comprising:

generating an indication of pose using the generated final indication of the one or more joints in the image and the generated indication of the one or more limbs in the image.

4. The computer-implemented method of claim 1 wherein generating the final indication of the one or more joints in the image comprises:

upsampling at least one initial indication of the one or more joints in the image to have a scale equivalent to a scale of a given initial indication of the one or more joints with a largest scale; and

adding together (i) the upsampled at least one initial indication of the one or more joints and (ii) the given initial indication of the one or more joints with the largest scale, to generate the final indication of the one or more joints in the image.

5. The method of claim 1 wherein unifying depth of the plurality of feature maps generated from the image comprises:

applying a respective convolutional layer to each of the plurality of feature maps generated from the image to create the plurality of feature maps each having the same depth.

6. The method of claim 1 wherein generating the initial indication of the one or more joints in the image for each of the plurality of feature maps having the same depth comprises:

applying a heatmap estimating layer to each of the plurality of feature maps having the same depth to generate each initial indication of the one or more joints in the image.

7. The method of claim 6 wherein the heatmap estimating layer is composed of a convolutional neural network.

8. The method of claim 7 wherein the image is a training image and the method further comprises:

training the convolutional neural network by comparing each generated initial indication of the one or more joints in the image to a respective ground-truth indication of the one or more joints in the training image to determine losses and back propagating the losses to the convolutional neural network, wherein each respective ground-truth indication of the one or more joints corresponds to a respective scale of a given feature map of the plurality of feature maps having the same depth.

9. The method of claim 1 further comprising:

generating the plurality of feature maps generated from the image by processing the image using a backbone neural network.

10. The method of claim 9 wherein processing the image using the backbone neural network comprises:

performing multi-scale feature extraction and multi-scale feature fusion to generate the plurality of feature maps from the image.

11. A computer system for identifying joints of a multi-limb body in an image, the computer system comprising:

a processor; and

a memory with computer code instructions stored thereon, the processor and the memory, with the computer code instructions, being configured to cause the system to:

create a plurality of feature maps each having a same depth by unifying depth of a plurality of feature maps generated from an image of a multi-limb body, wherein each feature map of the plurality of feature maps generated from the image has a respective scale;

for each of the plurality of feature maps having the same depth, generate an initial indication of one or more joints in the image, the one or more joints being located at an interconnection of a limb to the multi-limb body or at an interconnection of a limb to another limb; and

generate a final indication of the one or more joints in the image using each generated initial indication of the one or more joints.

12. The system of claim 11 wherein the processor and the memory, with the computer code instructions, are further configured to cause the system to:

from the generated final indication of the one or more joints in the image, generate an indication of one or more limbs in the image.

13. The system of claim 12 wherein the processor and the memory, with the computer code instructions, are further configured to cause the system to:

generate an indication of pose using the generated final indication of the one or more joints in the image and the generated indication of the one or more limbs in the image.

14. The system of claim 11 wherein, in generating the final indication of the one or more joints in the image, the processor and the memory, with the computer code instructions, are further configured to cause the system to:

upsample at least one initial indication of the one or more joints in the image to have a scale equivalent to a scale of a given initial indication of the one or more joints with a largest scale; and

add together (i) the upsampled at least one initial indication of the one or more joints and (ii) the given initial indication of the one or more joints with the largest scale, to generate the final indication of the one or more joints in the image.

15. The system of claim 11 wherein, in unifying depth of the plurality of feature maps generated from the image, the processor and the memory, with the computer code instructions, are further configured to cause the system to:

apply a respective convolutional layer to each of the plurality of feature maps generated from the image to create the plurality of feature maps each having the same depth.

16. The system of claim 11 wherein, in generating the initial indication of the one or more joints in the image for each of the plurality of feature maps having the same depth, the processor and the memory, with the computer code instructions, are further configured to cause the system to:

apply a heatmap estimating layer, composed of a convolutional neural network, to each of the plurality of feature maps having the same depth to generate each initial indication of the one or more joints in the image.

17. The system of claim 16 wherein the image is a training image and the processor and the memory, with the computer code instructions, are further configured to cause the system to:

train the convolutional neural network by comparing each generated initial indication of the one or more joints in the image to a respective ground-truth indication of the one or more joints in the training image to determine losses and back propagating the losses to the convolutional neural network, wherein each respective ground-truth indication of the one or more joints corresponds to a respective scale of a given feature map of the plurality of feature maps having the same depth.

18. The system of claim 11 wherein the processor and the memory, with the computer code instructions, are further configured to cause the system to:

generate the plurality of feature maps generated from the image by processing the image using a backbone neural network.

19. The system of claim 18 wherein, in processing the image using the backbone neural network, the processor and the memory, with the computer code instructions, are further configured to cause the system to:

perform multi-scale feature extraction and multi-scale feature fusion to generate the plurality of feature maps from the image.

20. A computer program product for identifying joints of a multi-limb body in an image, the computer program product comprising:

one or more non-transitory computer-readable storage devices and program instructions stored on at least one of the one or more storage devices, the program instructions, when loaded and executed by a processor, cause an apparatus associated with the processor to:

create a plurality of feature maps each having a same depth by unifying depth of a plurality of feature maps generated from an image of a multi-limb body, wherein each feature map of the plurality of feature maps generated from the image has a respective scale;

for each of the plurality of feature maps having the same depth, generate an initial indication of one or more joints in the image, the one or more joints being located at an interconnection of a limb to the multi-limb body or at an interconnection of a limb to another limb; and

generate a final indication of the one or more joints in the image using each generated initial indication of the one or more joints.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME AND EXECUTION DATES PREVIOUSLY RECORDED ON REEL 060741 FRAME 0351. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 10, 2022
From: FU, YUN; JIANG, SONGYAO; SUN, BIN
To: NORTHEASTERN UNIVERSITY
Reel/Frame 061141/0790 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2022
From: FU, YUN; JIANG, SONGYAO; SUN, BIN
To: NORTHEASTER UNIVERSITY
Reel/Frame 060741/0351 →
Continuity (2)
Provisional Application 62976099 · Feb 13, 2020
Related Publication 20230126178A1 · Apr 27, 2023
References Cited (28)
US 20100034457A1 · Berliner · 2010 [cited by examiner]
US 20190357615A1 · Koh · 2019 [cited by examiner]
CN 109271933A · 2019 [cited by applicant]
CN 109271933B · 2021 [cited by examiner]
WO 2021163103A1 · 2021 [cited by applicant]
Cluster-wise learning network for multi-person pose estimation. Zhao et al. (Year: 2020). [cited by examiner]
International Search Report and Written Opinion for PCT/US2021/017341 dated May 3, 2021 titled “Light-Weight Pose Estimation Network With Multi-Scale Heatmap Fusion”. [cited by applicant]
Zhao, Y. et al. “Cluster-wise learning network for multi-person pose estimation”, Pattern Recognition, Elsevier, GB, vol. 98, Oct. 3, 2019. [cited by applicant]
Gao, B. et al. “A Lightweight Network Pose Estimation”, Pattern Recognition. Image Analysis, Allen, Press, Lawrence, KS, US, vol. 29, No. 4, Oct. 1, 2019. [cited by applicant]
Zhu, X., et al. “Face alignment across large poses: A 3d solution.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2016. [cited by applicant]
Bulat, A., et al. “How far are we from solving the 2D & 3D Face Alignment problem? (and a dataset of 230,000 3D landmarks)”, Proceedings of the IEEE International Conference on Computer Vision. 2017. [cited by applicant]
Newell, A., et al. “Associative embedding: End-to-end learning for joint detection and grouping.” Advances in neural Information processing systems 30 (2017). [cited by applicant]
Pishchulin, L., et al. “Deepcut: Joint subset partition and labeling for multi person pose estimation.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2016. [cited by applicant]
Newell, A., et al. “Stacked hourglass networks for human pose estimation.” European conference on computer vision. Springer, Cham, 2016. [cited by applicant]
Sun, Ke, et al. “Deep high-resolution representation learning for human pose estimation.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019. [cited by applicant]
Sun, Bin, et al. “LRPRNet: Lightweight Deep Network by Low-Rank Pointwise Residual Convolution.” IEEE Transactions on Neural Networks and Learning Systems (2021). [cited by applicant]
He, K., et al. “Mask R-CNN”, Proceedings of the IEEE international conference on computer vision. 2017. [cited by applicant]
Tan, Mi., et al. “Mixconv: Mixed depthwise convolutional kernels.” arXiv preprint arXiv:1907.09595 (2019). [cited by applicant]
Howard, A.G., et al. “Mobilenets: Efficient convolutional neural networks for mobile vision applications.” arXiv preprint arXiv:1704.04861 (2017). [cited by applicant]
Sandler, M., et al. “MobileNetV2: Inverted Residuals and Linear Bottlenecks”, Proceedings of the IEEE conference on computer vision and pattern recognition. 2018. [cited by applicant]
Howard, A. et al. “Searching for mobilenetv3.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019. [cited by applicant]
Cao, Z., et al. “Realtime multi-person 2d pose estimation using part affinity fields.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2017. [cited by applicant]
Zhang, X., et al. “ShuffleNet: An extremely efficient convolutional neural network for mobile devices.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2018. [cited by applicant]
Ma, N., et al. “Shufflenet v2: Practical guidelines for efficient cnn architecture design.” Proceedings of the European conference on computer vision (ECCV). 2018. [cited by applicant]
Xiao, B. et al. “Simple baselines for human pose estimation and tracking.” Proceedings of the European conference on computer vision (ECCV). 2018. [cited by applicant]
Nie, X., et al. “Single-stage multi-person pose machines.” Proceedings of the IEEE/CVF international conference on computer vision. 2019. [cited by applicant]
Papandreou, G., et al. “Towards accurate multi-person pose estimation in the wild.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2017. [cited by applicant]
International Preliminary Report on Patentability for International Application No. PCT/US2021/017341 dated Aug. 25, 2022. [cited by applicant]
Cited By (1)
US 12,731,285