IP Library Granted Patent US 12,249,074
Granted Patent B2
US 12,249,074 · App. 17/256,307 · Granted Mar 11, 2025

Human pose analysis system and method

Inventors: Dongwook Cho (Montreal, CA); Maggie Zhang (Montreal, CA); Paul Kruszewski (Montreal, CA)
Assignee: Hinge Health, Inc.
G06T7/11G06F18/214G06N3/045G06V10/454G06V10/764G06V10/82G06V40/10G06V40/103G06V40/171
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,249,074
App. No.
17/256,307
Filed
Dec 28, 2020
Granted
Mar 11, 2025
Kind
B2
Art Unit
2663
USPC
382/103
Abstract

System and method for extracting human pose information from an image, comprising a feature extractor connected to a database, a convolutional neural network (CNN) with a plurality of CNN layers. Said system/method further comprising at least one of the following modules: a 2D body skeleton detector for determining 2D body skeleton information from the human-related image features; a body silhouette detector for determining body silhouette information from the human-related image features; a hand silhouette detector for determining hand silhouette detector from the human-related image features; a hand skeleton detector for determining hand skeleton from the human-related image features; a 3D body skeleton detector for determining 3D body skeleton from the human-related image features; and a facial keypoints detector for determining facial keypoints from the human-related image features.

Claims (46)

1. A system for extracting human pose information from an image, the system comprising:

a feature extractor for extracting features related to human joints and shapes from the image, the feature extractor being connectable to a database comprising a dataset of reference images and provided with a first convolutional neural network (CNN) architecture including a first plurality of CNN layers,

wherein each CNN layer of the first plurality of CNN layers applies a nonlinear activation function to its input data using trained kernel weights, and

wherein each CNN layer of the first plurality of CNN layers produces an output in the form of a tensor, such that the first CNN architecture outputs a plurality of tensors that are from different CNN layers and that collectively preserve the features;

a two-dimensional (2D) body skeleton detector to which the plurality of tensors are provided, as input, to obtain, for each of multiple joints, a heat map that visually indicates a location of that joint,

wherein the 2D body skeleton detector is provided with a second CNN architecture including a second plurality of CNN layers, and wherein the second CNN architecture accepts, as input, the plurality of tensors rather than the image; and

a facial keypoints detector to which at least one of the multiple heat maps produced by the 2D body skeleton detector for the multiple joints is provided, as input, to obtain locations of one or more facial keypoints.

2. The system of claim 1 , wherein the feature extractor comprises:

a low-level feature extractor for extracting low-level features from the image; and

an intermediate feature extractor for extracting intermediate features, the low-level features and the intermediate features forming together the features.

3. The system of claim 1 , wherein at least one of the first CNN architecture and the second CNN architecture comprises a deep CNN architecture.

4. The system of claim 1 , wherein one of the first and second pluralities of CNN layers comprises lightweight layers.

5. A method for extracting human pose information from an image, the method comprising:

receiving an image;

extracting human-related image features from the image using a feature extractor, the feature extractor being connectable to a database comprising a dataset of reference images and provided with a first convolutional neural network (CNN) architecture including a first plurality of CNN layers,

wherein each CNN layer of the first CNN architecture applies a convolutional operation to its input data using trained kernel weights to produce an output in the form of a tensor, and

wherein the first CNN architecture outputs a plurality of tensors that are from different CNN layers and that collectively preserve the human-related image features; and

determining the human pose information by:

providing the plurality of tensors to a two-dimensional (2D) body skeleton detector as input, so as to obtain, for each of multiple joints, a heat map that visually indicates a location of that joint,

wherein the 2D body skeleton is provided with a second CNN architecture including a second plurality of CNN layers, and wherein the second CNN architecture accepts, as input, the plurality of tensors, and

providing at least one of the multiple heat maps produced by the 2D body skeleton detector for the multiple joints to a facial keypoints detector as input, so as to obtain locations of one or more facial keypoints.

6. The method of claim 5 , wherein the feature extractor comprises:

a low-level feature extractor for extracting low-level features from the image; and

an intermediate feature extractor for extracting intermediate features, the low-level features and the intermediate features forming together the human-related image features.

7. The method of claim 5 , wherein at least one of the first CNN architecture and the second CNN architecture comprises a deep CNN architecture.

8. The method of claim 5 , wherein one of the first and second CNN architectures comprises lightweight layers.

9. A method comprising:

applying, to an image, a feature extractor that extracts human-related features in the form of a plurality of tensors,

wherein the feature extractor is provided with a first convolutional neural network (CNN) architecture having a first plurality of CNN layers, and

wherein each CNN layer applies a convolutional operation to its input using trained kernel weights to produce, as output, one of the plurality of tensors; and

determining, based on the plurality of tensors output by different CNN layers in the first CNN architecture, pose information of a human body in the image by:

providing the plurality of tensors to a two-dimensional (2D) body skeleton detector as input, so as to obtain, for each of multiple joints, a heat map that visually indicates a location of that joint, and

providing at least one of the multiple heat maps produced by the 2D body skeleton detector for the multiple joints to a facial keypoints detector as input, so as to obtain locations of one or more facial keypoints of the human body;

wherein the 2D body skeleton detector is provided with a second CNN architecture having a second plurality of CNN layers; and

wherein the facial keypoints detector is provided with a third CNN architecture having a third plurality of CNN layers.

10. The method of claim 9 , wherein the one or more facial keypoints include eyes, eyebrows, ears, nose, upper lip, lower lip, chin, or any combination thereof.

11. The method of claim 9 , wherein said determining further comprises:

providing the plurality of tensors to a hand silhouette detector as input, so as to obtain (i) a first mask for a left hand of the human body and (ii) a second mask for a right hand of the human body,

applying the first and second masks to the image, so as to obtain (i) a first masked image that includes the left hand of the human body and (ii) a second masked image that includes the right hand of the human body, and

providing the first and second masked images to a hand skeleton detector that produces, as output, (i) a first location of the left hand of the human body and (ii) a second location of the right hand of the human body.

12. The method of claim 9 , wherein a first subset of the first plurality of CNN layers is representative of a low-level feature extractor, and wherein a second subset of the first plurality of CNN layers is representative of an intermediate-level feature extractor.

13. The method of claim 12 ,

wherein the low-level feature extractor is configured for extracting low-level features that represent elemental characteristics of local regions in the image, and

wherein the intermediate-level feature extractor is configured for receiving the low-level features and determining intermediate features that correspond to high-level features by correlating the low-level features with shapes and/or relations of body parts.

14. The method of claim 13 , wherein the elemental characteristics include intensities, edges, gradients, curvatures, points, object shapes, or any combination thereof.

15. The system of claim 1 , wherein the nonlinear activation function is a rectified linear unit (ReLU).

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2021
From: WRNCH INC.
To: HINGE HEALTH, INC.
Reel/Frame 058364/0007 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2021
From: CHO, DONGWOOK; ZHANG, MAGGIE; KRUSZEWSKI, PAUL
To: WRNCH INC.
Reel/Frame 056414/0126 →
Continuity (2)
Provisional Application 62691818 · Jun 29, 2018
Related Publication 20210264144A1 · Aug 26, 2021
References Cited (23)
US 8437506B2 · Williams et al. · 2013 [cited by applicant]
US 20150347822A1 · Zhou · 2015 [cited by examiner]
US 20170270593A1 · Sherman · 2017 [cited by examiner]
US 20180018554A1 · Young · 2018 [cited by examiner]
US 20180024641A1 · Mao · 2018 [cited by examiner]
US 20180116620A1 · Chen · 2018 [cited by examiner]
US 20180150681A1 · Wang et al. · 2018 [cited by applicant]
US 20180276454A1 · Han · 2018 [cited by examiner]
US 20190251340A1 · Brown · 2019 [cited by examiner]
CN 105069423 · 2015 [cited by applicant]
CN 104346607 · 2017 [cited by applicant]
CN 104346607B · 2017 [cited by applicant]
JP 2008204384A · 2008 [cited by applicant]
JP 2017157138A · 2017 [cited by applicant]
Zhang et al., “Joint Face Detection and Alignment Using Multitask Cascaded Convolutional Networks,” Oct. 26, 2016, IEEE Signal Processing Letters, vol. 23, pp. 1499-1503. [cited by examiner]
Choi, Jinyoung , et al., “Human Body Orientation Estimation Using Convolutional Neural Network”, arXiv, the U.S., Cornell University; https://arxiv.org/abs/1609.01984, Sep. 7, 2016, pp. 1-5. [cited by applicant]
Kengo , “The accuracy estimation of the doubtful operation recognition technique of the person who utilized the machine learning based on attitude presumption technology”, IEICE Technical Report, Japan, vol. 117, No. 48… [cited by applicant]
Li, Sijin et al., “Heerogeneous Multi-task Learning for Human Pose Estimation with Deep Convolutional Neural Network”, 2014 IEEE Conference On Computer Vision and Pattern Recognition Workshops, IEEE, Jun. 23, 2014, pp. … [cited by applicant]
Ranjan, Rajeev et al., “HyperFace: A Deep Multi-Task Learning Framework for Face Detection, Landmark Localization, Pose Estimation, and Gender Recognition”, Retrieved from the Internet: URL:https://arxiv.org/pdf/1603.01… [cited by applicant]
Toshev, A. et al., “DeepPose: Human Pose stimation via Deep Neural Networks”, 2014 IEEE Conference On Computer Vision and Pattern Recognition, IEEE, Jun. 23, 2014 (Jun. 23, 2014), XP032649134, Jun. 23, 2014, pp. 1653-16… [cited by applicant]
Wang, Keze et al., “Human Pose Estimation from Depth Images via Inference Embedded Multi-task Learning”, Proceedings of The 2017 ACM On Conference On Information and Knowledge Management , CIKM '17, ACM Press, New York,… [cited by applicant]
Wei, S-E et al., “Convolutional Pose Machines”, 2016 IEEE Conference On Computer Vision and Pattern Recognition (CVPR), IEEE, XP033021664, Jun. 27, 2016, pp. 4724-4732. [cited by applicant]
Ranjan, Rajeev , et al., “HyperFace: A Deep Multi-Task Learning Framework for Face Detection, Landmark Localization, Pose Estimation, and Gender Recognition”, IEEE Transactions on Pattern Analysis and Machine Intelligen… [cited by applicant]
Cited By (2)
US 12,469,316 US 12,602,899