IP Library Granted Patent US 12,373,983
Granted Patent B2
US 12,373,983 · App. 18/518,175 · Granted Jul 29, 2025

Method and system for monocular depth estimation of persons

Inventors: Colin Brown (Montreal, CA); Louis Harbour (Montreal, CA)
Assignee: Hinge Health, Inc.
G06T7/73G06T3/40G06T7/50G06T2207/20081G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,983
App. No.
18/518,175
Granted
Jul 29, 2025
Kind
B2
Abstract

Systems and methods are provided for estimating the 3D joint location of skeleton joints from an image segment of an object and a 2D joint heatmaps comprising 2D locations of skeleton joints on the image segment. This includes applying the image segment and 2D joint heatmaps to a convolutional neural network containing at least one 3D convolutional layer block, wherein the 2D resolution is reduced at each 3D convolutional layer and the depth resolution is expanded to produce an estimated depth for each joint. Combining the 2D location of each kind of joint with the estimated depth of the kind of joint generates an estimated 3D joint position of the skeleton joint.

Claims (41)

1. A method for estimating a three-dimensional (3D) position of a joint of a person through analysis of an image that includes the joint of the person, the method comprising:

providing (i) at least a portion of the image and (ii) a two-dimensional (2D) heatmap that indicates a 2D location of the joint, as input, to a neural network that produces a one-dimensional (1D) heatmap that indicates an estimated depth range for the joint and then produces an estimated depth of the joint by aggregating confidence values at different depths within the estimated depth range,

wherein values in the estimated depth range are representative of confidence of the neural network in the joint being at that depth; and

establishing the 3D position of the joint by combining the 2D location indicated by the 2D heatmap with the estimated depth produced by the neural network.

2. The method of claim 1 ,

wherein the neural network is a convolutional neural network that contains a 3D convolutional layer, and

wherein at the 3D convolutional layer, 2D resolution is reduced while depth resolution is expanded to produce the estimated depth.

3. The method of claim 1 , wherein the joint is one of multiple joints for which estimated depths are produced by the neural network.

4. The method of claim 3 , wherein each joint of the multiple joints is associated with a corresponding one of multiple 2D heatmaps that are provided to the neural network as input.

5. The method of claim 3 , further comprising:

causing display of a visual representation of the multiple joints in 3D space.

6. The method of claim 5 , wherein the visual representation is a skeleton.

7. The method of claim 1 , wherein the 1D heatmap visually represents depth relative to a fixed point.

8. A non-transitory medium with instructions stored thereon that, when executed by a processor of a computing system, cause the computing system to perform operations comprising:

providing, to a neural network as input, (i) an image that includes joints of a person and (ii) a heatmap that indicates two-dimensional (2D) locations of the joints, so as to produce estimated depths of the joints,

wherein the neural network is trained using a set of images of persons in different poses, settings, and lighting conditions, and

wherein each of the images in the set is associated with at least one label indicating three-dimensional (3D) position of at least one joint;

determining 3D positions of the joints based on the 2D locations indicated by the heatmap and the estimated depths produced by the neural network; and

rendering a visual representation of the joints in 3D space.

9. The non-transitory medium of claim 8 , wherein the visual representation is rendered on an interface that includes the image and the heatmap.

10. The non-transitory medium of claim 8 , wherein the heatmap is representative of a combination of separate heatmaps, each of which is associated with a corresponding one of the joints.

11. The non-transitory medium of claim 8 , wherein the estimated depths are represented by heatmaps, each of which indicates a depth range for a corresponding one of the joints.

12. The non-transitory medium of claim 11 ,

wherein for each of the heatmaps, values within the depth range are representative of confidence of the neural network in the corresponding joint being at that depth, and

wherein for each of the joints, the estimated depth is derived by aggregating confidence values at different depths within the corresponding depth range.

13. The non-transitory medium of claim 11 , wherein the estimated depths are computed from the heatmaps using an argmax function.

14. A computing system comprising:

a processor; and

a memory with instructions that, when executed by the processor, cause the computing system to:

obtain images that include a person and that are arranged in temporal order; and

for each of the images,

obtain a two-dimensional (2D) heatmap that indicates 2D locations of joints, if any, that are visible in that image,

provide the 2D heatmap, as input, to a neural network that produces, for each joint that is visible in that image, a one-dimensional (1D) heatmap that indicates an estimated depth range for that joint and then produces an estimated depth for each joint that is visible in that image that joint by aggregating confidence values at different depths within the estimated depth range, and

determine a three-dimensional (3D) location of each joint that is visible in that image based on the estimated depth produced by the neural network.

15. The computing system of claim 14 , wherein the images are streamed to the computing system from a source external to the computing system.

16. The computing system of claim 14 , wherein the instructions further cause the computing system to:

for each of the images,

render, in 3D space, a visual representation of the joints visible in that image.

17. The computing system of claim 14 , wherein the neural network includes a 3D convolutional layer at which 2D resolution is reduced while depth resolution is expanded.

18. The computing system of claim 17 , wherein the 3D convolutional layer includes at least one 3D convolution, at least one rectified linear unit (ReLU), a max pooling layer, and a reshape layer.

19. The computing system of claim 17 , wherein the 3D convolutional layer is one of multiple 3D convolutional layers included in the neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2023
From: BROWN, COLIN; HARBOUR, LOUIS
To: HINGE HEALTH, INC.
Reel/Frame 065650/0679 →
Priority Claims (1)
CA 3046612 · Jun 14, 2019 · national
Continuity (4)
Continuation 17804909 · Jun 1, 2022
Continuation 17644221 · Dec 14, 2021
Continuation PCTIB2020052936 · Mar 27, 2020
Related Publication 20240087161A1 · Mar 14, 2024
References Cited (21)
US 10679046B1 · Black et al. · 2020 [cited by applicant]
US 11004230B2 · Pollefeys et al. · 2021 [cited by applicant]
US 11568533B2 · Anssari Moin et al. · 2023 [cited by applicant]
US 20190209046A1 · Addison · 2019 [cited by examiner]
US 20190278983A1 · Iqbal · 2019 [cited by examiner]
US 20200125877A1 · Phan · 2020 [cited by examiner]
US 20200160535A1 · Ali Akbarian · 2020 [cited by examiner]
US 20200273192A1 · Cheng · 2020 [cited by examiner]
US 20200302634A1 · Pollefeys et al. · 2020 [cited by applicant]
US 20220292714A1 · Brown et al. · 2022 [cited by applicant]
CN 108549876A · 2018 [cited by applicant]
WO 2017133009A1 · 2017 [cited by applicant]
WO 2018087933A1 · 2018 [cited by applicant]
International Search Report and Written Opinion mailed Jun. 23, 2020 for International Patent Application No. PCT/IB2020/052936 (7 pages). [cited by applicant]
Mazhar, Osama , et al., “Towards Real-time Physical Human-Robot Interaction using Skeleton Information and Hand Gestures”, Azhar Osama et al: “Towards Real-Time Physical Human-Robot Interaction Using Skeleton Informatio… [cited by applicant]
Tome, Denis , et al., “Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image”, Denis Tome et al: “Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image”, arxiv.org, Cornell Un… [cited by applicant]
Xu, Weipeng , et al., “Mo2Cap2: Real-time Mobile 3D Motion Capture with a Cap-mounted Fisheye Camera”, Xu Weipeng et al: “Mo2Cap2: Real-time Mobile 3D Motion Capture with a Cap-mounted Fisheye Camera”, IEEE Transactions… [cited by applicant]
Cao, Zhe et al., “Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”, Zhe Cao et al: “Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”, arxiv.org, Cornell University Library, 201 Ol… [cited by applicant]
Zhou, Xingyi , et al., “Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised Approach”, Zhou Xingyi et al: “Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised Approach”, 2017 IEEE Internation… [cited by applicant]
Muraoka, Taro , et al., “Motion Capturing System from Video Images”, Technical Report of the Institute of Electronics, Information and Communication Engineers, vol. 101, No. 737, Mar. 2002, pp. 97-104. [cited by applicant]
Popa, Alin-Ionut , et al., “Deep Multitask Architecture for Integrated 2D and 3D Human Sensing”, 2017 IEEE Conference on Computer Vision and Pattern Recognition, Jul. 21, 2017, pp. 4714-4723. [cited by applicant]