IP Library Granted Patent US 12,675,949
Granted Patent B2
US 12,675,949 · App. 18/431,126 · Granted Jul 7, 2026

Three-dimensional mesh generator based on two-dimensional image

Inventors: Sohail Zangenehpour (Beaconsfield, CA); Colin Joseph Brown (Saskatoon, CA); Paul Anthony Kruszewski (Westmount, CA)
Assignee: Hinge Health, Inc.
G06T17/20G06T7/12G06T2207/10024G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,675,949
App. No.
18/431,126
Filed
Feb 2, 2024
Granted
Jul 7, 2026
Kind
B2
Art Unit
2616
USPC
345/419
Abstract

An apparatus is provided. The apparatus includes a communications interface to receive raw data from an external source. The raw data includes a representation of an object. Furthermore, the apparatus includes a memory storage unit to store the raw data. The apparatus also includes a pre-processing engine to generate a coarse segmentation map and a joint heatmap from the raw data. The coarse segmentation map is to outline the object and the joint heatmap is to represent a point on the object. The apparatus further includes a neural network engine to receive the raw data, the coarse segmentation map, and the joint heatmap. The neural network engine is to generate a plurality of two-dimensional maps. Also, the apparatus includes a mesh creator engine to generate a three-dimensional mesh based on the plurality of two-dimensional maps.

Claims (53)

1 . A non-transitory medium with instructions stored thereon that, when executed by a processor, cause the processor to perform operations comprising:

obtaining a two-dimensional image of a person that is generated by a camera;

generating, based on the two-dimensional image, a coarse segmentation map that represents an outline of the person and a plurality of joint heatmaps, each of which represents a different one of a plurality of joints of the person;

applying a neural network to the two-dimensional image, the coarse segmentation map, and/or the plurality of joint heatmaps, so as to produce multiple two-dimensional maps,

wherein the multiple two-dimensional maps include

(i) a plurality of normal maps,

(ii) a distance map that provides, for each pixel in the two-dimensional image, distance to a reference plane, and

(iii) a thickness map that provides, for each pixel in the two-dimensional image, distance to a corresponding point along a back surface of a three-dimensional mesh; and

generating the three-dimensional mesh for the person by—

(a) isolating pixels in the two-dimensional image that form part of the person,

(b) forming a front surface of the three-dimensional mesh using at least one of the plurality of normal maps and the distance map, and

(c) forming the back surface of the three-dimensional mesh using at least one of the plurality of normal maps and the thickness map with reference to the front surface of the three-dimensional mesh.

2 . The non-transitory medium of claim 1 , wherein the multiple two-dimensional maps further include:

a first intensity map that represents, on a per-pixel basis, intensity of red in the two-dimensional image,

a second intensity map that represents, on a per-pixel basis, intensity of green in the two-dimensional image, and

a third intensity map that represents, on a per-pixel basis, intensity of blue in the two-dimensional image.

3 . The non-transitory medium of claim 1 , wherein said obtaining comprises:

receiving the two-dimensional image from a source external to a computing device of which the processor is a part via a communications interface.

4 . The non-transitory medium of claim 1 , wherein the operations further comprise:

processing the two-dimensional image so as to remove shadows and/or additional source lighting in the two-dimensional image.

5 . The non-transitory medium of claim 1 , wherein the operations further comprise:

setting the reference plane to be a plane other than a camera plane that corresponds to the camera.

6 . The non-transitory medium of claim 1 , wherein the camera is part of a same computing device as the processor.

7 . The non-transitory medium of claim 1 ,

wherein the plurality of normal maps includes:

a front red map that represents intensity of red across a front surface of the person,

a front green map that represents intensity of green across the front surface of the person,

a front blue map that represents intensity of blue across the front surface of the person,

a back red map that represents intensity of red across a back surface of the person,

a back green map that represents intensity of green across the back surface of the person, and

a back blue map that represents intensity of blue across the back surface of the person.

8 . The non-transitory medium of claim 1 , wherein the multiple two-dimensional maps further include a fine segmentation map that represents the outline of the person with greater accuracy than the coarse segmentation map, and wherein said isolating is based on the fine segmentation map such that pixels outside of a segmentation are discarded.

9 . The non-transitory medium of claim 1 , wherein the neural network is a fully convolutional network that includes two stacked U-nets.

10 . The non-transitory medium of claim 1

wherein the multiple two-dimensional maps further include:

a first intensity map that represents, on a per-pixel basis, intensity of red in the two-dimensional image,

a second intensity map that represents, on a per-pixel basis, intensity of green in the two-dimensional image, and

a third intensity map that represents, on a per-pixel basis, intensity of blue in the two-dimensional image, and

wherein the neural network is also applied to the first, second, and third intensity maps to produce the multiple two-dimensional maps.

11 . The non-transitory medium of claim 10 , further comprising:

using at least one of the multiple two-dimensional maps to improve a training process of the neural network.

12 . The non-transitory medium of claim 1 , wherein the plurality of joint heatmaps include separate heatmaps for at least two of a left eye, a right eye, a left shoulder, a right shoulder, a left elbow, a right elbow, a left wrist, a right wrist, a left hip, a right hip, a left knee, a right knee, a left ankle, a right ankle, a left toe, and a right toe.

13 . The non-transitory medium of claim 1 , wherein the plurality of normal maps includes:

for a first surface of the person,

a first red map that indicates intensity of red colors across the first surface,

a first green map that indicates intensity of green colors across the first surface, and

a first blue map that indicates intensity of blue colors across the first surface, and

for a second surface of the person,

a second red map that indicates intensity of red colors across the second surface,

a second green map that indicates intensity of green colors across the second surface, and

a second blue map that indicates intensity of blue colors across the second surface.

14 . The non-transitory medium of claim 13 , wherein the operations further comprise:

adding color to the three-dimensional mesh based on the first red map, the first green map, and the first blue map produced for the first surface of the person and the second red map, the second green map, and the second blue map produced for the second surface of the person.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2024
From: ZANGENEHPOUR, SOHAIL; BROWN, COLIN JOSEPH; KRUSZEWSKI, PAUL ANTHONY
To: HINGE HEALTH, INC.
Reel/Frame 067063/0501 →
Continuity (2)
Continuation 18250845 · Oct 29, 2020
Related Publication 20240169670A1 · May 23, 2024
References Cited (30)
US 8289318B1 · Hadap et al. · 2012 [cited by applicant]
US 10813715B1 · Chojnowski et al. · 2020 [cited by applicant]
US 11182924B1 · Akbas et al. · 2021 [cited by applicant]
US 11688139B1 · Karagoz et al. · 2023 [cited by applicant]
US 20190295318A1 · Levinson et al. · 2019 [cited by applicant]
US 20210161266A1 · Brown et al. · 2021 [cited by applicant]
US 20210392296A1 · Rabinovich et al. · 2021 [cited by applicant]
US 20220148296A1 · Brown et al. · 2022 [cited by applicant]
US 20230096013A1 · Agrawal et al. · 2023 [cited by applicant]
US 20230225832A1 · Cramer et al. · 2023 [cited by applicant]
KR 20180069786A · 2018 [cited by applicant]
Zhou K, Han X, Jiang N, Jia K, Lu J. Hemlets pose: Learning part-centric heatmap triplets for accurate 3d human pose estimation. InProceedings of the IEEE/CVF international conference on computer vision 2019 (pp. 2344-2… [cited by examiner]
Lassner C, Romero J, Kiefel M, Bogo F, Black MJ, Gehler PV. Unite the people: Closing the loop between 3d and 2d human representations. InProceedings of the IEEE conference on computer vision and pattern recognition 201… [cited by examiner]
Guidi G, Remondino F. 3D Modelling from Real Data. InModeling and Simulation in Engineering 2012 (pp. 69-102). InTech. [cited by examiner]
Yao Y, Schertler N, Rosales E, Rhodin H, Sigal L, Sheffer A. Front2back: Single view 3d shape reconstruction via front to back prediction. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit… [cited by examiner]
Ren Y, Li G, Liu S, Li TH. Deep Spatial Transformation for Pose-Guided Person Image Generation and Animation. arXiv preprint arXiv:2008.12606. Aug. 27, 2020. [cited by examiner]
Mehta D, Sotnychenko O, Mueller F, Xu W, Elgharib M, Fua P, Seidel HP, Rhodin H, Pons-Moll G, Theobalt C. XNect: Real-time multi-person 3D motion capture with a single RGB camera. Acm Transactions On Graphics (TOG). Jul… [cited by examiner]
Kato, H. , et al., “Neural 3d Mesh Renderer”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, USA, Jun. 18-22, 2018, pp. 3907-3916. [cited by applicant]
Kim, Y. , et al., “A CNN-based 3D human pose estimation based on projection of depth and ridge data”, Pattern Recognition, vol. 106, 107462, Oct. 2020. [cited by applicant]
Kniaz, W. , et al., “Single image 3D model reconstruction and photorealistic texturing”, In European Conference on Computer Vision, LNCS 12536, Aug. 23, 2020, pp. 595-611. [cited by applicant]
Moon, G. , et al., “I2L-meshnet: Image-to-lixel prediction network for accurate 3d human pose and mesh estimation from a single rgb image”, In Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, Aug. 23-28… [cited by applicant]
Natsume, Ryota, et al., “SiCloPe: Silhouette-Based Clothed People”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, XP033687199,, Jun. 15, 2019, pp. 4475-4485. [cited by applicant]
Pavlakos , et al., “Learning to Estimate 3D Human Pose and Shape from a Single Color Image”, Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, retrieved online from url:, Jun. 18, 2021,… [cited by applicant]
Tang , et al., “A Neural Network for Detailed Human Depth Estimation From a Single Image”, Proceedings of 2019 IEEE/CVF International Conference on Computer Vision (ICCV), retrieved online from url:, Oct. 27, 2019, pp. … [cited by applicant]
Varol , et al., “BodyNet: Volumetric Inference of 3D Human Body Shapes”, Proceedings of the 15th European Conference on Computer Vision—ECCV 2018; retrieved online from url: <http://www.ecva.net/papers/eccv _ 2018/paper… [cited by applicant]
Yang, L. , et al., “Bi hand: Recovering Hand Mesh with Multi-stage Bisected Hourglass Networks”, arXiv preprint arXiv2008.05079, Aug. 12, 2020, 14 pages. [cited by applicant]
Zhou , et al., “Learning to Reconstruct 3D Manhattan Wireframes from a Single Image, Learning to Reconstruct 3D Manhattan Wireframes from a Single Image”, Proceedings of2019 IEEE/CVF International Conference on Computer… [cited by applicant]
Kniaz, Vladimir V., et al., “StructureFromGAN: Single Image 3D Model Reconstruction and Photorealistic Texturing”, Computer Vision—ECCV 2020 Workshops. ECCV 2020. Lecture Notes in Computer Science, vol. 12536., Jan. 3, … [cited by applicant]
Natsume, Ryota , et al., “SiCloPe: Silhouette-Based Clothed People”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)., Apr. 10, 2019, pp. 1-11. [cited by applicant]
Pavlakos, Georgios , et al., “Learning to Estimate 3D Human Pose and Shape from a Single Color Image”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition., May 10, 2018, pp. 1-10. [cited by applicant]