IP Library › Granted Patent US 12,573,161
Granted Patent B2
US 12,573,161 · App. 18/252,118 · Granted Mar 10, 2026

Learning articulated shape reconstruction from imagery

Inventors: Deqing Sun (Cambridge, MA); Varun Jampani (Rockland, MA); Gengshan Yang (Pittsburgh, PA); Daniel Vlasic (Cambridge, MA); Huiwen Chang (Jersey City, NJ); Forrester H. Cole (Cambridge, MA); Ce Liu (Cambridge, MA); William Tafel Freeman (Acton, MA)
Assignee: GOOGLE LLC
G06T19/20G06T7/20G06T7/40G06T7/55G06T17/20G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30244G06T2219/2021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,161
App. No.
18/252,118
Granted
Mar 10, 2026
Kind
B2
Abstract

A computing system and method can be used to render a 3D shape from one or more images. In particular, the present disclosure provides a general pipeline for learning articulated shape reconstruction from images (LASR). The pipeline can reconstruct rigid or nonrigid 3D shapes. In particular, the pipeline can automatically decompose non-rigidly deforming shapes into rigid motions near rigid-bones. This pipeline incorporates an analysis-by-synthesis strategy and forward-renders silhouette, optical flow, and color images which can be compared against the video observations to adjust the internal parameters of the model. By inverting a rendering pipeline and incorporating optical flow, the pipeline can recover a mesh of a 3D model from the one or more images input by a user.

Claims (45)

1 . A computer-implemented method for determining 3D object shape from imagery, the method comprising:

obtaining, by a computing system comprising one or more computing devices, an input image that depicts an object and a current mesh model of the object, wherein the current mesh model of the object comprises a plurality of vertices, a plurality of joints, and a plurality of blend skinning weights for the plurality of joints relative to the plurality of vertices;

processing, by the computing system, the input image with a camera model to obtain camera parameters, articulation parameters, and object deformation data for the input image, wherein the camera parameters describe a camera pose for the input image, wherein the articulation parameters are based on the plurality of vertices, the plurality of joints, and the plurality of blend skinning weights, wherein the object deformation data describes one or more deformations of the current mesh model relative to a shape of the object shown in the input image, and wherein the processing comprises linear blend skinning based on the articulation parameters;

differentiably rendering, by the computing system, a rendered image of the object based on the camera parameters, the articulation parameters, the object deformation data, and the current mesh model;

evaluating, by the computing system, a loss function that compares one or more characteristics of the input image of the object with one or more characteristics of the rendered image of the object; and

modifying, by the computing system, one or more values of one or both of the camera model and the current mesh model based on a gradient of the loss function.

2 . The computer-implemented method of claim 1 , wherein evaluating the loss function comprises:

determining a first flow for the input image;

determining a second flow for the rendered image; and

evaluating the loss function based at least in part on a comparison of the first flow and the second flow.

3 . The computer-implemented method of claim 1 , wherein evaluating the loss function comprises:

determining a first silhouette for the input image;

determining a second silhouette for the rendered image; and

evaluating the loss function based at least in part on a comparison of the first silhouette and the second silhouette.

4 . The computer-implemented method of claim 1 , wherein evaluating the loss function comprises evaluating the loss function based at least in part on a comparison of first texture data associated with the input image and second texture data associated with the rendered image.

5 . The computer-implemented method of claim 1 , wherein the mesh model comprises a triangular mesh model.

6 . The computer-implemented method of claim 1 , wherein the camera parameters describes an object-to-camera transformation for the input image.

7 . The computer-implemented method of claim 1 , wherein the camera model comprises a convolutional neural network.

8 . The computer-implemented method of claim 1 , wherein the camera parameters further describe a focal length for the input image.

9 . The computer-implemented method of claim 1 , wherein the mesh model is initialized to a subdivided icosahedron projected to a sphere.

10 . The computer-implemented method of claim 1 , wherein one or both of the camera model and the current mesh model comprise both the camera model and the current mesh model.

11 . The computer-implemented method of claim 1 , wherein the plurality of joints and the plurality of blend skinning weights are learnable.

12 . The computer-implemented method of claim 1 , wherein differentiably rendering the rendered image of the object based on the camera parameters, the object deformation data, and the current mesh model comprises rendering the current mesh model deformed according to the object deformation data and from the camera pose according to the camera parameters.

13 . The computer-implemented method of claim 1 , further comprising performing the method of claim 1 on multiple images that depict multiple objects to build a library of shape models from the multiple images.

14 . The computer-implemented method of claim 1 , wherein obtaining the input image that depicts the object comprises manually choosing a canonical image frame from a video.

15 . The computer-implemented method of claim 1 , wherein obtaining the input image that depicts the object comprises:

selecting a number of candidate frames;

evaluating a loss for each of the candidate frames; and

selecting the candidate frame with a lowest final loss as a canonical frame.

16 . A computer system, comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store:

a machine-learned mesh model for an object, wherein the machine-learned mesh model for the object comprises a plurality of vertices, a plurality of joints, and a plurality of blend skinning weights for the plurality of joints relative to the plurality of vertices, wherein the machine-learned mesh model for the object has been learned jointly with a machine-learned camera model by minimizing a loss function that evaluates a difference between one or more input images of the object and one or more rendered images of the object, the one or more rendered images comprising images rendered based on the machine-learned mesh model, camera parameters, and linear blend skimming based on articulation parameters, wherein the articulation parameters and the camera parameters are generated by the machine-learned camera model from the one or more input images, and wherein the articulation parameters are based on the plurality of vertices, the plurality of joints, and the plurality of blend skinning weights; and

instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations, comprising:

receiving an additional set of camera parameters; and

rendering an additional rendered image of the object based on the machine-learned mesh model and the additional set of camera parameters.

17 . The computing system of claim 16 , wherein the loss function compares a first flow for each input image with a second flow for each rendered image.

18 . The computing system of claim 16 , wherein the loss function compares a first silhouette for each input image with a second silhouette for each rendered image.

19 . The computing system of claim 16 , wherein the loss function compares a first texture for each input image with a second texture for each rendered image.

20 . One or more non-transitory computer-readable media that collectively store instructions that, when executed by a computing system comprising one or more computing devices, cause the computing to perform operations, the operations comprising:

obtaining, by the computing system, an input image that depicts an object and a current mesh model of the object, wherein the current mesh model comprises a plurality of vertices, a plurality of joints, and a plurality of blend skinning weights for the plurality of joints relative to the plurality of vertices;

processing, by the computing system, the input image with a camera model to obtain camera parameters, articulation parameters, and object deformation data for the input image, wherein the camera parameters describe a camera pose for the input image, wherein the articulation parameters are based on the plurality of vertices, the plurality of joints, and the plurality of blend skinning weights, wherein the object deformation data describes one or more deformations of the current mesh model relative to a shape of the object shown in the image, and wherein the processing comprises linear blend skinning based on the articulation parameters;

differentiably rendering, by the computing system, a rendered image of the object based on the camera parameters, the articulation parameters, the object deformation data, and the current mesh model;

evaluating, by the computing system, a loss function that compares one or more characteristics of the input image of the object with one or more characteristics of the rendered image of the object; and

modifying, by the computing system, one or more values of one or both of the camera model and the current mesh model based on a gradient of the loss function.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2023
From: YANG, GENGSHAN; FREEMAN, WILLIAM TAFEL; SUN, DEQING; JAMPANI, VARUN; VLASIC, DANIEL; CHANG, HUIWEN; COLE, FORRESTER H.; LIU, CE
To: GOOGLE LLC
Reel/Frame 063568/0016 →
Continuity (1)
Related Publication 20240013497A1 · Jan 11, 2024
References Cited (31)
US 20130285908A1 · Kaplan · 2013 [cited by examiner]
US 20160223475A1 · Palamodov · 2016 [cited by examiner]
US 20200265294A1 · Kim · 2020 [cited by examiner]
US 20210012558A1 · Li · 2021 [cited by examiner]
US 20210125583A1 · Kaplanyan · 2021 [cited by examiner]
US 20210150287A1 · Baek · 2021 [cited by examiner]
US 20210158028A1 · Wu · 2021 [cited by examiner]
US 20210241522A1 · Guler · 2021 [cited by examiner]
US 20210256251A1 · Lin · 2021 [cited by examiner]
US 20220067940A1 · Heng · 2022 [cited by examiner]
US 20220245911A1 · Hu · 2022 [cited by examiner]
US 20220270402A1 · Cole et al. · 2022 [cited by applicant]
US 20220358770A1 · Guler · 2022 [cited by examiner]
US 20230070008A1 · Kulon · 2023 [cited by examiner]
CN 111563967 · 2020 [cited by applicant]
Yu et al., “Direct, dense, and deformable: Template-based non-rigid 3d reconstruction from rgb video,” Proceedings of the IEEE international conference on computer vision (Year: 2015). [cited by examiner]
Alldieck et al., “Learning to Reconstruct People in Clothing from a Single RGB Camera”, IEEE Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, XP033687140, pp. 1175-1186. [cited by applicant]
Henderson et al., “Learning to Generate and Reconstruct 3D Meshes with Only 2D Supervision”, arxiv.org, Jul. 24, 2018, XP081426760, 14 pages. [cited by applicant]
International Search Report for Application No. PCt/US2020/066305, mailed on Sep. 15, 2021, 2 pages. [cited by applicant]
Kato et al., “Differentiable Rendering: A Survey”, arxiv.org, Jul. 31, 2020, XP081726202, 20 pages. [cited by applicant]
Li et al., “Online Adaptation for Consistent Mesh Reconstruction in the Wild” arxiv.org, Dec. 6, 2020, XP081831342, 22 pages. [cited by applicant]
Zuffi et al., “Three-D Safar: Learning to Estimate Zebra Pose, Shape, and Texture from Images ‘In the Wild’”, 2019 IEEE International Conference on Computer Vision, Oct. 27, 2019, pp. 5358-5367, XP033724041. [cited by applicant]
Chen et al., “Deep Non-Rigid Structure from Motion.”, arXiv:1908.00052v2, Aug. 11, 2019, 10 pages. [cited by applicant]
Gotardo et al., “Non-Rigid Structure from Motion with Complementary Rank-3 Spaces.”, Twenty-fourth Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition, Colorado Spring… [cited by applicant]
Kanazawa et al., “Learning Category-Specific Mesh Reconstruction from Image Collections.”, arXiv:1803.07549v2, Jul. 30, 2018, 21 pages. [cited by applicant]
Loper et al., “SMPL: A Skinned Multi-Person Linear Model.”, Association for Computing Machinery Transactions on Graphics, vol. 34, No. 6, Article No. 248, Oct. 26, 2015, pp. 1-16. [cited by applicant]
Tomasi et al., “Shape and Motion from Image Streams Under Orthography: A Factorization Method.”, International Journal of Computer Vision, vol. 9, No. 2, pp. 137-154. [cited by applicant]
Yang et al., “LASR: Learning Articulated Shape Reconstruction from a Monocular Video.”, arXiv:2105.02976v1, May 6, 2021, 12 pages. [cited by applicant]
Zuffi et al., “Lions and Tigers and Bears: Capturing Non-Rigid, 3D, Articulated Shape from Images.”, 2018 Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition, Salt Lak… [cited by applicant]
International Preliminary Report on Patentability for Application No. PCT/US2020/066305, mailed Jun. 29, 2023, 11 pages. [cited by applicant]
Machine Translated Chinese Search Report Corresponding to Application No. 2020801023682 on Aug. 28, 2024. [cited by applicant]