IP Library › Granted Patent US 12,626,461
Granted Patent B2
US 12,626,461 · App. 18/242,380 · Granted May 12, 2026

Complete 3D object reconstruction from an incomplete image

Inventors: Jae Shin Yoon (San Jose, CA); Yangtuanfeng Wang (London, GB); Krishna Kumar Singh (San Jose, CA); Junying Wang (Los Angeles, CA); Jingwan Lu (Santa Clara, CA)
Assignee: Adobe Inc.
G06T17/10G06T5/77G06T7/70G06T2200/24G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,461
App. No.
18/242,380
Granted
May 12, 2026
Kind
B2
Abstract

A modeling system accesses a two-dimensional (2D) input image displayed via a user interface, the 2D input image depicting, at a first view, a first object. At least one region of the first object is not represented by pixel values of the 2D input image. The modeling system generates, by applying a 3D representation generation model to the 2D input image, a three-dimensional (3D) representation of the first object that depicts an entirety of the first object including the first region. The modeling system displays, via the user interface, the 3D representation, wherein the 3D representation is viewable via the user interface from a plurality of views including the first view.

Claims (60)

1 . A method performed by one or more computing devices associated with a modeling system, comprising:

accessing a two-dimensional (2D) input image displayed via a user interface, the 2D input image depicting, at a first view, a first object, wherein at least one region of the first object is not represented by pixel values of the 2D input image;

training a three-dimensional (3D) convolutional neural network of a 3D representation generation model using a 3D discriminator to produce generative volumetric features in the at least one region of the first object in the 2D input image;

combining fine-detailed surface normals in a multiview normal fusion framework to produce multiview surface normals based on pixel-aligned normal features;

combining, using the 3D representation generation model applied to the 2D input image, the multiview surface normals of the first object with the generative volumetric features to produce a 3D representation of the first object that depicts an entirety of the first object including the at least one region; and

displaying, via the user interface, the 3D representation, wherein the 3D representation is viewable via the user interface from a plurality of views including the first view.

2 . The method of claim 1 , wherein at the first view the at least one region is outside of an area of the 2D input image.

3 . The method of claim 1 , wherein at the first view the at least one region is occluded by a second object depicted in the 2D input image.

4 . The method of claim 1 , further comprising:

generating, based on the generative volumetric features determined based on the 2D input image, a coarse geometry for the first object using a coarse multilayer perceptron (MLP); and

generating, based on the coarse geometry and intermediate features generated by the coarse MLP, a fine geometry for the first object.

5 . The method of claim 4 , wherein the multiview surface normals are based on the coarse geometry.

6 . The method of claim 4 , further comprising:

generating an image feature volume for the 2D input image by extracting features of the 2D input image in a depth direction; and

determining concatenated image features by concatenating the image feature volume with a 3D pose of the first object recorded on the image feature volume, the 3D pose determined from a 3D object model.

7 . The method of claim 4 , further comprising applying, based on the fine geometry generated for the first object and the 2D input image, a progressive texture inpainting process to generate the 3D representation.

8 . The method of claim 1 , wherein the 3D representation is displayed at the first view and depicts the at least one region of the first object, and further comprising:

responsive to receiving an input via the user interface, displaying the 3D representation at a second view of the plurality of views that is different from the first view,

wherein the 3D representation displayed at the second view depicts the at least one region of the first object.

9 . A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device configured to perform operations comprising:

accessing a two-dimensional (2D) input image displayed via a user interface, the 2D input image depicting, at a first view, a first object, wherein at least one region of the first object is not represented by pixel values of the 2D input image;

training a three-dimensional (3D) convolutional neural network of a 3D representation generation model using a 3D discriminator to produce generative volumetric features in the at least one region of the first object in the 2D input image;

combining fine-detailed surface normals in a multiview normal fusion framework to produce multiview surface normals based on pixel-aligned normal features;

combining, using the 3D representation generation model applied to the 2D input image, the multiview surface normals of the first object with the generative volumetric features to produce a 3D representation of the first object that depicts an entirety of the first object including the at least one region; and

displaying, via the user interface, the 3D representation, wherein the 3D representation is viewable via the user interface from a plurality of views including the first view,

wherein the 3D representation is displayed at the first view and depicts the at least one region of the first object.

10 . The system of claim 9 , the operations further comprising:

responsive to receiving an input via the user interface, displaying the 3D representation at a second view of the plurality of views that is different from the first view,

wherein the 3D representation displayed at the second view depicts the at least one region of the first object.

11 . The system of claim 9 , wherein at the first view the at least one region is outside of an area of the 2D input image or is occluded by a second object depicted in the 2D input image.

12 . The system of claim 9 , the operations further comprising:

generating, based on the generative volumetric features determined based on the 2D input image, a coarse geometry for the first object using a coarse multilayer perceptron (MLP); and

generating, based on the coarse geometry and intermediate features generated by the coarse MLP, a fine geometry for the first object.

13 . The system of claim 12 , wherein the multiview surface normals are based on the coarse geometry.

14 . The system of claim 12 , the operations further comprising:

generating an image feature volume for the 2D input image by extracting features of the 2D input image in a depth direction; and

determining concatenated image features by concatenating the image feature volume with a 3D pose of the first object recorded on the image feature volume, the 3D pose determined from a 3D object model.

15 . The system of claim 12 , the operations further comprising applying, based on the fine geometry generated for the first object and the 2D input image, a progressive texture inpainting process to generate the 3D representation.

16 . The system of claim 12 , wherein the 3D representation is displayed at the first view and depicts the at least one region of the first object, the operations further comprising:

responsive to receiving an input via the user interface, displaying the 3D representation at a second view of the plurality of views that is different from the first view,

wherein the 3D representation displayed at the second view depicts the at least one region of the first object.

17 . A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

accessing a two-dimensional (2D) input image displayed via a user interface, the 2D input image depicting, at a first view, a first object, wherein at least one region of the first object is not represented by pixel values of the 2D input image, wherein the at least one region is outside of an area of the 2D input image or is occluded by a second object depicted in the 2D input image;

combining fine-detailed surface normals in a multiview normal fusion framework to produce multiview surface normals based on pixel-aligned normal features;

training a three-dimensional (3D) convolutional neural network of a 3D representation generation model using a 3D discriminator to produce generative volumetric features in the at least one region of the first object in the 2D input image;

combining, using the 3D representation generation model applied to the 2D input image, the multiview surface normals of the first object with the generative volumetric features to produce a 3D representation of the first object that depicts an entirety of the first object including the at least one region; and

displaying, via the user interface, the 3D representation, wherein the 3D representation is viewable via the user interface from a plurality of views including the first view,

wherein the 3D representation is displayed at the first view and depicts the at least one region of the first object.

18 . The non-transitory computer-readable medium of claim 17 , the operations further comprising:

generating, based on the generative volumetric features determined based on the 2D input image, a coarse geometry for the first object using a coarse multilayer perceptron (MLP);

generating, based on the coarse geometry and intermediate features generated by the coarse MLP, a fine geometry for the first object; and

applying, based on the fine geometry generated for the first object and the 2D input image, a progressive texture inpainting process to generate the 3D representation.

19 . The non-transitory computer-readable medium of claim 18 , the operations further comprising:

generating an image feature volume for the 2D input image by extracting features of the 2D input image in a depth direction; and

determining concatenated image features by concatenating the image feature volume with a 3D pose of the first object recorded on the image feature volume, the 3D pose determined from a 3D object model.

20 . The system of claim 12 , the operations further comprising:

responsive to receiving an input via the user interface, displaying the 3D representation at a second view of the plurality of views that is different from the first view,

wherein the 3D representation displayed at the second view depicts the at least one region of the first object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2023
From: YOON, JAE SHIN; WANG, YANGTUANFENG; SINGH, KRISHNA KUMAR; WANG, JUNYING; LU, JINGWAN
To: ADOBE INC.
Reel/Frame 064799/0711 →
Continuity (1)
Related Publication 20250078406A1 · Mar 6, 2025
References Cited (68)
US 20160314619A1 · Luo · 2016 [cited by examiner]
US 20210383115A1 · Alon · 2021 [cited by examiner]
US 20220157017A1 · Du · 2022 [cited by examiner]
US 20230060131A1 · Wang · 2023 [cited by examiner]
US 20230196617A1 · Zheng · 2023 [cited by examiner]
US 20230306686A1 · Zangenehpour · 2023 [cited by examiner]
US 20230334754A1 · Kirchmayer · 2023 [cited by examiner]
US 20230377182A1 · Shin · 2023 [cited by examiner]
US 20240104828A1 · Wang · 2024 [cited by examiner]
US 20240144595A1 · Ponjou Tasse · 2024 [cited by examiner]
Albahar et al., Pose with Style: Detail-Preserving Pose-Guided Image Synthesis with Conditional StyleGAN, ACM Transactions on Graphics (TOG), vol. 40, No. 6, Article 218, Available online at: https://arxiv.org/pdf/2109.… [cited by applicant]
Alldieck et al., Photorealistic Monocular 3D Reconstruction of Humans Wearing Clothing, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available online at: https://arxiv.org/pdf/22… [cited by applicant]
Alldieck et al., Tex2Shape: Detailed Full Human Body Geometry from a Single Image, in Proceedings of the IEEE/CVF International Conference on Computer Vision, Available online at: https://arxiv.org/pdf/1904.08645.pdf, S… [cited by applicant]
Cao et al., Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields, IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Available online at: https://arxiv.org/pdf/1611.08050.pdf, Apr. 14, 201… [cited by applicant]
Chen et al., gDNA: Towards Generative Detailed Neural Avatars, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available online at: https://arxiv.org/pdf/2201.04123.pdf, Apr. 13, 20… [cited by applicant]
Chen et al., Learning Implicit Fields for Generative Shape Modeling, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available online at: https://arxiv.org/pdf/1812.02822.pdf, Sep. … [cited by applicant]
Chen et al., Unpaired Pose Guided Human Image Generation, in Conference on Computer Vision and Pattern Recognition (CVPR 2019). Computer Vision Foundation (CVF), Available online at: https://arxiv.org/pdf/1901.02284.pdf… [cited by applicant]
Choi et al., Learning to Estimate Robust 3D Human Mesh from In-the-Wild Crowded Scenes, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available online at: https://arxiv.org/pdf/21… [cited by applicant]
Cicek et al., 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation, International Conference on Medical Image Computing and Computer-Assisted Intervention, Available online at: https://arxiv.org/pdf/1… [cited by applicant]
Fan et al., A Point Set Generation Network for 3D Object Reconstruction from a Single Image, Computer Vision and Pattern Recognition, Available Online at: https://arxiv.org/abs/1612.00603, Dec. 7, 2016, 12 pages. [cited by applicant]
Goodfellow et al., Generative Adversarial Networks, Communications of the ACM, vol. 63, No. 11, Available online at: https://dl.acm.org/doi/pdf/10.1145/3422622, Nov. 2020, pp. 139-144. [cited by applicant]
Grigorev et al., StylePeople: a Generative Model of Fullbody Human Avatars, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available online at: https://arxiv.org/pdf/2104.08363.pdf… [cited by applicant]
Groueix et al., A Papier-Mache Approach to Learning 3D Surface Generation, Cornell University, Computer Science; Computer Vision and Pattern Recognition, Available online at: https://arxiv.org/pdf/1802.05384.pdf, Jul. 2… [cited by applicant]
Guler et al., DensePose: Dense Human Pose Estimation in the Wild, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Available online at: https://arxiv.org/pdf/1802.00434.pdf, Feb. 1, 2018… [cited by applicant]
He et al., ARCH++: Animation-Ready Clothed Human Reconstruction Revisited, in Proceedings of the IEEE/CVF International Conference on Computer Vision, Available online at: https://arxiv.org/pdf/2108.07845.pdf, Aug. 17, … [cited by applicant]
He et al., Geo-PIFu: Geometry and Pixel Aligned Implicit Functions for Single-view Human Reconstruction, Advances in Neural Information Processing Systems, vol. 33, Available online at: https://arxiv.org/pdf/2006.08072.… [cited by applicant]
Ho et al., Denoising Diffusion Probabilistic Models, Advances in Neural Information Processing Systems, vol. 33, Available online at: https://arxiv.org/pdf/2006.11239.pdf, Dec. 16, 2020, 25 pages. [cited by applicant]
Hong et al., StereoPIFu: Depth Aware Clothed Human Digitization via Stereo Vision, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available online at: https://openaccess.thecvf.com… [cited by applicant]
Huang et al., ARCH: Animatable Reconstruction of Clothed Humans, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available online at: https://arxiv.org/pdf/2004.04572.pdf, Apr. 10, … [cited by applicant]
Isola et al., Image-to-Image Translation with Conditional Adversarial Networks, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Available Online at: http://gangw.cs.illinois.edu/class/c… [cited by applicant]
Jackson et al., 3D Human Body Reconstruction from a Single Image via Volumetric Regression, in Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Available online at: https://openaccess.thecvf.c… [cited by applicant]
Johnson et al., Perceptual Losses for Real-Time Style Transfer and Super-Resolution, in European Conference on Computer Vision, Available Online at: https://arxiv.org/pdf/1603.08155.pdf%7C, Mar. 27, 2016, 18 pages. [cited by applicant]
Kanazawa et al., End-to-end Recovery of Human Shape and Pose, Computer Vision and Pattern Recognition (CVPR), Available online at: https://arxiv.org/pdf/1712.06584.pdf, Jun. 23, 2018, 10 pages. [cited by applicant]
Kolotouros et al., Convolutional Mesh Regression for Single-Image Human Shape Reconstruction, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available online at: https://arxiv.org/… [cited by applicant]
Kolotouros et al., Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the Loop, in Proceedings of the IEEE International Conference on Computer Vision, Sep. 27, 2019, 10 pages. [cited by applicant]
Lassner et al., A Generative Model of People in Clothing, in Proceedings of the IEEE International Conference on Computer Vision, Available online at: https://arxiv.org/pdf/1705.04098.pdf, Jul. 31, 2017, 10 pages. [cited by applicant]
Lin et al., Learning Efficient Point Cloud Generation for Dense 3D Object Reconstruction, in proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, Available online at: https://arxiv.org/pdf/1706.07036.… [cited by applicant]
Liu et al., DIST: Rendering Deep Implicit Signed Distance Function with Differentiable Sphere Tracing, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available online at: https://a… [cited by applicant]
Loper et al., SMPL: a Skinned Multi-person Linear Model, ACM Transactions on Graphics, vol. 34, No. 6, Nov. 2, 2015, pp. 1-16. [cited by applicant]
Lorensen et al., Marching Cubes: a High-Resolution 3D Surface Construction Algorithm, ACM Siggraph Computer Graphics, vol. 21, No. 4, Jul. 1, 1987, pp. 163-169. [cited by applicant]
Ma et al., Pose Guided Person Image Generation, Advances in Neural Information Processing Systems, vol. 30, Available online at: https://arxiv.org/pdf/1705.09368.pdf, Jan. 28, 2018, 11 pages. [cited by applicant]
Mescheder et al., Occupancy Networks: Learning 3D Reconstruction in Function Space, Cornell University, Computer Science; Computer Vision and Pattern Recognition, Apr. 30, 2019, 11 pages. [cited by applicant]
Newell et al., Stacked Hourglass Networks for Human Pose Estimation, European conference on computer vision, Available Online at: https://arxiv.org/pdf/1603.06937.pdf, Jul. 26, 2016, 17 pages. [cited by applicant]
Nguyen-Phuoc et al., HoloGAN: Unsupervised Learning of 3D Representations from Natural Images, in Proceedings of the IEEE/CVF International Conference on Computer Vision, Available online at: https://arxiv.org/pdf/1904.… [cited by applicant]
Park et al., DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jan. 16, 2019, 19 pages. [cited by applicant]
Pavlakos et al., Expressive Body Capture: 3D Hands, Face, and Body from a Single Image, in Proceedings IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Apr. 11, 2019, 22 pages. [cited by applicant]
Pavlakos et al., Learning to Estimate 3D Human Pose and Shape from a Single Color Image, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Available online at: https://arxiv.org/pdf/1805.… [cited by applicant]
Pavlakos et al., TexturePose: Supervising Human Mesh Estimation with Texture Consistency, in Proceedings of the IEEE/CVF International Conference on Computer Vision, Available online at: https://arxiv.org/pdf/1910.11322… [cited by applicant]
Qi et al., PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Apr. 10, 2017, 19 pages. [cited by applicant]
Qi et al., PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space, 31st Conference on Neural Information Processing Systems, Jun. 7, 2017, 14 pages. [cited by applicant]
Saito et al., PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization, in Proceedings of the IEEE/CVF International Conference on Computer Vision, Dec. 3, 2019, pp. 2304-2314. [cited by applicant]
Saito et al., PIFuHD: Multi-Level Pixel-Aligned Implicit Function for High-Resolution 3D Human Digitization, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available online at: htt… [cited by applicant]
Sarkar et al., HumanGAN: a Generative Model of Humans Images, in 2021 International Conference on 3D Vision (3DV), Available online at: https://arxiv.org/pdf/2103.06902.pdf, Mar. 11, 2021, 21 pages. [cited by applicant]
Schwarz et al., GRAF: Generative Radiance Fields for 3D-Aware Image Synthesis, Advances in Neural Information Processing Systems, Available online at: https://arxiv.org/abs/2007.02442, Mar. 30, 2021, pp. 1-13. [cited by applicant]
Sohl-Dickstein et al., Deep Unsupervised Learning using Nonequilibrium Thermodynamics, in International Conference on Machine Learning, Available online at: https://arxiv.org/pdf/1503.03585.pdf, Nov. 18, 2015, 18 pages. [cited by applicant]
Torosdagli et al., Deep Geodesic Learning for Segmentation and Anatomical Landmarking, IEEE Transactions on Medical Imaging, vol. 38, No. 4, Apr. 2019, 39 pages. [cited by applicant]
Varol et al., BodyNet: Volumetric Inference of 3D Human Body Shapes, in Proceedings of the European Conference on Computer Vision (ECCV), Available online at: https://arxiv.org/pdf/1804.04875.pdf, Aug. 18, 2018, 27 page… [cited by applicant]
Wang et al., Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images, in Proceedings of the European Conference on Computer Vision (ECCV), Available online at: https://arxiv.org/pdf/1804.01654.pdf, Aug. 3, 2018, 16… [cited by applicant]
Wu et al., Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling, 29th Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain, Jan. 4, 2017, 11 pages. [cited by applicant]
Xiu et al., ICON: Implicit Clothed humans Obtained from Normals, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available online at: https://arxiv.org/pdf/2112.09127.pdf, Mar. 28, … [cited by applicant]
Yang, PointFlow: 3D Point Cloud Generation with Continuous Normalizing Flows, in Proceedings of the IEEE/CVF International Conference on Computer Vision, Available online at: https://arxiv.org/pdf/1906.12320.pdf, Sep. 2… [cited by applicant]
Yoon et al., Pose-Guided Human Animation from a Single Image in the Wild, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available online at: https://arxiv.org/pdf/2012.03796.pdf, … [cited by applicant]
Yu et al., Function4D: Real-time Human Volumetric Capture from Very Sparse Consumer RGBD Sensors, in IEEE Conference on Computer Vision and Pattern Recognition (CVPR2021), Available online at: https://arxiv.org/pdf/2105… [cited by applicant]
Zhang et al., PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback Loop, in Proceedings of the IEEE/CVF International Conference on Computer Vision, Apr. 1, 2021, pp. 11446-11456. [cited by applicant]
Zhang et al., PyMAF-X: Towards Well-aligned Full-body Model Regression from Monocular Images, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, No. 10, Available online at: http://export.arxiv.org… [cited by applicant]
Zheng et al., DeepHuman: 3D Human Reconstruction from a Single Image, in the IEEE International Conference on Computer Vision (ICCV), Available online at: https://arxiv.org/pdf/1903.06473.pdf, Mar. 28, 2019, 14 pages. [cited by applicant]
Zheng et al., DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview Cameras, in IEEE Conference on Computer Vision (ICCV 2021), Available online at: http://export.arxiv.org/pdf/2105.00261, Aug.… [cited by applicant]
Zheng et al., PaMIR: Parametric Model-Conditioned Implicit Representation for Image-Based Human Reconstruction, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, No. 6, Available online at: https:… [cited by applicant]