IP Library Granted Patent US 12,536,733
Granted Patent B2
US 12,536,733 · App. 17/551,046 · Granted Jan 27, 2026

Single-image inverse rendering

Inventors: Koki Nagano (Playa Vista, CA); Eric Ryan Chan (Alameda, CA); Sameh Khamis (Alameda, CA); Shalini De Mello (San Francisco, CA); Tero Tapani Karras (Helsinki, FI); Orazio Gallo (Santa Cruz, CA); Jonathan Tremblay (Redmond, WA)
Assignee: NVIDIA CORPORATION
G06T15/08G06N3/04G06N3/08G06T17/005G06T17/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,733
App. No.
17/551,046
Granted
Jan 27, 2026
Kind
B2
Abstract

A single two-dimensional (2D) image can be used as input to obtain a three-dimensional (3D) representation of the 2D image. This is done by extracting features from the 2D image by an encoder and determining a 3D representation of the 2D image utilizing a trained 2D convolutional neural network (CNN). Volumetric rendering is then run on the 3D representation to combine features within one or more viewing directions, and the combined features are provided as input to a multilayer perceptron (MLP) that predicts and outputs color (or multi-dimensional neural features) and density values for each point within the 3D representation. As a result, single-image inverse rendering may be performed using only a single 2D image as input to create a corresponding 3D representation of the scene in the single 2D image.

Claims (41)

1 . A method comprising, at a device:

training a two-dimensional (2D) convolutional neural network (CNN) to generate a three-dimensional (3D) representation of a given single 2D image by:

minimizing a reconstruction error between a 2D image selected from a dataset of single-view 2D images and a first rendering of a corresponding 3D representation produced for the 2D image, and

using an adversarial training objective to encourage one or more second renderings of the corresponding 3D representation to match a distribution of the single-view 2D images in the dataset.

2 . The method of claim 1 , wherein the 2D image and the first rendering are from a same viewpoint.

3 . The method of claim 2 , wherein the one or more second renderings are from a different viewpoint than the 2D image.

4 . The method of claim 3 , wherein the different viewpoint is arbitrary.

5 . The method of claim 1 , wherein the adversarial training objective combines GAN losses.

6 . The method of claim 5 , wherein the adversarial training objective ensures realism of a rendering the 3D representation from one or more viewpoints different from a viewpoint shown within the given single 2D image.

7 . The method of claim 1 , wherein the training is self-supervised.

8 . The method of claim 1 , wherein the corresponding 3D representation produced for the 2D image is a hybrid explicit-implicit model.

9 . The method of claim 8 , the corresponding 3D representation produced for the 2D image is a 3-plane representation.

10 . The method of claim 9 , wherein the 3-plane representation includes a plurality of sparse 2D neural planes oriented in 3D space in an axis-aligned manner.

11 . The method of claim 1 , wherein the corresponding 3D representation produced for the 2D image is a tree structure.

12 . The method of claim 1 , wherein the corresponding 3D representation produced for the 2D image is a dense voxel structure.

13 . The method of claim 1 , wherein the trained 2D CNN is a decoder.

14 . The method of claim 1 , wherein the trained 2D CNN is a generative adversarial network (GAN).

15 . The method of claim 1 , wherein the 2D CNN is trained to process features extracted from the given single 2D image to generate the 3D representation of a scene found within the given single 2D image.

16 . The method of claim 1 , further comprising, at the device:

outputting the trained 2D CNN.

17 . The method of claim 16 , wherein outputting the trained 2D CNN includes deploying the 2D CNN for use in generating the 3D representation of the given single 2D image.

18 . The method of claim 16 , wherein the 2D CNN is deployed to a local computing device.

19 . The method of claim 16 , wherein the 2D CNN is deployed to one or more networked or distributed computing devices.

20 . The method of claim 16 , wherein the 2D CNN is deployed to one or more cloud-computing devices.

21 . The method of claim 1 , wherein volumetric rendering is run on the corresponding 3D representation to aggregate neural features within one or more viewing directions.

22 . The method of claim 21 , wherein a signed distance function (SDF) is used to determine a 3D surface of the aggregated neural features, and an MLP is used to determine color values or multi-dimensional neural features for each point within the corresponding 3D representation.

23 . A system comprising:

a hardware processor of a device that is configured to:

train a two-dimensional (2D) convolutional neural network (CNN) to generate a three-dimensional (3D) representation of a given single 2D image by:

minimizing a reconstruction error between a 2D image selected from a dataset of single-view 2D images and a first rendering of a corresponding 3D representation produced for the 2D image, and

using an adversarial training objective to encourage one or more second renderings of the corresponding 3D representation to match a distribution of the single-view 2D images in the dataset.

24 . The system of claim 23 , wherein the 2D image and the first rendering are from a same viewpoint.

25 . The system of claim 23 , wherein the one or more second renderings are from a different viewpoint than the 2D image.

26 . The system of claim 25 , wherein the different viewpoint is arbitrary.

27 . A non-transitory computer-readable storage medium storing instructions that, when executed by a processor of a device, causes the processor to cause the device to:

train a two-dimensional (2D) convolutional neural network (CNN) to generate a three-dimensional (3D) representation of a given single 2D image by:

minimizing a reconstruction error between a 2D image selected from a dataset of single-view 2D images and a first rendering of a corresponding 3D representation produced for the 2D image, and

using an adversarial training objective to encourage one or more second renderings of the corresponding 3D representation to match a distribution of the single-view 2D images in the dataset.

28 . The non-transitory computer-readable storage medium of claim 27 , wherein the adversarial training objective combines GAN losses.

29 . The non-transitory computer-readable storage medium of claim 28 , wherein the adversarial training objective ensures realism of a rendering the 3D representation from one or more viewpoints different from a viewpoint shown within the given single 2D image.

30 . The non-transitory computer-readable storage medium of claim 27 , wherein the training is self-supervised.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2022
From: NAGANO, KOKI; CHAN, ERIC RYAN; KHAMIS, SAMEH; DE MELLO, SHALINI; KARRAS, TERO TAPANI; GALLO, ORAZIO; TREMBLAY, JONATHAN
To: NVIDIA CORPORATION
Reel/Frame 059325/0873 →
Continuity (2)
Provisional Application 63242993 · Sep 10, 2021
Related Publication 20230081641A1 · Mar 16, 2023
References Cited (28)
US 10297070B1 · Zhu · 2019 [cited by examiner]
US 10885707B1 · Fu · 2021 [cited by examiner]
US 20140009466A1 · Tzur · 2014 [cited by examiner]
US 20190261945A1 · Funka-Lea · 2019 [cited by examiner]
US 20200388071A1 · Grabner · 2020 [cited by examiner]
US 20210027536A1 · Fu · 2021 [cited by examiner]
US 20210103776A1 · Jiang · 2021 [cited by examiner]
US 20210150757A1 · Mustikovela · 2021 [cited by examiner]
US 20210241522A1 · Guler · 2021 [cited by examiner]
US 20210248811A1 · Shan · 2021 [cited by examiner]
US 20210279952A1 · Chen · 2021 [cited by examiner]
US 20210350620A1 · Bronstein · 2021 [cited by examiner]
US 20210383241A1 · Karras · 2021 [cited by examiner]
US 20220358770A1 · Guler · 2022 [cited by examiner]
Han et al., “Image-Based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Era”, IEEE transactions on pattern analysis and machine intelligence (Year: 2019). [cited by examiner]
Hiroharu Kato and Tatsuya Harada, “Self-supervised Learning of 3D Objects from Natural Images”, arXiv preprint arXiv:1911.08850, 2019 · arxiv.org (Year: 2019). [cited by examiner]
Park et al., “DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, ”CVPR, 2019, pp. 165-174, 2019 (Year: 2019). [cited by examiner]
Gafni et al., “Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction, ”CVPR, 2021, pp. 8649-8658 (Year: 2021). [cited by examiner]
Yu et al., “pixelNeRF: Neural Radiance Fields from One or Few Images,” CVPR, 2021, 20 pages, retrieved from https://arxiv.org/abs/2012.02190. [cited by applicant]
Saito et al., “PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization,” ICCV, 2019, pp. 2304-2314. [cited by applicant]
Mildenhall et al., “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” ECCV, 2020, pp. 1-25, retrieved from https://arxiv.org/abs/2003.08934. [cited by applicant]
Ramon et al., “H3D-Net: Few-Shot High-Fidelity 3D Head Reconstruction,” arXiv, 2021, 10 pages, retrieved from https://arxiv.org/abs/2107.12512. [cited by applicant]
Wang et al., “One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing,” CVPA, 2021, 16 pages, retrieved from https://arxiv.org/abs/2011.15126. [cited by applicant]
Zhou et al., “Rotate-and-Render: Unsupervised Photorealistic Face Rotation from Single-View Images,” CVPR preprint, 2020, 11 pages, retrieved from https://arxiv.org/abs/2003.08124. [cited by applicant]
Pan et al., “Do 2D GANs Know 3D Shape? Unsupervised 3D Shape Reconstruction from 2D Image GANs, ” ICLR, 2021, pp. 1-18. [cited by applicant]
Shi et al., “Lifting 2D StyleGAN for 3D-Aware Face Generation,” CVPR, 2021, 15 pages, retrieved from https://arxiv.org/abs/2011.13126. [cited by applicant]
Gafni et al., “Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction,” CVPR, 2021, pp. 8649-8658. [cited by applicant]
Guo et al., “AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis,” arXiv, 2021, 11 pages, retrieved from https://arxiv.org/abs/2103.11078. [cited by applicant]