IP Library › Granted Patent US 12,380,640
Granted Patent B2
US 12,380,640 · App. 18/071,821 · Granted Aug 5, 2025

3D generation of diverse categories and scenes

Inventors: Hsin-Ying Lee (San Jose, CA); Jian Ren (Hermosa Beach, CA); Aliaksandr Siarohin (Los Angeles, CA); Ivan Skorokhodov (Los Angeles, CA); Sergey Tulyakov (Santa Monica, CA); Yinghao Xu (Los Angeles, CA)
Assignee: Snap Inc.
G06T17/00G06T7/50G06T7/90G06V10/82G06T2207/10024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,640
App. No.
18/071,821
Granted
Aug 5, 2025
Kind
B2
Abstract

A three-dimensional (3D) scene is generated from non-aligned generic camera priors by producing a tri-plane representation for an input scene received in random latent code, obtaining a camera posterior including posterior parameters representing color and density data from the random latent code and from generic camera priors without alignment assumptions, and volumetrically rendering an image of the input scene from the color and density data to provide a scene having pixel colors and depth values from an arbitrary camera viewpoint. A depth adaptor processes depth values to generate an adapted depth map that bridges domains of rendered and estimated depth maps for the image of the input scene. The adapted depth map, color data, and scene geometry information from an external dataset are provided to a discriminator for selection of a 3D representation of the input scene.

Claims (121)

1. A method of generating a three-dimensional (3D) object or scene from non-aligned generic camera priors, the method comprising:

producing, by a 3D scene generator, a tri-plane representation for an input scene received in random latent code;

obtaining, by a camera generator, a camera posterior including posterior parameters representing color and density data from the random latent code and from generic camera priors without alignment assumptions of the generic camera priors;

volumetrically rendering, by a volume renderer, an image of the input scene from the color and density data to provide a scene having pixel colors and depth values from an arbitrary camera viewpoint;

processing, by a depth adaptor, the depth values to generate an adapted depth map that bridges domains of rendered and estimated depth maps for the image of the input scene;

providing the adapted depth map and color data to a discriminator;

providing external scene geometry information from an external dataset to the discriminator; and

selecting, by the discriminator, a 3D representation of the input scene based on the color data, adapted depth map, and external scene geometry information.

2. The method of claim 1 , wherein obtaining color and density data from the random latent code and generic camera priors comprises using a shallow 2-layer multi-layer perceptron (MLP) decoder to sample arbitrary camera viewpoints captured from ball-in-sphere camera parameterizations provided to the camera generator, the ball-in-sphere camera parameterization having four additional degrees of freedom including a field of view and pitch, yaw and radius of an inner sphere specifying a look-at point within an outer sphere of the ball-in-sphere camera parameterizations.

3. The method of claim 1 , further comprising learning the arbitrary camera viewpoint during training for each input dataset.

4. The method of claim 1 , further comprising pushing derivatives of predicted camera parameters with respect to prior camera parameters to either one or minus one to arrive at a camera gradient penalty L φi :

ℒ

φ

i

=

❘

"\[LeftBracketingBar]"

∂

φ

i

∂

φ

i

′

❘

"\[RightBracketingBar]"

+

❘

"\[LeftBracketingBar]"

∂

φ

i

∂

φ

i

′

❘

"\[RightBracketingBar]"

-

1

,

where φ′ i ∈ φ′ is a camera sampled from a prior camera distribution and φ i ∈ φ is produced by the camera generator.

5. The method of claim 1 , wherein volumetrically rendering the image of the input scene from the color and density data comprises rendering depths d by volumetric rendering as follows:

d

=

∫

t

n

t

f

T

⁡

(

t

)

⁢

σ

⁡

(

r

⁡

(

t

)

)

⁢

tdt

,

where t n and t f are near/far planes, T(t) is accumulated transmittance, and r(t) is a ray.

6. The method of claim 5 , wherein volumetrically rendering the image of the input scene from the color and density data further comprises shifting and scaling a depth d from a range of [t n , t f ] into [−1, 1] to obtain normalized depth d :

d

_

=

2

·

d

-

(

t

n

+

t

f

+

b

)

/

2

t

f

-

t

n

-

b

,

where b ∈ [0, (t n +t f )/2] is an additional learnable shift that accounts for empty space in front of a camera.

7. The method of claim 1 , wherein processing the depth values comprises producing the adapted depth map as a function of a normalized depth where the depth values are concatenated with RGB color data input and passed to the discriminator.

8. The method of claim 7 , wherein processing the depth values comprises using a convolutional network to generate separate depth maps with different levels of adaptation and the adapted depth map is randomly selected from the separate depth maps.

9. A system for generating a three-dimensional (3D) object or scene from non-aligned generic camera priors, comprising:

a 3D scene generator that produces a tri-plane representation for an input scene received in random latent code;

a camera generator that obtains a camera posterior including posterior parameters representing color and density data from the random latent code and from generic camera priors without alignment assumptions of the generic camera priors;

a volume renderer that volumetrically renders an image of the input scene from the color and density data to provide a scene having pixel colors and depth values from an arbitrary camera viewpoint;

a depth adaptor that processes the depth values to generate an adapted depth map that bridges domains of rendered and estimated depth maps for the image of the input scene; and

a discriminator that receives the adapted depth map, color data and external scene geometry information from an external dataset and selects a 3D representation of the input scene based on the color data, adapted depth map, and external scene geometry information.

10. The system of claim 9 , wherein the 3D scene generator comprises a mapping network, a synthesis network, and a tri-plane decoder.

11. The system of claim 10 , wherein the mapping network takes noise z ∈ 512 and class label c ∈ 0, . . . , K−1, where K is a number of classes and produces a style code w ∈ 512 , the mapping network comprising a 2-layer multi-layer perceptron (MLP) network with Leaky rectified linear unit (Leaky-ReLU) activations and 512 neurons in each layer.

12. The system of claim 10 , wherein the synthesis network comprises a decoder network that produces tri-plane features p=(p xy ), p yz , p xz ) ∈ 3x(512×512×32) wherein a feature vector f xyz ∈ 32 located (x, y, z) ∈ 3 is computed by projecting a coordinate back to the tri-plane representation, followed by bi-linearly interpolating nearby features and averaging features from different planes.

13. The system of claim 10 , wherein the tri-plane decoder comprises a two-layer multi-layer perceptron (MLP) network with Leaky-ReLU activations in a hidden layer that takes a tri-plane feature f xyz point as input and produces the color and density data in the tri-plane feature f xyz point.

14. The system of claim 9 , wherein the camera generator includes a learning system to adjust learnable posterior camera parameters and to provide six degrees of freedom to the learnable posterior camera parameters.

15. The system of claim 9 , wherein the camera generator avoids posterior collapse by reducing a Lipschitz constant for the camera generator.

16. The system of claim 9 , wherein the depth adaptor comprises a three layer convolutional neural network and a shared convolutional layer that converts outputs of the convolutional neural network into respective depth maps.

17. The system of claim 16 , wherein the depth adaptor normalizes an input depth image and applies the normalized input depth image to convolutional layers of the convolutional neural network to generate the respective depth maps obtained from different convolutional layers of the convolutional neural network and randomly selects one of the generated respective depth maps as the adapted depth map.

18. The system of claim 9 , wherein the discriminator receives distilled knowledge about the external scene geometry from a pretrained image source and a 3D representation of the input image from the depth adaptor.

19. The system of claim 9 , wherein the camera generator is conditioned on class labels of the input scene when generating a camera position and on random scene data when generating a look-at position and field-of-view.

20. A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a processor cause the processor to generate a three-dimensional (3D) object or scene from non-aligned generic camera priors by performing operations comprising:

producing a tri-plane representation for an input scene received in random latent code;

obtaining a camera posterior including posterior parameters representing color and density data from the random latent code and from generic camera priors without alignment assumptions of the generic camera priors;

volumetrically rendering an image of the input scene from the color and density data to provide a scene having pixel colors and depth values from an arbitrary camera viewpoint;

processing the depth values to generate an adapted depth map that bridges domains of rendered and estimated depth maps for the image of the input scene; and

selecting a 3D representation of the input scene based on the color data, adapted depth map, and external scene geometry information from an external dataset.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2025
From: LEE, HSIN-YING; REN, JIAN; SIAROHIN, ALIAKSANDR; SKOROKHODOV, IVAN; TULYAKOV, SERGEY; XU, YINGHAO
To: SNAP INC.
Reel/Frame 071486/0262 →
Continuity (1)
Related Publication 20240177414A1 · May 30, 2024
References Cited (10)
Chan, Eric R., et al. “Efficient geometry-aware 3d generative adversarial networks.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022. (Year: 2022). [cited by examiner]
Shi, Yichun, Divyansh Aggarwal, and Anil K. Jain. “Lifting 2d stylegan for 3d-aware face generation.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2021. (Year: 2021). [cited by examiner]
Zhao, Xiaoming, et al. “Generative multiplane images: Making a 2d gan 3d-aware.” European conference on computer vision. Cham: Springer Nature Switzerland, 2022. (Year: 2022). [cited by examiner]
International Search Report and Written Opinion for International Application No. PCT/US2024/081318, dated Apr. 11, 2024 (Apr. 11, 2024)—16 pages. [cited by applicant]
Ivan Skorokhodov et al.: “3D generation on ImageNet”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Mar. 2, 2023 (Mar. 2, 2023). [cited by applicant]
Ivan Skorokhodov et al: “EpiGRAF: Rethinking training of 3D GANs”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jun. 21, 2022 (Jun. 21, 2022). [cited by applicant]
Miguel Angel Bautista et al: “Gaudi: A Neural Architect for Immersive 3D Scene Generation”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jul. 27, 2022 (Jul. 27, 2022). [cited by applicant]
Niemeyer Michael et al: “Campari: Camera-Aware Decomposed Generative Neural Radiance Fields”, 2021 International Conference on 3D Vision (3DV), IEEE, Dec. 1, 2021 (Dec. 1, 2021), pp. 951-961. [cited by applicant]
Yin Wei et al: “Learning to Recover 3D Scene Shape from a Single Image”, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 20, 2021 (Jun. 20, 2021), pp. 204-213. [cited by applicant]
Zifan Shi et al: “Deep Generative Models on 3D Representations: A Survey”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 27, 2022 (Oct. 27, 2022). [cited by applicant]