IP Library › Granted Patent US 12,322,068
Granted Patent B1
US 12,322,068 · App. 17/940,637 · Granted Jun 3, 2025

Generating voxel representations using one or more neural networks

Inventors: Seung Wook Kim (Toronto, CA); Karsten Kreis (Vancouver, CA); Daiqing Li (Oakville, CA); Robin Rombach (Heidelberg, DE); Sanja Fidler (Toronto, CA); Antonio Torralba Barriuso (Somerville, MA); Bradley Brown (Oakvillle, CA)
Assignee: NVIDIA Corporation
G06T5/60G06T2207/20084G06T2210/61
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,322,068
App. No.
17/940,637
Filed
Sep 8, 2022
Granted
Jun 3, 2025
Kind
B1
Art Unit
2612
USPC
345/419
Abstract

Apparatuses, systems, and techniques are presented to generate digital content. In at least one embodiment, one or more neural networks are used to generate a three-dimensional voxel representation of a scene based, at least in part, upon a plurality of two-dimensional images of the scene.

Claims (40)

1. A processor, comprising:

one or more circuits to cause one or more neural networks to generate one or more hierarchical encodings each representing different types of scene information.

2. The processor of claim 1 , wherein the one or more circuits are further to cause a neural network to generate a three-dimensional voxel latent representation of a scene using a plurality of two-dimensional images of the scene and one or more parameters of one or more cameras used to capture the plurality of two-dimensional images, wherein the one or more hierarchical encodings are based on the three-dimensional voxel latent representation.

3. The processor of claim 2 , wherein the neural network is an autoencoder trained to generate the three-dimensional voxel latent representation based, at least in part, upon a density voxel representation and a feature voxel representation for this scene, generated using the plurality of two-dimensional images and the one or more parameters of the one or more cameras.

4. The processor of claim 2 , wherein the one or more neural networks used to generate one or more hierarchical encodings includes a hierarchical set of latent encoders to decompose the three-dimensional voxel latent representation of the scene into a set of smaller latent representations of the scene for use in learning a generative model.

5. The processor of claim 4 , wherein the generative model includes at least one diffusion neural network or a generative adversarial network.

6. The processor of claim 2 , wherein the neural network is trained to minimize a reconstruction loss and an adversarial loss in a three-dimensional space.

7. A system, comprising:

one or more processors to cause one or more neural networks to generate one or more hierarchical encodings each representing different types of scene information.

8. The system of claim 7 , wherein the one or more processors are further to cause a neural network to generate a three-dimensional voxel latent representation of a scene using a plurality of two-dimensional images of the scene and one or more parameters of one or more cameras used to capture the plurality of two-dimensional images, wherein the one or more hierarchical encodings are generated from the three-dimensional voxel latent representation.

9. The system of claim 8 , wherein the neural network is an autoencoder trained to generate the three-dimensional voxel latent representation based, at least in part, upon a density voxel representation and a feature voxel representation for this scene, generated using the plurality of two-dimensional images and the one or more parameters of the one or more cameras.

10. The system of claim 8 , wherein the one or more neural networks used to generate one or more hierarchical encodings includes a hierarchical set of latent encoders to decompose the three-dimensional voxel latent representation of the scene into a set of smaller latent representations of the scene for use in learning a generative model.

11. The system of claim 10 , wherein the generative model includes at least one diffusion neural network or a generative adversarial network.

12. The system of claim 8 , wherein the neural network is trained to minimize a reconstruction loss and an adversarial loss in a three-dimensional space.

13. A method, comprising:

causing, using one or more processors, one or more neural networks to generate one or more hierarchical encodings each representing different types of scene information.

14. The method of claim 13 , further comprising:

causing a neural network to generate a three-dimensional voxel latent representation of a scene using a plurality of two-dimensional images of the scene and one or more parameters of one or more cameras used to capture the plurality of two-dimensional images, wherein the one or more hierarchical encodings are generated from the three-dimensional voxel latent representation.

15. The method of claim 14 , wherein the neural network is an autoencoder trained to generate the three-dimensional voxel latent representation based, at least in part, upon a density voxel representation and a feature voxel representation for this scene, generated using the plurality of two-dimensional images and the one or more parameters of the one or more cameras.

16. The method of claim 14 , wherein the one or more neural networks used to generate one or more hierarchical encodings includes

a hierarchical set of latent encoders to decompose the three-dimensional voxel latent representation of the scene into a set of smaller latent representations of the scene for use in learning a generative model.

17. The method of claim 16 , wherein the generative model includes at least one diffusion neural network or a generative adversarial network.

18. The method of claim 14 , wherein the neural network is trained to minimize a reconstruction loss and an adversarial loss in a three-dimensional space.

19. A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

cause one or more neural networks to generate one or more hierarchical encodings each representing different types of scene information.

20. The non-transitory machine-readable medium of claim 19 , wherein the instructions, if performed, are further cause the one or more processors to cause a neural network to generate a three-dimensional voxel latent representation of a scene using a plurality of two-dimensional images of the scene and one or more parameters of one or more cameras used to capture the plurality of two-dimensional images, wherein the one or more hierarchical encodings are generated from the three-dimensional voxel latent representation.

21. The non-transitory machine-readable medium of claim 20 , wherein the neural network is an autoencoder trained to generate the three-dimensional voxel latent representation based, at least in part, upon a density voxel representation and a feature voxel representation for this scene, generated using the plurality of two-dimensional images and the one or more parameters of the one or more cameras.

22. The non-transitory machine-readable medium of claim 20 , wherein the one or more neural networks used to generate one or more hierarchical encodings includes

a hierarchical set of latent encoders to decompose the three-dimensional voxel latent representation of the scene into a set of smaller latent representations of the scene for use in learning a generative model.

23. The non-transitory machine-readable medium of claim 22 , wherein the generative model includes at least one diffusion neural network or a generative adversarial network.

24. The non-transitory machine-readable medium of claim 20 , wherein the neural network is trained to minimize a reconstruction loss and an adversarial loss in a three-dimensional space.

25. A scene generation system, comprising:

one or more processors to cause one or more neural networks to generate one or more hierarchical encodings each representing different types of scene information; and

memory for storing network parameters for the one or more neural networks.

26. The scene generation system of claim 25 , wherein the one or more processors are further to cause a neural network to generate a three-dimensional voxel latent representation of a scene using a plurality of two-dimensional images of the scene and one or more parameters of one or more cameras used to capture the plurality of two-dimensional images, wherein the one or more hierarchical encodings are generated from the three-dimensional voxel latent representation.

27. The scene generation system of claim 26 , wherein the neural network is an autoencoder trained to generate the three-dimensional voxel latent representation based, at least in part, upon a density voxel representation and a feature voxel representation for this scene, generated using the plurality of two-dimensional images and the one or more parameters of the one or more cameras.

28. The scene generation system of claim 26 , wherein the one or more neural networks used to generate one or more hierarchical encodings includes

a hierarchical set of latent encoders to decompose the three-dimensional voxel latent representation of the scene into a set of smaller latent representations of the scene for use in learning a generative model.

29. The scene generation system of claim 28 , wherein the generative model includes at least one diffusion neural network or a generative adversarial network.

30. The scene generation system of claim 26 , wherein the neural network is trained to minimize a reconstruction loss and an adversarial loss in a three-dimensional space.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2023
From: KIM, SEUNG WOOK; KREIS, KARSTEN; LI, DAIQING; ROMBACH, ROBIN; FIDLER, SANJA; TORRALBA BARRIUSO, ANTONIO; BROWN, BRADLEY
To: NVIDIA CORPORATION
Reel/Frame 064375/0552 →
References Cited (13)
US 11257298B2 · Kim · 2022 [cited by examiner]
US 20180089888A1 · Ondruska · 2018 [cited by examiner]
US 20230005217A1 · Chen · 2023 [cited by examiner]
Kucharski B., ‘Multi View Approach to 3D Generative Models’, [online, downloaded Feb. 29, 2024], https://bryonkucharski.github.io/files/MV3DVAEGAN_report.pdf, Jan. 9, (Year: 2021). [cited by examiner]
Su et al., ‘Multi-view Convolutional Neural Networks for 3D Shape Recognition’, arXiv:1505.00880v3 [cs.CV]. (Year: 2015). [cited by examiner]
Stigeborn P., ‘Generating 3D-objects using neural networks’, Royal Institute of Technology. (Year: 2018). [cited by examiner]
Wu et al., ‘Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling’, arXiv:1610.07584v2 [cs.CV]. (Year: 2017). [cited by examiner]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
International Electrotechnical Commission, “Functional safety of electrical/electronic/programmable electronic safety-related systems,” IEC Standard 61508-1, Apr. 2014, 23 pages. [cited by applicant]
International Organization for Standardization, “Road vehicles—Functional safety,” ISO Standard 26262, https://www.iso.org/obp/ui/#iso:std:iso:26262:-1:ed-1:v1:en, Nov. 11, 2011, 35 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Standard No. J3016-201806, dated Jun. … [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Ben Mildenhall, et al., “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”, arXiv:2003.08934v2, Aug. 3, 2020, pp. 1-25. [cited by applicant]
Cited By (8)
US 12,482,128 US 12,518,515 US 12,555,043 US 12,662,137 US 12,688,641 US 12,724,734 US 12,731,329 US 12,737,964