IP Library › Granted Patent US 12,633,034
Granted Patent B2
US 12,633,034 · App. 18/218,527 · Granted May 19, 2026

Generating images based on coordinates associated with objects depicted in the image using neural networks

Inventors: Valts Blukis (Seattle, WA); Taeyeop Lee (Daejeon, KR); Jonathan Tremblay (Redmond, WA); Bowen Wen (Bellevue, WA); Dieter Fox (Seattle, WA); Stanley Thomas Birchfield (Sammamish, WA)
Assignee: NVIDIA Corporation
G06T15/06G06T1/20G06T7/70G06T7/90G06V10/82H04N13/117G06T2207/10024G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,034
App. No.
18/218,527
Granted
May 19, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to generate an image of one or more objects. In at least one embodiment, an image of one or more objects is generated using a neural network based on, for example, a representation of a scene.

Claims (52)

1 . A system, comprising:

at least one processor; and

at least one memory comprising instructions that, in response to execution by the at least one processor, cause the system to at least:

generate a set of data indicating a plurality of rays that intersect with a plurality of objects within a scene based, at least in part, on a representation of the scene that comprises a plurality of latent codes comprising one or more latent codes for individual ones of the plurality of objects;

determine a plurality of coordinates associated with the plurality of objects in the scene based, at least in part, on where the plurality of rays intersect with the plurality of objects; and

generate an image based, at least in part, on a set of color and density values generated by at least one neural network based, at least in part, on the plurality of coordinates and the plurality of latent codes.

2 . The system of claim 1 , wherein the image is a first image depicting the plurality of objects in the scene from a first viewpoint associated with the plurality of rays, and the at least one memory comprises further instructions that, in response to execution by the at least one processor, cause the system to at least:

use at least the representation of the scene to generate a second image that depicts the plurality of objects in the scene from a second viewpoint.

3 . The system of claim 1 , wherein the at least one memory comprises further instructions that, in response to execution by the at least one processor, cause the system to at least:

calculate the plurality of latent codes based, at least in part, on one or more images of the scene, wherein the plurality of latent codes correspond to the plurality of objects.

4 . The system of claim 1 , wherein a first coordinate of the plurality of coordinates is along a first ray of the plurality of rays and is within a volume of a first object of the plurality of objects.

5 . The system of claim 1 , wherein the at least one memory comprises further instructions that, in response to execution by the at least one processor, cause the system to at least:

generate the plurality of rays in reference to a camera pose.

6 . The system of claim 1 , wherein the representation comprises a scene graph corresponding to the scene.

7 . A method, comprising:

obtaining a representation of a plurality of objects in a scene;

generating a set of matrices indicating a plurality of rays that intersect the plurality of objects in the scene based, at least in part, on where the plurality of rays intersect the representation, and wherein the plurality of rays are associated with a first viewpoint of the scene;

determining a set of coordinates along the plurality of rays, indicating where the plurality of rays intersect the plurality of objects based, at least in part, on the set of matrices; and

using a neural network to generate an image of the scene from the first viewpoint based, at least in part, on the set of coordinates and a plurality of latent codes associated with the plurality of objects, wherein the plurality of latent codes comprises one or more latent codes for individual ones of the plurality of objects.

8 . The method of claim 7 , further comprising:

generating a second set of matrices based, at least in part, on a second viewpoint in the scene;

determining a second set of coordinates based, at least in part, on the second set of matrices, the image being a first image; and

using the neural network to generate a second image of the scene from the second viewpoint based, at least in part, on the second set of coordinates and the plurality of latent codes.

9 . The method of claim 7 , further comprising:

determining the plurality of latent codes based, at least in part, on one or more images of the plurality of objects.

10 . The method of claim 7 , wherein the set of coordinates includes a plurality of coordinates that are within a plurality of volumes of the plurality of objects.

11 . The method of claim 7 , wherein the image is generated using one or more graphics processing units (GPUs).

12 . The method of claim 7 , wherein the representation indicates at least one or more poses of the plurality of objects in the scene.

13 . A non-transitory computer-readable medium comprising instructions that, when performed by at least one processor of a computing device, cause the computing device to at least:

obtain a representation of a scene indicating at least a plurality of first poses of a plurality of objects;

determine a plurality of rays that intersect the plurality of objects based, at least in part, on where the plurality of rays intersect the representation of the scene;

determine a set of coordinates based, at least in part, on the plurality of rays, wherein the set of coordinates indicate where the plurality of rays intersect the plurality of objects in the scene;

use a neural network to generate first data representing the plurality of objects based, at least in part, on the set of coordinates and a plurality of latent codes associated with the plurality of objects, wherein the plurality of latent codes comprises one or more latent codes determined for individual ones of the plurality of objects; and

generate a first image based at least in part on the first data, the first image comprising one or more first visual representations of the plurality of objects as viewed from a first viewpoint.

14 . The non-transitory computer-readable medium of claim 13 , comprising further instructions that when performed by the at least one processor of the computing device, cause the computing device to at least:

modify the representation of the scene to indicate at least one or more second poses of the plurality of objects;

use the neural network to generate second data representing the plurality of objects based, at least in part, on the representation of the scene; and

generate a second image based at least in part on the second data.

15 . The non-transitory computer-readable medium of claim 13 , comprising further instructions that when performed by the at least one processor of the computing device, cause the computing device to at least:

obtain a plurality of images depicting the plurality of objects; and

determine the plurality of latent codes based, at least in part, on the plurality of images and the neural network.

16 . The non-transitory computer-readable medium of claim 13 , comprising further instructions that when performed by the at least one processor of the computing device, cause the computing device to at least:

use at least the neural network to determine one or more pixel colors corresponding to the plurality of rays; and

generate the first image from at least the plurality of pixel colors.

17 . The non-transitory computer-readable medium of claim 13 , wherein the plurality of rays are indicated at least in part by one or more matrices.

18 . The non-transitory computer-readable medium of claim 13 , wherein each coordinate of the set of coordinates corresponds to a respective object of the plurality of objects.

19 . The non-transitory computer-readable medium of claim 13 , wherein the neural network is a decoder.

20 . The non-transitory computer-readable medium of claim 13 , wherein:

the plurality of objects are of a particular category; and

the neural network is trained in connection with one or more images depicting objects of one or more categories.

21 . The non-transitory computer-readable medium of claim 13 , comprising further instructions that when performed by the at least one processor of the computing device, cause the computing device to at least:

generate a second image based at least in part on the representation of the scene, the second image comprising one or more second visual representations of the plurality of objects as viewed from a second viewpoint.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2023
From: BLUKIS, VALTS; LEE, TAEYEOP; TREMBLAY, JONATHAN; WEN, BOWEN; FOX, DIETER; BIRCHFIELD, STANLEY THOMAS
To: NVIDIA CORPORATION
Reel/Frame 064201/0188 →
Continuity (2)
Provisional Application 63419661 · Oct 26, 2022
Related Publication 20240153196A1 · May 9, 2024
References Cited (38)
US 20220239844A1 · Lv · 2022 [cited by examiner]
US 20230252716A1 · Mcallister · 2023 [cited by examiner]
US 20230343014A1 · Ummenhofer · 2023 [cited by examiner]
Danny Driess et al., Reinforcement Learning with Neural Radiance Fields, arXiv:2206.01634v1 [cs.LG], Jun. 3, 2022, pp. 1-23. [cited by examiner]
Abou-Chakra et al., “Implicit Object Mapping with Noisy Data,” 2022, 7 pages. [cited by applicant]
Adamkiewicz et al., “Vision-Only Robot Navigation in a Neural Radiance World,” Jan. 4, 2022, 8 Pages. [cited by applicant]
Chan et al., “Efficient Geometry-aware 3D Generative Adversarial Networks,” CVPR, 2022, 11 Pages. [cited by applicant]
Downs et al., “Google Scanned Objects: A High-Quality Dataset of 3D Scanned Household Items,” 2022, 8 pages. [cited by applicant]
Driess et al., “Learning Multi-Object Dynamics with Compositional Neural Radiance Fields,” 6th Annual Conference on Robot Learning (CoRL), 2022, 14 Pages. [cited by applicant]
Driess et al., “Reinforcement Learning with Neural Radiance Fields,” 36th Conference on Neural Information Processing Systems (NeurIPS), 2022, 15 Pages. [cited by applicant]
Fridovivh-Keil et al., “Plenoxels: Radiance Fields without Neural Networks,” CVPR, 2022, 10 Pages. [cited by applicant]
Granskog et al., “Neural Scene Graph Rendering,” ACM Transactions on Graphics (TOG), Aug. 2021, 11 Pages. [cited by applicant]
Ichnowski et al., “Dex-NeRF: Using a Neural Radiance Field to Grasp Transparent Objects,” Oct. 27, 2021, 11 Pages. [cited by applicant]
IEEE “IEEE Standard for Floating-Point Arithmetic”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008, 70 pages. [cited by applicant]
Jang et al., “CodeNeRF: Disentangled Neural Radiance Fields for Object Categories,” IEEE/CVF International Conference on Computer Vision, 2021, 10 Pages. [cited by applicant]
Kay et al., “Ray Tracing Complex Scenes,” ACM SIGGRAPH Computer Graphics, 20(4): 1986, 10 pages. [cited by applicant]
Kerr et al., “Evo-neRF: Evolving neRF for Sequential Robot Grasping, ” 6th Annual Conference on Robot Learning, 2022, 15 pages. [cited by applicant]
Li et al., “3D Neural Scene Representations for Visuomotor Control,” 5th Conference on Robot Learning, CoRL, 2021. 12 Pages. [cited by applicant]
Lin et al., “MIRA: Mental Imagery for Robotic Affordances,” 6th Annual Conference on Robot Learning, 2022, 12 pages. [cited by applicant]
Mildenhall et al., “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” Proceedings of the European Conference on Computer Vision, Aug. 3, 2020, 25 pages. [cited by applicant]
Morrical et al., “NViSII: A Scriptable Tool for Photorealistic Image Generation,” May 28, 2021, 12 pages. [cited by applicant]
Muller et al., “Instant Neural Graphics Primitives with a Multiresolution Hash Encoding,” 2022, 15 pages. [cited by applicant]
Muller et al., “AutoRF: Learning 3D Object Radiance Fields from Single View Observations,” CVPR, 2022, 10 pages. [cited by applicant]
Orperel et al., “NVIDIAGameWorks/kaolin-wisp,” GitHub, retrieved from https://github.com/NVIDIAGameWorks/ kaolin-wisp, May 2, 2023, 5 pages. [cited by applicant]
Pantic et al., “Sampling-free obstacle gradients and reactive planning in Neural Radiance Fields,” May 3, 2022, 3 Pages. [cited by applicant]
Sitzmann et al., “Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations,” Advances in Neural Information Processing Systems, Jun. 4, 2019, 20 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Standard No. J3016-201806, dated Jun. … [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, Sep. 30, 2… [cited by applicant]
Sucar et al., “iMAP: Implicit Mapping and Positioning in Real-Time,” Proceedings of the International Conference on Computer Vision (ICCV), 2021, 10 Pages. [cited by applicant]
Takikawa et al., “Kaolin Wisp: A Pytorch Library and Engine for Neural Fields Research,” retrieved from https://github.com/NVIDIAGameWorks/kaolin-wisp, 2022, 5 Pages. [cited by applicant]
Turki et al., “Mega-NeRF: Scalable Construction of Large-Scale NeRFs for Virtual Fly-Throughs,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, 10 Pages. [cited by applicant]
Wang et al., “Image Quality Assessment: From Error Visibility to Structural Similarity,” IEEE, Transactions on Image Processing, Apr. 2004, 13(4): 14 pages. [cited by applicant]
Yang et al., “Learning Object-compositional Neural Radiance Field for Editable Scene Rendering,” International Conference on Computer Vision, Oct. 2021, 10 pages. [cited by applicant]
Yen-Chen, et al., “iNeRF: Inverting Neural Radiance Fields for Pose Estimation,” Aug. 10, 2021, 9 Pages. [cited by applicant]
Yen-Chn et al., “NeRF-Supervision: Learning Dense Object Descriptors from Neural Radiance Fields,” Apr. 27, 2022, 8 Pages. [cited by applicant]
Yu et al., “pixelNeRF: Neural Radiance Fields from One or Few Images,” Dec. 3, 2020, 20 Pages. [cited by applicant]
Zhang et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric,” CVPR, 2018, 10 pages. [cited by applicant]
Zhong et al., “Touching a NeRF: Leveraging Neural Radiance Fields for Tactile Sensory Data Generation,” 6th Annual Conference on Robot Learning (CoRL), 2022, 11 Pages. [cited by applicant]