IP Library › Granted Patent US 12,737,964
Granted Patent B2
US 12,737,964 · App. 18/513,105 · Granted Sep 15, 2026

Increasing levels of detail for neural fields using diffusion models

Inventors: Or Perel (Tel Aviv, IL); Maria Shugrina (Toronto, CA); Yoni Kasten (Hinanit, IL); Or Litany (Sunnyvale, CA); Gal Chechik (Ramat Hasharon, IL); Sanja Fidler (Toronto, CA)
Assignee: Nvidia Corporation
G06T15/08G06T15/20G06T2210/36
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,964
App. No.
18/513,105
Granted
Sep 15, 2026
Kind
B2
Abstract

Systems and methods of the present disclosure include providing higher levels of detail (LODs) for generated three-dimensional (3D) models, such as those represented by neural radiance fields (NeRFs). A 3D model may be presented to a user in which the user may request additional LODs, such as to zoom into the image or to receive information about features within the image. A request to generate finer levels of detail may include using one or more diffusion models to generate images at higher resolutions and/or to hallucinate finer details based on information extracted from the original image or text prompts. Newly generated images may then be added to a set of images associated with the 3D models to enable later model generation to have finer details.

Claims (89)

1 . A computer-implemented method, comprising:

determining a target level of detail for a three-dimensional (3D) volume;

generating, using an image generation network and based at least on a current view representing the 3D volume, an updated view representing a selected portion of the 3D volume at the target level of detail, wherein at least a portion of the 3D volume is removed in the updated view;

providing, responsive to the target level of detail, the updated view;

adding the updated view to a set of images associated with the 3D volume; and

updating the 3D volume, based at least on the updated view.

2 . The computer-implemented method of claim 1 , wherein the 3D volume is represented by a neural radiance field (NeRF).

3 . The computer-implemented method of claim 1 , wherein the image generation network is a diffusion model conditioned on both text and images.

4 . The computer-implemented method of claim 1 , further comprising:

receiving a prompt for the current view; and

providing, to a language model associated with the image generation network, the prompt.

5 . The computer-implemented method of claim 4 , wherein the language model is a large language model (LLM) configured to generate a hierarchy of information based, at least, on the prompt.

6 . The computer-implemented method of claim 1 , wherein the image generation network is a super-resolution model conditioned on an image having a resolution less than a threshold.

7 . The computer-implemented method of claim 1 , further comprising:

removing, upon receiving the updated view, one or more previous images from the set of images; and

updating a network associated with the 3D volume.

8 . The computer-implemented method of claim 1 , wherein the target level of detail is associated with an input command from a user of an interactive environment.

9 . The computer-implemented method of claim 1 , wherein the 3D volume is represented by a neural radiance field (NeRF), the method further comprising:

converting the NeRF to a mesh-based representation.

10 . The computer-implemented method of claim 1 , further comprising:

receiving, at an associated LLM, a prompt requesting a hierarchy of information for an object associated with the 3D volume;

determining a plurality of sub-levels for the 3D volume based, at least, on the hierarchy of information; and

establishing an ordering for the plurality of sub-levels associated with a respective level for each sub-level of the plurality of sub-levels.

11 . The computer-implemented method of claim 10 , further comprising:

storing the plurality of sub-levels;

providing, responsive to a first command, the object; and

providing, responsive to a second command, a sub-level of the plurality of sub-levels.

12 . A processor comprising:

one or more processing units to:

receive a request to generate an image using a neural radiance field (NeRF);

determine, from a prompt associated with the request, a target level of detail for the image corresponding to a target focal length of a digital camera associated with a representation of the NeRF;

determine that images generated using the NeRF will not meet the target level of detail;

generate, via one or more diffusion models, a new image at the target level of detail based at least on a current view representing the NeRF;

provide, responsive to the request, the new image corresponding to a selected portion of the NeRF at the target level of detail with the target focal length that is narrower than a first focal length associated with current view; and

add the new image to a set of images associated with the NeRF.

13 . The processor of claim 12 , wherein the prompt is a text prompt, and wherein the one or more processing units are further to:

provide the text prompt to a large language model (LLM);

receive, from the LLM, a command based, at least, on the text prompt; and

provide the command to the one or more diffusion models.

14 . The processor of claim 12 , wherein the one or more diffusion models are conditioned on both text and images.

15 . The processor of claim 12 , wherein at least one diffusion model of the one or more diffusion models include a super-resolution model conditioned on an image having a resolution less than a threshold.

16 . The processor of claim 12 , wherein the processor is comprised in at least one of:

a system for performing simulation operations;

a system for performing simulation operations to test or validate autonomous machine applications;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for rendering graphical output;

a system for performing deep learning operations;

a system implemented using an edge device;

a system for generating or presenting virtual reality (VR) content;

a system for generating or presenting augmented reality (AR) content;

a system for generating or presenting mixed reality (MR) content;

a system incorporating one or more Virtual Machines (VMs);

a system for performing operations for a conversational AI application;

a system for performing operations for a generative AI application;

a system for performing operations using a language model;

a system for performing one or more generative content operations using a large language model (LLM);

a system implemented at least partially in a data center;

a system for performing hardware testing using simulation;

a system for performing one or more generative content operations using a language model;

a system for synthetic data generation;

a collaborative content creation platform for 3D assets; or

a system implemented at least partially using cloud computing resources.

17 . A system, comprising:

one or more processors comprising processing circuitry to generate an output image with a finer level of detail (LOD) than an input image generated using a neural radiance field (NeRF), to update the NeRF using a set of images that includes the output image, and to render the output image at the finer LOD with at least a portion of the input image removed.

18 . The system of claim 17 , wherein the output image is generated by one or more diffusion models responsive to a request.

19 . The system of claim 17 , wherein the output image is at least one of a higher resolution image relative to the input image, or a hallucinated image.

20 . The system of claim 17 , wherein the system comprises at least one of:

a system for performing simulation operations;

a system for performing simulation operations to test or validate autonomous machine applications;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for rendering graphical output;

a system for performing deep learning operations;

a system implemented using an edge device;

a system for generating or presenting virtual reality (VR) content;

a system for generating or presenting augmented reality (AR) content;

a system for generating or presenting mixed reality (MR) content;

a system incorporating one or more Virtual Machines (VMs);

a system for performing operations for a conversational AI application;

a system for performing operations for a generative AI application;

a system for performing operations using a language model;

a system for performing one or more generative content operations using a large language model (LLM);

a system implemented at least partially in a data center;

a system for performing hardware testing using simulation;

a system for performing one or more generative content operations using a language model;

a system for synthetic data generation;

a collaborative content creation platform for 3D assets; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 27, 2023
From: PEREL, OR; SHUGRINA, MARIA; KASTEN, YONI; LITANY, OR; CHECHIK, GAL; FIDLER, SANJA
To: NVIDIA CORPORATION
Reel/Frame 065663/0803 →
Continuity (1)
Related Publication 20250166288A1 · May 22, 2025
References Cited (8)
US 11995803B1 · Karpman · 2024 [cited by examiner]
US 12322068B1 · Kim · 2025 [cited by examiner]
US 20210042991A1 · Wei · 2021 [cited by examiner]
US 20230412865A1 · Harviainen · 2023 [cited by examiner]
US 20240005604A1 · Kreis · 2024 [cited by examiner]
US 20240161327A1 · Chen · 2024 [cited by examiner]
US 20240185498A1 · Francis · 2024 [cited by examiner]
US 20240331356A1 · Wang · 2024 [cited by examiner]