Increasing levels of detail for neural fields using diffusion models
Systems and methods of the present disclosure include providing higher levels of detail (LODs) for generated three-dimensional (3D) models, such as those represented by neural radiance fields (NeRFs). A 3D model may be presented to a user in which the user may request additional LODs, such as to zoom into the image or to receive information about features within the image. A request to generate finer levels of detail may include using one or more diffusion models to generate images at higher resolutions and/or to hallucinate finer details based on information extracted from the original image or text prompts. Newly generated images may then be added to a set of images associated with the 3D models to enable later model generation to have finer details.
1 . A computer-implemented method, comprising:
determining a target level of detail for a three-dimensional (3D) volume;
generating, using an image generation network and based at least on a current view representing the 3D volume, an updated view representing a selected portion of the 3D volume at the target level of detail, wherein at least a portion of the 3D volume is removed in the updated view;
providing, responsive to the target level of detail, the updated view;
adding the updated view to a set of images associated with the 3D volume; and
updating the 3D volume, based at least on the updated view.
2 . The computer-implemented method of claim 1 , wherein the 3D volume is represented by a neural radiance field (NeRF).
3 . The computer-implemented method of claim 1 , wherein the image generation network is a diffusion model conditioned on both text and images.
4 . The computer-implemented method of claim 1 , further comprising:
receiving a prompt for the current view; and
providing, to a language model associated with the image generation network, the prompt.
5 . The computer-implemented method of claim 4 , wherein the language model is a large language model (LLM) configured to generate a hierarchy of information based, at least, on the prompt.
6 . The computer-implemented method of claim 1 , wherein the image generation network is a super-resolution model conditioned on an image having a resolution less than a threshold.
7 . The computer-implemented method of claim 1 , further comprising:
removing, upon receiving the updated view, one or more previous images from the set of images; and
updating a network associated with the 3D volume.
8 . The computer-implemented method of claim 1 , wherein the target level of detail is associated with an input command from a user of an interactive environment.
9 . The computer-implemented method of claim 1 , wherein the 3D volume is represented by a neural radiance field (NeRF), the method further comprising:
converting the NeRF to a mesh-based representation.
10 . The computer-implemented method of claim 1 , further comprising:
receiving, at an associated LLM, a prompt requesting a hierarchy of information for an object associated with the 3D volume;
determining a plurality of sub-levels for the 3D volume based, at least, on the hierarchy of information; and
establishing an ordering for the plurality of sub-levels associated with a respective level for each sub-level of the plurality of sub-levels.
11 . The computer-implemented method of claim 10 , further comprising:
storing the plurality of sub-levels;
providing, responsive to a first command, the object; and
providing, responsive to a second command, a sub-level of the plurality of sub-levels.
12 . A processor comprising:
one or more processing units to:
receive a request to generate an image using a neural radiance field (NeRF);
determine, from a prompt associated with the request, a target level of detail for the image corresponding to a target focal length of a digital camera associated with a representation of the NeRF;
determine that images generated using the NeRF will not meet the target level of detail;
generate, via one or more diffusion models, a new image at the target level of detail based at least on a current view representing the NeRF;
provide, responsive to the request, the new image corresponding to a selected portion of the NeRF at the target level of detail with the target focal length that is narrower than a first focal length associated with current view; and
add the new image to a set of images associated with the NeRF.
13 . The processor of claim 12 , wherein the prompt is a text prompt, and wherein the one or more processing units are further to:
provide the text prompt to a large language model (LLM);
receive, from the LLM, a command based, at least, on the text prompt; and
provide the command to the one or more diffusion models.
14 . The processor of claim 12 , wherein the one or more diffusion models are conditioned on both text and images.
15 . The processor of claim 12 , wherein at least one diffusion model of the one or more diffusion models include a super-resolution model conditioned on an image having a resolution less than a threshold.
16 . The processor of claim 12 , wherein the processor is comprised in at least one of:
a system for performing simulation operations;
a system for performing simulation operations to test or validate autonomous machine applications;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for rendering graphical output;
a system for performing deep learning operations;
a system implemented using an edge device;
a system for generating or presenting virtual reality (VR) content;
a system for generating or presenting augmented reality (AR) content;
a system for generating or presenting mixed reality (MR) content;
a system incorporating one or more Virtual Machines (VMs);
a system for performing operations for a conversational AI application;
a system for performing operations for a generative AI application;
a system for performing operations using a language model;
a system for performing one or more generative content operations using a large language model (LLM);
a system implemented at least partially in a data center;
a system for performing hardware testing using simulation;
a system for performing one or more generative content operations using a language model;
a system for synthetic data generation;
a collaborative content creation platform for 3D assets; or
a system implemented at least partially using cloud computing resources.
17 . A system, comprising:
one or more processors comprising processing circuitry to generate an output image with a finer level of detail (LOD) than an input image generated using a neural radiance field (NeRF), to update the NeRF using a set of images that includes the output image, and to render the output image at the finer LOD with at least a portion of the input image removed.
18 . The system of claim 17 , wherein the output image is generated by one or more diffusion models responsive to a request.
19 . The system of claim 17 , wherein the output image is at least one of a higher resolution image relative to the input image, or a hallucinated image.
20 . The system of claim 17 , wherein the system comprises at least one of:
a system for performing simulation operations;
a system for performing simulation operations to test or validate autonomous machine applications;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for rendering graphical output;
a system for performing deep learning operations;
a system implemented using an edge device;
a system for generating or presenting virtual reality (VR) content;
a system for generating or presenting augmented reality (AR) content;
a system for generating or presenting mixed reality (MR) content;
a system incorporating one or more Virtual Machines (VMs);
a system for performing operations for a conversational AI application;
a system for performing operations for a generative AI application;
a system for performing operations using a language model;
a system for performing one or more generative content operations using a large language model (LLM);
a system implemented at least partially in a data center;
a system for performing hardware testing using simulation;
a system for performing one or more generative content operations using a language model;
a system for synthetic data generation;
a collaborative content creation platform for 3D assets; or
a system implemented at least partially using cloud computing resources.