IP Library › Granted Patent US 12,608,873
Granted Patent B2
US 12,608,873 · App. 18/479,261 · Granted Apr 21, 2026

Interactive neural field editing in content generation systems and applications

Inventors: Karsten Julian Kreis (Vancouver, CA); Maria Shugrina (Toronto, CA); Ming-Yu Liu (San Jose, CA); Or Perel (Tel Aviv, IL); Sanja Fidler (Toronto, CA); Towaki Alan Takikawa (Toronto, CA); Tsung-Yi Lin (Sunnyvale, CA); Xiaohui Zeng (Toronto, CA)
Assignee: Nvidia Corporation
G06T15/06G06T15/005G06T2207/20084G06T2210/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,873
App. No.
18/479,261
Granted
Apr 21, 2026
Kind
B2
Abstract

Systems and methods of the present disclosure include interactive editing for generated three-dimensional (3D) models, such as those represented by neural radiance fields (NeRFs). A 3D model may be presented to a user in which the user may identify one or more localized regions for editing and/or modification. The localized regions may be selected and a corresponding 3D volume for that region may be provided to one or more generative networks, along with a prompt, to generate new content for the localized regions. Each of the original NeRF and the newly generated NeRF for the new content may then be combined into a single NeRF for a combined 3D representation with the original content and the localized modifications.

Claims (82)

1 . A computer-implemented method, comprising:

receiving an input corresponding to content for a selected portion of a 3D volume of a first neural radiance field (NeRF) corresponding to a scene;

receiving an interaction to the first NeRF, the interaction corresponding to the selected portion;

determining, based at least on a parameter of the interaction, a depth for the selected portion;

generating the selected portion using the depth;

generating a second NeRF of the selected portion based, at least, on the input; and

generating an output NeRF by blending the first NeRF and the second NeRF.

2 . The computer-implemented method of claim 1 , further comprising:

determining, for a region of the first NeRF, a feature value;

determining, for the region of the second NeRF, an overwriting feature value; and

replacing, in the output NeRF, the feature value with the overwriting feature value.

3 . The computer-implemented method of claim 1 , wherein the input is a text prompt.

4 . The computer-implemented method of claim 1 , wherein the output NeRF is generated by a diffusion model.

5 . The computer-implemented method of claim 1 , wherein the input is associated with a tool used to make the selection.

6 . The computer-implemented method of claim 1 , wherein the selected portion is a two-dimensional selection based on a first camera view, further comprising:

generating one or more rays from a virtual camera associated with the first camera view;

identifying one or more cells interacting with the one or more rays; and

determining a depth for the one or more cells.

7 . The computer-implemented method of claim 6 , wherein the depth is correlated to a duration of interaction with the selected portion.

8 . The computer-implemented method of claim 1 , further comprising:

generating a mask corresponding to the selected portion; and

providing the mask to a content generation pipeline.

9 . The computer-implemented method of claim 1 , further comprising:

rendering the content within an interaction environment; and

providing one or more tools, within the interaction environment, to identify the selected portion.

10 . The computer-implemented method of claim 9 , wherein both the scene and the content are generated objects using one or more content generation models.

11 . A processor comprising:

one or more processing units to:

generate, via one or more diffusion models, a scene representation as a neural radiance field (NeRF);

receive an input corresponding to a portion of the NeRF;

generate, based on an indication, a modified NeRF for the portion, the modified NeRF including one or more content portions different from the NeRF;

combine the modified NeRF and the NeRF to form an updated NeRF, wherein the portion of the NeRF is a three-dimensional (3D) representation generated from a two-dimensional (2D) input and one or more properties of the 2D input; and

provide a visual representation of the updated NeRF.

12 . The processor of claim 11 , wherein combining the modified NeRF and the NeRF corresponds to overwriting one or more feature values of the NeRF with modified feature values of the modified NeRF.

13 . The processor of claim 11 , wherein the indication is at least one of a textual prompt, an auditory prompt, or a feature associated with the input.

14 . The processor of claim 11 , wherein the processor is comprised in at least one of:

a system for performing simulation operations;

a system for performing simulation operations to test or validate autonomous machine applications;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for rendering graphical output;

a system for performing deep learning operations;

a system implemented using an edge device;

a system for generating or presenting virtual reality (VR) content;

a system for generating or presenting augmented reality (AR) content;

a system for generating or presenting mixed reality (MR) content;

a system incorporating one or more Virtual Machines (VMs);

a system for performing operations for a conversational AI application;

a system for performing operations for a generative AI application;

a system for performing operations using a language model;

a system for performing one or more generative content operations using a large language model (LLM);

a system implemented at least partially in a data center;

a system for performing hardware testing using simulation;

a system for performing one or more generative content operations using a language model;

a system for synthetic data generation;

a collaborative content creation platform for 3D assets; or

a system implemented at least partially using cloud computing resources.

15 . A system, comprising:

one or more processors comprising processing circuitry to generate an output representation of a three-dimensional (3D) scene as a combined neural radiance field (NeRF), the combined NeRF including an original volume and a modified volume added to the original volume, the modified volume being generated for a selected volume of the original volume based, at least, on an input corresponding to content for the modified volume, wherein the modified volume is computed from a second input corresponding to a two-dimensional area and a property of the second input.

16 . The system of claim 15 , wherein objects associated with both the original volume and the modified volume are formed using one or more diffusion models.

17 . The system of claim 15 , wherein the system comprises at least one of:

a system for performing simulation operations;

a system for performing simulation operations to test or validate autonomous machine applications;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for rendering graphical output;

a system for performing deep learning operations;

a system implemented using an edge device;

a system for generating or presenting virtual reality (VR) content;

a system for generating or presenting augmented reality (AR) content;

a system for generating or presenting mixed reality (MR) content;

a system incorporating one or more Virtual Machines (VMs);

a system for performing operations for a conversational AI application;

a system for performing operations for a generative AI application;

a system for performing operations using a language model;

a system for performing one or more generative content operations using a large language model (LLM);

a system implemented at least partially in a data center;

a system for performing hardware testing using simulation;

a system for performing one or more generative content operations using a language model;

a system for synthetic data generation;

a collaborative content creation platform for 3D assets; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2023
From: KREIS, KARSTEN JULIAN; SHUGRINA, MARIA; LIU, MING-YU; PEREL, OR; FIDLER, SANJA; TAKIKAWA, TOWAKI ALAN; LIN, TSUNG-YI; ZENG, XIAOHUI
To: NVIDIA CORPORATION
Reel/Frame 065097/0296 →
Continuity (1)
Related Publication 20250111588A1 · Apr 3, 2025
References Cited (6)
Song et al (NPL), (Sep. 2023), “Blending-NeRF: Text-Driven Localized Editing in Neural Radiance Fields.” (Year: 2023). [cited by examiner]
Haque et al (NPL), (Jun. 2023), “Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions.” (Year: 2023). [cited by examiner]
Mikaeili et al (NPL), (Aug. 2023), “SKED: Sketch-guided Text-based 3D Editing.” (Year: 2023). [cited by examiner]
Yuan et al (NPL), (2022), “NeRF-Editing: Geometry Editing of Neural Radiance Fields.” (Year: 2022). [cited by examiner]
Martin-Brualla et al (NPL), (2021), “NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections.” (Year: 2021). [cited by examiner]
Poole et al. (NPL), (Sep. 2022), “DreamFusion: Text-to-3D using 2D Diffusion.” (Year: 2022). [cited by examiner]