IP Library › Granted Patent US 12,646,253
Granted Patent B2
US 12,646,253 · App. 18/619,163 · Granted Jun 2, 2026

Neural semantic 3D capture with neural radiance fields

Inventor: Yangming Wen (Fremont, CA)
Assignee: Electronic Arts Inc.
G06T17/00G06V10/82G06V20/70G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,253
App. No.
18/619,163
Granted
Jun 2, 2026
Kind
B2
Abstract

A method of generating a three-dimensional (3D) model includes obtaining a set of two-dimensional (2D) images of a scene acquired by one or more cameras from a plurality of camera angles at a plurality of camera positions. Each 2D image corresponds to a respective camera angle and a respective camera position. The method further includes obtaining the respective camera angle and the respective camera position for each 2D image, and generating one or more semantic masks from the set of 2D images. Each semantic mask corresponds to a class of one or more objects in the scene. The method further includes training a neural radiance field (NeRF) model, using the set of 2D images and the one or more semantic masks as a training dataset, to obtain a trained NeRF model. The trained NeRF model is an implicit 3D model of the one or more objects in the scene.

Claims (43)

1 . A method of generating a three-dimensional (3D) model, the method comprising:

obtaining a set of two-dimensional (2D) images of a scene acquired by one or more cameras from a plurality of camera angles at a plurality of camera positions, wherein each 2D image in the set of 2D images corresponds to a respective camera angle and a respective camera position;

obtaining the respective camera angle and the respective camera position for each 2D image in the set of 2D images;

generating one or more semantic masks from the set of 2D images, wherein each semantic mask corresponds to a class of one or more objects in the scene and includes a respective group of associated voxels; and

training a neural radiance field (NeRF) model, using the set of 2D images, and the one or more semantic masks as a training dataset, to obtain a trained NeRF model, the trained NeRF model being an implicit 3D model of the one or more objects in the scene, wherein the training the NeRF model guides the trained NeRF model to fit to the one or more semantic masks.

2 . The method of claim 1 , further comprising postprocessing the implicit 3D model to generate a 3D mesh representing at least one of the one or more objects in the scene.

3 . The method of claim 2 , wherein the postprocessing is performed using a Marching Cubes algorithm.

4 . The method of claim 1 , wherein the training of the NeRF model comprises:

iteratively generating one or more predicted semantic masks;

wherein the NeRF model is trained by minimizing a loss function that relates to differences between the one or more predicted semantic masks and the one or more semantic masks.

5 . The method of claim 1 , wherein the generating the one or more semantic masks is performed using a machine-learning model.

6 . The method of claim 5 , wherein the machine-learning model comprises a convolutional neural network.

7 . The method of claim 5 , wherein the machine-learning model and the NeRF model are integrally trained together.

8 . The method of claim 1 , wherein the obtaining the respective camera angle and the respective camera position for each 2D image in the set of 2D images comprises:

preprocessing each 2D image in the set of 2D images using a Colmap algorithm to obtain the respective camera angle and the respective camera position for the 2D image.

9 . The method of claim 1 , wherein the one or more cameras comprise a mobile phone camera or a video camera.

10 . The method of claim 1 , wherein the scene comprises a casual world scene or a lightstage scene.

11 . A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause a computing device to generate a three-dimensional (3D) model, by performing the steps of:

obtaining a set of two-dimensional (2D) images of a scene acquired by one or more cameras from a plurality of camera angles at a plurality of camera positions, wherein each 2D image in the set of 2D images corresponds to a respective camera angle and a respective camera position;

obtaining the respective camera angle and the respective camera position for each 2D image in the set of 2D images;

generating one or more semantic masks from the set of 2D images, wherein each semantic mask corresponds to a class of one or more objects in the scene and includes a respective group of associated voxels; and

training a neural radiance field (NeRF) model, using the set of 2D images and the one or more semantic masks as a training dataset, to obtain a trained NeRF model, the trained NeRF model being an implicit 3D model of the one or more objects in the scene, wherein the training the NeRF model guides the trained NeRF model to fit to the one or more semantic masks.

12 . The non-transitory computer-readable storage medium of claim 11 , wherein the instructions, when executed by the one or more processors, further cause the computing device to perform the step of:

postprocessing the implicit 3D model to generate a 3D mesh representing at least one of the one or more objects in the scene.

13 . The non-transitory computer-readable storage medium of claim 11 , wherein the training of the NeRF model comprises:

iteratively generating one or more predicted semantic masks;

wherein the NeRF model is trained by minimizing a loss function that relates to differences between the one or more predicted semantic masks and the one or more semantic masks.

14 . The non-transitory computer-readable storage medium of claim 11 , wherein the generating the one or more semantic masks is performed using a machine-learning model.

15 . The non-transitory computer-readable storage medium of claim 14 , wherein the machine-learning model comprises a convolutional neural network.

16 . The non-transitory computer-readable storage medium of claim 14 , wherein the machine-learning model and the NeRF model are integrally trained together.

17 . A device for generating a three-dimensional (3D) model, the device comprising:

a memory storing instructions; and

one or more processors configured to execute the instructions to cause the device to:

obtain a set of two-dimensional (2D) images of a scene acquired by one or more cameras from a plurality of camera angles at a plurality of camera positions, wherein each 2D image in the set of 2D images corresponds to a respective camera angle and a respective camera position;

obtain the respective camera angle and the respective camera position for each 2D image in the set of 2D images;

generate one or more semantic masks from the set of 2D images, wherein each semantic mask corresponds to a class of one or more objects in the scene and includes a respective group of associated voxels; and

train a neural radiance field (NeRF) model, using the set of 2D images and the one or more semantic masks as a training dataset, to obtain a trained NeRF model, the trained NeRF model being an implicit 3D model of the one or more objects in the scene, wherein the training the NeRF model guides the trained NeRF model to fit to the one or more semantic masks.

18 . The device of claim 17 , wherein the instructions, when executed by the one or more processors, further cause the device to:

postprocess the implicit 3D model to generate a 3D mesh representing at least one of the one or more objects in the scene.

19 . The device of claim 17 , wherein the training of the NeRF model comprises:

iteratively generating one or more predicted semantic masks;

wherein the NeRF model is trained by minimizing a loss function that relates to differences between the one or more predicted semantic masks and the one or more semantic masks.

20 . The device of claim 17 , wherein the generating the one or more semantic masks is performed using a machine-learning model, and wherein the machine-learning model and the NeRF model are integrally trained together.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2024
From: WEN, YANGMING
To: ELECTRONIC ARTS INC.
Reel/Frame 066940/0419 →
Continuity (1)
Related Publication 20250308152A1 · Oct 2, 2025
References Cited (8)
US 12430934B2 · Sarkar · 2025 [cited by examiner]
US 20230023126A1 · Ansari · 2023 [cited by examiner]
US 20240104831A1 · Lin · 2024 [cited by examiner]
Metzer et al. “Latent-NeRF for Shape-Guided Generation of 3D Shapes and Textures,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12663-12673 (available at: https://arxiv.o… [cited by applicant]
Mildenhall et al. “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” Communications of the ACM, vol. 65, Iss. 1, pp. 99-106 (available at: https://arxiv.org/abs/2003.08934) (Dec. 17, 2021). [cited by applicant]
Kirillov et al. “Segment Anything,” Meta AI Research, FAIR (available at: https://arxiv.org/abs/2304.02643) (Apr. 5, 2023). [cited by applicant]
Zhang et al. “Dive into Deep Learning” (available at: https://d2l.ai/d2l-en.pdf) (Feb. 10, 2023). [cited by applicant]
Wang et al. “NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction,” 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Version 3 (available at: https://arxiv.or… [cited by applicant]