IP Library › Granted Patent US 12,555,343
Granted Patent B2
US 12,555,343 · App. 18/622,045 · Granted Feb 17, 2026

3D model generation using multimodal generative AI

Inventors: Cheng Xie (Toronto, CA); Jonathan Lorraine (Toronto, CA); Xiaohui Zeng (Toronto, CA); James Lucas (Royston, GB); Jun Gao (Toronto, CA); Sanja Fidler (Toronto, CA)
Assignee: NVIDIA Corporation
G06T19/20G06T17/20G06V20/647G06T2210/56G06T2219/2021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,343
App. No.
18/622,045
Granted
Feb 17, 2026
Kind
B2
Abstract

In various examples, systems and methods are disclosed relating to generating an output 3D latent representation by encoding, using a text encoder, a text prompt and encoding, using a 2D/3D encoder, a 2D image of an object or a 3D representation of the object. A 3D output is generated by applying the output 3D latent representation to a decoder. A reconstruction loss and a SDS loss are determined for the 3D output. At least one of the text encoder, the 2D/3D encoder, and the decoder is updated using the reconstruction loss and the SDS loss.

Claims (120)

1 . A system, comprising:

one or more circuits to:

generate an output 3D latent representation by:

encoding, using a first encoder, a text prompt; and

encoding, using a second encoder, at least one of a 2D image of an object or an input 3D representation of the object;

generate a 3D output by applying the output 3D latent representation to a decoder;

determine a reconstruction loss and a sampling loss for the 3D output; and

update at least one of the first encoder, the second encoder, or the decoder using the reconstruction loss and the sampling loss.

2 . The system of claim 1 , wherein the one or more circuits are to:

determine a first 3D latent representation by encoding, using the first encoder, the text prompt; and

determine a second 3D latent representation by encoding, using the second encoder, the 2D image of the object or the input 3D representation of the object; and

generate the output 3D latent representation by combining the first 3D latent representation and the second 3D latent representation.

3 . The system of claim 2 , wherein combining the first 3D latent representation and the second 3D latent representation comprises adding a first value at a first point of the first 3D latent representation to a second value at a second point of the second 3D latent representation to determine a value at a third point in the output 3D latent representation, wherein both the first point and the second point correspond to the third point in the output 3D latent representation.

4 . The system of claim 2 , wherein combining the first 3D latent representation and the second 3D latent representation comprises:

determining an adjusted first value at a first point of the first 3D latent representation by modifying a first value at the first point of the first 3D latent representation using a first blending parameter, wherein the first value is an output from the first encoder;

determining an adjusted second value at a second point of the second 3D latent representation by modifying a second value at the first point of the first 3D latent representation using a second blending parameter, wherein the second value is an output from the second encoder; and

adding the adjusted first value at the first point of the first 3D latent representation to the adjusted second value at the second point of the second 3D latent representation to determine a value at a third point in the output 3D latent representation, wherein both the first point and the second point correspond to the third point in the output 3D latent representation.

5 . The system of claim 1 , wherein

the first encoder is a text encoder that generates output parameters by encoding the text prompt; and

the second encoder is a 2D/3D encoder that encodes the 2D image of the object or the input 3D representation of the object based on the output parameters to generate the output 3D latent representation.

6 . The system of claim 1 , wherein

the first encoder is a text encoder that generates output parameters by encoding the text prompt;

determine adjusted output parameters by applying a blending parameter to each of the output parameters; and

the second encoder is a 2D/3D encoder that encodes the at least one of the 2D image of the object or the input 3D representation of the object based on the adjusted output parameters to generate the output 3D latent representation.

7 . The system of claim 1 , wherein the input 3D representation comprises at least one of a point cloud, a colored point cloud, an occupancy grid, a Signed Distance Fields (SDF) grid, or a 3D voxel representation.

8 . The system of claim 1 , wherein the 2D image comprises at least one of a multi-view image, a normal image, a depth image, a normal RGB image, or a depth RGB image.

9 . The system of claim 1 , wherein the 3D output comprises at least one of an occupancy field, a Signed Distance Fields (SDF) function, a texture field, a 3D field, a point cloud, a colored point cloud, a 3D mesh, or a 3D voxel.

10 . The system of claim 1 , wherein the one or more circuits generate the 3D output by:

determining a decoder output by applying the 3D latent representation to the decoder; and

generating an output 3D representation using the decoder output, the 3D output comprises at least one of the decoder output or the output 3D representation.

11 . The system of claim 10 , wherein

the decoder output comprises at least one of one or more implicit functions, one or more implicit values, one or more textures, an occupancy field, one or more Signed Distance Fields (SDF) function, a texture field, a 3D field, a point cloud, or a colored point cloud; and

the output 3D representation comprises at least one of a 3D field, a 3D mesh, or a 3D voxel.

12 . The system of claim 1 , wherein the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system implemented using a robot;

an aerial system;

a medical system;

a boating system,

a smart area monitoring system;

a system for performing deep learning operations;

a system for performing simulation operations;

a system for generating or presenting virtual reality (VR) content, augmented reality (AR) content, or mixed reality (MR) content;

a system for performing digital twin operations;

a system implemented using an edge device;

a system incorporating one or more virtual machines (VMs);

a system for generating synthetic data;

a system implemented at least partially in a data center;

a system for performing conversational artificial intelligence (AI) operations;

a system for performing generative AI operations;

a system implementing one or more language models;

a system implementing one or more large language models (LLMs);

a system implementing one or more vision language models (VLMs);

a system for hosting one or more real-time streaming applications;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets; or

a system implemented at least partially using cloud computing resources.

13 . A system, comprising:

one or more circuits to:

receive at least one of a first text prompt, a first 2D image, or a first input 3D representation of a first object; and

generate, using a machine learning model comprising a text encoder, a 2D/3D encoder, and a decoder, a first 3D output corresponding to the at least one of the first text prompt, the first 2D image, or the first input 3D representation, wherein

the machine learning model is updated by:

generating a first output 3D latent representation by:

encoding, using the text encoder, a second text prompt; and

encoding, using the 2D/3D encoder, a second 2D image of an object or a second input 3D representation of a second object;

generating a second 3D output by applying the output 3D latent representation to the decoder;

determining a reconstruction loss and a sampling loss for the second 3D output; and

updating at least one of the text encoder, the 2D/3D encoder, and the decoder using the reconstruction loss and the sampling loss.

14 . The system of claim 13 , wherein the one or more circuits are to:

determine a first 3D latent representation by encoding, using the text encoder, the first text prompt;

determine a second 3D latent representation by encoding, using the 2D/3D encoder, the first 2D image of the object or the first input 3D representation of the object; and

generate a second output 3D latent representation by combining the first 3D latent representation and the second 3D latent representation.

15 . The system of claim 14 , wherein combining the first 3D latent representation and the second 3D latent representation comprises adding a first value at a first point of the first 3D latent representation to a second value at a second point of the second 3D latent representation to determine a value at a third point in the second output 3D latent representation, wherein both the first point and the second point correspond to the third point in the second output 3D latent representation.

16 . The system of claim 14 , wherein combining the first 3D latent representation and the second 3D latent representation comprises:

determining an adjusted first value at a first point of the first 3D latent representation by modifying a first value at the first point of the first 3D latent representation using a first blending parameter, wherein the first value is an output from the text encoder;

determining an adjusted second value at a second point of the second 3D latent representation by modifying a second value at the first point of the first 3D latent representation using a second blending parameter, wherein the second value is an output from the 2D/3D encoder; and

adding the adjusted first value at the first point of the first 3D latent representation to the adjusted second value at the second point of the second 3D latent representation to determine a value at a third point in the second output 3D latent representation, wherein both the first point and the second point correspond to the third point in the second output 3D latent representation.

17 . The system of claim 13 , wherein

the text encoder generates output parameters by encoding the first text prompt; and

the 2D/3D encoder encodes the first 2D image of the object or the first input 3D representation of the object based on the output parameters to generate the second output 3D latent representation.

18 . The system of claim 13 , wherein

the text encoder generates output parameters by encoding the first text prompt;

determine adjusted output parameters by applying a blending parameter to each of the output parameters; and

the 2D/3D encoder encodes the 2D image of the object or the input 3D representation of the object based on the adjusted output parameters to generate the second output 3D latent representation.

19 . The system of claim 13 , wherein the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system implemented using a robot;

an aerial system;

a medical system;

a boating system,

a smart area monitoring system;

a system for performing deep learning operations;

a system for performing simulation operations;

a system for generating or presenting virtual reality (VR) content, augmented reality (AR) content, or mixed reality (MR) content;

a system for performing digital twin operations;

a system implemented using an edge device;

a system incorporating one or more virtual machines (VMs);

a system for generating synthetic data;

a system implemented at least partially in a data center;

a system for performing conversational artificial intelligence (AI) operations;

a system for performing generative AI operations;

a system implementing one or more language models;

a system implementing one or more large language models (LLMs);

a system implementing one or more vision language models (VLMs);

a system for hosting one or more real-time streaming applications;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets; or

a system implemented at least partially using cloud computing resources.

20 . A method, comprising:

receiving at least one of a first text prompt, first 2D image, or first input 3D representation of a first object; and

generating, using a machine learning model comprising a text encoder, a 2D/3D encoder, and a decoder, a first 3D output corresponding to the at least one of the first text prompt, the first 2D image, or the first input 3D representation, wherein

the machine learning model is updated by:

generating a first output 3D latent representation by:

encoding, using the text encoder, a second text prompt; and

encoding, using the 2D/3D encoder, a second 2D image of an object or a second input 3D representation of a second object;

generating a second 3D output by applying the output 3D representation to the decoder;

determining a reconstruction loss and a sampling loss for the second 3D output; and

updating at least one of the text encoder, the 2D/3D encoder, and the decoder using the reconstruction loss and the sampling loss.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2024
From: XIE, CHENG; LORRAINE, JONATHAN; ZENG, XIAOHUI; LUCAS, JAMES; GAO, JUN; FIDLER, SANJA
To: NVIDIA CORPORATION
Reel/Frame 067108/0216 →
Continuity (1)
Related Publication 20250308183A1 · Oct 2, 2025
References Cited (7)
US 11227448B2 · Lebaredian · 2022 [cited by examiner]
US 11361507B1 · Iqbal · 2022 [cited by examiner]
US 20230410397A1 · Wang et al. · 2023 [cited by applicant]
US 20240005604A1 · Kreis et al. · 2024 [cited by applicant]
US 20240153188A1 · Wang et al. · 2024 [cited by applicant]
US 20240161403A1 · Lin et al. · 2024 [cited by applicant]
US 20240273871A1 · Song · 2024 [cited by examiner]