IP Library › Granted Patent US 12,651,399
Granted Patent B2
US 12,651,399 · App. 18/486,046 · Granted Jun 9, 2026

Training data sampling for neural networks

Inventors: Jonathan Peter Lorraine (Toronto, CA); Cheng (Kevin) Xie (Toronto, CA); Xiaohui Zeng (Toronto, CA); Jun Gao (Toronto, CA); Sanja Fidler (Toronto, CA); James Lucas (Toronto, CA)
Assignee: NVIDIA Corporation
G06T15/005G06T7/97G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,399
App. No.
18/486,046
Granted
Jun 9, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to train one or more neural networks using stratified sampled training data parameters. In at least one embodiment, one or more stochastic training data parameters may be stratified sampled from one or more sampling ranges to compute a gradient for updating the one or more neural networks.

Claims (76)

1 . A processor comprising: one or more circuits to:

generate, using a text-to-3D model, one or more three-dimensional (3D) objects from a training text input;

sample, via stratified sampling over respective ranges, one or more stratified training data parameters that includes one or more camera orientation parameters, and at least one of a timestep length or a noise component;

render one or more images at least partially depicting the one or more 3D objects using the one or more camera orientation parameters;

generate, using a text-to-image model, through a plurality of diffusion steps each having the timestep length, a predicted noise component based on the training text input and at least one rendered image added with the noise component;

compute a gradient based on the predicted noise and stratified training data parameters; and

generate, from a text input, one or more 3D images using one or more neural networks trained based, at least in part, on the gradient computed with the stratified sampled training data parameters.

2 . The processor of claim 1 , wherein the stratified sampled training data parameters comprise one or more sampled values that are sampled for one or more stochastic training parameters.

3 . The processor of claim 2 , wherein the one or more circuits cause the gradient to be computed using the one or more sampled values and cause the one or more parameters of the one or more neural networks to be updated based on the gradient.

4 . The processor of claim 2 , wherein the one or more circuits cause the one or more neural networks to generate one or more three-dimensional images using a training text input, and wherein the one or more stochastic training parameters comprises one or more of:

one or more camera orientation parameters;

a timestep;

a noise sample;

an image type;

a data augmentation parameter; or

a prompt.

5 . The processor of claim 4 , wherein the one or more circuits cause the one or more three-dimensional images to be rendered into one or more images using the one or more camera orientation parameters,

cause an image model to generate a predicted noise based on the one or more rendered images and an injected noise, and

cause the gradient to be computed based on the predicted noise and the injected noise, and the one or more sampled values.

6 . The processor of claim 1 , wherein the one or more circuits cause the stratified sampled training data parameters to be obtained by:

evenly dividing a sampling range into a plurality of sub-ranges, and

randomly sampling at least one value from each of the plurality of sub-ranges.

7 . The processor of claim 1 , wherein the stratified sampled training data parameters are multi-dimensional and wherein the one or more circuits further cause the multi-dimensional stratified sampled training data parameters to be obtained by:

obtaining a respective sampled value through stratified sampling at each dimension,

and

aggregating respective stratified sampled values corresponding to the multiple dimensions to form the multi-dimensional stratified sampled training data parameters.

8 . A system comprising: one or more processors to:

generate, using a text-to-3D model, one or more three-dimensional (3D) objects from a training text input;

sample, via stratified sampling over respective ranges, one or more stratified training data parameters that includes one or more camera orientation parameters, and at least one of a timestep length or a noise component;

render one or more images at least partially depicting the one or more 3D objects using the one or more camera orientation parameters;

generate, using a text-to-image model, through a plurality of diffusion steps each having the timestep length, a predicted noise component based on the training text input and at least one rendered image added with the noise component;

compute a gradient based on the predicted noise and stratified training data parameters;

generate, from a text input, one or more 3D images using one or more 3D neural networks trained, at least in part, on the gradient computed with the stratified sampled training data parameters.

9 . The system of claim 8 , wherein the stratified sampled training data parameters comprises one or more sampled values that are sampled for one or more stochastic training parameters.

10 . The system of claim 9 , wherein the operations comprise causing the gradient to be computed using the one or more sampled values and cause the one or more parameters of the one or more 3D neural networks to be updated based on the gradient.

11 . The system of claim 9 , wherein the one or more stochastic training parameters comprises one or more of:

one or more camera orientation parameters;

a timestep;

a noise sample;

an image type;

a data augmentation parameter; or

a prompt.

12 . The system of claim 11 , wherein the operations include:

causing the one or more 3D objects to be rendered into one or more images using the one or more camera orientation parameters,

causing an image model to generate a predicted noise based on the one or more rendered images and an injected noise, and

causing the gradient to be computed based on the predicted noise and the injected noise, and the one or more sampled values.

13 . The system of claim 8 , wherein the stratified sampled training data parameters are obtained by:

evenly dividing a sampling range into a plurality of sub-ranges, and

randomly sampling at least one value from each of the plurality of sub-ranges.

14 . The system of claim 8 , wherein the stratified sampled training data parameters are multi-dimensional and wherein the multi-dimensional stratified sampled training data parameters are obtained by:

obtaining a respective sampled value through stratified sampling at each dimension,

and

aggregating respective stratified sampled values corresponding to the multiple dimensions to form the multi-dimensional stratified sampled training data parameters.

15 . A method for three-dimensional (3D) object generation using one or more 3D neural networks implemented on one or more processors, the method comprising:

generating, using a text-to-3D model, one or more 3D objects from a training text input;

sampling, via stratified sampling over respective ranges, one or more stratified training data parameters that includes one or more camera orientation parameters, and at least one of a timestep length or a noise component;

rendering one or more images at least partially depicting the one or more 3D objects using the one or more camera orientation parameters;

generating, using a text-to-image model, through a plurality of diffusion steps each having the timestep length, a predicted noise component based on the training text input and at least one rendered image added with the noise component;

computing a gradient based on the predicted noise and stratified training data parameters; and

generating, from a text input, one or more 3D images using one or more 3D neural networks trained based, at least in part, on the gradient computed with the stratified sampled training data parameters.

16 . The method of claim 15 , wherein the stratified sampled training data parameters comprises one or more sampled values that are sampled for one or more stochastic training parameters.

17 . The method of claim 16 , wherein the training the one or more 3D neural networks further comprises:

computing the gradient using the one or more sampled values; and

causing the one or more parameters of the one or more 3D neural networks to be updated based on the gradient.

18 . The method of claim 16 , wherein the one or more stochastic training parameters comprises any of:

one or more camera orientation parameters;

a timestep;

a noise sample;

an image type;

a data augmentation parameter; and

a prompt.

19 . The method of claim 16 , wherein the stratified sampled training data parameters are obtained by evenly dividing a sampling range into a plurality of sub-ranges, and randomly sampling at least one value from each of the plurality of sub-ranges.

20 . The method of claim 16 , wherein the stratified sampled training data parameters are multi-dimensional and wherein the multi-dimensional stratified sampled training data parameters are obtained by:

obtaining a respective sampled value through stratified sampling at each dimension,

and

aggregating respective stratified sampled values corresponding to the multiple dimensions to form the multi-dimensional stratified sampled training data parameters.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE'S ADDRESS PREVIOUSLY RECORDED AT REEL: 065225 FRAME: 0513. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT . Recorded Oct 17, 2023
From: LORRAINE, JONATHAN PETER; XIE, CHENG (KEVIN); ZENG, XIAOHUI; GAO, JUN; FIDLER, SANJA; LUCAS, JAMES
To: NVIDIA CORPORATION
Reel/Frame 065255/0765 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2023
From: LORRAINE, JONATHAN PETER; XIE, CHENG (KEVIN); ZENG, XIAOHUI; GAO, JUN; FIDLER, SANJA; LUCAS, JAMES
To: NVIDIA CORPORATION
Reel/Frame 065225/0513 →
Continuity (1)
Related Publication 20250124640A1 · Apr 17, 2025
References Cited (8)
US 20200134493A1 · Bhide · 2020 [cited by examiner]
US 20210110926A1 · Wang · 2021 [cited by examiner]
US 20210350621A1 · Bailey · 2021 [cited by examiner]
US 20210368157A1 · Overbeck · 2021 [cited by examiner]
US 20230055263A1 · Zhong · 2023 [cited by examiner]
Ben Poole, Ajay Jain, Jonathan T. Barron, Ben Mildenhall, “Dream Fusion: Text-to-3D using 2D Diffusion”, Sep. 29, 2022, arXiv.org, arXiv:2209.14988, https://arxiv.org/abs/2209.14988. [cited by examiner]
Zilong Chen, Feng Wang, Huaping Liu, “Text-to-3D using Gaussian Splatting”, Sep. 29, 2023, arXiv.org, Version 2, arXiv: 2309.16585v2, https://arxiv.org/abs/2309.16585v2. [cited by examiner]
Yunhao Ge, Harkirat Behl, Jiashu Xu, Suriya Gunasekar, Neel Joshi, Yale Song, Xin Wang, Laurent Itti, Vibhav Vineet, “Neural-Sim : Learning to Generate Training Data with NeRF”, Oct. 28, 2022, Springer, Computer Vision—… [cited by examiner]