Training data sampling for neural networks
Apparatuses, systems, and techniques to train one or more neural networks using stratified sampled training data parameters. In at least one embodiment, one or more stochastic training data parameters may be stratified sampled from one or more sampling ranges to compute a gradient for updating the one or more neural networks.
1 . A processor comprising: one or more circuits to:
generate, using a text-to-3D model, one or more three-dimensional (3D) objects from a training text input;
sample, via stratified sampling over respective ranges, one or more stratified training data parameters that includes one or more camera orientation parameters, and at least one of a timestep length or a noise component;
render one or more images at least partially depicting the one or more 3D objects using the one or more camera orientation parameters;
generate, using a text-to-image model, through a plurality of diffusion steps each having the timestep length, a predicted noise component based on the training text input and at least one rendered image added with the noise component;
compute a gradient based on the predicted noise and stratified training data parameters; and
generate, from a text input, one or more 3D images using one or more neural networks trained based, at least in part, on the gradient computed with the stratified sampled training data parameters.
2 . The processor of claim 1 , wherein the stratified sampled training data parameters comprise one or more sampled values that are sampled for one or more stochastic training parameters.
3 . The processor of claim 2 , wherein the one or more circuits cause the gradient to be computed using the one or more sampled values and cause the one or more parameters of the one or more neural networks to be updated based on the gradient.
4 . The processor of claim 2 , wherein the one or more circuits cause the one or more neural networks to generate one or more three-dimensional images using a training text input, and wherein the one or more stochastic training parameters comprises one or more of:
one or more camera orientation parameters;
a timestep;
a noise sample;
an image type;
a data augmentation parameter; or
a prompt.
5 . The processor of claim 4 , wherein the one or more circuits cause the one or more three-dimensional images to be rendered into one or more images using the one or more camera orientation parameters,
cause an image model to generate a predicted noise based on the one or more rendered images and an injected noise, and
cause the gradient to be computed based on the predicted noise and the injected noise, and the one or more sampled values.
6 . The processor of claim 1 , wherein the one or more circuits cause the stratified sampled training data parameters to be obtained by:
evenly dividing a sampling range into a plurality of sub-ranges, and
randomly sampling at least one value from each of the plurality of sub-ranges.
7 . The processor of claim 1 , wherein the stratified sampled training data parameters are multi-dimensional and wherein the one or more circuits further cause the multi-dimensional stratified sampled training data parameters to be obtained by:
obtaining a respective sampled value through stratified sampling at each dimension,
and
aggregating respective stratified sampled values corresponding to the multiple dimensions to form the multi-dimensional stratified sampled training data parameters.
8 . A system comprising: one or more processors to:
generate, using a text-to-3D model, one or more three-dimensional (3D) objects from a training text input;
sample, via stratified sampling over respective ranges, one or more stratified training data parameters that includes one or more camera orientation parameters, and at least one of a timestep length or a noise component;
render one or more images at least partially depicting the one or more 3D objects using the one or more camera orientation parameters;
generate, using a text-to-image model, through a plurality of diffusion steps each having the timestep length, a predicted noise component based on the training text input and at least one rendered image added with the noise component;
compute a gradient based on the predicted noise and stratified training data parameters;
generate, from a text input, one or more 3D images using one or more 3D neural networks trained, at least in part, on the gradient computed with the stratified sampled training data parameters.
9 . The system of claim 8 , wherein the stratified sampled training data parameters comprises one or more sampled values that are sampled for one or more stochastic training parameters.
10 . The system of claim 9 , wherein the operations comprise causing the gradient to be computed using the one or more sampled values and cause the one or more parameters of the one or more 3D neural networks to be updated based on the gradient.
11 . The system of claim 9 , wherein the one or more stochastic training parameters comprises one or more of:
one or more camera orientation parameters;
a timestep;
a noise sample;
an image type;
a data augmentation parameter; or
a prompt.
12 . The system of claim 11 , wherein the operations include:
causing the one or more 3D objects to be rendered into one or more images using the one or more camera orientation parameters,
causing an image model to generate a predicted noise based on the one or more rendered images and an injected noise, and
causing the gradient to be computed based on the predicted noise and the injected noise, and the one or more sampled values.
13 . The system of claim 8 , wherein the stratified sampled training data parameters are obtained by:
evenly dividing a sampling range into a plurality of sub-ranges, and
randomly sampling at least one value from each of the plurality of sub-ranges.
14 . The system of claim 8 , wherein the stratified sampled training data parameters are multi-dimensional and wherein the multi-dimensional stratified sampled training data parameters are obtained by:
obtaining a respective sampled value through stratified sampling at each dimension,
and
aggregating respective stratified sampled values corresponding to the multiple dimensions to form the multi-dimensional stratified sampled training data parameters.
15 . A method for three-dimensional (3D) object generation using one or more 3D neural networks implemented on one or more processors, the method comprising:
generating, using a text-to-3D model, one or more 3D objects from a training text input;
sampling, via stratified sampling over respective ranges, one or more stratified training data parameters that includes one or more camera orientation parameters, and at least one of a timestep length or a noise component;
rendering one or more images at least partially depicting the one or more 3D objects using the one or more camera orientation parameters;
generating, using a text-to-image model, through a plurality of diffusion steps each having the timestep length, a predicted noise component based on the training text input and at least one rendered image added with the noise component;
computing a gradient based on the predicted noise and stratified training data parameters; and
generating, from a text input, one or more 3D images using one or more 3D neural networks trained based, at least in part, on the gradient computed with the stratified sampled training data parameters.
16 . The method of claim 15 , wherein the stratified sampled training data parameters comprises one or more sampled values that are sampled for one or more stochastic training parameters.
17 . The method of claim 16 , wherein the training the one or more 3D neural networks further comprises:
computing the gradient using the one or more sampled values; and
causing the one or more parameters of the one or more 3D neural networks to be updated based on the gradient.
18 . The method of claim 16 , wherein the one or more stochastic training parameters comprises any of:
one or more camera orientation parameters;
a timestep;
a noise sample;
an image type;
a data augmentation parameter; and
a prompt.
19 . The method of claim 16 , wherein the stratified sampled training data parameters are obtained by evenly dividing a sampling range into a plurality of sub-ranges, and randomly sampling at least one value from each of the plurality of sub-ranges.
20 . The method of claim 16 , wherein the stratified sampled training data parameters are multi-dimensional and wherein the multi-dimensional stratified sampled training data parameters are obtained by:
obtaining a respective sampled value through stratified sampling at each dimension,
and
aggregating respective stratified sampled values corresponding to the multiple dimensions to form the multi-dimensional stratified sampled training data parameters.