Neural network image generation
Apparatuses, system, and techniques to process resources used to perform a neural network to generate one or more images. In at least one embodiment, a processor comprising circuitry uses one or more neural networks to generate one or more second images based, at least in part, on noise within a first image.
1 . A processor, comprising:
one or more circuits configured to;
use one or more diffusion model neural networks to generate video, based at least in part, on transforming a sequence of correlated noised images into a corresponding sequence of denoised video frames, wherein the correlated noised images comprise one or more subsequent images of the sequence containing noise based, at least in part, on noise within one or more earlier images of the sequence.
2 . The processor of claim 1 , wherein input to the one or more neural networks comprises a noisy image based, at least in part, on the noise within a first image.
3 . The processor of claim 1 , wherein the one or more circuits are to compute noise to be added to an input to the one or more neural networks based, at least in part, on noise of a first image and additional randomized noise.
4 . The processor of claim 1 , wherein the one or more neural networks are to generate the video based, at least in part, on a text input and an input image comprising correlated noise.
5 . The processor of claim 1 , wherein the one or more neural networks are to generate the video constrained by a text input.
6 . The processor of claim 1 , wherein the one or more neural networks comprise a diffusion model trained to denoise images in sequence based, at least in part, on one or more text-image pairs and one or more correlated noised images.
7 . A system, comprising:
one or more processors configured to;
use one or more diffusion model neural networks to generate video, based at least in part, on transforming a sequence of correlated noised images into a corresponding sequence of denoised video frames, wherein the correlated noised images comprise one or more subsequent images of the sequence containing noise based, at least in part, on noise within one or more earlier images of the sequence.
8 . The system of claim 7 , wherein input to the one or more neural networks comprises a noisy image based, at least in part, on the noise within a first image.
9 . The system of claim 7 , wherein the one or more processors are to compute noise to be added to an input to the one or more neural networks based, at least in part, on noise of a first image and additional randomized noise.
10 . The system of claim 7 , wherein the one or more neural networks are to generate the video based, at least in part, on a text input and an input image comprising correlated noise.
11 . The system of claim 7 , wherein the one or more neural networks are to generate the video constrained by a text input.
12 . The system of claim 7 , wherein the one or more neural networks comprise a diffusion model trained to denoise images in sequence based, at least in part, on one or more text-image pairs and one or more correlated noised images.
13 . A method, comprising:
using one or more neural networks, implemented using one or more processors, to generate video, based at least in part, on transforming a sequence of correlated noised images into a corresponding sequence of denoised video frames, wherein the correlated noised images comprise one or more subsequent images of the sequence containing noise based, at least in part, on noise within one or more earlier images of the sequence.
14 . The method of claim 13 , further comprising inputting, to the one or more neural networks to generate the one or more second images, a noisy image based, at least in part, on the noise within a first image.
15 . The method of claim 13 , further comprising computing noise to be added to an input to the one or more neural networks based, at least in part, on noise of a first image and additional randomized noise.
16 . The method of claim 13 , further comprising using the one or more neural networks to generate the video based, at least in part, on a text input and an input image comprising correlated noise.
17 . The method of claim 13 , further comprising using the one or more neural networks to generate the video based, at least in part, on data indicative of a textual input to the one or more neural networks.