IP Library Granted Patent US 12711583
Granted Patent B1
US 12711583 · App. 18/121,926 · Granted Aug 18, 2026

Neural network image generation

Inventors: Ming-Yu Liu (San Jose, CA); Yogesh Balaji (Mountain View, CA); Songwei Ge (Greenbelt, MD); Seungjun Nah (Santa Clara, CA)
Assignee: NVIDIA Corporation
G06T5/70G06T11/60G06T2207/10016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12711583
App. No.
18/121,926
Granted
Aug 18, 2026
Kind
B1
Abstract

Apparatuses, system, and techniques to process resources used to perform a neural network to generate one or more images. In at least one embodiment, a processor comprising circuitry uses one or more neural networks to generate one or more second images based, at least in part, on noise within a first image.

Claims (22)

1 . A processor, comprising:

one or more circuits configured to;

use one or more diffusion model neural networks to generate video, based at least in part, on transforming a sequence of correlated noised images into a corresponding sequence of denoised video frames, wherein the correlated noised images comprise one or more subsequent images of the sequence containing noise based, at least in part, on noise within one or more earlier images of the sequence.

2 . The processor of claim 1 , wherein input to the one or more neural networks comprises a noisy image based, at least in part, on the noise within a first image.

3 . The processor of claim 1 , wherein the one or more circuits are to compute noise to be added to an input to the one or more neural networks based, at least in part, on noise of a first image and additional randomized noise.

4 . The processor of claim 1 , wherein the one or more neural networks are to generate the video based, at least in part, on a text input and an input image comprising correlated noise.

5 . The processor of claim 1 , wherein the one or more neural networks are to generate the video constrained by a text input.

6 . The processor of claim 1 , wherein the one or more neural networks comprise a diffusion model trained to denoise images in sequence based, at least in part, on one or more text-image pairs and one or more correlated noised images.

7 . A system, comprising:

one or more processors configured to;

use one or more diffusion model neural networks to generate video, based at least in part, on transforming a sequence of correlated noised images into a corresponding sequence of denoised video frames, wherein the correlated noised images comprise one or more subsequent images of the sequence containing noise based, at least in part, on noise within one or more earlier images of the sequence.

8 . The system of claim 7 , wherein input to the one or more neural networks comprises a noisy image based, at least in part, on the noise within a first image.

9 . The system of claim 7 , wherein the one or more processors are to compute noise to be added to an input to the one or more neural networks based, at least in part, on noise of a first image and additional randomized noise.

10 . The system of claim 7 , wherein the one or more neural networks are to generate the video based, at least in part, on a text input and an input image comprising correlated noise.

11 . The system of claim 7 , wherein the one or more neural networks are to generate the video constrained by a text input.

12 . The system of claim 7 , wherein the one or more neural networks comprise a diffusion model trained to denoise images in sequence based, at least in part, on one or more text-image pairs and one or more correlated noised images.

13 . A method, comprising:

using one or more neural networks, implemented using one or more processors, to generate video, based at least in part, on transforming a sequence of correlated noised images into a corresponding sequence of denoised video frames, wherein the correlated noised images comprise one or more subsequent images of the sequence containing noise based, at least in part, on noise within one or more earlier images of the sequence.

14 . The method of claim 13 , further comprising inputting, to the one or more neural networks to generate the one or more second images, a noisy image based, at least in part, on the noise within a first image.

15 . The method of claim 13 , further comprising computing noise to be added to an input to the one or more neural networks based, at least in part, on noise of a first image and additional randomized noise.

16 . The method of claim 13 , further comprising using the one or more neural networks to generate the video based, at least in part, on a text input and an input image comprising correlated noise.

17 . The method of claim 13 , further comprising using the one or more neural networks to generate the video based, at least in part, on data indicative of a textual input to the one or more neural networks.