IP Library Granted Patent US 12,192,547
Granted Patent B2
US 12,192,547 · App. 18/181,729 · Granted Jan 7, 2025

High-resolution video generation using image diffusion models

Inventors: Karsten Julian Kreis (Vancouver, CA); Robin Rombach (Heidelberg, DE); Andreas Blattmann (Waldkirch, DE); Seung Wook Kim (Toronto, CA); Huan Ling (Toronto, CA); Sanja Fidler (Toronto, CA); Tim Dockhorn (Waterloo, CA)
Assignee: NVIDIA Corporation
H04N21/234363G06T9/00G06V10/24G06V10/25G06V10/82H04N7/0117
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,192,547
App. No.
18/181,729
Granted
Jan 7, 2025
Kind
B2
Abstract

In various examples, systems and methods are disclosed relating to aligning images into frames of a first video using at least one first temporal attention layer of a neural network model. The first video has a first spatial resolution. A second video having a second spatial resolution is generated by up-sampling the first video using at least one second temporal attention layer of an up-sampler neural network model, wherein the second spatial resolution is higher than the first spatial resolution.

Claims (52)

1. A processor, comprising:

one or more circuits to:

align a plurality of images into frames of a first video using a neural network model comprising a latent diffusion model (LDM), wherein the first video has a first spatial resolution, the LDM comprises:

an encoder to map an input from an image space to a latent space; and

a decoder to map latent encoding from the latent space to the image space; and

generate a second video having a second spatial resolution by up-sampling the first video using an up-sampler neural network model, wherein the second spatial resolution is higher than the first spatial resolution, wherein the decoder is updated according to one or more temporal incoherencies in mapping the latent encoding from the latent space to the image space.

2. The processor of claim 1 , wherein the neural network model is modified from an image diffusion model by adding at least one first temporal attention layer into the image diffusion model.

3. The processor of claim 1 , wherein the up-sampler neural network model comprises a first diffusion model for video generation, and the first diffusion model is modified from a second diffusion model for image generation by adding at least one second temporal attention layer into the first diffusion model for image generation.

4. The processor of claim 1 , wherein the plurality of images are consecutive frames of the first video.

5. The processor of claim 1 , wherein:

the neural network model comprises a first diffusion model and a second diffusion model;

the first diffusion model is to generate a third video; and

the second diffusion model is to generate the first video by generating at least one frame between two consecutive frames of the third video.

6. The processor of claim 1 , wherein the neural network model is to:

generate a third video; and

generate the first video by generating at least one frame between two consecutive frames of the third video according to relative time step embedding.

7. The processor of claim 1 , wherein the first video is generated according to at least one of:

one or more text prompts;

one or more bounding boxes; or

one or more conditioning signals.

8. A processor, comprising:

one or more circuits to:

update a neural network model to align a plurality of images into frames of a first video by updating at least one layer of the neural network model, wherein the first video has a first spatial resolution, wherein the neural network model comprises a Latent Diffusion Model (LDM) comprising:

an encoder to map an input from an image space to a latent space; and

a decoder to map latent encoding from the latent space to the image space,

wherein to update the neural network model, the one or more processing units are to update the decoder according to one or more temporal incoherencies in mapping the latent encoding from the latent space to the image space; and

update an up-sampler neural network model to generate a second video having a second spatial resolution via up-sampling the first video by updating at least one layer of the up-sampler neural network model, wherein the second spatial resolution is higher than the first spatial resolution.

9. The processor of claim 8 , wherein the one or more circuits are to:

insert the at least one neural network layer into a denoising neural network of an image diffusion model; and

update the at least one neural network layer to align the plurality of images into the frames of the first video.

10. The processor of claim 8 , wherein the one or more circuits are to:

update at least one other layer of the neural network model using a plurality of sample images; and

update the at least one layer using at least one sample sequence of images, while maintaining parameters of the at least one other layer unchanged.

11. The processor of claim 8 , wherein the plurality of images are consecutive frames of the first video.

12. The processor of claim 8 , wherein:

the neural network model comprises a first diffusion model and a second diffusion model;

the first diffusion model is updated to generate a third video; and

the second diffusion model is updated to generate the first video by generating at least one frame between two consecutive frames of the third video.

13. The processor of claim 8 , wherein:

the neural network model is updated to generate a third video; and

the neural network model is updated to generate the first video by generating at least one frame between two consecutive frames of the third video according to relative time step embedding.

14. The processor of claim 8 , wherein the one or more circuits are to update the neural network model to generate the first video according to at least one of:

one or more text prompts;

one or more bounding boxes; or

one or more conditioning signals.

15. The processor of claim 8 , wherein the one or more circuits are to update the neural network model to generate the first video according to modified conditioning signals for classifier-free guidance.

16. A method, comprising:

updating a neural network model to align a plurality of images into frames of a first video by updating at least one first temporal attention layer of the neural network model, wherein the first video has a first spatial resolution, wherein the neural network model comprises a Latent Diffusion Model (LDM) comprising:

an encoder to map an input from an image space to a latent space; and

a decoder to map latent encoding from the latent space to the image space,

wherein updating the neural network model comprises updating the decoder according to one or more temporal incoherencies in mapping the latent encoding from the latent space to the image space; and

updating an up-sampler neural network model to generate a second video having a second spatial resolution via up-sampling the first video, by updating at least one second temporal attention layer of the up-sampler neural network model, wherein the second spatial resolution is higher than the first spatial resolution.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2023
From: ROMBACH, ROBIN
To: NVIDIA CORPORATION
Reel/Frame 063200/0239 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2023
From: KREIS, KARSTEN JULIAN; ROMBACH, ROBIN; BLATTMANN, ANDREAS; KIM, SEUNG WOOK; LING, HUAN; FIDLER, SANJA; DOCKHORN, TIM
To: NVIDIA CORPORATION
Reel/Frame 063035/0519 →
Continuity (2)
Provisional Application 63426037 · Nov 16, 2022
Related Publication 20240171788A1 · May 23, 2024
References Cited (9)
US 11335048B1 · Lee · 2022 [cited by examiner]
US 11436787B2 · Wang · 2022 [cited by examiner]
US 20210366082A1 · Xiao · 2021 [cited by examiner]
US 20230351558A1 · Chen · 2023 [cited by examiner]
D. Zhou et al., “MagicVideo: Efficient Video Generation With Latent Diffusion Models,” arXiv, Nov. 20, 2022, https://arxiv.org/abs/2211.11018. [cited by applicant]
T. Hoppe et al., “Diffusion Models for Video Prediction and Infilling,” Published in Transactions on Machine Learning Research, Nov. 14, 2022, https://arxiv.org/abs/2206.07696. [cited by applicant]
U. Singer et al., “Make-A-Video: Text-to-Video Generation without Text-Video Data,” arXiv, Sep. 29, 2022, https://makeavideo.studio/. [cited by applicant]
V. Voleti et al., “MCVD: Masked Conditional Video Diffusion for Prediction, Generation, and Interpolation,” Mila, University of Montreal, Canada, Oct. 12, 2022, https://arxiv.org/abs/2205.09853. [cited by applicant]
W. Harvey et al., “Flexible Diffusion Modeling of Long Videos,” Department of Computer Science, University of British Columbia, Vancouver, Canada, Sep. 15, 2022, https://arxiv.org/abs/2205.11495. [cited by applicant]