IP Library › Granted Patent US 12,506,897
Granted Patent B2
US 12,506,897 · App. 16/812,058 · Granted Dec 23, 2025

System and method for content and motion controlled action video generation

Inventors: Ming-Yu Liu (Sunnyvale, CA); Xiaodong Yang (San Jose, CA); Jan Kautz (Lexington, MA); Sergey Tulyakov (Santa Clara, CA)
Assignee: NVIDIA Corporation
H04N19/521G06N3/044G06N3/045G06N3/047G06N3/08G06T13/40G06V20/64G06V40/171G06T2207/20081G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,506,897
App. No.
16/812,058
Granted
Dec 23, 2025
Kind
B2
Abstract

A method, computer readable medium, and system are disclosed for action video generation. The method includes the steps of generating, by a recurrent neural network, a sequence of motion vectors from a first set of random variables and receiving, by a generator neural network, the sequence of motion vectors and a content vector sample. The sequence of motion vectors and the content vector sample are sampled by the generator neural network to produce a video clip.

Claims (37)

1 . A computer-implemented method, comprising:

using a first neural network to generate a plurality of motion vectors; and

using a second neural network to use the plurality of motion vectors to produce content, wherein the second neural network comprises a generator neural network.

2 . The computer-implemented method of claim 1 , wherein the first neural network comprises a recurrent neural network.

3 . The computer-implemented method of claim 1 , further comprising:

using a third neural network to use image frames from the content to generate updated information for the second neural network.

4 . The computer-implemented method of claim 3 , wherein the third neural network comprises a discriminative neural network.

5 . The computer-implemented method of claim 3 , further comprising:

using the third neural network to use sets of sequential frames from the content to generate updated information for the first and the second neural networks.

6 . The computer-implemented method of claim 1 , wherein the content comprises a sequence of image frames.

7 . The computer-implemented method of claim 1 , further comprising:

passing a first set of variables to the first neural network to generate the plurality of motion vectors; and

passing a second set of variables to the first neural network to generate a second set of a plurality of motion vectors that is different from the plurality of vectors to generate additional content.

8 . A processor, comprising:

one or more arithmetic logic units (ALUs) to use a first neural network to generate a plurality of motion vectors and a second neural network to use the plurality of motion vectors to produce content, wherein the second neural network comprises a generator neural network.

9 . The processor of claim 8 , wherein the first neural network comprises a recurrent neural network.

10 . The processor of claim 8 , further comprising one or more ALUs to use a third neural network to use image frames from the content to generate updated information for the second neural network.

11 . The processor of claim 10 , wherein the third neural network comprises a discriminative neural network.

12 . The processor of claim 10 , further comprising one or more ALUs to use the third neural network to use sets of sequential frames from the content to generate updated information for the first and the second neural networks.

13 . The processor of claim 8 , wherein the content comprises a sequence of video frames.

14 . The processor of claim 8 , further comprising one or more ALUs to:

use the first neural network generate additional plurality of motion vectors using different input used to generate the plurality of motion vectors; and

use the second neural network to generate additional content by using the additional plurality of vectors.

15 . A system, comprising:

one or more computers having one or more processors to use a first neural network to generate a plurality of motion vectors and a second neural network to use the plurality of motion vectors to produce content, wherein the second neural network comprises a generator neural network.

16 . The system of claim 15 , wherein the first neural network comprises a recurrent neural network.

17 . The system of claim 15 , further comprising one or more computers having one or more processors to use a third neural network to use image frames from the content to generate updated information for the second neural network.

18 . The system of claim 17 , wherein the third neural network comprises a discriminative neural network.

19 . The system of claim 18 , further comprising one or more computers having one or more processors to use the third neural network to use sets of sequential frames from the content to generate updated information for the first and the second neural networks.

20 . The system of claim 15 , wherein the content comprises a sequence of image frames.

21 . The system of claim 15 , further comprising one or more computers having one or more processors to pass input to the first neural network to generate additional plurality of motion vectors; and use the second neural network to generate additional content using the additional plurality of motion vectors.

22 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to use a first neural network to generate a plurality of motion vectors and a second neural network to use the plurality of motion vectors to produce content, wherein the second neural network comprises a generator neural network.

23 . The non-transitory machine-readable medium of claim 22 , wherein the first neural network comprises a recurrent neural network.

24 . The non-transitory machine-readable medium of claim 22 , having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to further use a third neural network to use image frames from the content to generate updated information for the second neural network.

25 . The non-transitory machine-readable medium of claim 24 , wherein the third neural network comprises a discriminative neural network.

26 . The non-transitory machine-readable medium of claim 24 , having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to further use the third neural network to use sets of sequential frames from the content to generate updated information for the first and the second neural networks.

27 . The non-transitory machine-readable medium of claim 24 , wherein the content comprises a sequence of image frames.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2020
From: LIU, MING-YU; YANG, XIAODONG; KAUTZ, JAN; TULYAKOV, SERGEY
To: NVIDIA CORPORATION
Reel/Frame 052043/0395 →
Continuity (3)
Continuation 15939098 · Mar 28, 2018
Provisional Application 62480094 · Mar 31, 2017
Related Publication 20200204822A1 · Jun 25, 2020
References Cited (67)
US 6385245B1 · De Haan et al. · 2002 [cited by applicant]
US 7636662B2 · Dimtrova et al. · 2009 [cited by applicant]
US 7797259B2 · Jiang et al. · 2010 [cited by applicant]
US 9015093B1 · Commons · 2015 [cited by examiner]
US 9020239B2 · Graepel et al. · 2015 [cited by applicant]
US 9053562B1 · Rabin et al. · 2015 [cited by applicant]
US 9449412B1 · Rogers et al. · 2016 [cited by applicant]
US 9524582B2 · Ma et al. · 2016 [cited by applicant]
US 9715496B1 · Sapoznik · 2017 [cited by examiner]
US 9786084B1 · Bhat et al. · 2017 [cited by applicant]
US 10867236B2 · Lee · 2020 [cited by examiner]
US 20130124206A1 · Rezvani et al. · 2013 [cited by applicant]
US 20160232440A1 · Gregor et al. · 2016 [cited by applicant]
US 20170127016A1 · Yu et al. · 2017 [cited by applicant]
US 20170278135A1 · Majumdar et al. · 2017 [cited by applicant]
US 20170285732A1 · Daly et al. · 2017 [cited by applicant]
US 20170293815A1 · Cosatto et al. · 2017 [cited by applicant]
US 20170351935A1 · Liu · 2017 [cited by examiner]
US 20180067605A1 · Lin · 2018 [cited by examiner]
US 20180157902A1 · Tu · 2018 [cited by examiner]
US 20180176576A1 · Rippel · 2018 [cited by examiner]
US 20180240257A1 · Li · 2018 [cited by examiner]
US 20180276533A1 · Ding · 2018 [cited by examiner]
CN 102054270B · 2013 [cited by applicant]
Holden et al. (A Deep Learning Framework for Character Motion Synthesis and Editing; Jul. 2016) (Year: 2016). [cited by examiner]
Goodfellow et al.; Generative Adversarial Nets; Dec. 2, 2014 (Year: 2014). [cited by examiner]
IEEE Computer Society, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” The Institute of Electrical and Electronics Engineers, Inc., Aug. 29, 2008, 70 pages. [cited by applicant]
Li et al., “Video Generation From Text,” arXiv preprint arXiv:1710.00421, 2017, pp. 1-8. [cited by applicant]
Saito et al., “Temporal Generative Adversarial Nets,” arXiv:1611.06624v1, Nov. 21, 2017, pp. 1-10. [cited by applicant]
Villegas et al., “Decomposing motion and content for natural video sequence prediction,” arXiv preprint arXiv:1706.08033, 2017, pp. 1-22. [cited by applicant]
Vondrick et al., “Generating Videos with Scene Dynamics,” 29th Conference on Neural Information Processing Systems (NIPS), 2016, pp. 1-10. [cited by applicant]
Aifanti et al., “The Mug Facial Expression Database,” In Image Analysis for Multimedia Interactive Services, 2010, 4 pages. [cited by applicant]
Amos et al., “OpenFace: A General-Purpose Face Recognition Library with Mobile Applications,” Technical Report, CMU School of Computer Science, Jun. 2016, 20 pages. [cited by applicant]
Arjovsky et al., “Wasserstein GAN,” Dec. 6, 2017, 32 pages. [cited by applicant]
Chen et al., “Infogan: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets,” In Advances in Neural Information Processing Systems, 2016, 14 pages. [cited by applicant]
Chung et al., “Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling,” Dec. 11, 2014, 9 pages. [cited by applicant]
Denton et al., “Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks,” Advances in Neural Information Processing Systems, 2015, 9 pages. [cited by applicant]
Doretto et al., “Dynamic Textures,” International Journal of Computer Vision, 51(2): 2003, 19 pages. [cited by applicant]
Durugkar et al., “Generative Multl-Adversarial Networks,” In International Conference on Learning Representations. 2016, 14 pages. [cited by applicant]
Finn et al., “Unsupervised Learning for Physical Interaction Through Video Prediction,” Advances in Neural Information Processing Systems, 2016, 9 pages. [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets,” In Advances in Neural Information Processing Systems, Jun. 10, 2014, 9 pages. [cited by applicant]
Gorelick et al., “Actions as Space-Time Shapes,” PAMI, 29(12): Dec. 2007, 7 pages. [cited by applicant]
Gregor et al., “DRAW: A Recurrent Neural Network for Image Generation,” International Conference on Machine Learning, 2015, 10 pages. [cited by applicant]
Hochreiter et al., “Long Short-Term Memory,” Neural Computation, 9(8): 1997, pp. 1735-1780. [cited by applicant]
Im et al., “Generating Images with Recurrent Adversarial Networks,” Dec. 13, 2016, 20 pages. [cited by applicant]
Kalchbrenner et al., “Video Pixel Networks,” Oct. 3, 2016, 16 pages. [cited by applicant]
Kingma et al., “Adam: A Method for Stochastic Optimization,” ICLR, 2015, 15 pages. [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes,” Dec. 27, 2013, 14 pages. [cited by applicant]
Li et al., “Generative Moment Matching Networks,” International Conference on Machine Learning, 2015, 10 pages. [cited by applicant]
Liu et al., “Coupled Generative Adversarial Networks,” 30th Conference on Neural Information Processing, Dec. 5, 2016, 9 pages. [cited by applicant]
Mathieu et al., “Deep Multi-Scale Video Prediction Beyond Mean Square Error,” Nov. 23, 2015, 11 pages. [cited by applicant]
Nguyen et al., “Plug & Play Generative Networks: Conditional Iterative Generation of Images in Latent Space,” Nov. 30, 2016, 33 pages. [cited by applicant]
Oh et al., “Action-Conditional Video Prediction using Deep Networks in Atari Games,” Advances in Neural Information Processing Systems, 2015, 9 pages. [cited by applicant]
Oord et al., “Conditional Image Generation with Pixelcnn Decoders,” Advances in Neural Information Processing Systems, 2016, 9 pages. [cited by applicant]
Radford et al., “Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks,” Jan. 7, 2016, 16 pages. [cited by applicant]
Reed et al., “Generative Adversarial Text-to-Image Synthesis,” ICML, Jun. 5, 2016, 10 pages. [cited by applicant]
Rezende et al., “Stochastic Backpropagation and Approximate Inference in Deep Generative Models,” In International Conference on Machine Learning, 2014, 9 pages. [cited by applicant]
Salimans et al., “Improved Techniques for Training GANs,” Advances in Neural Information Processing Systems, 2016, 9 pages. [cited by applicant]
Srivastava et al., “Unsupervised Learning of Video Representations using LSTMs,” International Conference on Machine Learning, 2015, 10 pages. [cited by applicant]
Szummer et al., “Temporal Texture Modeling,” International Conference on Image Processing, 1996, 57 pages. [cited by applicant]
Taigman et al., “DeepFace: Closing the Gap to Human-Level Performance in Face Verification,” CVPR, 2014, 8 pages. [cited by applicant]
Theis et al., “A Note on the Evaluation of Generative Models,” International Conference on Learning Representations, Apr. 24, 2016, 10 pages. [cited by applicant]
Tokui et al., “Chainer: A Next-Generation Open Source Framework for Deep Learning,” Workshop on Machine Learning Systems on Advances in Neural Information Processing Systems, 2015, 6 pages. [cited by applicant]
Van Amersfoort et al., “Transformation-Based Models of Video Sequences,” Apr. 24, 2017, 11 pages. [cited by applicant]
Wei et al., “Fast Texture Synthesis using Tree-structured Vector Quantization,” Proceedings of the 27th Annual Conference on Computer Graphics and Interactive Techniques, 2000, 11 pages. [cited by applicant]
Xue et al., “Probabilistic Modeling of Future Frames from a Single Image,” Advances In Neural Information Processing Systems, 2016, 9 pages. [cited by applicant]
Zhang et al., “StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks,” ICCV, 2017, 9 pages. [cited by applicant]