IP Library Granted Patent US 10,595,039
Granted Patent B2
US 10,595,039 · App. 15/939,098 · Granted Mar 17, 2020

System and method for content and motion controlled action video generation

Inventors: Ming-Yu Liu (Sunnyvale, CA); Xiaodong Yang (San Jose, CA); Jan Kautz (Lexington, MA); Sergey Tulyakov (Santa Clara, CA)
Assignee: NVIDIA Corporation
H04N19/521G06K9/00201G06K9/00281G06N3/0445G06N3/0454G06N3/0472G06N3/08G06T13/40G06T2207/20081G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,595,039
App. No.
15/939,098
Granted
Mar 17, 2020
Kind
B2
Abstract

A method, computer readable medium, and system are disclosed for action video generation. The method includes the steps of generating, by a recurrent neural network, a sequence of motion vectors from a first set of random variables and receiving, by a generator neural network, the sequence of motion vectors and a content vector sample. The sequence of motion vectors and the content vector sample are sampled by the generator neural network to produce a video clip.

Claims (50)

1. A computer-implemented method, comprising:

generating, by a recurrent neural network, a sequence of motion vectors from a first set of random variables;

receiving, by a generator neural network, the sequence of motion vectors and a content vector sample; and

processing the sequence of motion vectors and the content vector sample by the generator neural network to produce a video clip.

2. The computer-implemented method of claim 1 , further comprising:

generating, by the recurrent neural network, an additional sequence of motion vectors from a second set of random variables; and

processing the additional sequence of motion vectors and the content vector sample by the generator neural network to produce an additional video clip.

3. The computer-implemented method of claim 2 , wherein a number of frames in the video clip differs from a number of frames in the additional video clip.

4. The computer-implemented method of claim 1 , further comprising:

receiving, by the generator neural network, an additional content vector sample; and

processing the first sequence of motion vectors and the additional content vector sample by the generator neural network to produce an additional video clip.

5. The computer-implemented method of claim 1 , further comprising generating, by an encoder, the content vector sample based on identified content.

6. The computer-implemented method of claim 1 , further comprising sampling a Gaussian distribution of content to produce the content vector sample.

7. The computer-implemented method of claim 1 , further comprising:

sampling the video clip to produce image frames; and

processing the image frames by a discriminative neural network configured to distinguish real images from generated images to generate updated parameters for the generator neural network.

8. The computer-implemented method of claim 1 , further comprising:

sampling the video clip to produce sets of sequential frames; and

processing the sets of sequential frames by a discriminative neural network configured to distinguish real video clips from generated video clips to generate updated parameters for the generator neural network and the recurrent neural network.

9. The computer-implemented method of claim 1 , further comprising, prior to generating the sequence of motion vectors, combining an action label associated with an action category with the first set of random variables.

10. The computer-implemented method of claim 9 , wherein the action category represents facial expression.

11. The computer-implemented method of claim 9 , wherein the action category represents motion directions.

12. A system, comprising:

a parallel processing unit configured to implement a recurrent neural network and a generator network, wherein

the recurrent neural network is configured to generate a sequence of motion vectors from a first set of random variables,

the generator neural network receives the sequence of motion vectors and a content vector sample, and

the generator neural network processes the sequence of motion vectors and the content vector sample to produce a video clip.

13. The system of claim 12 , wherein

the recurrent neural network is further configured to generate an additional sequence of motion vectors from a second set of random variables; and

the generator neural network is further configured to process the additional sequence of motion vectors and the content vector sample by to produce an additional video clip.

14. The system of claim 13 , wherein a number of frames in the video clip differs from a number of frames in the additional video clip.

15. The system of claim 12 , wherein

the generator neural network is further configured to receive an additional content vector sample; and

the generator neural network is further configured to process the first sequence of motion vectors and the additional content vector sample to produce an additional video clip.

16. The system of 12 , further comprising an encode configured to generate the content vector sample based on identified content.

17. The system of claim 12 , further comprising sampling a Gaussian distribution of content to produce the content vector sample.

18. The system of claim 12 , further comprising:

an image sampler configured to sample the video clip to produce image frames; and

a discriminative neural network configured to:

process the image frames, distinguishing real images from the image frames; and

generate updated parameters for the generator neural network.

19. The system of claim 12 , further comprising:

a video sampler configured to sample the video clip to produce sets of sequential frames; and

a discriminative neural network configured to:

process the sets of sequential frames, distinguishing real video clips from the sets of sequential frames; and

generate updated parameters for the generator neural network and the recurrent neural network.

20. A non-transitory computer-readable media storing computer instructions for translating images that, when executed by a processor, cause the processor to perform the steps of:

generating, by a recurrent neural network, a sequence of motion vectors from a first set of random variables; and

receiving, by a generator neural network, the sequence of motion vectors and a content vector sample; and

processing the sequence of motion vectors and the content vector sample by the generator neural network to produce a video clip.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2018
From: LIU, MING-YU; YANG, XIAODONG; KAUTZ, JAN; TULYAKOV, SERGEY
To: NVIDIA CORPORATION
Reel/Frame 046104/0450 →
Continuity (2)
Provisional Application 62480094 · Mar 31, 2017
Related Publication 20180288431A1 · Oct 4, 2018