IP Library Granted Patent US 11,580,395
Granted Patent B2
US 11,580,395 · App. 17/069,449 · Granted Feb 14, 2023

Generative adversarial neural network assisted video reconstruction

Inventors: Tero Tapani Karras (Helsinki, FI); Samuli Matias Laine (Vantaa, FI); David Patrick Luebke (Charlottesville, VA); Jaakko T. Lehtinen (Helsinki, FI); Miika Samuli Aittala (Helsinki, FI); Timo Oskari Aila (Tuusula, FI); Ming-Yu Liu (San Jose, CA); Arun Mohanray Mallya (San Jose, CA); Ting-Chun Wang (Santa Clara, CA)
Assignee: NVIDIA Corporation
G06N3/08G06T5/003G06T7/73G06T9/002G06V40/168H04N7/157H04N19/20G06T2207/20081G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,580,395
App. No.
17/069,449
Granted
Feb 14, 2023
Kind
B2
Abstract

A latent code defined in an input space is processed by the mapping neural network to produce an intermediate latent code defined in an intermediate latent space. The intermediate latent code may be used as appearance vector that is processed by the synthesis neural network to generate an image. The appearance vector is a compressed encoding of data, such as video frames including a person's face, audio, and other data. Captured images may be converted into appearance vectors at a local device and transmitted to a remote device using much less bandwidth compared with transmitting the captured images. A synthesis neural network at the remote device reconstructs the images for display.

Claims (29)

1. A computer-implemented method, comprising:

obtaining data for replicating a style specific to a subject, wherein the data is determined by training a generator neural network to produce images of the subject that are compared with captured images of the subject; configuring a neural network to apply the data to a sequence of vectors generated for a human face captured in a first sequence of images to modify at least one attribute according to the style specific to the subject; receiving, through a communication network, at least a first vector in the sequence of vectors, wherein the first vector encodes attributes of the human face captured in a first image in the first sequence of images; and processing, by the neural network, the sequence of vectors to reconstruct a second sequence of images of the human face including the at least one attribute that is modified based on the style.

2. The computer-implemented method of claim 1 , wherein the first vector is a compressed encoding of the human face.

3. The computer-implemented method of claim 1 , wherein the first image in the first sequence of images is a frame of video and further comprising receiving vector adjustment values for each additional frame of the video.

4. The computer-implemented method of claim 3 , further comprising successively applying each vector adjustment value to the first vector to reconstruct additional images of the human face including the at least one attribute that is modified based on the style.

5. The computer-implemented method of claim 1 , wherein the attributes comprise head pose and facial expression.

6. The computer-implemented method of claim 1 , wherein the first vector encodes at least one additional attribute associated with clothing, hairstyle, or lighting.

7. The computer-implemented method of claim 1 , further comprising displaying the second sequence of images of the human face in a viewing environment, wherein the neural network reconstructs the second sequence of images according to lighting in the viewing environment instead of different lighting associated with the first sequence of images and that is encoded in the first vector.

8. The computer-implemented method of claim 1 , further comprising receiving encoded background image data that is combined with the second sequence of images of the human face.

9. The computer-implemented method of claim 1 , wherein the first vector comprises an abstract latent code.

10. The computer-implemented method of claim 9 , wherein the abstract latent code is computed by a remote mapping neural network and transmitted to the neural network through the communication network.

11. The computer-implemented method of claim 9 , wherein the abstract latent code is computed by transforming facial landmark points that delineate positions of key points on the human face according to a learned or optimized matrix.

12. The computer-implemented method of claim 1 , wherein the first vector is transmitted to the neural network during a videoconferencing session.

13. The computer-implemented method of claim 1 , wherein the subject is a real or synthetic character and the human face is a different human compared with the subject.

14. The computer-implemented method of claim 1 , wherein the subject is a real or synthetic character and the human face corresponds to the subject.

15. The computer-implemented method of claim 1 , further comprising interpolating a third vector and a second vector corresponding to two frames in a video to produce the first vector in the sequence of vectors, wherein the first image is between the two frames.

16. The computer-implemented method of claim 1 , further comprising receiving audio data, wherein the audio data is used to reconstruct the second sequence of images of the human face.

17. The computer-implemented method of claim 1 , wherein the first vector comprises a first portion corresponding to a first frame in a video and a second portion corresponding to a second frame in the video, wherein the human face is more blurry in the first frame compared to the second frame.

18. The computer-implemented method of claim 17 , wherein the processing combines the first portion and the second portion to reconstruct an image in the second sequence of images with the human face by using the first portion to control coarse scale styles and the second portion to control fine scale styles.

19. The computer-implemented method of claim 1 , wherein the human face captured in the first image is blurry and the processing reconstructs an image in the second sequence of images with the human face by using the first vector to control coarse scale styles and the data to control fine scale styles.

20. The computer-implemented method of claim 1 , further comprising displaying the second sequence of images of the human face in a viewing environment, wherein the neural network reconstructs the second sequence of images according to a gaze location within each image in the second sequence of images that is intersected by a gaze direction of a viewer observing the image as sensed in the viewing environment.

21. The computer-implemented method of claim 20 , wherein a gaze direction of the human face in each image in the second sequence of images is modified to appear directed towards the gaze location.

22. The computer-implemented method of claim 1 , wherein the first vector includes a gaze location corresponding to an image viewed by the human face and a gaze direction of a reconstructed image of the human face in a viewing environment is towards the image that is also reconstructed and displayed in the viewing environment.

23. The computer-implemented method of claim 1 , wherein the steps of obtaining, receiving, and processing are performed on a virtual machine comprising a portion of a graphics processing unit.

24. The computer-implemented method of claim 1 , wherein the first sequence of images or the second sequence of images is used for training, testing, or certifying a neural network employed in a machine, robot, or autonomous vehicle.

25. A system, comprising a processor configured to:

obtain data for replicating a style specific to a subject, wherein the data is determined by training a generator neural network to produce images of the subject that are compared with captured images of the subject; and implement a neural network that is configured to: apply the data to a sequence of vectors generated for a human face captured in a first sequence of images to modify at least one attribute according to the style specific to the subject; receive, through a communication network, at least a first vector in the sequence of vectors, wherein the first vector encodes attributes of the human face captured in a first image in the first sequence of images; and process the sequence of vectors to reconstruct a second sequence of images of the human face including the at least one attribute that is modified based on the style.

26. A non-transitory, computer-readable storage medium storing instructions that, when executed by a processing unit, cause the processing unit to: obtain data for replicating a style specific to a subject, wherein the data is determined by training a generator neural network to produce images of the subject that are compared with captured images of the subject; and

implement a neural network that is configured to: apply the data to multiple vectors to modify at least one attribute according to the style specific to the subject; receive, through a communication network, a at least a first vector in the sequence of vectors, wherein the first vector encodes attributes of the human face captured in a first image in the first sequence of images; and process the sequence of vectors to reconstruct a second sequence of images of the human face including the at least one attribute that is modified based on the style.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2020
From: KARRAS, TERO TAPANI; LAINE, SAMULI MATIAS; LUEBKE, DAVID PATRICK; LEHTINEN, JAAKKO T.; AITTALA, MIIKA SAMULI; AILA, TIMO OSKARI; LIU, MING-YU; MALLYA, ARUN MOHANRAY; WANG, TING-CHUN
To: NVIDIA CORPORATION
Reel/Frame 054077/0641 →
Continuity (5)
Continuation In Part 16418317 · May 21, 2019
Provisional Application 63010511 · Apr 15, 2020
Provisional Application 62767985 · Nov 15, 2018
Provisional Application 62767417 · Nov 14, 2018
Related Publication 20210049468A1 · Feb 18, 2021
Cited By (4)
US 12,506,886 US 12,548,274 US 12,651,459 US 12,659,428