IP Library Granted Patent US 11,114,086
Granted Patent B2
US 11,114,086 · App. 16/509,370 · Granted Sep 7, 2021

Text and audio-based real-time face reenactment

Inventors: Pavel Savchenkov (Sochi, RU); Maxim Lukin (Sochi, RU); Aleksandr Mashrabov (Sochi, RU)
Assignee: Snap Inc.
G10L13/00G06K9/00281G06T13/40G10L13/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,114,086
App. No.
16/509,370
Granted
Sep 7, 2021
Kind
B2
Abstract

Provided are systems and methods for text and audio-based real-time face reenactment. An example method includes receiving an input text and a target image, the target image including a target face; generating, based on the input text, a sequence of sets of acoustic features representing the input text; determining, based on the sequence of sets of acoustic features, a sequence of sets of scenario data indicating modifications of the target face for pronouncing the input text; generating, based on the sequence of sets of scenario data, a sequence of frames, wherein each of the frames includes the target face modified based on at least one of the sets of scenario data; generating, based on the sequence of frames, an output video; and synthesizing, based on the sequence of sets of acoustic features, an audio data and adding the audio data to the output video.

Claims (68)

1. A method for text and audio-based real-time face reenactment, the method comprising:

receiving, by a computing device, an input text and a target image, the target image including a target face;

generating, by the computing device and based on the input text, a sequence of sets of acoustic features representing the input text;

generating, by the computing device and based on the sequence of sets of acoustic features, a sequence of sets of scenario data, the sets of scenario data indicating modifications of facial key points of a model face pronouncing the input text, wherein the generating the sequence of sets of scenario data includes:

generating, based on the sequence of sets of acoustic features, a sequence of sets of mouth key points of the model face by providing the sets of acoustic features to a first model, the first model being trained to generate, based on a set of acoustic features, a set of mouth key points, the first model being trained based on first images of a first actor pronouncing first training input texts; and

generating, based on the sequence of sets of mouth key points, a sequence of sets of the facial key points of the model face by providing the sets of mouth key points to a second model, the second model being trained to generate, based on the set of mouth key points, a set of the facial key points, the second model being trained based on second images of a second actor pronouncing second training input texts;

generating, by the computing device and based on the sequence of sets of scenario data and the target image, a sequence of frames, wherein each of the frames includes the target face modified based on at least one set of scenario data of the sequence of sets of scenario data; and

generating, by the computing device and based on the sequence of frames, an output video.

2. The method of claim 1 , further comprising:

synthesizing, by the computing device and based on the sequence of sets of acoustic features, an audio data representing the input text; and

adding, by the computing device, the audio data to the output video.

3. The method of claim 1 , wherein the acoustic features include Mel- frequency cepstral coefficients.

4. The method of claim 1 , wherein the sequence of sets of acoustic features is generated by a neural network.

5. The method of claim 1 , wherein:

the generating the sequence of frames includes:

determining, based on a sequence of sets of the facial key points, a sequence of sets of two-dimensional (2D) deformations; and

applying each set of 2D deformations of the sequence of the sets of 2D deformations to the target image to obtain the sequence of frames.

6. The method of claim 5 , wherein:

the sequence of sets of mouth key points is generated by a neural network; and

at least one set of the sequence of sets of mouth key points is generated based on a pre-determined number of sets preceding the at least one set in the sequence of sets of mouth key points.

7. The method of claim 6 , wherein:

the at the least one set of the sequence of sets of mouth key points corresponds to at least one set (S) of the sequence of sets of acoustic features; and

the at least one set of the sequence of sets of mouth key points is generated based on a first pre-determined number of sets of acoustic features preceding the S in the sequence of sets of acoustic features and a second pre-determined number sets of acoustic features succeeding the S in the sequence of sets of acoustic features.

8. The method of claim 5 , wherein:

the sequence of sets of facial key points is generated by a neural network; and

at least one set of the sequence of sets of facial key points is determined based on a pre-determined number of sets preceding the at least one set in the sequence of sets of facial key points.

9. The method of claim 5 , further comprising:

generating, by the computing device and based on the sequence of sets of mouth key points, a sequence of mouth texture images; and

inserting, by the computing device, each of the sequence of mouth texture images in a corresponding frame of the sequence of the frames.

10. The method of claim 9 , wherein each mouth texture image of the sequence of mouth texture images is generated by a neural network based on a first pre-determined number of mouth texture images preceding the mouth region image in the sequence of mouth region images.

11. A system for text and audio-based real-time face reenactment, the system comprising at least one processor, a memory storing processor-executable codes, wherein the at least one processor is configured to implement the following operations upon executing the processor-executable codes:

receiving an input text and a target image, the target image including a target face;

generating, based on the input text, a sequence of sets of acoustic features representing the input text;

generating, based on the sequence of sets of acoustic features, a sequence of sets of scenario data, the sets of scenario data indicating modifications of facial key points of a model face pronouncing the input text, wherein the

generating the sequence of sets of scenario data includes:

generating, based on the sequence of sets of acoustic features, a sequence of sets of mouth key points of the model face by providing the sets of acoustic features to a first model, the first model being trained to generate, based on a set of acoustic features, a set of mouth key points, the first model being trained based on first images of a first actor pronouncing first training input texts; and

generating, based on the sequence of sets of mouth key points, a sequence of sets of the facial key points of the model face by providing the sets of mouth key points to a second model, the second model being trained to generate, based on the set of mouth key points, a set of the facial key points, the second model being trained based on second images of a second actor pronouncing second training input texts;

generating, based on the sequence of sets of scenario data and the target image, a sequence of frames, wherein each of the frames includes the target face modified based on at least one set of scenario data of the sequence of sets of scenario data; and

generating, based on the sequence of frames, an output video.

12. The system of claim 11 , further comprising:

synthesizing, based on the sequence of sets of acoustic features, an audio data representing the input text; and

adding the audio data to the output video.

13. The system of claim 11 , wherein the acoustic features include Mel-frequency cepstral coefficients.

14. The system of claim 11 , wherein the sequence of sets of acoustic features is generated based on a neural network.

15. The system of claim 11 , wherein:

the generating the sequence of frames includes:

determining, based on a sequence of sets of the facial key points, a sequence of sets of 2D deformations; and

applying each set of two-dimensional (2D) deformations of the sequence of the sets of 2D deformations to the target image to obtain the sequence of frames.

16. The system of claim 15 , wherein:

the sequence of sets of mouth key points is generated by a neural network; and

at least one set of the sequence of sets of mouth key points is generated based on a pre-determined number of sets preceding the at least one set in the sequence of sets of mouth key points.

17. The system of claim 16 , wherein:

the at the least one set of the sequence of sets of mouth key points corresponds to at least one set (S) of the sequence of sets of acoustic features; and

the at least one set of the sequence of sets of mouth key points is generated based on a first pre-determined number of sets of acoustic features preceding the S in the sequence of sets of acoustic features and a second pre-determined number sets of acoustic features succeeding the S in the sequence of sets of acoustic features.

18. The system of claim 15 , wherein:

the sequence of sets of facial key points is generated by a neural network; and

at least one set of the sequence of sets of facial key points is generated based on a pre-determined number of sets preceding the at least one set in the sequence of sets of facial key points.

19. The system of claim 15 , further comprising:

generating, by the computing device and based on the sequence of sets of mouth key points, a sequence of mouth texture images; and

inserting, by the computing device, each of the sequence of mouth texture images in a corresponding frame of the sequence of the frames.

20. A non-transitory processor-readable medium having instructions stored thereon, which when executed by one or more processors, cause the one or more processors to implement a method for text and audio- based real-time face reenactment, the method comprising:

receiving an input text and a target image, the target image including a target face;

generating, based on the input text, a sequence of sets of acoustic features representing the input text;

generating, based on the sequence of sets of acoustic features, a sequence of sets of scenario data, the sets of scenario data indicating modifications of facial key points of a model face pronouncing the input text, wherein the generating the sequence of sets of scenario data includes:

generating, based on the sequence of sets of acoustic features, a sequence of sets of mouth key points of the model face by providing the sets of acoustic features to a first model, the first model being trained to generate, based on a set of acoustic features, a set of mouth key points, the first model being trained based on first images of a first actor pronouncing first training input texts; and

generating, based on the sequence of sets of mouth key points, a sequence of sets of the facial key points of the model face by providing the sets of mouth key points to a second model, the second model being trained to generate, based on the set of mouth key points, a set of the facial key points, the second model being trained based on second images of a second actor pronouncing second training input texts;

generating, based on the sequence of sets of scenario data and the target image, a sequence of frames, wherein each of the frames includes the target face modified based on at least one set of scenario data of the sequence of sets of scenario data; and

generating, based on the sequence of frames, an output video.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2020
From: AI FACTORY, INC.
To: SNAP INC.
Reel/Frame 051789/0760 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2019
From: SAVCHENKOV, PAVEL; LUKIN, MAXIM; MASHRABOV, ALEKSANDR
To: AI FACTORY, INC.
Reel/Frame 049737/0359 →
Continuity (2)
Continuation In Part 16251436 · Jan 18, 2019
Related Publication 20200234690A1 · Jul 23, 2020