IP Library Granted Patent US 11,288,880
Granted Patent B2
US 11,288,880 · App. 16/661,086 · Granted Mar 29, 2022

Template-based generation of personalized videos

Inventors: Victor Shaburov (Ocean Village, GI); Alexander Mashrabov (Sochi, RU); Dmitriy Matov (Saratov, RU); Sofia Savinova (Sochi, RU); Alexey Pchelnikov (Sochi, RU); Roman Golobokov (Sochi, RU)
Assignee: Snap Inc.
G06T19/20G06K9/00228G06K9/00281G06K9/00288G06K9/00302G06T13/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,288,880
App. No.
16/661,086
Granted
Mar 29, 2022
Kind
B2
Abstract

Disclosed are systems and methods for template-based generation of personalized videos. An example method may commence with receiving video configuration data including a sequence of frame images, a sequence of face area parameters defining positions of a face area in the frame images, and a sequence of facial landmark parameters defining positions of facial landmarks in the frame images. The method may continue with receiving an image of a source face. The method may further include generating an output video. The generation of the output video may include modifying a frame image of the sequence of frame images. Specifically, the image of the source face may be modified to obtain a further image featuring the source face adopting a facial expression corresponding to the facial landmark parameters. The further image may be inserted into the frame image at a position determined by face area parameters corresponding to the frame image.

Claims (80)

1. A method for template-based generation of personalized videos, the method comprising:

receiving, by a computing device, video configuration data including:

a sequence of frame images featuring images of an actor;

a sequence of face area parameters corresponding to positions of a face area of the actor in the frame images; and

a sequence of facial landmark parameters corresponding to the frame images, wherein the facial landmark parameters are generated based on further frame images featuring a face of a face synchronization actor absent from the frame images;

receiving, by the computer device, an image of a source face absent from the frame images and the further frame images; and

generating, by the computing device, an output video, wherein the generating the output video includes modifying a frame image of the sequence of frame images by:

modifying, based on facial landmark parameters corresponding to the frame image, the image of the source face to obtain a further face image featuring the source face adopting a facial expression corresponding to the facial landmark parameters; and

inserting the further face image into the frame image at a position determined by face area parameters corresponding to the frame image.

2. The method of claim 1 , wherein the sequence of frame images is generated based on one of the following: an animation video and a live action video.

3. The method of claim 1 , wherein the sequence of facial landmark parameters is generated based on a live action video featuring the face of the face synchronization actor.

4. The method of claim 1 , wherein:

the video configuration data include a sequence of skin masks defining a skin area of a body of the actor featured in the frame images or a skin area of 2D/3D animation of a further body; and

the generating the output video includes:

determining color data associated with the source face; and

recoloring, based on the color data, the skin area in the frame image.

5. The method of claim 1 , wherein:

the video configuration data further include a sequence of mouth region images, each of the mouth region images corresponding to at least one of the frame images; and

the generating the output video includes inserting, into the frame image, a mouth region corresponding to the frame image.

6. The method of claim 1 , wherein:

the video configuration data further include a sequence of eye parameters corresponding to the frame images and determined based on positions of an iris in a sclera of the face synchronization actor featured in the further frame images; and

the generating the output video includes:

generating, based on the eye parameters corresponding to the frame, an image of eyes region; and

inserting the image of eyes region in the frame image.

7. The method of claim 1 , wherein:

the video configuration data include a sequence of head parameters defining one or more of a rotation, a turn, a position, and a scale of a head.

8. The method of claim 1 , wherein:

the generating the output video includes:

determining, based on the image of the source face, a hair mask;

generating, based on the hair mask, a hair image; and

inserting the hair image into the frame image.

9. The method of claim 1 , wherein:

the video configuration data include a sequence of animated object images, wherein each of the animated object images corresponds to at least one of the frame images; and

the generating the output video includes inserting, into the frame image, an animated object image corresponding to the frame image.

10. The method of claim 1 , wherein:

the video configuration data include a soundtrack; and

the generating the output video further includes adding the soundtrack to the output video.

11. A system for template-based generation of personalized videos, the system comprising at least one processor and a memory storing processor-executable codes, wherein the at least one processor is configured to implement the following operations upon executing the processor-executable codes:

receiving, by a computing device, video configuration data including:

a sequence of frame images featuring images of an actor;

a sequence of face area parameters corresponding to positions of a face area of the actor in the frame images; and

a sequence of facial landmark parameters corresponding to the frame images, wherein the facial landmark parameters are generated based on further frame images featuring a face of a face synchronization actor absent from the frame images;

receiving, by the computer device, an image of a source face absent from the frame images and the further frame images; and

generating, by the computing device, an output video, wherein the generating the output video includes modifying a frame image of the sequence of frame images by:

modifying, based on facial landmark parameters corresponding to the frame image, the image of the source face to obtain a further face image featuring the source face adopting a facial expression corresponding to the facial landmark parameters; and

inserting the further face image into the frame image at a position determined by face area parameters corresponding to the frame image.

12. The system of claim 11 , wherein the sequence of frame images is generated based on one of the following: an animation video and a live action video.

13. The system of claim 11 , wherein the sequence of facial landmark parameters is generated based on a live action video featuring the face of the face synchronization actor.

14. The system of claim 11 , wherein:

the video configuration data include a sequence of skin masks defining a skin area of a body of the actor featured in the frame images or a skin area of 2D/3D animation of a further body; and

the generating the output video includes:

determining color data associated with the source face; and

recoloring, based on the color data, the skin area in the frame image.

15. The system of claim 11 , wherein:

the video configuration data further include a sequence of mouth region images, each of the mouth region images corresponding to at least one of the frame images; and

the generating the output video includes inserting, into the frame image, a mouth region corresponding to the frame image.

16. The system of claim 11 , wherein:

the video configuration data further include a sequence of eye parameters corresponding to the frame images and determined based on positions of an iris in a sclera of the face synchronization actor featured in the further frame images; and

the generating the output video includes:

generating, based on the eye parameters corresponding to the frame, an image of eyes region; and

inserting the image of eyes region in the frame image.

17. The system of claim 11 , wherein:

the video configuration data include a sequence of head parameters defining one or more of a rotation, a turn, a position, and a scale of a head.

18. The system of claim 11 , wherein:

the generating the output video includes:

determining, based on the source face image, a hair mask;

generating, based on the hair mask, a hair image; and

inserting the hair image into the frame image.

19. The system of claim 11 , wherein:

the video configuration data include a sequence of animated object images, wherein each of the animated object images corresponds to at least one of the frame images; and

the generating the output video includes inserting, into the frame image, an animated object image corresponding to the frame image.

20. A non-transitory processor-readable medium having instructions stored thereon, which when executed by one or more processors, cause the one or more processors to implement a method for template-based generation of personalized videos, the method comprising:

receiving, by a computing device, video configuration data including:

a sequence of frame images featuring images of an actor;

a sequence of face area parameters corresponding to positions of a face area of the actor in the frame images; and

a sequence of facial landmark parameters corresponding to the frame images, wherein the facial landmark parameters are generated based on further frame images featuring a face of a face synchronization actor absent from the frame images;

receiving, by the computer device, an image of a source face absent from the frame images and the further frame images; and

generating, by the computing device, an output video, wherein the generating the output video includes modifying a frame image of the sequence of frame images by:

modifying, based on facial landmark parameters corresponding to the frame image, the image of the source face to obtain a further face image featuring the source face adopting a facial expression corresponding to the facial landmark parameters; and

inserting the further face image into the frame image at a position determined by face area parameters corresponding to the frame image.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2020
From: AI FACTORY, INC.
To: SNAP INC.
Reel/Frame 051789/0760 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2019
From: SHABUROV, VICTOR; MASHRABOV, ALEXANDER; MATOV, DMITRIY; SAVINOVA, SOFIA; PCHELNIKOV, ALEXEY; GOLOBKOV, ROMAN
To: AI FACTORY, INC.
Reel/Frame 050800/0546 →
Continuity (11)
Continuation In Part 16594771 · Oct 7, 2019
Continuation In Part 16251436 · Jan 18, 2019
Continuation In Part 16661086
Continuation In Part 16594690 · Oct 7, 2019
Continuation In Part 16251436 · Jan 18, 2019
Continuation In Part 16661086
Continuation In Part 16551756 · Aug 27, 2019
Continuation In Part 16434185 · Jun 7, 2019
Continuation In Part 16661086
Continuation In Part 16251472 · Jan 18, 2019
Related Publication 20200234508A1 · Jul 23, 2020
Cited By (1)
US 12,705,840