IP Library Granted Patent US 12694593
Granted Patent B2
US 12694593 · App. 18/388,785 · Granted Jul 28, 2026

Creating real-time interactive videos

Inventors: Peilin Li (Los Angeles, CA); Guoxian Song (Los Angeles, CA)
Assignee: Lemon Inc.
G06T13/00G06T7/248G06V10/77G06V10/945G06V40/176G06T2200/24G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12694593
App. No.
18/388,785
Granted
Jul 28, 2026
Kind
B2
Abstract

The present disclosure describes techniques for creating a real-time interactive video. A source image is generated by a first machine learning model based on capturing an image of a user. The image comprises a face of the user. One or more facial images of the user are captured. The one or more facial images depict one or more facial expressions. The source image and information extracted from the one or more facial images are input into a second machine learning model. The second machine learning model is configured and trained to transfer facial expressions of creators to machine-generated images in real-time. The real-time interactive video is created by dynamically driving the source image based on the one or more facial expressions.

Claims (76)

1 . A method of creating a real-time interactive video, comprising:

capturing an image of a user, wherein the image comprises a face of the user;

displaying an interface by a device during a process of generating a source image based on the captured image, wherein the interface comprises an indication that the process of generating the source image using a first machine learning model is in progress, and wherein the source image comprises a machine-generated avatar of the user;

displaying a message by the device in response to determining that the process of generating the source image is complete, wherein the message prompts the user to make one or more facial expressions;

capturing one or more facial images of the user by the device, wherein the one or more facial images depict the one or more facial expressions;

inputting, via the device, the source image and information extracted from the one or more facial images into a second machine learning model, wherein the second machine learning model is configured and trained to transfer facial expressions of creators to machine-generated images in real-time;

causing to display the one or more facial expressions on the machine-generated avatar of the user by the device; and

creating, by the device, the real-time interactive video by dynamically driving the machine-generated avatar of the user based on the one or more facial expressions.

2 . The method of claim 1 , further comprising:

causing to display an interface configured to guide the user to position the face at a predetermined location; and

generating the source image by the first machine learning model based on scanning the face positioned at the predetermined location.

3 . The method of claim 1 , further comprising:

causing to display the source image generated by the first machine learning model; and

causing to display information configured to prompt the user to show a facial expression.

4 . The method of claim 1 , further comprising:

extracting facial landmark data from the one or more facial images; and

inputting the facial landmark data into the second machine learning model.

5 . The method of claim 4 , further comprising:

detecting key points indicative of one or more motion fields associated with the one or more facial expressions by a first sub-model of the second machine learning model.

6 . The method of claim 5 , further comprising:

generating a deformation file based on the key points by a second sub-model of the second machine learning model; and

refining the one or more motion fields and generating occlusion maps by the second sub-model of the second machine learning model.

7 . The method of claim 6 , further comprising:

deforming the source image based on the deformation file by the second sub-model of the second machine learning model.

8 . The method of claim 7 , further comprising:

inputting the occlusion maps and the deformed source image into a third sub-model of the second machine learning model; and

generating one or more images that comprises the one or more facial expressions on the source image by the third sub-model of the second machine learning model.

9 . The method of claim 1 , wherein the method is implemented by a mobile computing device.

10 . A system of creating a real-time interactive video, comprising:

at least one processor; and

at least one memory communicatively coupled to the at least one processor and comprising computer-readable instructions that upon execution by the at least one processor cause the at least one processor to perform operations comprising:

capturing an image of a user, wherein the image comprises a face of the user;

displaying an interface by a device during a process of generating a source image based on the captured image, wherein the interface comprises an indication that the process of generating the source image using a first machine learning model is in progress, and wherein the source image comprises a machine-generated avatar of the user;

displaying a message by the device in response to determining that the process of generating the source image is complete, wherein the message prompts the user to make one or more facial expressions;

capturing one or more facial images of the user by the device, wherein the one or more facial images depict the one or more facial expressions;

inputting, via the device, the source image and information extracted from the one or more facial images into a second machine learning model, wherein the second machine learning model is configured and trained to transfer facial expressions of creators to machine-generated images in real-time;

causing to display the one or more facial expressions on the machine-generated avatar of the user by the device; and

creating, by the device, the real-time interactive video by dynamically driving the machine-generated avatar of the user based on the one or more facial expressions.

11 . The system of claim 10 , the operations further comprising:

causing to display an interface configured to guide the user to position the face at a predetermined location; and

generating the source image by the first machine learning model based on scanning the face positioned at the predetermined location.

12 . The system of claim 10 , the operations further comprising:

causing to display the source image generated by the first machine learning model; and

causing to display information configured to prompt the user to show a facial expression.

13 . The system of claim 10 , the operations further comprising:

extracting facial landmark data from the one or more facial images; and

inputting the facial landmark data into the second machine learning model.

14 . The system of claim 13 , the operations further comprising:

detecting key points indicative of one or more motion fields associated with the one or more facial expressions by a first sub-model of the second machine learning model.

15 . The system of claim 14 , the operations further comprising:

generating a deformation file based on the key points by a second sub-model of the second machine learning model;

deforming the source image based on the deformation file by the second sub-model of the second machine learning model; and

refining the one or more motion fields and generating occlusion maps by the second sub-model of the second machine learning model.

16 . The system of claim 15 , the operations further comprising:

inputting the occlusion maps and the deformed source image into a third sub-model of the second machine learning model; and

generating one or more images that comprises the one or more facial expressions on the source image by the third sub-model of the second machine learning model.

17 . A non-transitory computer-readable storage medium, storing computer-readable instructions that upon execution by a processor cause the processor to implement operations comprising:

capturing an image of a user, wherein the image comprises a face of the user;

displaying an interface by a device during a process of generating a source image based on the captured image, wherein the interface comprises an indication that the process of generating the source image using a first machine learning model is in progress, and

wherein the source image comprises a machine-generated avatar of the user;

displaying a message by the device in response to determining that the process of generating the source image is complete, wherein the message prompts the user to make one or more facial expressions;

capturing one or more facial images of the user by the user device, wherein the one or more facial images depict the one or more facial expressions;

inputting, via the device, the source image and information extracted from the one or more facial images into a second machine learning model, wherein the second machine learning model is configured and trained to transfer facial expressions of creators to machine-generated images in real-time;

causing to display the one or more facial expressions on the machine-generated avatar of the user by the device; and

creating, by the device, the real-time interactive video by dynamically driving the machine-generated avatar of the user based on the one or more facial expressions.

18 . The non-transitory computer-readable storage medium of claim 17 , the operations further comprising:

extracting facial landmark data from the one or more facial images; and

inputting the facial landmark data into the second machine learning model.

19 . The non-transitory computer-readable storage medium of claim 18 , the operations further comprising:

detecting key points indicative of one or more motion fields associated with the one or more facial expressions by a first sub-model of the second machine learning model;

generating a deformation file based on the key points by a second sub-model of the second machine learning model;

deforming the source image based on the deformation file by the second sub-model of the second machine learning model; and

refining the one or more motion fields and generating occlusion maps by the second sub-model of the second machine learning model.

20 . The non-transitory computer-readable storage medium of claim 19 , the operations further comprising:

inputting the occlusion maps and the deformed source image into a third sub-model of the second machine learning model; and

generating one or more images that comprises the one or more facial expressions on the source image by the third sub-model of the second machine learning model.