Creating real-time interactive videos
The present disclosure describes techniques for creating a real-time interactive video. A source image is generated by a first machine learning model based on capturing an image of a user. The image comprises a face of the user. One or more facial images of the user are captured. The one or more facial images depict one or more facial expressions. The source image and information extracted from the one or more facial images are input into a second machine learning model. The second machine learning model is configured and trained to transfer facial expressions of creators to machine-generated images in real-time. The real-time interactive video is created by dynamically driving the source image based on the one or more facial expressions.
1 . A method of creating a real-time interactive video, comprising:
capturing an image of a user, wherein the image comprises a face of the user;
displaying an interface by a device during a process of generating a source image based on the captured image, wherein the interface comprises an indication that the process of generating the source image using a first machine learning model is in progress, and wherein the source image comprises a machine-generated avatar of the user;
displaying a message by the device in response to determining that the process of generating the source image is complete, wherein the message prompts the user to make one or more facial expressions;
capturing one or more facial images of the user by the device, wherein the one or more facial images depict the one or more facial expressions;
inputting, via the device, the source image and information extracted from the one or more facial images into a second machine learning model, wherein the second machine learning model is configured and trained to transfer facial expressions of creators to machine-generated images in real-time;
causing to display the one or more facial expressions on the machine-generated avatar of the user by the device; and
creating, by the device, the real-time interactive video by dynamically driving the machine-generated avatar of the user based on the one or more facial expressions.
2 . The method of claim 1 , further comprising:
causing to display an interface configured to guide the user to position the face at a predetermined location; and
generating the source image by the first machine learning model based on scanning the face positioned at the predetermined location.
3 . The method of claim 1 , further comprising:
causing to display the source image generated by the first machine learning model; and
causing to display information configured to prompt the user to show a facial expression.
4 . The method of claim 1 , further comprising:
extracting facial landmark data from the one or more facial images; and
inputting the facial landmark data into the second machine learning model.
5 . The method of claim 4 , further comprising:
detecting key points indicative of one or more motion fields associated with the one or more facial expressions by a first sub-model of the second machine learning model.
6 . The method of claim 5 , further comprising:
generating a deformation file based on the key points by a second sub-model of the second machine learning model; and
refining the one or more motion fields and generating occlusion maps by the second sub-model of the second machine learning model.
7 . The method of claim 6 , further comprising:
deforming the source image based on the deformation file by the second sub-model of the second machine learning model.
8 . The method of claim 7 , further comprising:
inputting the occlusion maps and the deformed source image into a third sub-model of the second machine learning model; and
generating one or more images that comprises the one or more facial expressions on the source image by the third sub-model of the second machine learning model.
9 . The method of claim 1 , wherein the method is implemented by a mobile computing device.
10 . A system of creating a real-time interactive video, comprising:
at least one processor; and
at least one memory communicatively coupled to the at least one processor and comprising computer-readable instructions that upon execution by the at least one processor cause the at least one processor to perform operations comprising:
capturing an image of a user, wherein the image comprises a face of the user;
displaying an interface by a device during a process of generating a source image based on the captured image, wherein the interface comprises an indication that the process of generating the source image using a first machine learning model is in progress, and wherein the source image comprises a machine-generated avatar of the user;
displaying a message by the device in response to determining that the process of generating the source image is complete, wherein the message prompts the user to make one or more facial expressions;
capturing one or more facial images of the user by the device, wherein the one or more facial images depict the one or more facial expressions;
inputting, via the device, the source image and information extracted from the one or more facial images into a second machine learning model, wherein the second machine learning model is configured and trained to transfer facial expressions of creators to machine-generated images in real-time;
causing to display the one or more facial expressions on the machine-generated avatar of the user by the device; and
creating, by the device, the real-time interactive video by dynamically driving the machine-generated avatar of the user based on the one or more facial expressions.
11 . The system of claim 10 , the operations further comprising:
causing to display an interface configured to guide the user to position the face at a predetermined location; and
generating the source image by the first machine learning model based on scanning the face positioned at the predetermined location.
12 . The system of claim 10 , the operations further comprising:
causing to display the source image generated by the first machine learning model; and
causing to display information configured to prompt the user to show a facial expression.
13 . The system of claim 10 , the operations further comprising:
extracting facial landmark data from the one or more facial images; and
inputting the facial landmark data into the second machine learning model.
14 . The system of claim 13 , the operations further comprising:
detecting key points indicative of one or more motion fields associated with the one or more facial expressions by a first sub-model of the second machine learning model.
15 . The system of claim 14 , the operations further comprising:
generating a deformation file based on the key points by a second sub-model of the second machine learning model;
deforming the source image based on the deformation file by the second sub-model of the second machine learning model; and
refining the one or more motion fields and generating occlusion maps by the second sub-model of the second machine learning model.
16 . The system of claim 15 , the operations further comprising:
inputting the occlusion maps and the deformed source image into a third sub-model of the second machine learning model; and
generating one or more images that comprises the one or more facial expressions on the source image by the third sub-model of the second machine learning model.
17 . A non-transitory computer-readable storage medium, storing computer-readable instructions that upon execution by a processor cause the processor to implement operations comprising:
capturing an image of a user, wherein the image comprises a face of the user;
displaying an interface by a device during a process of generating a source image based on the captured image, wherein the interface comprises an indication that the process of generating the source image using a first machine learning model is in progress, and
wherein the source image comprises a machine-generated avatar of the user;
displaying a message by the device in response to determining that the process of generating the source image is complete, wherein the message prompts the user to make one or more facial expressions;
capturing one or more facial images of the user by the user device, wherein the one or more facial images depict the one or more facial expressions;
inputting, via the device, the source image and information extracted from the one or more facial images into a second machine learning model, wherein the second machine learning model is configured and trained to transfer facial expressions of creators to machine-generated images in real-time;
causing to display the one or more facial expressions on the machine-generated avatar of the user by the device; and
creating, by the device, the real-time interactive video by dynamically driving the machine-generated avatar of the user based on the one or more facial expressions.
18 . The non-transitory computer-readable storage medium of claim 17 , the operations further comprising:
extracting facial landmark data from the one or more facial images; and
inputting the facial landmark data into the second machine learning model.
19 . The non-transitory computer-readable storage medium of claim 18 , the operations further comprising:
detecting key points indicative of one or more motion fields associated with the one or more facial expressions by a first sub-model of the second machine learning model;
generating a deformation file based on the key points by a second sub-model of the second machine learning model;
deforming the source image based on the deformation file by the second sub-model of the second machine learning model; and
refining the one or more motion fields and generating occlusion maps by the second sub-model of the second machine learning model.
20 . The non-transitory computer-readable storage medium of claim 19 , the operations further comprising:
inputting the occlusion maps and the deformed source image into a third sub-model of the second machine learning model; and
generating one or more images that comprises the one or more facial expressions on the source image by the third sub-model of the second machine learning model.