IP Library › Granted Patent US 12,450,810
Granted Patent B2
US 12,450,810 · App. 18/190,684 · Granted Oct 21, 2025

Animated facial expression and pose transfer utilizing an end-to-end machine learning model

Inventor: Cameron Smith (Santa Clara, CA)
Assignee: Adobe Inc.
G06T13/40G06T7/251G06V10/95G06V40/176G06T2200/24G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,810
App. No.
18/190,684
Granted
Oct 21, 2025
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify digital images via scene-based editing using image understanding facilitated by artificial intelligence. For example, in one or more embodiments the disclosed systems utilize generative machine learning models to create modified digital images portraying human subjects. In particular, the disclosed systems generate modified digital images by performing infill modifications to complete a digital image or human inpainting for portions of a digital image that portrays a human. Moreover, in some embodiments, the disclosed systems perform reposing of subjects portrayed within a digital image to generate modified digital images. In addition, the disclosed systems in some embodiments perform facial expression transfer and facial expression animations to generate modified digital images or animations.

Claims (79)

1. A computer-implemented method comprising:

extracting, utilizing a first three-dimensional encoder, a first target facial expression animation embeddings for a first resolution from a first frame of a target digital video portraying a target animation of a face;

extracting, utilizing the first three-dimensional encoder, a first target pose animation embeddings for the first resolution from the first frame of the target digital video;

identifying a static source digital image portraying a source face having a source shape and facial expression;

generating, utilizing a second three-dimensional encoder, a first source shape embedding for the first resolution from the static source digital image;

generating a first combined embeddings by concatenating the first target facial expression animation embeddings for the first resolution from the first frame of the target digital video, the first target pose animation embeddings for the first resolution from the first frame of the target digital video, and the first source shape embedding for the first resolution from the static source digital image;

extracting, utilizing the first three-dimensional encoder, a second target facial expression animation embedding for a second resolution from the first frame of the target digital video portraying the target animation of the face;

extracting, utilizing the first three-dimensional encoder, a second target pose animation embedding for the second resolution from the first frame of the target digital video;

generating, utilizing the second three-dimensional encoder, a second source shape embedding for the second resolution from the static source digital image;

generating a second combined embedding by concatenating the second target facial expression animation embedding for the second resolution from the first frame of the target digital video, the second target pose animation embedding for the second resolution from the first frame of the target digital video, and the second source shape embedding for the second resolution from the static source digital image; and

generating, utilizing a facial animation generative adversarial neural network comprising a first layer corresponding to the first resolution and a second layer corresponding to the second resolution, an animation by conditioning the first layer of the facial animation generative adversarial neural network with the first combined embeddings and conditioning the second layer of the facial animation generative adversarial neural network with the second combined embedding, wherein the animation portrays the source face animated according to the target animation from the target digital video.

2. The computer-implemented method of claim 1 , wherein generating the animation comprises:

utilizing a comodulated generative adversarial neural network as the facial animation generative adversarial neural network to generate the animation by:

generating, utilizing a first modulation layer of the first layer of the comodulated generative adversarial neural network, an intermediate vector from the static source digital image by conditioning the first modulation layer according to the first combined embedding; and

generating, utilizing a second modulation layer of the second layer of the comodulated generative adversarial neural network, an additional intermediate vector from the intermediate vector by conditioning the second modulation layer according to the second combined embedding.

3. The computer-implemented method of claim 2 , further comprising generating, utilizing the comodulated generative adversarial neural network, the animation from the intermediate vector, the additional intermediate vector, and the static source digital image.

4. The computer-implemented method of claim 1 , further comprising generating, utilizing a three-dimensional morphable machine learning model as the first three-dimensional encoder, a third target facial expression animation embedding for the first resolution from a second frame of the target digital video and a third target pose animation embedding for the first resolution from the second frame of the target digital video.

5. The computer-implemented method of claim 4 , further comprising:

generating a third combined embedding by concatenating the third target facial expression animation embedding for the first resolution from the second frame, the third target pose animation embedding for the first resolution from the second frame, and the first source shape embedding for the first resolution from the static source digital image; and

generating, utilizing the facial animation generative adversarial neural network, the animation by conditioning the first layer of the facial animation generative adversarial neural network with the first combined embedding and the third combined embedding.

6. The computer-implemented method of claim 1 , further comprising:

training the facial animation generative adversarial neural network by:

accessing a digital video from a digital video dataset, wherein the digital video comprises a first training frame, a second training frame, and a third training frame;

generating a training source shape embedding from the first training frame, a training target pose embedding from the second training frame, and a training target facial expression embedding from the second training frame;

generating, a training combined embedding by concatenating the training source shape embedding, the training target pose embedding, and the training target facial expression embedding;

generating, utilizing the facial animation generative adversarial neural network, a training modified source digital image; and

comparing the training modified source digital image with the second training frame of the digital video to determine a measure of loss and modify parameters of the facial animation generative adversarial neural network.

7. The computer-implemented method of claim 1 , further comprising:

providing, for display via a user interface of a client device, a source image selection element and a target animation selection element; and

based on user interaction with the source image selection element and the target animation selection element, identify the static source digital image portraying the source face and extract the first target facial expression animation embeddings and the second target facial expression animation embedding.

8. The computer-implemented method of claim 7 , wherein providing, for display via the user interface of the client device, the target animation selection element comprises:

receiving a set of target digital videos from a client device;

extracting target facial expression data from the set of target digital videos;

generating a plurality of pre-defined target animations from the target facial expression data; and

providing the plurality of pre-defined target animations for display via a user interface of the client device, wherein the plurality of pre-defined target animations comprise a plurality of text descriptions of pre-defined target faces provided for display on the client device.

9. The computer-implemented method of claim 7 , further comprising identifying the target animation from the target digital video obtained from a camera roll of the client device based on user interaction with the target animation selection element.

10. A system comprising:

one or more memory devices comprising a target digital video, a static source digital image, and a facial animation generative neural network; and

one or more processors configured to cause the system to:

based on a user interaction with a source image selection element and a target animation selection element, extract, utilizing a first three-dimensional encoder, a first target facial expression animation embeddings for a first resolution and a first target pose animation embeddings for the first resolution from a first frame of the target digital video portraying a target animation of a face;

extract, utilizing a second three-dimensional encoder, a first source shape embedding for the first resolution from the static source digital image portraying a source face having a source shape and facial expression;

generate a first combined embedding by concatenating the first target facial expression animation embedding for the first resolution from the first frame of the target digital video, the first target pose animation embedding for the first resolution from the first frame of the target digital video, and the first source shape embedding for the first resolution from the static source digital image;

generate, utilizing a first denoising neural network, a first denoising representation from a diffusion noise representation by conditioning the first denoising neural network with the first combined embedding;

generate a second combined embedding by concatenating a second target facial expression animation embedding for a second resolution from the first frame of the target digital video, a second target pose animation embedding for the second resolution from the first frame of the target digital video, and a second source shape embedding for the first resolution from the static source digital image;

generate, utilizing one or more additional denoising neural networks, a final denoised representation from the first denoising representation by conditioning the one or more additional denoising neural networks with the second combined embedding;

generate, utilizing a decoder, an animation that portrays the source face animated according to the target animation of a target face from the final denoised representation; and

provide the animation for display via a user interface of a client device.

11. The system of claim 10 , wherein the one or more processors are configured to cause the system to provide, for display via the user interface of the client device, an option to select from a plurality of pre-defined target animations.

12. The system of claim 10 , wherein the one or more processors are configured to cause the system to provide, for display via the user interface of the client device, an option to select a digital image from a camera roll of the client device.

13. The system of claim 12 , wherein the one or more processors are configured to cause the system to identify the static source digital image portraying the source face by receiving a selection of the digital image from the camera roll of the client device.

14. The system of claim 10 , wherein the one or more processors are configured to cause the system to: generate a third target facial expression animation embedding for the first resolution from a second frame of the target digital video and a third target pose animation embedding for the first resolution from the second frame of the target digital video.

15. The system of claim 14 , wherein the one or more processors are configured to cause the system to:

generate a third combined embedding by concatenating the third target facial expression animation embedding for the first resolution from the second frame, the third target pose animation embedding for the first resolution from the second frame, and the first source shape embedding for the first resolution from the static source digital image; and

generate, utilizing the one or more additional denoising neural networks, the final denoised representation from the first denoising representation by conditioning the one or more additional denoising neural networks with the second combined embedding and the third combined embedding.

16. A non-transitory computer-readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

extracting, utilizing a first three-dimensional encoder, a first target pose animation embeddings and a first target facial expression animation embeddings for a first resolution from a first frame of a target digital video portraying a target animation of a face;

extracting, utilizing a second three-dimensional encoder, a first source shape embedding for the first resolution from a static source digital image that portrays a source face;

generating a first combined embeddings by concatenating the first target pose animation embeddings, the first target facial expression animation embeddings, and the first source shape embedding;

extracting, utilizing the first three-dimensional encoder, a second target facial expression animation embedding for a second resolution from the first frame of the target digital video portraying the target animation of the face;

extracting, utilizing the first three-dimensional encoder, a second target pose animation embedding for the second resolution from the first frame of the target digital video;

generating, utilizing the second three-dimensional encoder, a second source shape embedding for the second resolution from the static source digital image;

generating a second combined embedding by concatenating the second target facial expression animation embedding for the second resolution from the first frame of the target digital video, the second target pose animation embedding for the second resolution from the first frame of the target digital video, and the second source shape embedding for the second resolution from the static source digital image;

generating, utilizing a facial animation generative adversarial neural network comprising a first style block corresponding to the first resolution and a second style block corresponding to the second resolution, an animation that portrays the source face animated according to an animation selection element by conditioning the first style blocks of the facial animation generative adversarial neural network with the first combined embeddings and conditioning the second style block of the facial animation generative adversarial neural network with the second combined embedding; and

providing the animation for display via a user interface of a client device.

17. The non-transitory computer-readable medium of claim 16 , wherein generating the animation further comprises:

utilizing a comodulated generative adversarial neural network as the facial animation generative adversarial neural network to generate the animation by:

generating, utilizing a first modulation layer of the first style block of the comodulated generative adversarial neural network, an intermediate vector from the static source digital image by conditioning the first modulation layer according to the first combined embedding;

generating, utilizing a second modulation layer of the second style block of the comodulated generative adversarial neural network, an additional intermediate vector from the intermediate vector by conditioning the second modulation layer according to the second combined embedding; and

generating, utilizing the comodulated generative adversarial neural network, the animation from the intermediate vector, the additional intermediate vector, and the static source digital image.

18. The non-transitory computer-readable medium of claim 16 , the operations further comprising:

receiving the static source digital image comprising the source face and a plurality of additional source faces;

generating a recommendation to animate the source face in the static source digital image by transferring a different facial expression to replace the source face;

providing, for display via a user interface of a client device, the recommendation to animate the source face; and

in response to a selection of the recommendation, generating the animation that portrays the source face animated according to the animation selection element.

19. The non-transitory computer-readable medium of claim 18 ,

wherein extracting the first source shape embedding and the second source shape embedding from the static source digital image that portrays the source face comprises identifying the static source digital image based on user interaction with a source image selection element comprising an option to select a digital image from a camera roll of the client device.

20. The non-transitory computer-readable medium of claim 16 , the operations further comprising:

generating a third combined embedding by concatenating a third target facial expression animation embedding for the first resolution from a second frame, a third target pose animation embedding for the first resolution from the second frame, and the first source shape embedding for the first resolution from the static source digital image; and

generating, utilizing the facial animation generative adversarial neural network, the animation by conditioning the first style block of the facial animation generative adversarial neural network with the first combined embedding and the third combined embedding.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: SMITH, CAMERON
To: ADOBE INC.
Reel/Frame 063130/0465 →
Continuity (1)
Related Publication 20240331247A1 · Oct 3, 2024
References Cited (50)
US 11462040B2 · Lin et al. · 2022 [cited by applicant]
US 12026845B2 · Pardeshi · 2024 [cited by applicant]
US 20190114748A1 · Lin et al. · 2019 [cited by applicant]
US 20200234480A1 · Volkov · 2020 [cited by examiner]
US 20200394828A1 · Shukla · 2020 [cited by examiner]
US 20210056348A1 · Berlin et al. · 2021 [cited by applicant]
US 20210264207A1 · Smith et al. · 2021 [cited by applicant]
US 20220068037A1 · Pardeshi · 2022 [cited by applicant]
US 20220207262A1 · Jeong et al. · 2022 [cited by applicant]
US 20220392133A1 · Volkov · 2022 [cited by examiner]
US 20230037339A1 · Villegas · 2023 [cited by examiner]
US 20230110206A1 · Karras et al. · 2023 [cited by applicant]
US 20230123820A1 · Wang · 2023 [cited by examiner]
US 20230252687A1 · Zhang et al. · 2023 [cited by applicant]
US 20230319223A1 · Naruniec · 2023 [cited by examiner]
US 20230410447A1 · Cheng · 2023 [cited by examiner]
US 20240135511A1 · Singh et al. · 2024 [cited by applicant]
US 20240135513A1 · Singh et al. · 2024 [cited by applicant]
US 20240153047A1 · Smith et al. · 2024 [cited by applicant]
US 20240169624A1 · Brandt et al. · 2024 [cited by applicant]
US 20240169701A1 · Kulal · 2024 [cited by examiner]
US 20240171848A1 · Figueroa et al. · 2024 [cited by applicant]
US 20240249459A1 · Bradley · 2024 [cited by examiner]
US 20240331322A1 · Smith · 2024 [cited by examiner]
CN 113240613B · 2021 [cited by applicant]
CN 114862697A · 2022 [cited by applicant]
CN 114943656A · 2022 [cited by applicant]
GB 2606253A · 2022 [cited by applicant]
WO 2022083504A1 · 2022 [cited by applicant]
Qiao, Fengchun, et al. “Geometry-contrastive gan for facial expression transfer.” arXiv preprint arXiv:1802.01822 (2018). (Year: 2018). [cited by examiner]
Chen, Yajing, et al. “Self-supervised learning of detailed 3d face reconstruction.” IEEE Transactions on Image Processing 29 (2020): 8696-8705. (Year: 2020). [cited by examiner]
Screen captures from YouTube video clip entitled “How to Use xpression camera—For Video Chat, Vlogging, Live Streaming, Content Creation, Gaming,” 4 pages, uploaded on Feb. 8, 2023 by user “EmbodyMe”. Retrieved from Int… [cited by examiner]
Michail Christos Doukas, Stefanos Zafeiriou, Viktorija Sharmanska HeadGAN: One-Shot Neural Head Synthesis and Editing , Aug. 23, 2021. [cited by applicant]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer—High-Resolution Image Synthesis with Latent Diffusion Models arXiv:2112.10752 Wed, Apr. 13, 2022. [cited by applicant]
Badour AlBahar, Jingwan Lu, Jimei Yang, Zhixin Shu, Eli Shechtman, Jia-Bin Huang—Pose with Style: Detail-Preserving Pose-Guided Image Synthesis with Conditional StyleGAN Badour et al., Siggraph Asia 2021. [cited by applicant]
Artur Grigorev, Artem Sevastopolsky, Alexander Vakhiov, and Victor Lempitsky. Coordinate-based texture inpainting for pose-guided image generation. arXiv preprint arXiv:1811.11459, 2018. [cited by applicant]
Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of IEEE Conference on Computer Vision and Pattern Recogn… [cited by applicant]
Kripasindhu Sarkar, Vladislav Golyanik, Lingjie Liu, and Christian Theobalt. Style and pose control for image synthesis of humans from a single monocular view, 2021. [cited by applicant]
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. [cited by applicant]
Combined Search and Examination Report as received in GB2318853.5 dated Jun. 12, 2024. [cited by applicant]
Combined Search and Examination Report as received in GB2319660.3 dated Jun. 14, 2024. [cited by applicant]
Combined Search and Examination Report as received in GB2319084.6 dated Jun. 25, 2024. [cited by applicant]
Wiles, 0., Koepke, A and Zisserman, A, 2018. “X2face: A network for controlling face generation using images, audio, and pose codes.” in Proceedings of the European conference on computer vision (ECCV) (pp. 690-706). [cited by applicant]
U.S. Appl. No. 18/190,544, filed Nov. 20, 2024, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,556, filed Dec. 13, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,500, filed Feb. 26, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,500, filed Apr. 15, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,556, filed Mar. 12, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,673, filed Mar. 13, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,673, May 23, 2025, Office Action. [cited by applicant]