IP Library Granted Patent US 12,731,319
Granted Patent B2
US 12,731,319 · App. 18/348,856 · Granted Sep 8, 2026

Generating facial representations

Inventors: Gaurav Bharaj (Santa Monica, CA); Qingju Liu (London, GB)
Assignee: Flawless Holdings Limited
G06T13/40G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,319
App. No.
18/348,856
Granted
Sep 8, 2026
Kind
B2
Abstract

A method of generating facial representations using a machine learned model is provided. A text encoder generates a representation of an input text segment. An aligner determines a time alignment between the input text segment and an input audio signal. A decoder generates a facial representation based at least in part on the representation of the input text segment, the time alignment, and on target style data, the target style data representing a target audio style, wherein the facial representation comprises a sequence of facial expressions corresponding to the input text segment. A method of training such a machine learning model is also provided.

Claims (50)

1 . A computer-implemented method of generating a facial representation, the method comprising:

receiving, at a machine learned model, an input text segment, an input audio signal corresponding to the input text segment, and target style data, wherein the target style data represents a target audio style;

generating, by a text encoder of the machine learned model, a representation of the input text segment;

determining, by an aligner of the machine learned model, a time alignment between the input text segment and the input audio signal; and

generating, by a decoder of the machine learned model, the facial representation based at least in part on the representation of the input text segment, the time alignment, and the target style data, wherein the facial representation comprises a sequence of facial expressions corresponding to the input text segment;

wherein the machine learned model is trained by:

performing a first training operation comprising training the model based at least in part on:

generating, by a first configuration of the decoder, an output audio representation based on a first training text segment and a corresponding first training audio signal; and

updating the model so as to reduce a deviation between the output audio representation and an audio representation of the first training audio signal; and

performing a second training operation comprising training the model based at least in part on:

generating, by a second configuration of the decoder, an output facial representation based at least in part on a second training text segment and a corresponding second training audio signal; and

updating the model so as to reduce a deviation between the output facial representation and a training facial representation corresponding to the second training text segment and second training audio signal.

2 . The method of claim 1 , comprising:

receiving, at the machine learned model, a reference audio signal exhibiting the target audio style; and

generating, by a style encoder of the machine learned model, the target style data based at least in part on the reference audio signal.

3 . The method of claim 1 , comprising:

receiving input video data comprising target footage of a target human face; and

generating output video data based at least in part on the input video data and on the facial representation, wherein in the output video data the target human face exhibits the sequence of facial expressions corresponding to the input text segment.

4 . The method of claim 1 , wherein a weight associated with the text encoder, energy, or pitch is held fixed during the second training operation.

5 . The method of claim 1 , wherein during the first training operation the model is trained with a first duration of training audio signals, and during the second training operation the model is trained with a second duration of training audio signals, wherein the first duration is at least five times the second duration, or at least ten times the second duration.

6 . The method of claim 1 , wherein generating the time alignment comprises determining, from the input audio signal, a ground truth duration of phonemes represented in the input text segment, and wherein the facial representation is generated based at least in part on the ground truth duration of the phonemes.

7 . The method of claim 1 , wherein the target style data is an embedded representation of the target audio style, and wherein the method further comprises predicting, from the target style data, an energy parameter, a pitch parameter, and a residual parameter.

8 . The method of claim 1 , wherein the target style data specifies an energy parameter, a pitch parameter, and a residual parameter.

9 . The method of claim 7 , wherein the method further comprises generating, by a variance adaptor of the machine learned model, a combined signal representing a combination of the representation of the input text segment, the energy parameter, the pitch parameter, and the residual parameter; and

wherein the facial representation is generated by the decoder based at least in part on the combined signal.

10 . The method of claim 1 , wherein the facial representation comprises a blendshape.

11 . A computer-implemented method of training a machine learning model for generating a facial representation, the method comprising:

initializing the machine learning model, wherein the machine learning model comprises:

a text encoder configured to generate, from a received text segment, a representation of the received text segment;

an aligner configured to determine a time alignment between the received text segment and a received audio signal corresponding to the received text segment; and

a decoder configured to generate an output based at least in part on the representation of the received text segment, the time alignment, and received target style data representing an audio style;

performing a first training operation comprising training the machine learning model based at least in part on:

generating, by a first configuration of the decoder, an output audio representation based on a first training text segment and a corresponding first training audio signal; and

updating the machine learning model so as to reduce a deviation between the output audio representation and an audio representation of the first training audio signal; and

performing a second training operation comprising training the model based at least in part on:

generating, by a second configuration of the decoder, an output facial representation based at least in part on a second training text segment and a corresponding second training audio signal; and

updating the machine learning model so as to reduce a deviation between the output facial representation and a training facial representation corresponding to the second training text segment and second training audio signal; and

outputting, based at least in part on the first training operation and the second training operation, a machine learned model, wherein the machine learned model comprises the decoder in the second configuration.

12 . A system comprising one or more processors and one or more non-transient storage media storing machine readable instructions which, when executed by the one or more processors, cause the one or more processors to carry out a method comprising:

receiving, at a machine learned model, an input text segment, an input audio signal corresponding to the input text segment, and target style data, wherein the target style data represents a target audio style;

generating, by a text encoder of the machine learned model, a representation of the input text segment;

determining, by an aligner of the machine learned model, a time alignment between the input text segment and the input audio signal; and

generating, by a decoder of the machine learned model, the facial representation based at least in part on the representation of the input text segment, the time alignment, and the target style data, wherein the facial representation comprises a sequence of facial expressions corresponding to the input text segment;

wherein the machine learned model is trained by:

performing a first training operation comprising training the model based at least in part on:

generating, by a first configuration of the decoder, an output audio representation based on a first training text segment and a corresponding first training audio signal; and

updating the model so as to reduce a deviation between the output audio representation and an audio representation of the first training audio signal; and

performing a second training operation comprising training the model based at least in part on:

generating, by a second configuration of the decoder, an output facial representation based at least in part on a second training text segment and a corresponding second training audio signal; and

updating the model so as to reduce a deviation between the output facial representation and a training facial representation corresponding to the second training text segment and second training audio signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2023
From: BHARAJ, GAURAV; LIU, QINGJU
To: FLAWLESS HOLDINGS LIMITED
Reel/Frame 064822/0769 →
Continuity (1)
Related Publication 20250014253A1 · Jan 9, 2025
References Cited (43)
US 9548048B1 · Solh · 2017 [cited by examiner]
US 11049308B2 · del Val Santos · 2021 [cited by examiner]
US 11238885B2 · Mittal · 2022 [cited by examiner]
US 11398255B1 · Mann et al. · 2022 [cited by applicant]
US 11562521B2 · del Val Santos · 2023 [cited by examiner]
US 11562597B1 · Kim · 2023 [cited by applicant]
US 11847727B2 · del Val Santos · 2023 [cited by examiner]
US 12033259B2 · Kwatra · 2024 [cited by examiner]
US 12315057B2 · Beith · 2025 [cited by examiner]
US 12322018B2 · Ume · 2025 [cited by examiner]
US 12374014B2 · Starke · 2025 [cited by examiner]
US 12406419B1 · Villanueva Aylagas · 2025 [cited by examiner]
US 20200135226A1 · Mittal et al. · 2020 [cited by applicant]
US 20200302667A1 · del Val Santos · 2020 [cited by examiner]
US 20210319610A1 · del Val Santos · 2021 [cited by examiner]
US 20230123486A1 · del Val Santos · 2023 [cited by examiner]
US 20230154089A1 · Bradley · 2023 [cited by examiner]
US 20230154090A1 · Bradley · 2023 [cited by examiner]
US 20230335107A1 · Zhao · 2023 [cited by examiner]
US 20230394734A1 · Marquis Bolduc · 2023 [cited by examiner]
US 20240013462A1 · Seol · 2024 [cited by examiner]
US 20240037824A1 · Biswas · 2024 [cited by examiner]
US 20240038271A1 · de Juan · 2024 [cited by examiner]
US 20240212249A1 · Ume · 2024 [cited by examiner]
US 20240312095A1 · Wei · 2024 [cited by examiner]
WO 2023219752A1 · 2023 [cited by applicant]
Cong Gaoxiang et al.,: “Learning to Dub Movies via Hierarchical Prosody Models”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, pp. 14687-14697, XP034401043, Jun. 17, 2023. [cited by applicant]
International Search Report and Written Opinion dated Oct. 1, 2024 for International PCT Application No. PCT/GB2024/051777. [cited by applicant]
Daniel Cudeiro et al., “Capture, learning, and synthesis of 3d speaking styles”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10101-10111, 2019, May 8, 2019. [cited by applicant]
Yingruo Fan et al., “Faceformer: Speech-driven 3d facial animation with transformers”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18770-18780, Dec. 10, 2021. [cited by applicant]
Tero Karras et al., “Audio-driven facial animation by joint end-to-end learning of pose and emotion”, ACM Transactions on Graphics (TOG), 36(4):1-12, Jul. 30, 2017. [cited by applicant]
K R Prajwal et al., “A lip sync expert is all you need for speech to lip generation in the wild”, In Proceedings of the 28th ACM International Conference on Multimedia, pp. 484-492, 2020, Aug. 23, 2020. [cited by applicant]
Alexander Richard et al., “Meshtalk: 3d face animation from speech using cross-modality disentanglement”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1173-1182, Apr. 16, 2021. [cited by applicant]
Yi Ren et al., “Fastspeech 2: Fast and high-quality end-to-end text to speech”, arXiv preprint arXiv:2006.04558, Jun. 8, 2020. [cited by applicant]
Michael McAuliffe et al., “Montreal forced aligner: Trainable text-speech alignment using Kaldi”, In Interspeech, vol. 2017, pp. 498-502, Aug. 20, 2017. [cited by applicant]
Dongchan Min et al., “Meta-stylespeech: Multi-speaker adaptive text-to-speech generation”, In International Conference on Machine Learning, pp. 7748-7759. PMLR, Jun. 6, 2021. [cited by applicant]
Ralph Gross et al., “Multi-pie”, Dec. 1, 2013. [cited by applicant]
Yinghao Aaron Li et al., “Styletts: A style-based generative model for natural and diverse text-to-speech synthesis”, arXiv preprint arXiv:2205.15439, May 30, 2022. [cited by applicant]
Ashish Vaswani et al., “Attention is all you need”, Advances in neural information processing systems, 30, Jun. 12, 2017. [cited by applicant]
Kun Zhou et al., “Emotional Voice Conversion: Theory, Databases and ESD”, Speech Communication, 137:1-18, May 31, 2021. [cited by applicant]
Yuxuan Wang et al., “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis”, arXiv:1803.09017, Mar. 23, 2018. [cited by applicant]
Kyubyong Park et al., “g2pE: A Simple Python Module for English Grapheme To Phoneme Conversion”, https://github.com/Kyubyong/g2p, Dec. 31, 2019. [cited by applicant]
Daniel Jurafsky et al., “Speech and language processing” (3rd draft ed.), Oct. 16, 2019. [cited by applicant]