IP Library › Granted Patent US 12,548,229
Granted Patent B2
US 12,548,229 · App. 18/686,587 · Granted Feb 10, 2026

Robust facial animation from video and audio

Inventors: Iñaki Navarro (Zurich, CH); Dario Kneubuehler (Zurich, CH); Tijmen Verhulsdonck (Gothenburg, SE); Eloi Du Bois (Austin, TX); William Welch (San Francisco, CA); Charles Shang (San Mateo, CA); Ian Sachs (Corte Madera, CA); Kiran Bhat (San Francisco, CA)
Assignee: Roblox Corporation
G06T13/40G06T7/246G06T7/73G06T13/205G06V10/82G06V40/161G06V40/171G06T2207/20081G06T2207/20084G06T2207/30201G06V10/774G06V10/776
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,229
App. No.
18/686,587
Granted
Feb 10, 2026
Kind
B2
Abstract

Implementations described herein relate to methods, systems, and computer-readable media to generate animations for a 3D avatar from input video and audio captured at a client device. A camera may capture video of a face while a trained face detection model and a trained regression model output a set of video FACS weights, head poses, and facial landmarks to be translated into the animations of the 3D avatar. Additionally, a microphone may capture audio uttered by a user while a trained facial movement detection model and a trained regression model output a set of audio FACS weights. Additionally, a blending term is provided for identification of lapses in audio. A modularity mixing component fuses the video FACS weights and the audio FACS weights based on the blending term to create final FACS weights for animating the user's avatar, a character rig, or another animation-capable construct.

Claims (43)

1 . A computer-implemented method, comprising:

receiving input video frames;

receiving input audio frames and a blending term, wherein the input audio frames include audio associated with the input video frames;

obtaining video facial action coding system (FACS) weights from a first trained machine learning model based on the input video frames, wherein the first trained machine learning model comprises at least one encoder and at least three task-specific decoders, and wherein:

a first task-specific decoder of the at least three task-specific decoders is configured to output a predicted headpose,

a second task-specific decoder of the at least three task-specific decoders is configured to output a probability that a face is visible in an input video frame of the received input video frames, and

a third task-specific decoder of the at least three task-specific decoders is configured to output facial landmarks;

obtaining audio FACS weights from a second trained machine learning model based on the input audio frames;

combining the video FACS weights and the audio FACS weights to obtain final FACS weights, wherein the combining is based at least in part on the blending term; and

outputting the final FACS weights to drive facial animation of a three-dimensional (3D) model.

2 . The computer-implemented method of claim 1 , wherein the first trained machine learning model and the second trained machine learning model are trained in a two-stage semi-supervised training process.

3 . The computer-implemented method of claim 1 , wherein at least one of the at least three task-specific decoders comprises causal convolution layers applied over a time dimension.

4 . The computer-implemented method of claim 1 , wherein the second trained machine learning model comprises at least one encoder and at least two task-specific decoders.

5 . The computer-implemented method of claim 4 , wherein a first task-specific decoder of the at least two task-specific decoders of the second trained machine learning model is configured to output the audio FACS weights and a second task-specific decoder of the at least two task-specific decoders of the second trained machine learning model is configured to output the blending term.

6 . A system, comprising:

a memory with instructions stored thereon; and

a processing device, coupled to the memory, wherein the processing device is configured to access the memory and execute the instructions, and wherein the instructions cause the processing device to perform operations comprising:

receiving input video frames;

receiving input audio frames and a blending term, wherein the input audio frames include audio associated with the input video frames;

obtaining video facial action coding system (FACS) weights from a first trained machine learning model based on the input video frames, wherein the first trained machine learning model comprises at least one encoder and at least three task-specific decoders, and wherein:

a first task-specific decoder of the at least three task-specific decoders is configured to output a predicted headpose,

a second task-specific decoder of the at least three task-specific decoders is configured to output a probability that a face is visible in an input video frame of the received input video frames, and

a third task-specific decoder of the at least three task-specific decoders is configured to output facial landmarks;

obtaining audio FACS weights from a second trained machine learning model based on the input audio frames;

combining the video FACS weights and the audio FACS weights to obtain final FACS weights, wherein the combining is based at least in part on the blending term; and

outputting the final FACS weights to drive facial animation of a three-dimensional (3D) model.

7 . The system of claim 6 , wherein the first trained machine learning model and the second trained machine learning model are trained in a two-stage semi-supervised training process.

8 . The system of claim 6 , wherein at least one of the at least three task-specific decoders comprises causal convolution layers applied over a time dimension.

9 . The system of claim 6 , wherein the second trained machine learning model comprises at least one encoder and at least two task-specific decoders.

10 . The system of claim 9 , wherein a first task-specific decoder of the at least two task-specific decoders of the second trained machine learning model is configured to output the audio FACS weights and a second task-specific decoder of the at least two task-specific decoders of the second trained machine learning model is configured to output the blending term.

11 . A non-transitory computer-readable medium with instructions stored thereon that, responsive to execution by a processing device, causes the processing device to perform operations comprising:

receiving input video frames;

receiving input audio frames and a blending term, wherein the input audio frames include audio associated with the input video frames;

obtaining video facial action coding system (FACS) weights from a first trained machine learning model based on the input video frames, wherein the first trained machine learning model comprises at least one encoder and at least three task-specific decoders, and wherein:

a first task-specific decoder of the at least three task-specific decoders is configured to output a predicted headpose,

a second task-specific decoder of the at least three task-specific decoders is configured to output a probability that a face is visible in an input video frame of the received input video frames, and

a third task-specific decoder of the at least three task-specific decoders is configured to output facial landmarks;

obtaining audio FACS weights from a second trained machine learning model based on the input audio frames;

combining the video FACS weights and the audio FACS weights to obtain final FACS weights, wherein the combining is based at least in part on the blending term; and

outputting the final FACS weights to drive facial animation of a three-dimensional (3D) model.

12 . The non-transitory computer-readable medium of claim 11 , wherein the first trained machine learning model and the second trained machine learning model are trained in a two-stage semi-supervised training process.

13 . The non-transitory computer-readable medium of claim 11 , wherein at least one of the at least three task-specific decoders comprises causal convolution layers applied over a time dimension.

14 . The non-transitory computer-readable medium of claim 11 , wherein the second trained machine learning model comprises at least one encoder and at least two task-specific decoders, and wherein a first task-specific decoder of the at least two task-specific decoders of the second trained machine learning model is configured to output the audio FACS weights and a second task-specific decoder of the at least two task-specific decoders of the second trained machine learning model is configured to output the blending term.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 29, 2024
From: NAVARRO, IÑAKI; KNEUBUEHLER, DARIO; VERHULSDONCK, TIJMEN; DU BOIS, ELOI; WELCH, WILLIAM; SHANG, CHARLES; SACHS, IAN; BHAT, KIRAN
To: ROBLOX CORPORATION
Reel/Frame 066601/0636 →
Continuity (2)
Provisional Application 63440993 · Jan 25, 2023
Related Publication 20240428492A1 · Dec 26, 2024
References Cited (52)
US 9786084B1 · Bhat et al. · 2017 [cited by applicant]
US 10062198B2 · Bhat et al. · 2018 [cited by applicant]
US 10169905B2 · Bhat et al. · 2019 [cited by applicant]
US 10198845B1 · Bhat et al. · 2019 [cited by applicant]
US 10521946B1 · Roche et al. · 2019 [cited by applicant]
US 10559111B2 · Sachs et al. · 2020 [cited by applicant]
US 20130044183A1 · Jeon et al. · 2013 [cited by applicant]
US 20170243387A1 · Li et al. · 2017 [cited by applicant]
US 20190356894A1 · Oh · 2019 [cited by examiner]
US 20200242383A1 · el Kaliouby · 2020 [cited by examiner]
US 20210004976A1 · Guizilini et al. · 2021 [cited by applicant]
US 20210027511A1 · Shang et al. · 2021 [cited by applicant]
Anonymous, Multimodal input and deep learning for robust real-time facial animation and avatar lip sync, ACM Trans. Graph., vol. 1, No. 1,Jan. 2022, 11 pages. [cited by applicant]
Anonymous, “Audiovisual inputs for robust, real-time facial animation with lip sync”, ACM Trans. Graph., vol. 1, No. 1, Jan. 2023, 13 pages. [cited by applicant]
“Preliminary Report on Patentability in International Application No. PCT/US2024/012900”, Aug. 7, 2025, 9 Pages. [cited by applicant]
Aaron Van Den Oord, et al., “Wavenet: A generative model for raw audio”, 2016. arXiv preprint arXiv:1609.03499 (2016)., Sep. 19, 2016. [cited by applicant]
Abhinav Kumar, et al., “Luvli face alignment: Estimating landmarks' location, uncertainty, and visibility likelihood”, 2020. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18 pages… [cited by applicant]
Adrian Bulat, et al., “How far are we from solving the 2D & 3D Face Alignment problem? (and a dataset of 230,000 3D facial landmarks)”, In 2017 IEEE International Conference on Computer Vision (ICCV). IEEE Computer Soci… [cited by applicant]
Adrian Bulat, et al., “Subpixel Heatmap Regression for Facial Landmark Localization”, arXiv preprint arXiv:2111.02360 (2021). [cited by applicant]
Ahmed Hussen Abdelaziz, et al., “Modality Dropout for Improved Performance-driven Talking Faces”, In Proceedings of the 2020 International Conference on Multimodal Interaction (Virtual Event, Netherlands) (ICMI '20). As… [cited by applicant]
Alex Graves, et al., “Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks”, 2006, In Proceedings of the 23rd international conference on Machine learning. 369-376. [cited by applicant]
Alexander Richard, et al., “MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement.”, 2021. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV). 1153-1162. https://doi.org/10.1109/I… [cited by applicant]
Andrew L Maas, et al., “Rectifier nonlinearities improve neural network acoustic models”, 2013. In Proc. icml, vol. 30. Atlanta, Georgia, USA, 6 pages. [cited by applicant]
Awni Hannun, et al., “Deep speech: Scaling up end-to-end speech recognition”, arXiv preprint arXiv:1412.5567 (2014)., Dec. 19, 2014. [cited by applicant]
Daniel Cudeiro, et al., “Capture, Learning, and Synthesis of 3D Speaking Styles”, In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR). 10101-10111. http://voca.is.tue.mpg.de/, 2019. [cited by applicant]
David I Perrett, et al., “Symmetry and Human Facial Attractiveness”, 1999. Evolution and Human Behavior 20, 5 (1999), 295-307. https://doi.org/10.1016/S1090-5138(99)00014-8. [cited by applicant]
Erroll Wood, “3D Face Reconstruction with Dense Landmarks”, In Computer Vision—ECCV 2022, Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner (Eds.). Springer Nature Switzerland, Cha… [cited by applicant]
Ivan Grishchenko, “Attention Mesh: High-fidelity Face Mesh Prediction in Real-time”, 2020. Attention Mesh: High-fidelity Face Mesh Prediction in Real-time. arXiv:2006.10962 [cs.CV], Jun. 19, 2020. [cited by applicant]
Jiyuan Zhang, et al., “High Performance Zero-Memory Overhead Direct Convolutions”, CoRR abs/1809.10170 (2018). arXiv:1809.10170 http://arxiv.org/abs/1809.10170, Sep. 20, 2018. [cited by applicant]
Jort F Gemmeke, et al., “Audio Set: An Ontology and Human-Labeled Dataset for Audio Events”, 2017. Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE international conference on acoustics, s… [cited by applicant]
Joseph Roth, et al., “Ava active speaker: An audio-visual dataset for active speaker detection”, 2020. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, May 25,… [cited by applicant]
Kaipeng Zhang, et al., “Joint Face Detection and Alignment using Multi-task Cascaded Convolutional Networks”, CoRR abs/1604.02878 (2016). arXiv:1604.02878 http://arxiv.org/abs/1604.02878. [cited by applicant]
Lisha Chen, et al., “Face Alignment with Kernel Density Deep Neural Network”, In 2019 IEEE/CVF International Conference on Computer Vision (ICCV). 6991-7001. https://doi.org/10.1109/ICCV.2019.00709. [cited by applicant]
Mark Sandler, et al., “Inverted Residuals and Linear Bottlenecks:”, Mobile Networks for Classification, Detection and Segmentation, CoRR abs/1801.04381 (2018). arXiv:1801.04381 http://arxiv.org/abs/1801.04381, Mar. 21, … [cited by applicant]
Michael Mcauliffe, et al., “Montreal Forced Aligner: trainable text-speech alignment using Kaldi”, 2017. In Interspeech, vol. 2017. 498-502. [cited by applicant]
Ross Girshick, “Fast R-CNN”, In Proceedings of the IEEE international conference on computer vision. 1440-1448., Sep. 27, 2015. [cited by applicant]
Samuli Laine, et al., “Production-Level Facial Performance Capture Using Deep Convolutional Neural Networks”, In Proceedings of the ACM SIGGRAPH /Eurographics Symposium on Computer Animation (Los Angeles, California) (S… [cited by applicant]
Sergey Ioffe, et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”, In International conference on machine learning. PMLR,, Mar. 2, 2015, 11 pages. [cited by applicant]
Sina Honari, et al., “Improving Landmark Localization with Semi-Supervised Learning”, 2018. arXiv:1709.01591 [cs.CV], Oct. 28, 2018. [cited by applicant]
Supasorn Suwajanakorn, et al., “Synthesizing Obama: Learning Lip Sync from Audio”, 2017. ACM Trans. Graph. 36, 4, Article 95 (Jul. 2017), 13 pages. https://doi.org/10.1145/3072959.3073640. [cited by applicant]
Tero Karras, et al., “Audio-Driven Facial Animation by Joint End-to-End Learning of Pose and Emotion”, ACM Trans. Graph. 36, 4, Article 94 (Jul. 2017), 12 pages. https://doi.org/10.1145/3072959. 3073658. [cited by applicant]
Tianye Li, et al., “Learning a model of facial shape and expression from 4D scans”, ACM Trans. Graph. 36, 6, Article 194 (Nov. 2017), 17 pages. https://doi.org/10.1145/3130800.3130813. [cited by applicant]
Vassil Panayotov, et al., “Librispeech: An ASR Corpus Based on Public Domain Audio Books”, 2015. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 5206-5210. [cited by applicant]
Xiaojie Guo, et al., “PFLD: A Practical Facial Landmark Detector”, CoRR abs/1902.10859 (2019). arXiv:1902.10859 http://arxiv.org/abs/1902.10859, Mar. 3, 2019. [cited by applicant]
Xin Chen, et al., “Joint Audio-Video Driven Facial Animation”, In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 3046-3050. https://doi.org/10.1109/ICASSP.2018.8461502. [cited by applicant]
Yang Zhou, et al., “VisemeNet: Audio-Driven Animator-Centric Speech Animation”, ACM Transactions on Graphics (TOG) 37, 4 (2018), 1-10., May 24, 2018. [cited by applicant]
Yao Feng, et al., “Joint 3D Face Reconstruction and Dense Alignment with Position Map Regression Network”, In Proceedings of the European conference on computer vision (ECCV). 534-551. arXiv:1803.07835, Mar. 21, 2018. [cited by applicant]
Yao Feng, et al., “Learning an Animatable Detailed 3D Face Model From In-The-Wild Images”, ACM Trans. Graph. 40, 4, Article 88 (Jul. 2021), 13 pages. https://doi.org/10.1145/3450626.3459936, Jun. 2, 2021. [cited by applicant]
Yufei Xu, et al., “ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation”, In Advances in Neural Information Processing Systems., Oct. 13, 2022. [cited by applicant]
Richard, “Audio- and Gaze-driven Facial Animation of Codec Avatars”, Retrieved from Internet: https://arxiv.org/abs/2008.05023, 2020, 10 pages. [cited by applicant]
USPTO, International Search Report for International Patent Application No. PCT/US2024/012900, Apr. 25, 2024, 2 pages. [cited by applicant]
USPTO, Written Opinion for International Patent Application No. PCT/US2024/012900, Apr. 25, 2024, 7 pages. [cited by applicant]