IP Library Granted Patent US 12,406,419
Granted Patent B1
US 12,406,419 · App. 18/128,997 · Granted Sep 2, 2025

Generating facial animation data from speech audio

Inventors: Monica Villanueva Aylagas (Sundbyberg, SE); Mattias Teye (Sundbyberg, SE); Hector Leon (Malmö, SE)
Assignee: ELECTRONIC ARTS INC.
G06T13/205G06T13/40G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,419
App. No.
18/128,997
Granted
Sep 2, 2025
Kind
B1
Abstract

This specification provides a system comprising: or more computing devices; and one or more storage devices communicatively coupled to the one or more computing devices. The one or more storage devices store instructions that, when executed by the one or more computing devices, cause the one or more computing devices to perform operations comprising: receiving input data derived from speech audio; generating facial animation data, comprising processing the input data and a conditioning input using a machine-learned generative model; generating further animation data, comprising processing the input data using a further machine-learned generative model; and generating animation data for at least a face in a video game using the facial animation data and the further animation data, wherein the animation data animates at least the face in the video game in accordance with speech sounds of the speech audio.

Claims (47)

1. A computer-implemented method, the method comprising:

obtaining a trained generative machine-learning model, the trained generative machine-learning model configured to process (i) input data derived from speech audio and (ii) a conditioning input representing a particular facial expression to generate facial animation data corresponding to the speech audio and the particular facial expression;

obtaining input data derived from speech audio for processing by the trained generative machine-learning model;

determining a conditioning input representing a particular facial expression from a set of reference speech animation examples, each reference speech animation example comprising data derived from speech audio and corresponding ground-truth facial animation data having the particular facial expression, wherein determining the conditioning input comprises:

initializing the conditioning input;

processing, using the trained generative machine-learning model: (i) the conditioning input, and (ii) the data derived from speech audio of one or more reference speech animation examples from the set of reference speech animation examples;

generating, as output of the trained generative machine learning model, predicted facial animation data for each reference speech animation example;

determining a loss for each reference speech animation example, wherein the loss for a reference speech animation example is dependent on the predicted facial animation data and the ground truth facial animation data of the reference speech animation example; and

updating the conditioning input based on the losses of the speech animation examples whilst the weights of the trained generative machine-learning model are held frozen;

processing, by the trained generative machine-learning model, (i) the input data derived from speech audio for processing and (ii) the determined conditioning input representing a particular facial expression from the set of reference speech animation examples to generate facial animation data corresponding to the speech audio and the particular facial expression.

2. The method of claim 1 , wherein the facial animation data animates a face in a video game.

3. The method of claim 1 , wherein updating the conditioning input comprises:

updating the conditioning input using a gradient-based optimization procedure and the losses of the reference speech animation examples.

4. The method of claim 1 , wherein determining the loss for a reference speech animation example comprises performing a comparison between the predicted facial animation data and the ground-truth facial animation data of the reference speech animation example.

5. The method of claim 1 , wherein the one or more reference speech animation examples comprises a plurality of speech animation examples generated using speech audio associated with a plurality of speakers.

6. A system comprising:

one or more computing devices; and

one or more storage devices communicatively coupled to the one or more computing devices, wherein the one or more storage devices store instructions that, when executed by the one or more computing devices, cause the one or more computing devices to perform operations comprising:

obtaining a trained generative machine-learning model, the trained generative machine-learning model configured to process (i) input data derived from speech audio and (ii) a conditioning input representing a particular facial expression to generate facial animation data corresponding to the speech audio and the particular facial expression;

obtaining input data derived from speech audio for processing by the trained generative machine-learning model;

determining a conditioning input representing a particular facial expression from a set of reference speech animation examples, each reference speech animation example comprising data derived from speech audio and corresponding ground-truth facial animation data having the particular facial expression, wherein determining the conditioning input comprises:

initializing the conditioning input;

processing, using the trained generative machine-learning model: (i) the conditioning input, and (ii) the data derived from speech audio of one or more reference speech animation examples from the set of reference speech animation examples;

generating, as output of the trained generative machine-learning model, predicted facial animation data for each reference speech animation example;

determining a loss for each reference speech animation example, wherein the loss for a reference speech animation example is dependent on the predicted facial animation data and the ground-truth facial animation data of the reference speech animation example; and

updating the conditioning input based on the losses of the speech animation examples whilst the weights of the trained generative machine-learning model are held frozen;

processing, by the trained generative machine-learning model, (i) the input data derived from speech audio for processing and (ii) the determined conditioning input representing a particular facial expression from the set of reference speech animation examples to generate facial animation data corresponding to the speech audio and the particular facial expression.

7. The system of claim 6 , wherein the facial animation data animates a face in a video game.

8. The system of claim 6 , wherein updating the conditioning input comprises:

updating the conditioning input using a gradient-based optimization procedure and the losses of the reference speech animation examples.

9. The system of claim 6 , wherein determining the loss for a reference speech animation example comprises performing a comparison between the predicted facial animation data and the ground-truth facial animation data of the reference speech animation example.

10. The system of claim 6 , wherein the one or more reference speech animation examples comprises a plurality of speech animation examples generated using speech audio associated with a plurality of speakers.

11. One or more non-transitory computer storage media storing instructions that when executed by one or more computing devices cause the one or more computing devices to perform operations comprising:

obtaining a trained generative machine learning model, the trained generative machine learning model configured to process (i) input data derived from speech audio and (ii) a conditioning input representing a particular facial expression to generate facial animation data corresponding to the speech audio and the particular facial expression;

obtaining input data derived from speech audio for processing by the trained generative machine learning model;

determining a conditioning input representing a particular facial expression from a set of reference speech animation examples, each reference speech animation example comprising data derived from speech audio and corresponding ground-truth facial animation data having the particular facial expression, wherein determining the conditioning input comprises:

initializing the conditioning input;

processing, using the trained generative machine-learning model: (i) the conditioning input, and (ii) the data derived from speech audio of one or more reference speech animation examples from the set of reference speech animation examples;

generating, as output of the trained generative machine-learning model, predicted facial animation data for each reference speech animation example;

determining a loss for each reference speech animation example, wherein the loss for a reference speech animation example is dependent on the predicted facial animation data and the ground-truth facial animation data of the reference speech animation example; and

updating the conditioning input based on the losses of the speech animation examples whilst the weights of the trained generative machine learning model are held frozen;

processing, by the trained generative machine learning model, (i) the input data derived from speech audio for processing and (ii) the determined conditioning input representing a particular facial expression from the set of reference speech animation examples to generate facial animation data corresponding to the speech audio and the particular facial expression.

12. The non-transitory computer storage media of claim 11 , wherein the facial animation data animates a face in a video game.

13. The non-transitory computer storage media of claim 11 , wherein updating the conditioning input comprises:

updating the conditioning input using a gradient-based optimization procedure and the losses of the reference speech animation examples.

14. The non-transitory computer storage media of claim 11 , wherein determining the loss for a reference speech animation example comprises performing a comparison between the predicted facial animation data and the ground-truth facial animation data of the reference speech animation example.

15. The non-transitory computer storage media of claim 11 , wherein the one or more reference speech animation examples comprises a plurality of speech animation examples generated using speech audio associated with a plurality of speakers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2023
From: VILLANUEVA AYLAGAS, MONICA; TEYE, MATTIAS; LEON, HECTOR
To: ELECTRONIC ARTS INC.
Reel/Frame 063504/0606 →
Continuity (2)
Provisional Application 63327647 · Apr 5, 2022
Provisional Application 63327633 · Apr 5, 2022
References Cited (103)
US 20180336464A1 · Karras · 2018 [cited by examiner]
US 20200234690A1 · Savchenkov · 2020 [cited by examiner]
US 20200302667A1 · del Val Santos · 2020 [cited by examiner]
US 20210027511A1 · Shang · 2021 [cited by examiner]
US 20210375260A1 · Yu · 2021 [cited by examiner]
US 20220020196A1 · Kuta · 2022 [cited by examiner]
US 20220068001A1 · Kaushik · 2022 [cited by examiner]
Abdal, Rameen, Yipeng Qin, and Peter Wonka, “Image2stylegan++: How to edit the embedded images?,” In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 8296-8305, 2020. [cited by applicant]
Abrevaya, Victoria Fernández, Adnane Boukhayma, Philip HS Torr, and Edmond Boyer, “Cross-modal deep face normals with deactivable skip connections,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Patte… [cited by applicant]
Baevski, Alexei, Henry Zhou, Abdelrahman Mohamed, and Michael Auli, “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” arXiv preprint arXiv:2006.11477 Oct. 22, 2020. [cited by applicant]
Bailey, Stephen W., Dalton Omens, Paul Dilorenzo, and James F. O'brien, “Fast and deep facial deformations,” ACM Transactions on Graphics (TOG) 39, No. 4: 94-1, Jul. 2020. [cited by applicant]
Bakker, Iris, Theo Van Der Voordt, Peter Vink, and Jan De Boon, “Pleasure, arousal, dominance: Mehrabian and Russell revisited,” Current Psychology 33: 405-421, Jun. 11, 2014. [cited by applicant]
Benzeghiba, Mohamed, Renato De Mori, Olivier Deroo, Stephane Dupont, Teodora Erbes, Denis Jouvet, Luciano Fissore et al. “Automatic speech recognition and speech variability: A review,” Speech communication 49, No. 10-1… [cited by applicant]
Bhat, Chitralekha, and Sunil Kopparapu, “Viseme comparison based on phonetic cues for varying speech accents,” In Sixteenth Annual Conference of the International Speech Communication Association, Sep. 6, 2015. [cited by applicant]
Bishop, Chris M., “Training with noise is equivalent to Tikhonov regularization,” Neural computation 7, No. 1: 108-116, 1995. [cited by applicant]
Botha, Johnny, and Heloise Pieterse, “Fake news and deepfakes: A dangerous threat for 21st century information security,” In ICCWS 2020 15th International Conference on Cyber Warfare and Security, Academic Conferences a… [cited by applicant]
Brand, Matthew, “Voice puppetry,” In Proceedings of the 26th annual conference on Computer graphics and interactive techniques, pp. 21-28, 1999. [cited by applicant]
Bregler, Christoph, Michele Covell, and Malcolm Slaney, “Video rewrite: Driving visual speech with audio,” In Proceedings of the 24th annual conference on Computer graphics and interactive techniques, pp. 353-360, 1997. [cited by applicant]
Brown, Tom B., Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan et al. “Language Models are Few-Shot Learners,” arXiv preprint arXiv:2005.14165, Jul. 22, 2020. [cited by applicant]
Burgess, Christopher P., Irina Higgins, Arka Pal, Loic Matthey, Nick Watters, Guillaume Desjardins, and Alexander Lerchner, “Understanding disentangling in $\beta $-VAE,” arXiv preprint arXiv:1804.03599, Apr. 10, 2018. [cited by applicant]
Busso, Carlos, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” Language… [cited by applicant]
Cao, Houwei, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova, and Ragini Verma, “Crema-d: Crowd-sourced emotional multimodal actors dataset,” IEEE transactions on affective computing 5, No. 4: 377-390, 2… [cited by applicant]
Chai, Yujin, Yanlin Weng, Lvdi Wang, and Kun Zhou, “Speech-driven facial animation with spectral gathering and temporal attention,” Frontiers of Computer Science 16: 1-10, Sep. 23, 2020. [cited by applicant]
Chen, Mingyi, Xuanji He, Jing Yang, and Han Zhang, “3-D convolutional recurrent neural networks with attention model for speech emotion recognition,” IEEE Signal Processing Letters 25, No. 10: 1440-1444, Oct. 10, 2018. [cited by applicant]
Chung, Joon Son, Amir Jamaludin, and Andrew Zisserman, “You said that?,” arXiv preprint arXiv:1705.02966, Jul. 18, 2017. [cited by applicant]
Chung, Joon Son, and Andrew Zisserman, “Lip reading in the wild,” In Computer Vision—ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, Nov. 20-24, 2016, Revised Selected Papers, Part II 13, pp. 87-103… [cited by applicant]
Cootes, Timothy F., Gareth J. Edwards, and Christopher J. Taylor, “Active appearance models,” IEEE Transactions on pattern analysis and machine intelligence 23, No. 6: 681-685, 2001. [cited by applicant]
Cudeiro, Daniel, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J. Black, “Capture, learning, and synthesis of 3D speaking styles,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec… [cited by applicant]
Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, Oct. 11, 2018. [cited by applicant]
Diederik P. Kingma, , And Jimmy Ba, “ADAM: A Method for Stochastic Optimization,” In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, Conference Track Proceedings, May 7, 2015. [cited by applicant]
Doersch, Carl, “Tutorial on variational autoencoders,” arXiv preprint arXiv:1606.05908, Aug. 13, 2016. [cited by applicant]
Dossou, Bonaventure FP, and Yeno KS Gbenou, “FSER: Deep Convolutional Neural Networks for Speech Emotion Recognition,” arXiv preprint arXiv:2109.07916, 2021. [cited by applicant]
Ekman, Paul, “Facial expressions of emotion: an old controversy and new findings,” Philosophical Transactions of the Royal Society of London. Series B: Biological Sciences 335, No. 1273: 63-69, 1992. [cited by applicant]
Engwall, Olov, and Jonas Beskow, “Resynthesis of 3D tongue movements from facial data,” In Eighth European Conference on Speech Communication and Technology, 2003. [cited by applicant]
Ezzat, Tony, Gadi Geiger, and Tomaso Poggio, “Trainable videorealistic speech animation,” ACM Transactions on Graphics (TOG) 21, No. 3: 388-398, Jun. 6, 2002. [cited by applicant]
Fabre, Diandra, Thomas Hueber, Laurent Girin, Xavier Alameda-Pineda, and Pierre Badin, “Automatic animation of an articulatory tongue model from ultrasound images of the vocal tract,” Speech Communication 93: 63-75, Sep… [cited by applicant]
FFX, FaceFX, URL: https://facefx.com/, Retrieved on: Sep. 3, 2021. [cited by applicant]
Gidaris, Spyros, Praveer Singh, and Nikos Komodakis, “Unsupervised representation learning by predicting image rotations,” arXiv preprint arXiv:1803.07728, Mar. 21, 2018. [cited by applicant]
Hahn, Fabian, Bernhard Thomaszewski, Stelian Coros, Robert W. Sumner, and Markus Gross, “Efficient simulation of secondary motion in rig-space,” In Proceedings of the 12th ACM SIGGRAPH/eurographics symposium on computer… [cited by applicant]
Hahn, Fabian, Sebastian Martin, Bernhard Thomaszewski, Robert Sumner, Stelian Coros, and Markus Gross, “Rig-space physics.” ACM transactions on graphics (TOG) 31, No. 4: 1-8, 2012. [cited by applicant]
Han, Kun, Dong Yu, and Ivan Tashev, “Speech emotion recognition using deep neural network and extreme learning machine,” In Interspeech, 2014. [cited by applicant]
Hannun, Awni, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger et al., “Deep speech: Scaling up end-to-end speech recognition,” arXiv preprint arXiv:1412.5567, Dec. 19, 2014. [cited by applicant]
Higgins, Irina, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner, “beta-vae: Learning basic visual concepts with a constrained variational framework,”… [cited by applicant]
Holden, Daniel, Jun Saito, and Taku Komura, “Learning inverse rig mappings by nonlinear regression,” IEEE transactions on visualization and computer graphics 23, No. 3: 1167-1178, 2016. [cited by applicant]
Huang, Xuedong, and Kai-Fu Lee, “On speaker-independent, speaker-dependent, and speaker-adaptive speech recognition,” IEEE Transactions on Speech and Audio processing 1, No. 2: 150-157, Apr. 1993. [cited by applicant]
James, Jesin, Li Tian, and Catherine Watson, “An open source emotional speech corpus for human robot interaction applications,” Interspeech, Sep. 6, 2018. [cited by applicant]
Jonell, Patrik, Taras Kucherenko, Gustav Eje Henter, and Jonas Beskow, “Let's face it: Probabilistic multi-modal interlocutor-aware generation of facial gestures in dyadic settings,” In Proceedings of the 20th ACM Inter… [cited by applicant]
Karras, Tero, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila, “Alias-Free Generative Adversarial Networks,” arXiv preprint arXiv:2106.12423, Mar. 13, 2021. [cited by applicant]
Karras, Tero, Samuli Laine, and Timo Aila, “A style-based generator architecture for generative adversarial networks,” In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4401-4410,… [cited by applicant]
Karras, Tero, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila, “Analyzing and Improving the Image Quality of StyleGAN,” In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition … [cited by applicant]
Karras, Tero, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen, “Audio-driven facial animation by joint end-to-end learning of pose and emotion,” ACM Transactions on Graphics (TOG) 36, No. 4: 1-12, 2017. [cited by applicant]
Kingma, D. P., and M. Welling, “Auto-encoding variational Bayes. 2nd international conference on learning representations (ICLR2014),” Preprint, submitted Dec. 23, 2014: arXiv: http://arxiv.org/abs/1312.6114v10, May 1, … [cited by applicant]
Kingma, Diederik P., and Max Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114, Dec. 20, 2013. [cited by applicant]
Korn, Oliver, Lukas Stamm, and Gerd Moeckl, “Designing authentic emotions for non-human characters: A study evaluating virtual affective behavior,” In Proceedings of the 2017 conference on designing interactive systems,… [cited by applicant]
Lewis, John P., and Ken-Ichi Anjyo, “Direct manipulation blendshapes,” IEEE Computer Graphics and Applications 30, No. 4: 42-50, Jul. 2010. [cited by applicant]
Lewis, John P., Ken Anjyo, Taehyun Rhee, Mengjie Zhang, Frederic H. Pighin, and Zhigang Deng, “Practice and theory of blendshape facial models,” Eurographics (State of the Art Reports) 1, No. 8: 2, 2014. [cited by applicant]
Li, Zhi, Anne Aaron, Ioannis Katsavounidis, Anush Moorthy, and Megha Manohara, “Toward a practical perceptual video quality metric,” The Netflix Tech Blog 6, No. 2: 2, Jun. 6, 2016. [cited by applicant]
Lin, Ji, Richard Zhang, Frieder Ganz, Song Han, and Jun-Yan Zhu, “Anycost gans for interactive image synthesis and editing,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1498… [cited by applicant]
Livingstone, Steven R., and Frank A. Russo, “The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,” PloS one 13, N… [cited by applicant]
Luo, Changwei, Jun Yu, Xian Li, and Leilei Zhang, “HMM based speech-driven 3D tongue animation,” In 2017 IEEE International Conference On Image Processing (ICIP), pp. 4377-4381, IEEE, Sep. 2017. [cited by applicant]
News, Guinness World Records. Star Wars: The Old Republic Recognised Guinness World Records 2012 Gamer's Edition, URL: https://www.guinnessworldrecords.com, Retrieved on: Dec. 17, 2021. [cited by applicant]
Nwe, Tin Lay, Say Wei Foo, and Liyanage C. De Silva, “Speech emotion recognition using hidden Markov models,” Speech communication 41, No. 4: 603-623, Nov. 2003. [cited by applicant]
Orvalho, Verónica, Pedro Bastos, Frederic I. Parke, Bruno Oliveira, and Xenxo Alvarez, “A Facial Rigging Survey,” Eurographics (State of the Art Reports): 183-204, 2012. [cited by applicant]
Paszke, Adam, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen et al, “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processi… [cited by applicant]
Patamia, Rutherford Agbeshi, Wu Jin, Kingsley Nketia Acheampong, Kwabena Sarpong, and Edwin Kwadwo Tenagyei, “Transformer based multimodal speech emotion recognition with improved neural networks,” In 2021 IEEE 2nd Inte… [cited by applicant]
Pelachaud, Catherine, Cornelius Wam Van Overveld, and Chin Seah, “Modeling and animating the human tongue during speech production,” In Proceedings of Computer Animation'94, pp. 40-49. IEEE, May 1994. [cited by applicant]
Pham, Hai Xuan, Yuting Wang, and Vladimir Pavlovic, “End-to-end learning for 3d facial animation from speech,” In Proceedings of the 20th ACM International Conference on Multimodal Interaction, pp. 361-365, 2018. [cited by applicant]
Richard, Alexander, Colin Lea, Shugao Ma, Jurgen Gall, Fernando De La Torre, and Yaser Sheikh, “Audio-and gaze-driven facial animation of codec avatars,” In Proceedings of the IEEE/CVF winter conference on applications … [cited by applicant]
Richard, Alexander, Michael Zollhoefer, Yandong Wen, Fernando De La Torre, and Yaser Sheikh, “MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement,” arXiv preprint arXiv:2104.08223, Apr. 16, 2021. [cited by applicant]
Schneider, Steffen, Alexei Baevski, Ronan Collobert, and Michael Auli, “wav2vec: Unsupervised pre-training for speech recognition,” arXiv preprint arXiv:1904.05862, Sep. 11, 2019. [cited by applicant]
Schwartz, Roy, Jesse Dodge, Noah A. Smith, and Oren Etzioni, “Green AI,” arXiv preprint arXiv:1907.10597, Aug. 13, 2019. [cited by applicant]
SG, Speech Graphics, URL: https://www.speech-graphics.com/, Retrieved on: Sep. 3, 2021. [cited by applicant]
Si, Shijing, Jianzong Wang, Xiaoyang Qu, Ning Cheng, Wenqi Wei, Xinghua Zhu, and Jing Xiao, “Speech2video: Cross-modal distillation for speech to video generation,” arXiv preprint arXiv:2107.04806, Jul. 10, 2021. [cited by applicant]
Sohn, Kihyuk, Honglak Lee, and Xinchen Yan, “Learning structured output representation using deep conditional generative models,” Advances in neural information processing systems 28, 2015. [cited by applicant]
Taylor, Sarah, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews, “A deep learning approach for generalized speech animation,” ACM Transactions On Graphic… [cited by applicant]
Theis, Lucas, Aäron Van Den Oord, and Matthias Bethge, “A note on the evaluation of generative models,” arXiv preprint arXiv:1511.01844, Nov. 5, 2015. [cited by applicant]
Tzirakis, Panagiotis, Athanasios Papaioannou, Alexandros Lattas, Michail Tarasiou, Björn Schuller, and Stefanos Zafeiriou, “Synthesising 3D facial motion from “in-the-wild” speech,” In 2020 15th IEEE International Confe… [cited by applicant]
Vaessen, Nik, and David A. Van Leeuwen, “Fine-tuning wav2vec2 for speaker recognition,” arXiv preprint arXiv:2109.15053, Sep. 30, 2021. [cited by applicant]
Verma, Ashish, Nitendra Rajput, and L. Venkata Subramaniam, “Using viseme based acoustic models for speech driven lip synthesis,” In 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003, … [cited by applicant]
Wang, Chengyi, Yu Wu, Sanyuan Chen, Shujie Liu, Jinyu Li, Yao Qian, and Zhenglu Yang, “Self-supervised learning for speech recognition with intermediate layer supervision,” arXiv preprint arXiv:2112.08778, Dec. 16, 2021. [cited by applicant]
Yang, Lin, Yi Shen, Yue Mao, and Longjun Cai, “Hybrid Curriculum Learning for Emotion Recognition in Conversation,” arXiv preprint arXiv:2112.11718 Dec. 22, 2021. [cited by applicant]
Yannakakis, Georgios N., Roddy Cowie, and Carlos Busso, “The ordinal nature of emotions: An emerging approach,” IEEE Transactions on Affective Computing 12, No. 1: 16-35, 2018. [cited by applicant]
Zachary C. Lipton, and Subarna Tripathi, “Precise Recovery of Latent Vectors from Generative Adversarial Networks”, 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, Apr. 24-26, 2017, … [cited by applicant]
Zhang, Yuanyuan, Jun Du, Zirui Wang, and Jianshu Zhang, “Attention Based Fully Convolutional Network for Speech Emotion Recognition,” arXiv preprint arXiv:1806.01506, Jun. 5, 2018. [cited by applicant]
Zhou, Yang, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh, “Visemenet: Audio-driven animator-centric speech animation,” ACM Transactions on Graphics (TOG) 37, No. 4: 1-10, Aug. 2018. [cited by applicant]
Zhou, Yang, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh, “VisemeNet: Audio-Driven Animator-Centric Speech Animation,” arXiv preprint arXiv:1805.09488, May 24, 2018. [cited by applicant]
Zhu, Lixing, Gabriele Pergola, Lin Gui, Deyu Zhou, and Yulan He, “Topic-Driven and Knowledge-Aware Transformer for Dialogue Emotion Detection,” In Proceedings of the 59th Annual Meeting of the Association for Computatio… [cited by applicant]
Abdelaziz, Ahmed Hussen, et al., “Audiovisual Speech Synthesis using Tacotron2,” arXiv preprint arXiv:2008.00620 Aug. 3, 2020. [cited by applicant]
Peng, Ziqiao, et al., “EmoTalk: Speech-driven emotional disentanglement for 3D face animation,” arXiv preprint arXiv:2303.11089 , Mar. 20, 2023. [cited by applicant]
Eskimez, Sefik Emre, et al., “Speech driven talking face generation from a single image and an emotion condition,” IEEE Transactions on Multimedia 24: 3480-3490 Jul. 21, 2021. [cited by applicant]
Wang, Kaisiyuan, et al., “Mead: A large-scale audio-visual dataset for emotional talking-face generation.” Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, Aug. 23-28, 2020, Proceedings, Part XXI. Cham:… [cited by applicant]
Zeng, Dan et al., “Talking face generation with expression-tailored generative adversarial network,” Proceedings of the 28th ACM International Conference on Multimedia, Supplemental Material Video Retrieved from: https:… [cited by applicant]
Sadiq, Rizwan, et al., “Emotion Dependent Facial Animation from Affective Speech.” 2020 IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP). IEEE, 2020. [cited by applicant]
Bolduc, Maquis et al., “Rig Inversion by Training a Differentiable Rig Function,” SIGGRAPH Asia 2022 Technical Communications, 1-4 2022. [cited by applicant]
Medina, Salvador, et al., Speech Driven Tongue Animations, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. [cited by applicant]
Fan, Yingruo, et al., “Faceformer: Speech-driven 3d facial animation with transformers,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Bao, Linchao et al., “Learning Audio-Driven Viseme Dynamics for 3D Face Animation,” arXiv preprint arXiv:2301.06059, Jan. 15, 2023. [cited by applicant]
Tarantino, L., et al., Self-Attention for Speech Emotion Recognition. Proc. Interspeech 2019, 2578-2582, doi: 10.21437/Interspeech.2019-2822, Sep. 15, 2019. [cited by applicant]
Li, Y., et al., ) Improved End-to-End Speech Emotion Recognition Using Self Attention Mechanism and Multitask Learning. Proc. Interspeech 2019, 2803-2807, doi: 10.21437/Interspeech.2019-2594, Sep. 15, 2019. [cited by applicant]
Jali, Jali Research Inc., URL: http://jaliresearch.com, 5 pages, visited on Sep. 9, 2021. [cited by applicant]
Lithgow, K. and Edge, J, “Surrey AudioVisual Expressed Emotion (SAVEE) Database,” URL: http://kahlan.eps.surrey.ac.uk/savee/. 7 pages, Apr. 2, 2015. [cited by applicant]
NVIDIA, Omniverse Audio 2Face, https://www.nvidia.com/en-us/omiverse/apps/audio2face/, 1 page, Retrieved on: May 22, 2023. [cited by applicant]
Speech Graphics, SGX, https://www.speech-graphics.com/sgx-production-audio-to-face-animation-software/; 2pages, Retrieved on: May 22, 2023. [cited by applicant]
Cited By (1)
US 12,731,319