IP Library Granted Patent US 12,592,018
Granted Patent B1
US 12,592,018 · App. 18/128,995 · Granted Mar 31, 2026

Generating expressive facial animation data from speech audio using speech emotion recognition

Inventors: Monica Villanueva Aylagas (Sundbyberg, SE); Mattias Teye (Sundbyberg, SE); Hector Leon (Malmö, SE); Damian Valle (Stockholm, SE)
Assignee: ELECTRONIC ARTS INC.
G06T13/40G10L15/16G10L25/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,592,018
App. No.
18/128,995
Granted
Mar 31, 2026
Kind
B1
Abstract

This specification provides a computer-implemented method for generating facial animation data. The facial animation data animates a face in a video game in accordance with speech sounds and speech emotion of speech audio. Speech audio data representing the speech audio is processed using a machine-learned speech emotion recognition model. One or more emotion representations that represent the emotion of the speech audio are generated. For each of the one or more emotion representations, a respective conditioning input for a facial animation generative model is determined. Facial animation data for the speech audio data is generated, comprising processing, using the facial animation generative model: (i) input data derived from the speech audio data, and (ii) the one or more conditioning inputs.

Claims (56)

1 . A computer-implemented method for generating facial animation data that animates a face in a video game in accordance with speech sounds and speech emotion of speech audio, the method comprising:

generating one or more emotion representations that represent the emotion of the speech audio, comprising processing, using a machine-learned speech emotion recognition model, speech audio data representing the speech audio;

wherein each of the one or more emotion representations comprises a respective score for each emotion classification of a set of emotion classifications;

wherein each emotion classification of the set of emotion classifications is associated with a pre-determined conditioning input;

determining, for each of the one or more emotion representations, a respective conditioning input for a facial animation generative model comprising:

for each emotion classification of the set of emotion classifications, weighting the pre-determined conditioning input associated with the emotion classification based on the respective score for the emotion classification generated by the machine-learned speech emotion recognition model; and

summing together the weighted pre-determined conditioning inputs to generate the respective conditioning input for the facial animation generative model; and

generating facial animation data for the speech audio data, comprising processing, using the facial animation generative model: (i) input data derived from the speech audio data, and (ii) the one or more conditioning inputs.

2 . The method of claim 1 , wherein determining, for each of the one or more emotion representations, the respective conditioning input for the facial animation generative model comprises selecting the pre-determined conditioning input associated with a highest-scoring emotion classification.

3 . The method of claim 1 , wherein:

generating the one or more emotion representations comprises generating a global emotion representation that represents the emotion of the entire speech audio; and

determining, for each of the one or more emotion representations, the respective conditioning input comprises determining a global conditioning input for the entire speech audio.

4 . The method of claim 1 , wherein:

generating the one or more emotion representations comprises generating an emotion representation for each time step of a plurality of time steps of the speech audio data, comprising processing speech audio data of the time step using the machine-learned speech emotion recognition model; and

determining, for each of the one or more emotion representations, the respective conditioning input comprises determining a conditioning input for each time step using the emotion representation generated for the time step.

5 . The method of claim 1 , wherein generating the facial animation data for the speech audio data comprises:

generating, for each time step of a plurality of time steps of the speech audio data, mesh data for a facial mesh.

6 . The method of claim 1 , wherein generating the facial animation data for the speech audio data comprises:

generating, for each time step of a plurality of time steps of the speech audio data, rig parameters for a facial animation rig.

7 . The method of claim 1 , wherein the facial animation generative model comprises at least one of:

a decoder of a conditional variational autoencoder;

a generator neural network of a conditional generative adversarial neural network;

a conditional flow-based model, which is optionally a conditional normalizing flow model; and

a conditional denoising diffusion model.

8 . A computing system for generating facial animation data that animates a face in a video game in accordance with speech sounds and emotion of speech audio, the computing system comprising one or more computing devices configured to:

generate one or more emotion representations that represent the emotion of the speech audio, comprising processing, using a machine-learned speech emotion recognition model, speech audio data representing the speech audio;

wherein each of the one or more emotion representations comprises a respective score for each emotion classification of a set of emotion classifications;

wherein each emotion classification of the set of emotion classifications is associated with a pre-determined conditioning input;

determine, for each of the one or more emotion representations, a respective conditioning input for a facial animation generative model comprising:

for each emotion classification of the set of emotion classifications, weighting the pre-determined conditioning input associated with the emotion classification based on the respective score for the emotion classification generated by the machine-learned speech emotion recognition model; and

summing together the weighted pre-determined conditioning inputs to generate the respective conditioning input for the facial animation generative model; and

generate facial animation data for the speech audio data, comprising processing, using the facial animation generative model: (i) input data derived from the speech audio, and (ii) the one or more conditioning inputs.

9 . The computing system of claim 8 , wherein the input data derived from the speech audio data comprises acoustic features for each time step of a plurality of time steps of the speech audio data.

10 . The computing system of claim 9 , wherein the acoustic features comprise spectrogram parameters.

11 . The computing system of claim 8 , wherein the input data derived from the speech audio data comprises outputs generated by one or more neural network layers of a speech transcription neural network, wherein the outputs are generated by processing the speech audio data using the speech transcription neural network.

12 . The computing system of claim 8 , wherein the facial animation generative model comprises at least one of:

a decoder of a conditional variational autoencoder;

a generator neural network of a conditional generative adversarial neural network;

a conditional flow-based model, which is optionally a conditional normalizing flow model; and

a conditional denoising diffusion model.

13 . A non-transitory computer-readable medium storing instructions, which when executed by one or more processors, cause the processor to:

generate one or more emotion representations for speech audio, comprising processing, using a machine-learned speech emotion recognition model, speech audio data representing the speech audio;

wherein each of the one or more emotion representations comprises a respective score for each emotion classification of a set of emotion classifications;

wherein each emotion classification of the set of emotion classifications is associated with a pre-determined conditioning input;

determine, for each of the one or more emotion representations, a respective conditioning input for a facial animation generative model comprising:

for each emotion classification of the set of emotion classifications, weighting the pre-determined conditioning input associated with the emotion classification based on the respective score for the emotion classification generated by the machine-learned speech emotion recognition model; and

summing together the weighted pre-determined conditioning inputs to generate the respective conditioning input for the facial animation generative model; and

generate facial animation data for the speech audio data, comprising processing, using the facial animation generative model: (i) input data derived from the speech audio, and (ii) the one or more conditioning inputs.

14 . The non-transitory computer-readable medium of claim 13 , wherein the input data derived from the speech audio data comprises acoustic features for each time step of a plurality of time steps of the speech audio data.

15 . The non-transitory computer-readable medium of claim 14 , wherein the acoustic features comprise spectrogram parameters.

16 . The non-transitory computer-readable medium of claim 13 , wherein the input data derived from the speech audio data comprises outputs generated by one or more neural network layers of a speech transcription neural network, wherein the outputs are generated by processing the speech audio data using the speech transcription neural network.

17 . The non-transitory computer-readable medium of claim 16 , wherein the facial animation generative model comprises at least one of:

a decoder of a conditional variational autoencoder;

a generator neural network of a conditional generative adversarial neural network;

a conditional flow-based model, which is optionally a conditional normalizing flow model; and

a conditional denoising diffusion model.

Continuity (2)
Provisional Application 63327647 · Apr 5, 2022
Provisional Application 63327633 · Apr 5, 2022
References Cited (104)
US 10176619B2 · Jiao · 2019 [cited by examiner]
US 11349989B2 · Murali · 2022 [cited by examiner]
US 11756251B2 · Kaushik · 2023 [cited by examiner]
US 20210319610A1 · del val Santos · 2021 [cited by examiner]
CN 113393832A · 2021 [cited by examiner]
A. Van den Oord et al., “WaveNet: A Generative Model for Raw Audio” Sep. 19, 2016, Speech Synthesis Workshop, 15 pages (Year: 2016). [cited by examiner]
X. Wu et al., “F [cited by examiner]
Karras et al, “Audio-driven facial animation by joint end-to-end learning of pose and emotion” (Year: 2017). [cited by examiner]
Li, Zhi, Anne Aaron, Ioannis Katsavounidis, Anush Moorthy, and Megha Manohara, “Toward a practical perceptual video quality metric,” The Netflix Tech Blog 6, No. 2: 2, Jun. 6, 2016. [cited by applicant]
Lin, Ji, Richard Zhang, Frieder Ganz, Song Han, and Jun-Yan Zhu, “Anycost gans for interactive image synthesis and editing,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1498… [cited by applicant]
Lithgow, K. and Edge, J, “Surrey AudioVisual Expressed Emotion (SAVEE) Database,” URL: http://kahlan.eps.surrey.ac.uk/savee/. Apr. 2, 2015. [cited by applicant]
Livingstone, Steven R., and Frank A. Russo, “The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,” PloS one 13, N… [cited by applicant]
Luo, Changwei, Jun Yu, Xian Li, and Leilei Zhang, “Hmm based speech-driven 3D tongue animation,” In 2017 IEEE International Conference On Image Processing (ICIP), pp. 4377-4381, IEEE, Sep. 2017. [cited by applicant]
News, Guinness World Records. Star Wars: The Old Republic Recognised Guinness World Records 2012 Gamer's Edition, URL: https://www.guinnessworldrecords.com, Retrieved on: Dec. 17, 2021. [cited by applicant]
Nwe, Tin Lay, Say Wei Foo, and Liyanage C. De Silva, “Speech emotion recognition using hidden Markov models,” Speech communication 41, No. 4: 603-623, Nov. 2003. [cited by applicant]
Orvalho, Verónica, Pedro Bastos, Frederic I. Parke, Bruno Oliveira, and Xenxo Alvarez, “A Facial Rigging Survey,” Eurographics (State of the Art Reports): 183-204, 2012. [cited by applicant]
Paszke, Adam, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen et al, “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processi… [cited by applicant]
Patamia, Rutherford Agbeshi, Wu Jin, Kingsley Nketia Acheampong, Kwabena Sarpong, and Edwin Kwadwo Tenagyei, “Transformer based multimodal speech emotion recognition with improved neural networks,” In 2021 IEEE 2nd Inte… [cited by applicant]
Pelachaud, Catherine, Cornelius Wam Van Overveld, and Chin Seah, “Modeling and animating the human tongue during speech production,” In Proceedings of Computer Animation'94, pp. 40-49. IEEE, May 1994. [cited by applicant]
Pham, Hai Xuan, Yuting Wang, and Vladimir Pavlovic, “End-to-end learning for 3d facial animation from raw waveforms of speech.” arXiv preprint arXiv:1710.00920 Dec. 7, 2017. [cited by applicant]
Richard, Alexander, Colin Lea, Shugao Ma, Jurgen Gall, Fernando De La Torre, and Yaser Sheikh, “Audio-and gaze-driven facial animation of codec avatars,” In Proceedings of the IEEE/CVF winter conference on applications … [cited by applicant]
Richard, Alexander, Michael Zollhoefer, Yandong Wen, Fernando De La Torre, and Yaser Sheikh, “MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement,” arXiv preprint arXiv:2104.08223, Apr. 16, 2021. [cited by applicant]
Schneider, Steffen, Alexei Baevski, Ronan Collobert, and Michael Auli, “wav2vec: Unsupervised pre-training for speech recognition,” arXiv preprint arXiv:1904.05862, Sep. 11, 2019. [cited by applicant]
Schwartz, Roy, Jesse Dodge, Noah A. Smith, and Oren Etzioni, “Green AI,” arXiv preprint arXiv:1907.10597, Aug. 13, 2019. [cited by applicant]
SG, Speech Graphics, URL: https://www.speech-graphics.com/, Retrieved on: Sep. 3, 2021. [cited by applicant]
Si, Shijing, Jianzong Wang, Xiaoyang Qu, Ning Cheng, Wenqi Wei, Xinghua Zhu, and Jing Xiao, “Speech2video: Cross-modal distillation for speech to video generation,” arXiv preprint arXiv:2107.04806, Jul. 10, 2021. [cited by applicant]
Sohn, Kihyuk, Honglak Lee, and Xinchen Yan, “Learning structured output representation using deep conditional generative models,” Advances in neural information processing systems 28, 2015. [cited by applicant]
Taylor, Sarah, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews, “A deep learning approach for generalized speech animation,” ACM Transactions On Graphic… [cited by applicant]
Theis, Lucas, Aäron Van Den Oord, and Matthias Bethge, “A note on the evaluation of generative models,” arXiv preprint arXiv:1511.01844, Nov. 5, 2015. [cited by applicant]
Tzirakis, Panagiotis, Athanasios Papaioannou, Alexandros Lattas, Michail Tarasiou, Björn Schuller, and Stefanos Zafeiriou, “Synthesising 3D facial motion from “in-the-wild” speech,” In 2020 15th IEEE International Confe… [cited by applicant]
Vaessen, Nik, and David A. Van Leeuwen, “Fine-tuning wav2vec2 for speaker recognition,” arXiv preprint arXiv:2109.15053, Sep. 30, 2021. [cited by applicant]
Verma, Ashish, Nitendra Rajput, and L. Venkata Subramaniam, “Using viseme based acoustic models for speech driven lip synthesis,” In 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003, … [cited by applicant]
Wang, Chengyi, Yu Wu, Sanyuan Chen, Shujie Liu, Jinyu Li, Yao Qian, and Zhenglu Yang, “Self- supervised learning for speech recognition with intermediate layer supervision,” arXiv preprint arXiv:2112.08778, Dec. 16, 202… [cited by applicant]
Yang, Lin, Yi Shen, Yue Mao, and Longjun Cai, “Hybrid Curriculum Learning for Emotion Recognition in Conversation,” arXiv preprint arXiv:2112.11718 Dec. 22, 2021. [cited by applicant]
Yannakakis, Georgios N., Roddy Cowie, and Carlos Busso, “The ordinal nature of emotions: An emerging approach,” IEEE Transactions on Affective Computing 12, No. 1: 16-35, 2018. [cited by applicant]
Zachary C. Lipton, and Subarna Tripathi, “Precise Recovery of Latent Vectors from Generative Adversarial Networks”, 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, Apr. 24-26, 2017, … [cited by applicant]
Zhang, Yuanyuan, Jun Du, Zirui Wang, and Jianshu Zhang, “Attention Based Fully Convolutional Network for Speech Emotion Recognition,” arXiv preprint arXiv:1806.01506, Jun. 5, 2018. [cited by applicant]
Zhou, Yang, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh, “Visemenet: Audio-driven animator-centric speech animation,” ACM Transactions on Graphics (TOG) 37, No. 4: Aug. 1-10, 2018. [cited by applicant]
Zhou, Yang, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh, “VisemeNet: Audio-Driven Animator-Centric Speech Animation,” arXiv preprint arXiv:1805.09488, May 24, 2018. [cited by applicant]
Zhu, Lixing, Gabriele Pergola, Lin Gui, Deyu Zhou, and Yulan He, “Topic-Driven and Knowledge-Aware Transformer for Dialogue Emotion Detection,” In Proceedings of the 59th Annual Meeting of the Association for Computatio… [cited by applicant]
Nvidia, Omniverse Audio 2Face, https://www.nvidia.com/en-us/omniverse/apps/audio2face/, Retrieved on: May 22, 2023. [cited by applicant]
Speech Graphics, SGX, https://www.speech-graphics.com/sgx-production-audio-to-face-animation-software/; Retrieved on: May 22, 2023. [cited by applicant]
Abdelaziz, Ahmed Hussen, et al., “Audiovisual Speech Synthesis using Tacotron2,” arXiv preprint arXiv:2008.00620 Aug. 3, 2020. [cited by applicant]
Peng, Ziqiao, et al., “EmoTalk: Speech-driven emotional disentanglement for 3D face animation,” arXiv preprint arXiv:2303.11089 , Mar. 20, 2023. [cited by applicant]
Eskimez, Sefik Emre, et al., “Speech driven talking face generation from a single image and an emotion condition,” IEEE Transactions on Multimedia 24: 3480-3490 Jul. 21, 2021. [cited by applicant]
Wang, Kaisiyuan, et al., “Mead: A large-scale audio-visual dataset for emotional talking-face generation.” Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, Aug. 23-28, 2020, Proceedings, Part XXI. Cham:… [cited by applicant]
Zeng, Dan et al., “Talking face generation with expression-tailored generative adversarial network,” Proceedings of the 28th ACM International Conference on Multimedia, Supplemental Material Video Retrieved from: https:… [cited by applicant]
Sadiq, Rizwan, et al., “Emotion Dependent Facial Animation from Affective Speech.” 2020 IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP). IEEE, 2020. [cited by applicant]
Bolduc, Maquis et al., “Rig Inversion by Training a Differentiable Rig Function,” SIGGRAPH Asia 2022 Technical Communications, Jan. 4, 2022. [cited by applicant]
Medina, Salvador, et al., Speech Driven Tongue Animations, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. [cited by applicant]
Fan, Yingruo, et al., “Faceformer: Speech-driven 3d facial animation with transformers,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Bao, Linchao et al., “Learning Audio-Driven Viseme Dynamics for 3D Face Animation,” arXiv preprint arXiv:2301.06059, Jan. 15, 2023. [cited by applicant]
Tarantino, L., et al., Self-Attention for Speech Emotion Recognition. Proc. Interspeech 2019, 2578-2582, doi: 10.21437/Interspeech.2019-2822, Sep. 15, 2019. [cited by applicant]
Li, Y., et al., ) Improved End-to-End Speech Emotion Recognition Using Self Attention Mechanism and Multitask Learning. Proc. Interspeech 2019, 2803-2807, doi: 10.21437/Interspeech.2019-2594, Sep. 15, 2019. [cited by applicant]
Abdal, Rameen, Yipeng Qin, and Peter Wonka, “Image2stylegan++: How to edit the embedded images?,” In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 8296-8305, 2020. [cited by applicant]
Abrevaya, Victoria Fernández, Adnane Boukhayma, Philip HS Torr, and Edmond Boyer, “Cross-modal deep face normals with deactivable skip connections,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Patte… [cited by applicant]
Baevski, Alexei, Henry Zhou, Abdelrahman Mohamed, and Michael Auli, “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” arXiv preprint arXiv:2006.11477 Oct. 22, 2020. [cited by applicant]
Bailey, Stephen W., Dalton Omens, Paul Dilorenzo, and James F. O'Brien, “Fast and deep facial deformations,” ACM Transactions on Graphics (TOG) 39, No. 4: 94-1, Jul. 2020. [cited by applicant]
Bakker, Iris, Theo Van Der Voordt, Peter Vink, and Jan De Boon, “Pleasure, arousal, dominance: Mehrabian and Russell revisited,” Current Psychology 33: 405-421, Jun. 11, 2014. [cited by applicant]
Benzeghiba, Mohamed, Renato De Mori, Olivier Deroo, Stephane Dupont, Teodora Erbes, Denis Jouvet, Luciano Fissore et al. “Automatic speech recognition and speech variability: A review,” Speech communication 49, No. 10-1… [cited by applicant]
Bhat, Chitralekha, and Sunil Kopparapu, “Viseme comparison based on phonetic cues for varying speech accents,” In Sixteenth Annual Conference of the International Speech Communication Association, Sep. 6, 2015. [cited by applicant]
Bishop, Chris M., “Training with noise is equivalent to Tikhonov regularization,” Neural computation 7, No. 1: 108-116, 1995. [cited by applicant]
Botha, Johnny, and Heloise Pieterse, “Fake news and deepfakes: A dangerous threat for 21st century information security,” In ICCWS 2020 15th International Conference on Cyber Warfare and Security, Academic Conferences a… [cited by applicant]
Brand, Matthew, “Voice puppetry,” In Proceedings of the 26th annual conference on Computer graphics and interactive techniques, pp. 21-28, 1999. [cited by applicant]
Bregler, Christoph, Michele Covell, and Malcolm Slaney, “Video rewrite: Driving visual speech with audio,” In Proceedings of the 24th annual conference on Computer graphics and interactive techniques, pp. 353-360, 1997. [cited by applicant]
Brown, Tom B., Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan et al. “Language Models are Few-Shot Learners,” arXiv preprint arXiv:2005.14165, Jul. 22, 2020. [cited by applicant]
Burgess, Christopher P., Irina Higgins, Arka Pal, Loic Matthey, Nick Watters, Guillaume Desjardins, and Alexander Lerchner, “Understanding disentangling in $\beta $-VAE,” arXiv preprint arXiv:1804.03599, Apr. 10, 2018. [cited by applicant]
Busso, Carlos, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” Language… [cited by applicant]
Cao, Houwei, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova, and Ragini Verma, “Crema-d: Crowd-sourced emotional multimodal actors dataset,” IEEE transactions on affective computing 5, No. 4: 377-390, 2… [cited by applicant]
Chai, Yujin, Yanlin Weng, Lvdi Wang, and Kun Zhou, “Speech-driven facial animation with spectral gathering and temporal attention,” Frontiers of Computer Science 16: 1-10, Sep. 23, 2020. [cited by applicant]
Chen, Mingyi, Xuanji He, Jing Yang, and Han Zhang, “3-D convolutional recurrent neural networks with attention model for speech emotion recognition,” IEEE Signal Processing Letters 25, No. 10: 1440-1444, Oct. 10, 2018. [cited by applicant]
Chung, Joon Son, Amir Jamaludin, and Andrew Zisserman, “You said that?,” arXiv preprint arXiv:1705.02966, Jul. 18, 2017. [cited by applicant]
Chung, Joon Son, and Andrew Zisserman, “Lip reading in the wild,” In Computer Vision—ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, Nov. 20-24, 2016, Revised Selected Papers, Part II 13, pp. 87-103… [cited by applicant]
Cootes, Timothy F., Gareth J. Edwards, and Christopher J. Taylor, “Active appearance models,” IEEE Transactions on pattern analysis and machine intelligence 23, No. 6: 681-685, 2001. [cited by applicant]
Cudeiro, Daniel, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J. Black, “Capture, learning, and synthesis of 3D speaking styles,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec… [cited by applicant]
Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, Oct. 11, 2018. [cited by applicant]
Diederik P. Kingma, , and Jimmy BA, “ADAM: A Method for Stochastic Optimization,” In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, Conference Track Proceedings, May 7, 2015. [cited by applicant]
Doersch, Carl, “Tutorial on variational autoencoders,” arXiv preprint arXiv:1606.05908, Aug. 13, 2016. [cited by applicant]
Dossou, Bonaventure FP, and Yeno KS Gbenou, “FSER: Deep Convolutional Neural Networks for Speech Emotion Recognition,” arXiv preprint arXiv:2109.07916, 2021. [cited by applicant]
Ekman, Paul, “Facial expressions of emotion: an old controversy and new findings,” Philosophical Transactions of the Royal Society of London. Series B: Biological Sciences 335, No. 1273: 63-69, 1992. [cited by applicant]
Engwall, Olov, and Jonas Beskow, “Resynthesis of 3D tongue movements from facial data,” In Eighth European Conference on Speech Communication and Technology, 2003. [cited by applicant]
Ezzat, Tony, Gadi Geiger, and Tomaso Poggio, “Trainable videorealistic speech animation,” ACM Transactions on Graphics (TOG) 21, No. 3: 388-398, Jun. 6, 2002. [cited by applicant]
Fabre, Diandra, Thomas Hueber, Laurent Girin, Xavier Alameda-Pineda, and Pierre Badin, “Automatic animation of an articulatory tongue model from ultrasound images of the vocal tract,” Speech Communication 93: 63-75, Sep… [cited by applicant]
FFX, FaceFX, URL: https://facefx.com/, Retrieved on: Sep. 3, 2021. [cited by applicant]
Gidaris, Spyros, Praveer Singh, and Nikos Komodakis, “Unsupervised representation learning by predicting image rotations,” arXiv preprint arXiv:1803.07728, Mar. 21, 2018. [cited by applicant]
Hahn, Fabian, Bernhard Thomaszewski, Stelian Coros, Robert W. Sumner, and Markus Gross, “Efficient simulation of secondary motion in rig-space,” In Proceedings of the 12th ACM SIGGRAPH/eurographics symposium on computer… [cited by applicant]
Hahn, Fabian, Sebastian Martin, Bernhard Thomaszewski, Robert Sumner, Stelian Coros, and Markus Gross, “Rig-space physics.” ACM transactions on graphics (TOG) 31, No. 4: Jan. 8, 2012. [cited by applicant]
Han, Kun, Dong Yu, and Ivan Tashev, “Speech emotion recognition using deep neural network and extreme learning machine,” In Interspeech, 2014. [cited by applicant]
Hannun, Awni, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger et al, “Deep speech: Scaling up end-to-end speech recognition,” arXiv preprint arXiv:1412.5567, Dec. 19, 2014. [cited by applicant]
Higgins, Irina, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner, “beta-vae: Learning basic visual concepts with a constrained variational framework,”… [cited by applicant]
Holden, Daniel, Jun Saito, and Taku Komura, “Learning inverse rig mappings by nonlinear regression,” IEEE transactions on visualization and computer graphics 23, No. 3: 1167-1178, 2016. [cited by applicant]
Huang, Xuedong, and Kai-Fu Lee, “On speaker-independent, speaker-dependent, and speaker-adaptive speech recognition,” IEEE Transactions on Speech and Audio processing 1, No. 2: 150-157, Apr. 1993. [cited by applicant]
Jali, Jali Research Inc. 2021, URL: http://jaliresearch.com, Retrieved on: Sep. 3, 2021. [cited by applicant]
James, Jesin, Li Tian, and Catherine Watson, “An open source emotional speech corpus for human robot interaction applications,” Interspeech, Sep. 2, 2018. [cited by applicant]
Jonell, Patrik, Taras Kucherenko, Gustav Eje Henter, and Jonas Beskow, “Let's face it: Probabilistic multi-modal interlocutor-aware generation of facial gestures in dyadic settings,” In Proceedings of the 20th ACM Inter… [cited by applicant]
Karras, Tero, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila, “Alias-Free Generative Adversarial Networks,” arXiv preprint arXiv:2106.12423, Mar. 13, 2021. [cited by applicant]
Karras, Tero, Samuli Laine, and Timo Aila, “A style-based generator architecture for generative adversarial networks,” In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4401-4410,… [cited by applicant]
Karras, Tero, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila, “Analyzing and Improving the Image Quality of StyleGAN,” In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition … [cited by applicant]
Karras, Tero, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen, “Audio-driven facial animation by joint end-to-end learning of pose and emotion,” ACM Transactions on Graphics (TOG) 36, No. 4: Jan. 12, 2017. [cited by applicant]
Kingma, D. P., and M. Welling, “Auto-encoding variational Bayes. 2nd international conference on learning representations (ICLR2014),” Preprint, submitted Dec. 23 (2014): arXiv: http://arxiv.org/abs/1312.6114v10, May 1,… [cited by applicant]
Kingma, Diederik P., and Max Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114, Dec. 20, 2013. [cited by applicant]
Korn, Oliver, Lukas Stamm, and Gerd Moeckl, “Designing authentic emotions for non-human characters: A study evaluating virtual affective behavior,” In Proceedings of the 2017 conference on designing interactive systems,… [cited by applicant]
Lewis, John P., and Ken-Ichi Anjyo, “Direct manipulation blendshapes,” IEEE Computer Graphics and Applications 30, No. 4: 42-50, Jul. 2010. [cited by applicant]
Lewis, John P., Ken Anjyo, Taehyun Rhee, Mengjie Zhang, Frederic H. Pighin, and Zhigang Deng, “Practice and theory of blendshape facial models,” Eurographics (State of the Art Reports) 1, No. 8: Feb. 2014. [cited by applicant]