IP Library Granted Patent US 12,592,247
Granted Patent B2
US 12,592,247 · App. 17/859,660 · Granted Mar 31, 2026

Inferring emotion from speech in audio data using deep learning

Inventors: Ilia Federov (Moscow, RU); Dmitry Aleksandrovich Korobchenko (Moscow, RU)
Assignee: Nvidia Corporation
G10L25/63G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,592,247
App. No.
17/859,660
Granted
Mar 31, 2026
Kind
B2
Abstract

A deep neural network can be trained to infer emotion data from input audio. The network can be a transformer-based network that can infer probability values for a set of emotions or emotion classes. The emotion probability values can be modified using one or more heuristics, such as to provide for smoothing of emotion determinations over time, or via a user interface, where a user can modify emotion determinations as appropriate. A user may also provide prior emotion values to be blended with these emotion determination values. Determined emotion values can be provided as input to an emotion-based operation, such as to provide audio-driven speech animation.

Claims (58)

1 . A computer-implemented method, comprising:

computing, based at least on a neural network processing audio data representative of speech, a plurality of values indicative of emotions from a set of emotions in the speech;

selecting, from the set of emotions, a subset of one or more emotions, the selecting being based at least on respective values from the plurality of values corresponding to the subset of one or more emotions exceeding a threshold; and

rendering, based at least on a second neural network processing the audio data, data indicating the subset of one or more emotions, and one or more weighted values indicating one or more respective emotional strength values for the subset of the one or more emotions, an animation of a face such that one or more portions of the face move within the animation according to the speech and the subset of one or more emotions.

2 . The computer-implemented method of claim 1 , wherein the plurality of values include respective one or more probability values for each emotion of the set of emotions, and the plurality of values are normalized and summed to an absolute value.

3 . The computer-implemented method of claim 1 , wherein the set of emotions include at least one of anger, disgust, fear, joy, sadness, or a neutral emotion.

4 . The computer-implemented method of claim 1 , further comprising:

providing an interface to receive user input corresponding to one or more adjustments of one or more values of the plurality of values.

5 . The computer-implemented method of claim 4 , further comprising:

receiving, via the interface, the one or more respective emotion strength values for use in weighting the subset of one or more emotions with respect to at least one other emotion of the subset of one or more emotions determined to correspond to the speech.

6 . The computer-implemented method of claim 4 , further comprising:

receiving one or more first prior values corresponding to a first emotional state separate from the audio data;

blending the one or more first prior values with the one or more values to generate one or more first blended values;

receiving, through the interface, one or more second prior values corresponding to the subset of one or more emotions; and

blending the one or more second prior values with the one or more values to generate one or more second blended values,

wherein the neural network is a transformer-based neural network and determining that at least one emotion of the subset of one or more emotions corresponds to the speech is based at least in part on the one or more first blended values and the one or more second blended values.

7 . The computer-implemented method of claim 6 , further comprising:

receiving one or more prior emotion strength values corresponding to the subset of one or more emotions, the one or more prior emotion strength values indicating one or more weights to be used in the blending of the one or more second prior values with the one or more values.

8 . The computer-implemented method of claim 1 , further comprising:

determining one or more probability values for the subset of one or more emotions for each keyframe of a set of keyframes in the audio data, the one or more probability values being determined using a sliding window of audio data for a given audio segment.

9 . The computer-implemented method of claim 8 , further comprising:

smoothing the one or more values corresponding to the subset of one or more emotions across a plurality of iterations.

10 . The computer-implemented method of claim 1 , wherein the audio data is represented using an audio file format.

11 . A processor comprising:

one or more processing units to:

provide audio data in an audio file format as input to a neural network;

provide at least one style vector including instructions for rendering one or more facial components for one or more emotions;

compute, based at least on the neural network processing the audio data, a plurality of values indicative of emotions from a set of emotions in the audio data; and

render, based at least on a second neural network processing the audio data, data indicating a selected subset of one or more emotions, and one or more weighted values indicating one or more respective emotional strength values for the selected subset of the one or more emotions, an animation of the one or more facial components according to the audio data, the subset of one or more emotions, the weighted values indicating the one or more respective emotional strength values, and the at least one style vector.

12 . The processor of claim 11 , wherein the set of emotions include a predetermined set of emotions, wherein the predetermined set of emotions includes at least anger, disgust, fear, joy, sadness, or neutral.

13 . The processor of claim 11 , wherein the one or more processing units are further to:

weight the plurality of values based at least in part on the one or more respective emotional strength values corresponding to respective emotions of the selected subset of one or more emotions.

14 . The processor of claim 11 , wherein the one or more processing units are further to:

receive one or more prior values corresponding to the selected subset of one or more emotions; and

blend the one or more prior values with the respective values of the plurality of values to generate one or more blended values,

wherein a determination that the at least one emotion of the selected subset of one or more emotions corresponds to the audio data is based at least in part on the one or more blended values.

15 . The processor of claim 11 , wherein the audio file format includes at least one of an uncompressed audio file format, a lossless compression audio file format, or a lossy compression audio file format.

16 . A system comprising:

one or more processing units to:

compute, based at least on one or more neural networks processing audio data representative of speech, a plurality of values indicating a probability that emotions from a set of emotions correspond to the speech;

compute, based at least on the one or more neural networks processing a selected subset of first values having respective values exceeding a threshold and having one or more weighted emotional strength values, a plurality of second values corresponding to a style of animation, and the audio data, a plurality of third values indicating one or more positions of one or more feature points corresponding to a virtual object; and

render the virtual object based at least in part on the one or more third values.

17 . The system of claim 16 , wherein the audio data corresponds to an audio file format.

18 . The system of claim 16 , wherein the audio data is processed using a transformer neural network of the one or more networks in an audio file format and the audio data is processed using the one or more neural networks in an image file format.

19 . The system of claim 16 , wherein the one or more feature points correspond to one or more facial features or one or more body features of the virtual object.

20 . The system of claim 16 , wherein the system comprises at least one of:

a system for performing simulation operations;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets;

a system for performing deep learning operations;

a system implemented using an edge device;

a system implemented using a robot;

a system for performing conversational AI operations;

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2022
From: FEDEROV, ILIA; KOROBCHENKO, DMITRY ALEKSANDROVICH
To: NVIDIA CORPORATION
Reel/Frame 060871/0339 →
Continuity (1)
Related Publication 20240013802A1 · Jan 11, 2024
References Cited (34)
US 11521639B1 · Shon · 2022 [cited by examiner]
US 11551393B2 · Shang · 2023 [cited by examiner]
US 20040095344A1 · Dojyun · 2004 [cited by examiner]
US 20080052080A1 · Narayanan · 2008 [cited by applicant]
US 20100036660A1 · Bennett · 2010 [cited by applicant]
US 20100094634A1 · Park · 2010 [cited by examiner]
US 20120130717A1 · Xu · 2012 [cited by examiner]
US 20150206543A1 · Lee · 2015 [cited by examiner]
US 20180144761A1 · Amini · 2018 [cited by examiner]
US 20200250278A1 · Liu · 2020 [cited by examiner]
US 20200410739A1 · Shin · 2020 [cited by examiner]
US 20210027512A1 · Park · 2021 [cited by examiner]
US 20210050033A1 · Bui · 2021 [cited by examiner]
US 20210193169A1 · Faizakof · 2021 [cited by examiner]
US 20210233299A1 · Zhou · 2021 [cited by examiner]
US 20220020196A1 · Kuta · 2022 [cited by examiner]
US 20230154487A1 · Huang · 2023 [cited by examiner]
US 20230186906A1 · Gordon · 2023 [cited by examiner]
US 20230230303A1 · Yang · 2023 [cited by examiner]
US 20230351662A1 · Sinha · 2023 [cited by examiner]
US 20240105207A1 · Kruk · 2024 [cited by examiner]
CN 103810994A · 2014 [cited by applicant]
WO WO2007098560A1 · 2007 [cited by examiner]
WO WO2008087621A1 · 2008 [cited by examiner]
Siriwardhana, Shamane, et al. “Jointly fine-tuning” bert-like “self supervised models to improve multimodal speech emotion recognition.” arXiv preprint arXiv:2008.06682 (2020). (Year: 2020). [cited by examiner]
Boigne, Jonathan, Biman Liyanage, and Ted Östrem. “Recognizing more emotions with less data using self-supervised transfer learning.” arXiv preprint arXiv:2011.05585 (2020). (Year: 2020). [cited by examiner]
Macary, Manon, et al. “On the use of self-supervised pre-trained acoustic and linguistic features for continuous speech emotion recognition.” 2021 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2021. (Year: 2021). [cited by examiner]
Pepino, Leonardo, Pablo Riera, and Luciana Ferrer. “Emotion recognition from speech using wav2vec 2.0 embeddings.” arXiv preprint arXiv:2104.03502 (2021). (Year: 2021). [cited by examiner]
Fan, Yingruo, et al. “Faceformer: Speech-driven 3d facial animation with transformers.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022. (Year: 2022). [cited by examiner]
Wang, Xianfeng, et al. “A novel end-to-end speech emotion recognition network with stacked transformer layers.” ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2… [cited by examiner]
Sadiq, Rizwan, Sasan Asadiabadi, and Engin Erzin. “Emotion Dependent Facial Animation from Affective Speech.” 2020 IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP). IEEE, 2020. (Year: 2020). [cited by examiner]
Pham, Hai X., Yuting Wang, and Vladimir Pavlovic. “Learning continuous facial actions from speech for real-time animation.” IEEE Transactions on Affective Computing 13.3 (2020): 1567-1580. (Year: 2020). [cited by examiner]
Patamia, Rutherford Agbeshi, et al. “Transformer based multimodal speech emotion recognition with improved neural networks.” 2021 IEEE 2nd International Conference on Pattern Recognition and Machine Learning (PRML). IEE… [cited by examiner]
International Search Report and Written Opinion issued in PCT Application No. PCT/RU2022/000220, dated Apr. 6, 2023. [cited by applicant]