IP Library Granted Patent US 12,236,975
Granted Patent B2
US 12,236,975 · App. 17/526,810 · Granted Feb 25, 2025

Bi-directional recurrent encoders with multi-hop attention for speech emotion recognition

Inventors: Trung Bui (San Jose, CA); Subhadeep Dey (Martigny, CH); Seunghyun Yoon (Seoul, KR)
Assignee: Adobe Inc.
G10L25/63G06F17/16G06F17/18G06N3/047G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,975
App. No.
17/526,810
Granted
Feb 25, 2025
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer readable media for determining speech emotion. In particular, a speech emotion recognition system generates an audio feature vector and a textual feature vector for a sequence of words. Further, the speech emotion recognition system utilizes a neural attention mechanism that intelligently blends together the audio feature vector and the textual feature vector to generate attention output. Using the attention output, which includes consideration of both audio and text modalities for speech corresponding to the sequence of words, the speech emotion recognition system can apply attention methods to one of the feature vectors to generate a hidden feature vector. Based on the hidden feature vector, the speech emotion recognition system can generate a speech emotion probability distribution of emotions among a group of candidate emotions, and then select one of the candidate emotions as corresponding to the sequence of words.

Claims (47)

1. A system comprising:

one or more memory devices comprising:

an audio bi-directional recurrent encoder that generates an audio feature vector for one or more words in an acoustic sequence;

a textual bi-directional recurrent encoder that generates a textual feature vector for the one or more words in a textual sequence corresponding to the acoustic sequence;

a multi-hop neural attention model that generates an attention output at each hop that alternates from utilizing the textual feature vector and the audio feature vector as context; and

a hidden feature vector generator that generates a hidden feature vector based on the attention output and one or more of the audio feature vector and the textual feature vector; and

one or more processors configured to cause the system to determine an emotion of the acoustic sequence based on the hidden feature vector.

2. The system of claim 1 , wherein the multi-hop neural attention model further comprises a first-hop attention model that generates a first attention output based on the audio feature vector at a final state of the acoustic sequence and the textual feature vector at each state of the textual sequence.

3. The system of claim 2 , wherein the hidden feature vector generator generates the hidden feature vector as a first-hop hidden feature vector based on the textual feature vector at each state of the textual sequence and the first attention output.

4. The system of claim 1 , wherein the multi-hop neural attention model further comprises a second-hop attention model that generates a second attention output based on the audio feature vector at each state of the acoustic sequence.

5. The system of claim 4 , wherein the second-hop attention model generates the second attention output based on the hidden feature vector.

6. The system of claim 4 , wherein the hidden feature vector generator generates an additional hidden feature vector as a second-hop hidden feature vector based on the second attention output and the audio feature vector at each state of the acoustic sequence.

7. The system of claim 6 , wherein the multi-hop neural attention model further comprises a third-hop attention model that generates a third attention output based on the textual feature vector at each state of the textual sequence and the second-hop hidden feature vector.

8. The system of claim 7 , wherein the hidden feature vector generator generates another hidden feature vector as a third-hop hidden feature vector based on the third attention output and the textual feature vector at each state of the textual sequence.

9. The system of claim 1 , wherein the one or more processors are configured to cause the system to determine the emotion of the acoustic sequence based on two or more of a first-hop hidden feature vector, a second-hop hidden feature vector, or a third-hop feature vector.

10. A system comprising:

one or more memory devices comprising:

an audio encoder that generates an audio feature vector for one or more words in an acoustic sequence;

a textual encoder that generates a textual feature vector for the one or more words in a textual sequence corresponding to the acoustic sequence;

a first neural attention model that generates a first attention output by applying attention to the textual feature vector using the audio feature vector as context;

a first hidden feature vector generator that generates a first hidden feature vector based on the first attention output;

a second neural attention model that generates a second attention output by applying attention to the audio feature vector using the first hidden feature vector as context; and

a second hidden feature vector generator that generates a second hidden feature vector based on the second attention output and the audio feature vector; and

one or more processors configured to cause the system to determine an emotion of the acoustic sequence based on the first hidden feature vector and the second hidden feature vector.

11. The system of claim 10 , wherein:

the first neural attention model generates the first attention output based on the audio feature vector at a final state of the acoustic sequence and the textual feature vector at each state of the textual sequence; and

the first hidden feature vector generator generates the first hidden feature vector based on first attention output and the textual feature vector at each state of the textual sequence.

12. The system of claim 10 , wherein:

the second neural attention model generates the second attention output based on the audio feature vector at each state of the acoustic sequence; and

the second hidden feature vector generator generates the second hidden feature vector based on the audio feature vector at each state of the acoustic sequence.

13. The system of claim 10 , wherein the one or more memory devices further comprise:

a third neural attention model that generates a third attention output based on the textual feature vector at each state of the textual sequence and the second hidden feature vector; and

a third hidden feature vector generator that generates a third hidden feature vector based on the third attention output and the textual feature vector at each state of the textual sequence.

14. The system of claim 13 , wherein the one or more processors are configured to cause the system to determine the emotion of the acoustic sequence based on the third hidden feature vector.

15. The system of claim 10 , wherein:

the audio encoder comprises a bi-directional recurrent encoder; and

the textual encoder comprises a bi-directional recurrent encoder.

16. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to:

generate, utilizing an audio bi-directional recurrent encoder, an audio feature vector for one or more words in an acoustic sequence;

generate, utilizing a textual bi-directional recurrent encoder, a textual feature vector for the one or more words in a textual sequence corresponding to the acoustic sequence;

generate, utilizing a neural attention model, an attention output by applying attention to the audio feature vector using the textual feature vector as a context vector;

generate, utilizing a hidden feature vector generator, a hidden feature vector based on the attention output and the audio feature vector; and

determine an emotion of the acoustic sequence based on the hidden feature vector.

17. The non-transitory computer-readable medium of claim 16 , further comprising instructions that, when executed by the at least one processor, cause the computing device to generate, utilizing the neural attention model, an additional attention output by applying attention to the textual feature vector using the audio feature vector as a context vector.

18. The non-transitory computer-readable medium of claim 17 , further comprising instructions that, when executed by the at least one processor, cause the computing device to generate, utilizing the hidden feature vector generator, an additional hidden feature vector based on the additional attention output and the textual feature vector.

19. The non-transitory computer-readable medium of claim 18 , further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the emotion of the acoustic sequence based on the additional hidden feature vector, the emotion corresponding to an emotion category of sad, happy, angry, or neutral.

20. The non-transitory computer-readable medium of claim 16 , further comprising instructions that, when executed by the at least one processor, cause the computing device to generate a transcription of the acoustic sequence including the one or more words in the textual sequence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2021
From: BUI, TRUNG; DEY, SUBHADEEP; YOON, SEUNGHYUN
To: ADOBE INC.
Reel/Frame 058116/0687 →
Continuity (2)
Continuation 16543342 · Aug 16, 2019
Related Publication 20220076693A1 · Mar 10, 2022
References Cited (37)
US 9984682B1 · Tao et al. · 2018 [cited by applicant]
US 11205444B2 · Bui · 2021 [cited by examiner]
US 20190341025A1 · Omote et al. · 2019 [cited by applicant]
US 20190371298A1 · Hannun et al. · 2019 [cited by applicant]
S. Yoon, S. Byun, S. Dey and K. Jung, “Speech Emotion Recognition Using Multi-hop Attention Mechanism,” ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK, 2… [cited by examiner]
Kurpukdee, Nattapong, et al. “Speech emotion recognition using convolutional long short-term memory neural network and support vector machines.” 2017 Asia-Pacific Signal and Information Processing Association Annual Sum… [cited by examiner]
Sun, Licai, et al. “Multimodal cross-and self-attention network for speech emotion recognition.” ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021. (Year: 202… [cited by examiner]
Rosalind W Picard, “Affective computing: challenges,” International Journal of Human-Computer Studies, vol. 59, No. 1-2, pp. 55-64, 2003. [cited by applicant]
Carlos Busso, Murtaza Bulut, Shrikanth Narayanan, J Gratch, and S Marsella, “Toward effective automatic recognition systems of emotion in speech,” Social emotions in nature and artifact: emotions in human and human-comp… [cited by applicant]
Agata Kołakowska, Agnieszka Landowska, Mariusz Szwoch, Wioleta Szwoch, and Michal R Wrobel, “Emotion recognition and its applications,” in Human-Computer Systems Inter-action: Backgrounds and Applications 3, pp. 51-62. … [cited by applicant]
Kun Han, Dong Yu, and Ivan Tashev, “Speech emotion recognition using deep neural network and extreme learning machine,” in Fifteenth annual conference of the international speech communication association, 2014. [cited by applicant]
Bjorn Schuller, Gerhard Rigoll, and Manfred Lang, “Speech emotion recognition combining acoustic features and linguistic information in a hybrid support vector machine-belief network architecture,” in Acoustics, Speech,… [cited by applicant]
Qin Jin, Chengxin Li, Shizhe Chen, and Huimin Wu, “Speech emotion recognition with acoustic and lexical features,” in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on. IEEE, 2015, … [cited by applicant]
Seunghyun Yoon, Seokhyun Byun, and Kyomin Jung, “Multi-modal speech emotion recognition using audio and text,” arXiv preprint arXiv:1810.04635, 2018. [cited by applicant]
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan, “lemocap: Interactive emotional dyadic motion capture database,” Language re… [cited by applicant]
Thapanee Seehapoch and Sartra Wongthanavasu, “Speech emotion recognition using support vector machines,” in Knowledge and Smart Technology (KST), 2013 5th International Conference on. IEEE, 2013, pp. 86-91. [cited by applicant]
Bjorn Schuller, Gerhard Rigoll, and Manfred Lang, “Hidden markov model-based speech emotion recognition,” in Multi-media and Expo, 2003. ICME'03. Proceedings. 2003 International Conference on. IEEE, 2003, vol. 1, pp. I-… [cited by applicant]
Chi-Chun Lee, Emily Mower, Carlos Busso, Sungbok Lee, and Shrikanth Narayanan, “Emotion recognition using a hierarchical binary decision tree approach,” Speech Communication, vol. 53, No. 9-10, pp. 1162-1171, 2011. [cited by applicant]
Dario Bertero and Pascale Fung, “A first look into a convolutional neural network for speech emotion detection,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on. IEEE, 2017, pp… [cited by applicant]
Abdul Malik Badshah, Jamil Ahmad, Nasir Rahim, and Sung Wook Baik, “Speech emotion recognition from spectrograms with deep convolutional neural network,” in Platform Technology and Service (PlatCon), 2017 International … [cited by applicant]
Zakaria Aldeneh and Emily Mower Provost, “Using regional saliency for speech emotion recognition,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on. IEEE, 2017, pp. 2741-2745. [cited by applicant]
Aharon Satt, Shai Rozenberg, and Ron Hoory, “Efficient emotion recognition from speech using deep learning on spectrograms,” Proc. Interspeech 2017, pp. 1089-1093, 2017. [cited by applicant]
Pengcheng Li, Yan Song, Ian McLoughlin, Wu Guo, and Lirong Dai, “An attention pooling based representation learning method for speech emotion recognition,” Proc. Interspeech 2018, pp. 3087-3091, 2018. [cited by applicant]
Seyedmahdad Mirsamadi, Emad Barsoum, and Cha Zhang, “Automatic speech emotion recognition using recurrent neural networks with local attention,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE Internation… [cited by applicant]
Jaejin Cho, Raghavendra Pappagari, Purva Kulkarni, Jesus Villalba, Yishay Carmiel, and Najim Dehak, “Deep neural networks for emotion recognition combining audio and transcripts,” Proc. Interspeech 2018, pp. 247-251, 20… [cited by applicant]
Yun Wang, Leonardo Neves, and Florian Metze, “Audio-based multimedia event detection using deep recurrent neural networks,” in Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on. IEE… [cited by applicant]
Yiren Wang and Fei Tian, “Recurrent residual learning for sequence classification,” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 2016, pp. 938-943. [cited by applicant]
Samarth Tripathi and Homayoon Beigi, “Multi-modale motion recognition on iemocap dataset using deep learning,” arXiv preprint arXiv:1804.05788, 2018. [cited by applicant]
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., “The kaldi speech recognition toolkit,” Tech. Rep., IEEE Sig… [cited by applicant]
Florian Eyben, Felix Weninger, Florian Gross, and Bjorn Schuller, “Recent developments in opensmile, the munich open-source multimedia feature extractor,” in Proceedings of the 21st ACM international conference on Multi… [cited by applicant]
Google, “Cloud speech-to-text,” http://cloud.google.com/speech-to-text/, 2018. [cited by applicant]
Michael Neumann and Ngoc Thang Vu, “Attentive convolutional neural network based speech emotion recognition: A study on the impact of input features, signal length, and acted speech,” Proc. Interspeech 2017, pp. 1263-12… [cited by applicant]
Diederik Kingma and Jimmy Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980, 2014. [cited by applicant]
Lee, J. & Tashev, I. High-level feature representation using recurrent neural network for speech emotion recognition. In Sixteenth Annual Conference of the International Speech Communication Association (2015). [cited by applicant]
Lorenzo-Trueba, J., Henter, G. E., Takaki, S., Yamagishi, J., Morino, Y., & Ochiai, Y. (2018). Investigating different representations for modeling and controlling multiple emotions in DNN-based speech synthesis. Speech… [cited by applicant]
U.S. Appl. No. 16/543,342, Apr. 15, 2021, Office Action. [cited by applicant]
U.S. Appl. No. 16/543,342, Aug. 13, 2021, Notice of Allowance. [cited by applicant]