IP Library › Granted Patent US 12,340,788
Granted Patent B2
US 12,340,788 · App. 18/656,983 · Granted Jun 24, 2025

Generating expressive speech audio from text data

Inventors: Siddharth Gururani (Santa Clara, CA); Kilol Gupta (Redwood City, CA); Dhaval Shah (Redwood City, CA); Zahra Shakeri (Newark, CA); Jervis Pinto (Toronto, CA); Mohsen Sardari (Burlingame, CA); Navid Aghdaie (San Jose, CA); Kazi Zaman (Foster City, CA)
Assignee: ELECTRONIC ARTS INC.
G10L13/00A63F13/60G06N3/044G06N3/08A63F13/63A63F2300/6018
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,788
App. No.
18/656,983
Granted
Jun 24, 2025
Kind
B2
Abstract

A system for use in video game development to generate expressive speech audio comprises a user interface configured to receive user-input text data and a user selection of a speech style. The system includes a machine-learned synthesizer comprising a text encoder, a speech style encoder and a decoder. The machine-learned synthesizer is configured to generate one or more text encodings derived from the user-input text data, using the text encoder of the machine-learned synthesizer; generate a speech style encoding by processing a set of speech style features associated with the selected speech style using the speech style encoder of the machine-learned synthesizer; combine the one or more text encodings and the speech style encoding to generate one or more combined encodings; and decode the one or more combined encodings with the decoder of the machine-learned synthesizer to generate predicted acoustic features. The system includes one or more modules configured to process the predicted acoustic features, the one or more modules comprising a machine-learned vocoder configured to generate a waveform of the expressive speech audio.

Claims (40)

1. A system for use in video game development for generating a speech audio waveform, the system comprising:

a user interface configured to receive user-input text data and a user selection of a speech style; and

a machine-learned synthesizer comprising a text encoder, a speech style encoder and a decoder, the machine-learned synthesizer being configured to:

generate one or more text encodings derived from the user-input text data, using the text encoder of the machine-learned synthesizer;

generate a speech style encoding by processing a set of speech style features associated with the selected speech style using the speech style encoder of the machine-learned synthesizer;

combine the one or more text encodings and the speech style encoding to generate one or more combined encodings; and

decode the one or more combined encodings with the decoder of the machine-learned synthesizer to generate a compressed representation of a speech audio waveform.

2. The system of claim 1 , wherein the compressed representation comprises a sequence of vectors, each vector representing acoustic information for a respective time period.

3. The system of claim 1 , wherein the compressed representation comprises one or a combination of a log-mel spectrogram representation, a Mel-Frequency Cepstral Coefficients (MFCC) representation, a log fundamental frequency (LFO) representation, or a band aperiodicity (bap) representation.

4. The system of claim 1 , wherein the compressed representation comprises spectrogram magnitudes.

5. The system of claim 4 , wherein the spectrogram magnitudes comprise linear spectrogram magnitudes or log-transformed mel-spectrogram magnitudes.

6. The system of claim 1 , wherein the set of speech style features comprises prosodic features determined from the selected speech style.

7. The system of claim 1 , wherein the compressed representation comprises a representation of one or more of frequency, magnitude or phase.

8. The system of claim 1 , wherein:

the user interface is further configured to receive a user selection of an instance of speech audio; and

the system further comprises a prosody analyzer configured to process the selected instance of speech audio to determine the prosodic features.

9. The system of claim 1 , wherein:

the user interface is further configured to receive a user selection of speaker attribute information; and

the set of speech style features further comprises the speaker attribute information.

10. The system of claim 1 further comprising a machine-learned vocoder configured to generate the speech audio waveform.

11. A computer-implemented method comprising:

receiving user-input text data and a user selection of a speech style;

generating one or more text encodings derived from the user-input text data, using a text encoder of a machine-learned synthesizer;

generating a speech style encoding by processing a set of speech style features associated with the selected speech style using a speech style encoder of the machine-learned synthesizer;

combining one or more text encodings and the speech style encoding to generate one or more combined encodings; and

decoding the one or more combined encodings with a decoder of the machine-learned synthesizer to generate a compressed representation of a speech audio waveform.

12. The method of claim 11 , wherein the compressed representation comprises a sequence of vectors, each vector representing acoustic information for a respective time period.

13. The method of claim 11 , wherein the compressed representation comprises one or a combination of a log-mel spectrogram representation, a Mel-Frequency Cepstral Coefficients (MFCC) representation, a log fundamental frequency (LFO) representation, or a band aperiodicity (bap) representation.

14. The method of claim 11 , wherein the compressed representation comprises linear spectrogram magnitudes or log-transformed mel-spectrogram magnitudes.

15. The method of claim 11 , wherein the compressed representation comprises a representation of one or more of frequency, magnitude or phase.

16. A non-transitory computer readable medium storing instructions, which when executed by one or more processors, cause one or more of the processors to:

receive user-input text data and a user selection of a speech style;

generate one or more text encodings derived from the user-input text data, using a text encoder of a machine-learned synthesizer;

generate a speech style encoding by processing a set of speech style features associated with the selected speech style using a speech style encoder of the machine-learned synthesizer;

combine one or more text encodings and the speech style encoding to generate one or more combined encodings; and

decode the one or more combined encodings with a decoder of the machine-learned synthesizer to generate a compressed representation of a speech audio waveform.

17. The non-transitory computer readable medium of claim 16 , wherein the compressed representation comprises a sequence of vectors, each vector representing acoustic information for a respective time period.

18. The non-transitory computer readable medium of claim 16 , wherein the compressed representation comprises one or a combination of a log-mel spectrogram representation, a Mel-Frequency Cepstral Coefficients (MFCC) representation, a log fundamental frequency (LFO) representation, or a band aperiodicity (bap) representation.

19. The non-transitory computer readable medium of claim 16 , wherein the compressed representation comprises spectrogram magnitudes.

20. The non-transitory computer readable medium of claim 16 , wherein the spectrogram magnitudes comprise linear spectrogram magnitudes or log-transformed mel-spectrogram magnitudes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2024
From: GURURANI, SIDDHARTH; GUPTA, KILOL; SHAH, DHAVAL; SHAKERI, ZAHRA; PINTO, JERVIS; SARDARI, MOHSEN; AGHDAIE, NAVID; ZAMAN, KAZI
To: ELECTRONIC ARTS INC.
Reel/Frame 067351/0886 →
Continuity (4)
Continuation 17682206 · Feb 28, 2022
Continuation 16840070 · Apr 3, 2020
Provisional Application 62936249 · Nov 15, 2019
Related Publication 20240290316A1 · Aug 29, 2024
References Cited (75)
US 5729694A · Holzrichter · 1998 [cited by applicant]
US 7113909B2 · Nukaga · 2006 [cited by examiner]
US 7249021B2 · Morio · 2007 [cited by applicant]
US 7831420B2 · Sinder · 2010 [cited by applicant]
US 8527276B1 · Senior · 2013 [cited by examiner]
US 9159329B1 · Agiomyrgiannakis · 2015 [cited by applicant]
US 9697820B2 · Jeon · 2017 [cited by applicant]
US 9905220B2 · Fructuoso · 2018 [cited by applicant]
US 9916825B2 · Edrenkin · 2018 [cited by applicant]
US 9934775B2 · Raitio · 2018 [cited by applicant]
US 10192541B2 · Mairano · 2019 [cited by applicant]
US 10249289B2 · Chun · 2019 [cited by applicant]
US 10431188B1 · Nelson · 2019 [cited by applicant]
US 10692484B1 · Merritt · 2020 [cited by applicant]
US 10699695B1 · Nadolski · 2020 [cited by applicant]
US 10706837B1 · Chicote · 2020 [cited by applicant]
US 10741169B1 · Trueba · 2020 [cited by applicant]
US 10902841B2 · Liu · 2021 [cited by applicant]
US 10911596B1 · Do · 2021 [cited by applicant]
US 11017761B2 · Peng · 2021 [cited by applicant]
US 11069335B2 · Pollet · 2021 [cited by applicant]
US 11488576B2 · Park · 2022 [cited by examiner]
US 11715485B2 · Park · 2023 [cited by applicant]
US 20020188449A1 · Nukaga · 2002 [cited by examiner]
US 20040054537A1 · Morio · 2004 [cited by applicant]
US 20070233472A1 · Sinder · 2007 [cited by examiner]
US 20070298886A1 · Aguilar, Jr. · 2007 [cited by applicant]
US 20150186359A1 · Fructuoso · 2015 [cited by examiner]
US 20160093289A1 · Pollet · 2016 [cited by examiner]
US 20160140951A1 · Agiomyrgiannakis · 2016 [cited by examiner]
US 20170092259A1 · Jeon · 2017 [cited by examiner]
US 20170110110A1 · Pollet · 2017 [cited by applicant]
US 20180096677A1 · Pollet · 2018 [cited by applicant]
US 20180268806A1 · Chun et al. · 2018 [cited by applicant]
US 20190108830A1 · Pollet · 2019 [cited by applicant]
US 20190189109A1 · Yuan · 2019 [cited by applicant]
US 20200005764A1 · Chae · 2020 [cited by examiner]
US 20200066253A1 · Peng · 2020 [cited by applicant]
US 20200265829A1 · Liu · 2020 [cited by applicant]
US 20200288014A1 · Huet · 2020 [cited by applicant]
US 20210097976A1 · Chicote · 2021 [cited by applicant]
US 20210151029A1 · Gururani · 2021 [cited by applicant]
US 20210193112A1 · Cui · 2021 [cited by applicant]
US 20220208170A1 · Gururani · 2022 [cited by applicant]
EP 3376497A1 · 2018 [cited by applicant]
Arik, Sercan O., et al. “Neural Voice Cloning with a Few Samples.” arXiv preprint arXiv:1802.06006, Retrieved from https://arxiv.org/abs/1802.06006v3, dated 2018. [cited by applicant]
Jia, Ye, et al. “Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis.” arXiv preprint arXiv:1806.04558, Retrieved from https://arxiv.org/abs/1806.04558, dated 2018. [cited by applicant]
Polyak, Adam, et al. “TTS skins: Speaker conversion via ASR.” arXiv preprint arXiv:1904.08983, Retrieved from https://arxiv.org/abs/1904.08983, dated 2019. [cited by applicant]
Sun, Lifa, et al. “Voice conversion using deep bidirectional long short-term memory based recurrent neural networks.” 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2015. ,… [cited by applicant]
Kameoka, Hirokazu, et al. “StarGAN-VC: Non-parallel many-to-many voice conversion with star generative adversarial networks.” arXiv preprint arXiv:1806.02169, Retrieved from https://arxiv.org/abs/1806.02169, dated 2018. [cited by applicant]
Liu, Li-Juan, et al. “WaveNet Vocoder with Limited Training Data for Voice Conversion.” Interspeech, Retrieved from https://www.isca-speech.org/archive/Interspeech_2018/pdfs/1190.pdf, dated 2018. [cited by applicant]
Narayanan, Praveen, et al. “Hierarchical sequence to sequence voice conversion with limited data.” arXiv preprint arXiv:1907.07769, Retrieved from https://arxiv.org/abs/1907.07769, dated 2019. [cited by applicant]
Zhang, Mingyang, et al. “Joint training framework for text-to-speech and voice conversion using multi-source tacotron and wavenet.” arXiv preprint arXiv:1903.12389 , Retrieved from https://arxiv.org/abs/1903.12389, date… [cited by applicant]
Huang, Wen-Chin, et al. “Voice transformer network: Sequence-to-sequence voice conversion using transformer with text-to-speech pretraining.” arXiv preprint arXiv:1912.06813, Retrieved from https://arxiv.org/abs/1912.06… [cited by applicant]
Luong, Hieu-Thi, and Junichi Yamagishi. “Bootstrapping non-parallel voice conversion from speaker-adaptive text-to-speech.” arXiv preprint arXiv:1909.06532, Retrieved from https://arxiv.org/abs/1909.06532, dated 2019. [cited by applicant]
Kim, Tae-Ho, et al. “Emotional Voice Conversion using multitask learning with Text-to-speech.” arXiv preprint arXiv:1911.06149 Retrieved from: https://arxiv.org/abs/1911.06149, dated 2019. [cited by applicant]
Ren, Yi, et al. “FastSpeech 2: Fast and High-Quality End-to-End Text-to-Speech.” arXiv preprint arXiv:2006.04558, Retrieved from: https://arxiv.org/abs/2006.04558, dated 2020. [cited by applicant]
Skerry-Ryan, R. J., et al. “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron.” arXiv preprint arXiv: 1803.09047, Retrieved from: https://arxiv.org/abs/1803.09047, dated 2018. [cited by applicant]
Theune, Mariët, et al. “Generating expressive speech for storytelling applications.” IEEE Transactions on Audio, Speech, and Language Processing 14.4: 1137-1144., Retrieved from: https://ieeexplore.ieee.org/abstract/doc… [cited by applicant]
Gibiansky, Andrew, et al. “Deep voice 2: Multi-speaker neural text-to-speech.” Advances in neural information processing systems, Retrieved from: http://papers.nips.cc/paper/6889-deep-voice-2-multi-speaker-neural-text-t… [cited by applicant]
Shen, Jonathan, et al. “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions.” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, Retrieved from: https:… [cited by applicant]
Wu, Xixin, et al. “Rapid Style Adaptation Using Residual Error Embedding for Expressive Speech Synthesis.” Interspeech, Retrieved from: http://www1.se.cuhk.edu.hk/˜hccl/publications/pub/2018_201809_INTERSPEECH_XixinWU.p… [cited by applicant]
Zen, Heiga, et al. “The HMM-based speech synthesis system (HTS) version 2.0.” SSW., Retrieved from: https://www.cs.cmu.edu/afs/cs/Web/People/awb/papers/ssw6/ssw6_294.pdf, 6 pages, dated 2007. [cited by applicant]
Kalchbrenner, Nal, et al. “Efficient Neural Audio Synthesis.” International Conference on Machine Learning, Retrieved from: http://proceedings.mlr.press/v80/kalchbrenner18a.html, dated 2018. [cited by applicant]
Eyben, Florian, et al. “Unsupervised clustering of emotion and voice styles for expressive TTS.” 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, Retrieved from: https://ieee… [cited by applicant]
Wang, Yuxuan, et al. “Tacotron: Towards end-to-end speech synthesis.” arXiv preprint arXiv:1703.10135, Retrieved from: https://arxiv.org/abs/1703.10135, dated 2017. [cited by applicant]
Wang, Yuxuan, et al. “Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis.” ICML., Retrieved from: https://openreview.net/forum?id=Hy4s90-OZr, dated 2018. [cited by applicant]
Akuzawa, Kei, Yusuke Iwasawa, and Yutaka Matsuo. “Expressive Speech Synthesis via Modeling Expressions with Variational Autoencoder.” Proc. Interspeech 2018: 3067-3071, Retrieved from: https://www.isca-speech.org/archiv… [cited by applicant]
Taigman, Yaniv, et al. “VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop.” International Conference on Learning Representations, Retrieved from: https://openreview.net/forum?id=SkFAWax0-&noteld=SkFAWax0-, … [cited by applicant]
Hsu, Wei-Ning, et al. “Hierarchical Generative Modeling for Controllable Speech Synthesis.” International Conference on Learning Representations.Retrieved from: https://openreview.net/forum?id=rygkk305YQ, dated 2018. [cited by applicant]
Kenter, Tom, et al. “CHiVE: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network.” International Conference on Machine Learning, Retrieved from: http://pr… [cited by applicant]
Lee, Younggun, and Taesu Kim. “Robust and fine-grained prosody control of end-to-end speech synthesis.” ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, Retrieved… [cited by applicant]
Klimkov, Viacheslav, et al. “Fine-Grained Robust Prosody Transfer for Single-Speaker Neural Text-To-Speech.” Proc. Interspeech 2019: 4440-4444, Retrieved from: https://www.isca-speech.org/archive/Interspeech_2019/abstra… [cited by applicant]
Daniel, Povey, et al. “The Kaldi speech recognition toolkit.” IEEE 2011 workshop on automatic speech recognition and understanding. No. EPFL-CONF-192584. Retrieved from: https://www.fit.vut.cz/research/product/304/, dat… [cited by applicant]
Ghahremani, Pegah, et al. “A pitch extraction algorithm tuned for automatic speech recognition.” 2014 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, Retrieved from: https://ieee… [cited by applicant]