IP Library Granted Patent US 12,573,370
Granted Patent B2
US 12,573,370 · App. 17/984,590 · Granted Mar 10, 2026

Synthetic speech generation

Inventors: Subhankar Ghosh (Santa Clara, CA); Boris Ginsburg (Sunnyvale, CA)
Assignee: NVIDIA Corporation
G10L13/047G10L13/08G10L25/18G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,370
App. No.
17/984,590
Granted
Mar 10, 2026
Kind
B2
Abstract

Disclosed are apparatuses, systems, and techniques that may use machine learning for generating artificial speech. The techniques include obtaining a synthetic embedding using learned embeddings associated with different speakers. At least one learned embedding may be generated using a multi-stage training of a machine learning model (MLM) with progressively increasing quality of training speech utterances. The techniques may further include using the MLM and the synthetic embedding to generate synthetic audio data.

Claims (59)

1 . A method comprising:

obtaining a synthetic embedding using two or more learned embeddings associated with different speakers, at least one of the two or more learned embeddings being generated using a multi-stage training of a machine learning model (MLM) that was based at least on:

a first plurality of training utterances of a first quality during a first stage of the multi-stage training; and

a second plurality of training utterances of a second quality during a second stage of the multi-stage training, the second quality being higher than the first quality; and

generating audio data corresponding to a text representation based at least on the MLM processing the text representation and the synthetic embedding.

2 . The method of claim 1 , wherein the first quality of the first plurality of training utterances is characterized by a lower signal-to-noise ratio than the second quality of the second plurality of training utterances, wherein the first plurality of training utterances are associated with a first plurality of speakers and the second plurality of training utterances are associated with a second plurality of speakers, a number of the first plurality of speakers being larger than a number of the second plurality of speakers.

3 . The method of claim 1 , wherein the synthetic embedding is obtained, at least, by computing a weighted combination of the two or more learned embeddings.

4 . The method of claim 3 , wherein weights in the weighted combination of the two or more learned embeddings are selected randomly.

5 . The method of claim 1 , wherein the MLM comprises at least one transformer neural subnetwork with one or more attention layers.

6 . The method of claim 1 , wherein the MLM comprises:

a first subnetwork to associate units of the audio data with respective units of the text representation; and

a second subnetwork to determine durations of the units of the audio data.

7 . The method of claim 6 , wherein the first subnetwork and the second subnetwork comprise one or more convolutional layers and one or more fully connected layers.

8 . The method of claim 1 , wherein the text representation comprises a text embedding, and the text embedding is applied to the MLM in combination with the synthetic embedding.

9 . A method comprising:

obtaining a plurality of sets of training data, two or more sets of training data of the plurality of sets of training data being associated with a different audio quality (AQ) index characterizing audio quality of a corresponding set of training data, at least one set of training data of the plurality of sets of training data comprising:

a training input comprising a batch of text representations, and

a target output comprising a batch of audio data; and

training a machine learning model (MLM) using a plurality of training stages, at least one training stage of the plurality of training stages comprising applying the at least one set of training data to the MLM to generate learned embeddings corresponding to respective speakers associated with the at least one set of training data.

10 . The method of claim 9 , wherein at least one of:

the plurality of training stages are performed in an order of decreasing number of speakers associated with the at least one set of the training data; or

the plurality of training stages are performed in an order of increasing AQ index associated with the at least one set of training data.

11 . The method of claim 9 , wherein the MLM comprises at least one transformer neural subnetwork having one or more attention layers.

12 . The method of claim 9 , wherein one or more of the plurality of training stages comprise:

selecting a text representation from the batch of text representations of the training input for a corresponding training stage of the one or more training stages;

selecting audio data from the batch of audio data of the target output for the corresponding training stage;

training a first subnetwork of the MLM to associate units of the selected audio data with correct units of the selected text representation; and

training a second subnetwork of the MLM to determine duration of the units of the selected audio data.

13 . The method of claim 12 , wherein the units of the selected audio data comprise speech spectrograms.

14 . The method of claim 12 , wherein the first subnetwork and the second subnetwork comprise one or more convolutional layers of neurons and one or more fully connected layers of neurons.

15 . The method of claim 14 , wherein the one or more training stages of the plurality of training stages further comprise:

obtaining an output of the MLM comprising synthetic audio data generated for the selected text representation, the target speaker ID, and the embedding for the target speaker; and modifying the embedding for the target speaker based on a difference between the synthetic audio data and the selected audio data.

16 . The method of claim 9 , wherein one or more training stages of the plurality of training stages comprise:

selecting a text representation from the batch of text representations of the training input for a corresponding training stage of the one or more training stages;

selecting audio data from the batch of audio data of the target output for the corresponding training stage;

obtaining a target speaker identification (ID) identifying a target speaker associated with the selected audio data; and

applying, to the MLM, at least:

the selected text representation,

the target speaker ID, and

an embedding for the target speaker.

17 . The method of claim 16 , wherein the one or more training stages of the plurality of training stages further comprise:

obtaining an output of the MLM comprising synthetic audio data generated for the selected text representation, the target speaker ID, and the embedding for the target speaker; and modifying parameters of the MLM based on a difference between the synthetic audio data and the selected audio data.

18 . A system comprising:

one or more processing units to cause presentation of synthetic speech generated based at least on one or more machine learning models (MLMs) processing a synthetic embedding and an associated textual representation, the synthetic embedding generated based at least on combining two or more learned embeddings corresponding to two or more different speakers, wherein at least one MLM of the one or more MLMs is trained using a multi-stage training process where respective stages include different audio quality (AQ) indexes associated with respective sets of training data corresponding to the respective stage.

19 . The system of claim 18 , wherein the system is comprised in at least one of:

an in-vehicle infotainment system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets;

a system for performing deep learning operations;

a system implemented using an edge device;

a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content;

a system implemented using a robot;

a system for performing conversational AI operations;

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2022
From: GHOSH, SUBHANKAR; GINSBURG, BORIS
To: NVIDIA CORPORATION
Reel/Frame 061795/0792 →
Continuity (1)
Related Publication 20240161728A1 · May 16, 2024
References Cited (90)
US 11158307B1 · Ghias · 2021 [cited by examiner]
US 11605388B1 · Gupta · 2023 [cited by examiner]
US 11900947B2 · Ghaemmaghami et al. · 2024 [cited by applicant]
US 12100383B1 · Ezzerg et al. · 2024 [cited by applicant]
US 12233338B1 · Zhong · 2025 [cited by examiner]
US 20080270133A1 · Tian · 2008 [cited by examiner]
US 20180211649A1 · Li · 2018 [cited by examiner]
US 20200058290A1 · Chae et al. · 2020 [cited by applicant]
US 20200302223A1 · Dutta · 2020 [cited by examiner]
US 20200394997A1 · Trueba · 2020 [cited by examiner]
US 20200402497A1 · Semenov · 2020 [cited by examiner]
US 20210142782A1 · Wolf · 2021 [cited by examiner]
US 20210174791A1 · Shen · 2021 [cited by examiner]
US 20210224319A1 · Ingel et al. · 2021 [cited by applicant]
US 20210256960A1 · Daido · 2021 [cited by examiner]
US 20210304769A1 · Ye · 2021 [cited by examiner]
US 20210390668A1 · Ren · 2021 [cited by examiner]
US 20220013106A1 · Deng · 2022 [cited by examiner]
US 20220028367A1 · Shekhar et al. · 2022 [cited by applicant]
US 20220029863A1 · Wu · 2022 [cited by examiner]
US 20220051654A1 · Finkelstein · 2022 [cited by examiner]
US 20220068256A1 · Jia · 2022 [cited by examiner]
US 20220068257A1 · Biadsy · 2022 [cited by examiner]
US 20220084166A1 · Navarrete Michelini · 2022 [cited by examiner]
US 20220092274A1 · Arivazhagan · 2022 [cited by examiner]
US 20220093106A1 · Mosayyebpour Kaskari et al. · 2022 [cited by applicant]
US 20220147876A1 · Dalli · 2022 [cited by examiner]
US 20220223144A1 · Sun et al. · 2022 [cited by applicant]
US 20220319018A1 · Gervais et al. · 2022 [cited by applicant]
US 20220343903A1 · Mostafazadeh · 2022 [cited by examiner]
US 20230011337A1 · Qian · 2023 [cited by examiner]
US 20230107450A1 · Chang · 2023 [cited by examiner]
US 20230137652A1 · Khoury · 2023 [cited by examiner]
US 20230150498A1 · Seong et al. · 2023 [cited by applicant]
US 20230206898A1 · Stanton · 2023 [cited by examiner]
US 20230386506A1 · Shor · 2023 [cited by examiner]
US 20230410789A1 · Sharma · 2023 [cited by examiner]
US 20230410814A1 · Yin · 2023 [cited by examiner]
US 20240037316A1 · Mohanty et al. · 2024 [cited by applicant]
US 20240046161A1 · Pham · 2024 [cited by examiner]
US 20240104055A1 · McAnallen · 2024 [cited by applicant]
US 20240105289A1 · Khan et al. · 2024 [cited by applicant]
US 20240170007A1 · Qian · 2024 [cited by examiner]
CN 111696572A · 2020 [cited by applicant]
JP 2015102667A · 2015 [cited by applicant]
WO 2022063758A1 · 2022 [cited by applicant]
Garcia-Romero, D. et al., “Speaker Diarization Using Deep Neural Network Embeddings”, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4930-4934, Mar. 5, 2017. [cited by applicant]
Han, W. et al., “ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context”, arXiv preprint arXiv:2005.03191 (2020). [cited by applicant]
Kingma, D. et al., “Improved Variational Inference with Inverse Autoregressive Flow”, 30th Conference on Neural Information Processing Systems, 9 pages, 2016. [cited by applicant]
Koluguri, N. et al., “Titanet: Neural Model for Speaker Representation With 1D Depth-Wise Separable Convolutions and Global Context”, International Conference on Acoustics, Speech and Signal Processing, pp. 8102-8106, M… [cited by applicant]
Lancucki, Adrian, “Fastpitch: Parallel Text-to-Speech with Pitch Prediction”, In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6588-6592. IEEE, 2021. [cited by applicant]
Muller, T. et al., “Neural Importance Sampling”, ACM Transactions on Graphics (ToG), vol. 38, No. 5, 2019, pp. 1-19. [cited by applicant]
Park, T. et al., “Auto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengap”, IEEE Signal Processing Letters, vol. 27, pp. 381-385, 2019. [cited by applicant]
Park, T. et al., “Multi-Scale Speaker Diarization with Neural Affinity Score Fusion”, In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7173-7177. IEEE, 2021. [cited by applicant]
Peng, K. et al., “Non-Autoagressive Neural Text-to-Speech”, International Conference on Machine Learning, pp. 7586-7598, 2020. [cited by applicant]
Ren, Y. et al., “FastSpeech: Fast, Robust and Controllable Text to Speech”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 13 pages. [cited by applicant]
Sasirekha, D. et al., “Text-to-Speech: A Simple Tutorial”, International Journal of Soft Computing and Engineering, vol. 2, Issue 1, 4 pages, Mar. 2012. [cited by applicant]
Shih, K. et al., “Generative Modeling for Low Dimensional Speech Attributes with Neural Spline Flows”, arXiv preprint arXiv:2203.01786 (2022). [cited by applicant]
Valle, R., “Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Sythesis”, arXiv preprint arXiv:2005.05957. May 12, 2020. [cited by applicant]
Chen, J., et al., “VFlow: More Expressive Generative Flows with Variational Data Augmentation,” Proceedings of the 37th International Conference on Machine Learning, PMLR, 2020, vol. 119, pp. 1660-1669. Retrived from th… [cited by applicant]
Doty C., et al., “What Is Speaker Diarization,” Deepgram, Aug. 2022, 9 Pages Retrieved from internet URL: https://deepgram.com/learn/what-is-speaker-diarization. [cited by applicant]
Dumoulin, V., et al., “A Learned Representation for Artistic Style,” Computer Vision and Pattern Recognition, ICLR, 2017, 26 Pages. Retrived from the Internet: [https://arxiv.org/abs/1610.07629]. [cited by applicant]
Dupont, E., et al., “Augmented Neural ODEs,” 33rd Conference on Neural Information Processing Systems, NeurIPS, 2019, vol. 32, pp. 3140-3150. [cited by applicant]
Durkan, C., et al., “Neural Spline Flows,” Advances in Neural Information Processing Systems, 2019, vol. 32, pp. 7511-7522. [cited by applicant]
Huang, C.W., et al., “Augmented Normalizing Flows: Bridging the Gap Between Generative Flows and Latent Variable Models,” Machine Learning, 2022, 27 Pages. Retrived from the Internet: [https://arxiv.org/abs/2002.07101]. [cited by applicant]
Ito, K., et al., “The LJ Speech Dataset,” BibSonomy, 2017. Retrived from the Internet: [https://keithito.com/LJ-Speech-Dataset/]. [cited by applicant]
Jeong, M., et al., “Diff-TTS: A Denoising Diffusion Model for Text-to-Speech,” Department of Electrical and Computer Engineering and INMC, 2021, 5 Pages. Retrived from the Internet: [https://arxiv.org/pdf/2104.01409.pdf… [cited by applicant]
Kawahara, H., et al., “Nearly Defect-free Fo Trajectory Extraction for Expressive Speech Modifications Based on Straight,” In Ninth European Conference on Speech Communication and Technology, 2005, 5 Pages. [cited by applicant]
Kim, H., et al., “Softflow: Probabilistic Framework for Normalizing Flow on Manifolds,” Advances in Neural Information Processing Systems, 2020, 10 Pages. [cited by applicant]
Kim, J., et al., “Glow-TTS: A Generative Flow for Text-to-speech via Monotonic Alignment Search,” Advances in Neural Information Processing Systems, 2020b, 11 Pages. [cited by applicant]
Kingma, D.P., et al., “Glow: Generative flow with invertible 1×1 Convolutions,” Advances in Neural Information Processing Systems, 2018, vol. 31, 10 Pages. [cited by applicant]
Kong, J., et al., “Hifi-gan: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,” Advances in Neural Information Processing Systems, 2020, pp. 1-14. [cited by applicant]
Kong, J., et al., “Hifi-gan: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,” 2022, Retrived from the Internet: [https://github. com/jik876/hifi-gan,]. [cited by applicant]
Mauch, M., et al., “Pyin: A Fundamental Frequency Estimator Using Probabilistic Threshold Distributions,” In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014, pp. 659-663. [cited by applicant]
Mccree, A.V., et al., “A Mixed Excitation Ipc Vocoder Model for Low Bit Rate Speech Coding,” IEEE Transactions on Speech and Audio Processing, 1995, vol. 3(4), pp. 242-250. [cited by applicant]
Miao, C., et al., “Flow-TTS: A Non-autoregressive Network for Text to Speech Based on Flow,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, 5 Pages. [cited by applicant]
Moog, B., “MIDI: Musical Instrument Digital Interface,” Journal of the Audio Engineering Society, 1986, vol. 34(5), pp. 394-404. [cited by applicant]
Nakatani, T., et al., “A Method for Fundamental Frequency Estimation and Voicing Decision: Application to Infant Utterances Recorded in Real Acoustical Environments,” Speech Communication, Mar. 2008, vol. 50(3), pp. 203… [cited by applicant]
Park T.J., et al., “Multi-Scale Speaker Diarization With Neural Affinity Score Fusion.” IEEE, May 2021, pp. 7173-7177 Retrieved from Internet URL: https://ieeexplore.ieee.org/stamp/stamp.jsptp=&arnumber=9414578. [cited by applicant]
Ping, W., et al., “Waveflow: A Compact Flow-based Model for Raw Audio,” In International Conference on Machine Learning, pp. 7706-7716. [cited by applicant]
Popov, V., et al., “Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech,” Arxiv Preprint, 2021, 10 Pages. [cited by applicant]
Ren, Y., et al., “FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,” Aug. 8, 2022, 15 Pages. Retrived from the Internet: [https://speechresearch.github.io/fastspeech2/]. [cited by applicant]
Ren, Y., et al., “FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,” arXiv:2006.04558v1 [eess.AS] Jun. 8, 2020. [cited by applicant]
Shih, K.J., et al., “RAD-TTS: Parallel Flowbased Tts With Robust Alignment Learning and Diverse Synthesis,” In ICML Workshop on Invertible Neural Networks, Normalizing Flows, and Explicit Likelihood Models, 2021. [cited by applicant]
Suni, et al., Wavelets for Intonation Modeling in Hmm Speech Synthesis, ISCA, 2013, 7 Pages. [cited by applicant]
Valle, R., et al., “Flowtron: An Autoregressive Flow-based Generative Network for Text-to-speech Synthesis,” In International Conference on Learning Representations, 2020b. [cited by applicant]
Valle, R., et al., “Mellotron: Multispeaker Expressive Voice Synthesis by Conditioning on Rhythm, Pitch and Global Style Tokens,” In ICASSP 2020-2020 Ieee International Conference on Acoustics, Speech and Signal Process… [cited by applicant]
Wang, Q., et al., “Speaker Diarization With LSTM,” ArXiv, Jan. 2022, 5 Pages Retrieved from internet URL: https://arxiv.org/pdf/1710.10468. [cited by applicant]
Zhang Y., et al., “Audio segmentation based on multi-scale audio classification.” IEEE, Aug. 2004, vol. 4, pp. 349-352 Retrieved from Internet URL: https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=1326835&am… [cited by applicant]
Raj D., et al., “Dover-Lap: A Method for Combining Overlap-Aware Diarization Outputs,” arXiv, Nov. 2020, 8 Pages. [cited by applicant]