IP Library › Granted Patent US 12,451,120
Granted Patent B2
US 12,451,120 · App. 17/927,865 · Granted Oct 21, 2025

Speech synthesis and speech recognition

Inventors: Xu Tan (Redmond, WA); Tao Qin (Beijing, CN); Junwei Gan (Redmond, WA); Sheng Zhao (Redmond, WA); Tieyan Liu (Beijing, CN)
Assignee: Microsoft Technology Licensing, LLC
G10L15/063G10L13/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,451,120
App. No.
17/927,865
Granted
Oct 21, 2025
Kind
B2
Abstract

Implementations of the subject matter described herein provide a solution for speech synthesis and speech recognition. In this solution, a Text to Speech (TTS) model and an Automatic Speech Recognition (ASR) model supporting at least one language are obtained. The TTS model and the ASR model are adjusted, based on a first set of paired data in a target language, to support the target language. The TTS model is optimized based on the first set of paired data and a first set of synthesized paired data in the target language generated by the ASR model while the ASR model is optimized based on the first set of paired data and a second set of synthesized paired data in the target language generated by the TTS model. As such, the solution can provide TTS and ASR models with high accuracy for languages lacking training data by using less training data.

Claims (41)

1. A computer-implemented method, comprising:

training a Text to Speech (TTS) model and an Automatic Speech Recognition (ASR) model using paired speech and text data in a first language to develop speech-text alignment capabilities;

adjusting the trained TTS model and the ASR model to support a target language different from the first language by initializing the models for the target language using pre-trained parameters from the first language while updating phoneme embeddings, character embeddings and speaker embeddings for the target language, wherein the adjusting is based on a first set of paired data comprising speech data in the target language from multiple speakers and corresponding text data; and

optimizing the TTS model based on the first set of paired data and a first set of synthesized paired data in the target language while optimizing the ASR model based on the first set of paired data and a second set of synthesized paired data in the target language, wherein the first set of synthesized paired data comprises a first set of speech data from multiple speakers and a first set of text data generated by the ASR model based on the first set of speech data, and the second set of synthesized paired data comprises a second set of text data and a second set of speech data of multiple speakers generated by the TTS model based on the second set of text data.

2. The method of claim 1 , wherein training the TTS model and the ASR model comprises:

training, based on a second set of paired data in the first language, the TTS model and the ASR model, wherein the second set of paired data comprises speech data in the first language from multiple speakers and corresponding text data.

3. The method of claim 1 , further comprising:

training a target TTS model and a target ASR model based on the first set of paired data and a plurality of sets of synthesized paired data in the target language generated by the optimized TTS model and the optimized ASR model.

4. The method of claim 3 , wherein training the target TTS model comprises:

obtaining, from the first set of paired data, a third set of paired data associated with a target speaker in a plurality of speakers, wherein the third set of paired data comprises speech data in the target language from the target speaker and corresponding text data;

generating, using the optimized TTS model, a third set of synthesized paired data in the target language, wherein the third set of synthesized paired data comprises a third set of text data and a third set of speech data of the target speaker generated by the optimized TTS model based on the third set of text data; and

training the target TTS model based on the third set of paired data and the third set of synthesized paired data, such that the target TTS model can generate, based on text data in the target language, speech data of the target speaker corresponding to the text data.

5. The method of claim 4 , wherein training the target TTS model based on the third set of paired data and the third set of synthesized paired data comprises:

obtaining a fourth set of synthesized paired data by removing unqualified speech data from the third set of speech data and removing text data corresponding to the unqualified speech data from the third set of text data; and

training the target TTS model based on the third set of paired data and the fourth set of synthesized paired data.

6. The method of claim 5 , wherein the unqualified speech data comprises at least one of:

speech data with missing words;

speech data with repeated words; and

incomprehensible speech data.

7. The method of claim 5 , wherein removing the unqualified speech data comprises:

removing, from the third set of speech data, speech data with a Word Coverage Rate (WCR) lower than a predetermined threshold, wherein the WCR is inversely correlated with a possibility that missing words or repeated words exist in the speech data.

8. The method of claim 5 , wherein removing the unqualified speech data comprises:

removing, from the third set of speech data, speech data with an Attention Diagonal Ratio (ADR) lower than a predetermined threshold, wherein the ADR indicates an alignment degree between the speech data and text data used to generate the speech data in the third set of text data.

9. The method of claim 3 , wherein training the target ASR model comprises:

generating, using the optimized TTS model, a fifth set of synthesized paired data in the target language, wherein the fifth set of synthesized paired data comprises a third set of text data and a fourth set of speech data of multiple speakers generated by the optimized TTS model based on the third set of text data;

generating, using the optimized ASR model, a sixth set of synthesized paired data in the target language, wherein the sixth set of synthesized paired data comprises a fifth set of speech data from multiple speakers and a fourth set of text data generated by the optimized ASR model based on the fifth set of speech data; and

training the target ASR model based on the first set of paired data, the fifth set of synthesized paired data and the sixth set of synthesized paired data, such that the target ASR model can generate, based on speech data in the target language from multiple speakers, text data corresponding to the speech data.

10. An electronic device, comprising:

a processing unit; and

a memory coupled to the processing unit and having instructions stored thereon, the instructions when executed by the processing unit causing the electronic device to perform acts comprising:

training a Text to Speech (TTS) model and an Automatic Speech Recognition (ASR) model using paired speech and text data in a first language to develop speech-text alignment capabilities;

adjusting the trained TTS model and the ASR model to support a target language different from the first language by initializing the models for the target language using pre-trained parameters from the first language while updating phoneme embeddings, character embeddings and speaker embeddings for the target language, wherein the adjusting is based on a first set of paired data comprising speech data in the target language from multiple speakers and corresponding text data; and

optimizing the TTS model based on the first set of paired data and a first set of synthesized paired data in the target language while optimizing the ASR model based on the first set of paired data and a second set of synthesized paired data in the target language, wherein the first set of synthesized paired data comprises a first set of speech data from multiple speakers and a first set of text data generated by the ASR model based on the first set of speech data, and the second set of synthesized paired data comprises a second set of text data and a second set of speech data of multiple speakers generated by the TTS model based on the second set of text data.

11. The electronic device of claim 10 , wherein training the TTS model and the ASR model comprises:

training, based on a second set of paired data in the first language, the TTS model and the ASR model, wherein the second set of paired data comprises speech data in the first language from multiple speakers and corresponding text data.

12. The electronic device of claim 10 , wherein the acts further comprise:

training a target TTS model and a target ASR model based on the first set of paired data and a plurality of sets of synthesized paired data in the target language generated by the optimized TTS model and the optimized ASR model.

13. A computer program product stored tangibly in a computer storage medium and including machine-executable instructions which, when executed by a device, cause the device to perform acts comprising:

training a Text to Speech (TTS) model and an Automatic Speech Recognition (ASR) model using paired speech and text data in a first language to develop speech-text alignment capabilities;

adjusting the trained TTS model and the ASR model to support a target language different from the first language by initializing the models for the target language using pre-trained parameters from the first language while updating phoneme embeddings, character embeddings and speaker embeddings for the target language, wherein the adjusting is based on a first set of paired data comprising speech data in the target language from multiple speakers and corresponding text data; and

optimizing the TTS model based on the first set of paired data and a first set of synthesized paired data in the target language while optimizing the ASR model based on the first set of paired data and a second set of synthesized paired data in the target language, wherein the first set of synthesized paired data comprises a first set of speech data from multiple speakers and a first set of text data generated by the ASR model based on the first set of speech data, and the second set of synthesized paired data comprises a second set of text data and a second set of speech data of multiple speakers generated by the TTS model based on the second set of text data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2022
From: TAN, XU; QIN, TAO; GAN, JUNWEI; ZHAO, SHENG; LIU, TIEYAN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 061976/0829 →
Priority Claims (1)
CN 202010620533.5 · Jun 30, 2020 · national
Continuity (1)
Related Publication 20230298567A1 · Sep 21, 2023
References Cited (77)
US 20170092258A1 · Edrenkin · 2017 [cited by applicant]
US 20180254034A1 · Li · 2018 [cited by examiner]
US 20200082806A1 · Kim · 2020 [cited by examiner]
US 20200152184A1 · Steedman Henderson · 2020 [cited by applicant]
US 20210304769A1 · Ye · 2021 [cited by examiner]
US 20210312906A1 · Kuo · 2021 [cited by examiner]
CN 102360543A · 2013 [cited by applicant]
CN 109117483A · 2019 [cited by applicant]
CN 110428818A · 2019 [cited by applicant]
KR 20200048620A · 2020 [cited by applicant]
WO 2019139428A1 · 2019 [cited by applicant]
Baskar, et al., “Self-supervised sequence-to-sequence ASR using unpaired speech and text”, arXiv preprint arXiv:1905.01152, 2019, 6 pages. [cited by applicant]
Communication pursuant to Article 94(3) Received in European Patent Application No. 21731622.3, mailed on Jun. 24, 2024, 6 pages. [cited by applicant]
Nakayama, et al., “Zero-Shot Code-Switching ASR and TTS with Multilingual Machine Speech Chain”, IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Dec. 18, 2019, pp. 964-971. [cited by applicant]
Ren, et al., “Almost Unsupervised Text to Speech and Automatic Speech Recognition”, Proceedings of Machine Learning Research, 2019, pp. 5410-5419. [cited by applicant]
“Open Source Database—Open Source Data Products”, Retrieved From: https://web.archive.org/web/20200502011204/https://www.data-baker.com/open_source.html, May 2, 2020, 2 Pages. [cited by applicant]
Artetxe, et al., “On the Cross-lingual Transferability of Monolingual Representations”, In Repository of arXiv:1910.11856v1, Oct. 25, 2019, 15 Pages. [cited by applicant]
Baevski, et al., “Effectiveness of Self-supervised Pre-training for Speech Recognition”, In Repository of arXiv:1911.03912v1, Nov. 10, 2019, 10 Pages. [cited by applicant]
Bahdanau, et al., “Neural Machine Translation by Jointly Learning to Align and Translate”, In Repository of arXiv:1409.0473v1, Sep. 1, 2014, 15 Pages. [cited by applicant]
Bruguier, et al., “Dictionary Augmented Sequence-to-Sequence Neural Network for Grapheme to Phoneme prediction”, In Proceedings of 19th Annual Conference of the International Speech Communication Association, Sep. 2, 20… [cited by applicant]
Bu, et al., “AISHELL-1: An Open-Source Mandarin Speech Corpus and a Speech Recognition Baseline”, In Proceedings of 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases… [cited by applicant]
Chan, et al., “Listen, Attend and Spell: A Neural Network for Large Vocabulary Conversational Speech Recognition”, In Proceedings of International Conference on Acoustics, Speech and Signal Processing, Mar. 20, 2016, pp… [cited by applicant]
Chen, et al., “End-to-end Text-to-speech for Low-resource Languages by Cross-Lingual Transfer Learning”, In Proceedings of 20th Annual Conference of the International Speech Communication Association, Sep. 15, 2019, pp.… [cited by applicant]
Chen, et al., “Extensible Cross-Modal Hashing”, In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, Aug. 10, 2019, pp. 2109-2115. [cited by applicant]
Chen, et al., “Towards Unsupervised Automatic Speech Recognition Trained by Unaligned Speech and Text Only”, In Repository of arXiv:1803.10952v1, Mar. 29, 2018, 5 Pages. [cited by applicant]
Chiu, et al., “State-of-the-Art Speech Recognition with Sequence-to-Sequence Models”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 15, 2018, pp. 4774-4778. [cited by applicant]
Chorowski, et al., “End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results”, In Repository of arXiv:1412.1602v1, Dec. 4, 2014, 10 Pages. [cited by applicant]
Chung, et al., “Semi-supervised Training for Improving Data Efficiency in End-to-end Speech Synthesis”, In Proceedings of International Conference on Acoustics, Speech and Signal Processing, May 12, 2019, pp. 6940-6944. [cited by applicant]
Cooper, et al., “Characteristics of Text-to-Speech and Other Corpora”, In Proceedings of Speech Prosody, Jun. 13, 2018, pp. 690-694. [cited by applicant]
Cooper, Erica, “Text-to-Speech Synthesis Using Found Data for Low-Resource Languages”, In Dissertation Submitted to Graduate School of Arts and Sciences, Jan. 29, 2019, 150 Pages. [cited by applicant]
Hori, et al., “Cycle-consistency Training for End-to-end Speech Recognition”, In Proceedings of International Conference on Acoustics, Speech and Signal Processing, May 12, 2019, pp. 6271-6275. [cited by applicant]
Ito, Keith, “The LJ Speech Dataset”, Retrieved From: https://keithito.com/LJ-Speech-Dataset/, 2017, 5 Pages. [cited by applicant]
Kaiser, et al., “Tensor2Tensor”, Retrieved From: https://web.archive.org/web/20190130121118/https://github.com/tensorflow/tensor2tensor, Mar. 4, 2020, 6 Pages. [cited by applicant]
Karrer, Tony, “Text-to-Speech Costs—Licensing and Pricing”, Retrieved From: http://elearningtech.blogspot.com/2010/11/text-to-speech-costs-licensing-and.html, Nov. 18, 2010, 3 Pages. [cited by applicant]
Kim, et al., “Sequence-Level Knowledge Distillation”, In Proceedings of Conference on Empirical Methods in Natural Language Processing, Nov. 1, 2016, pp. 1317-1327. [cited by applicant]
Kuhl, et al., “Phonetic Learning as a Pathway to Language: New Data and Native Language Magnet Theory Expanded (NLM-E)”, In Journal of Philosophical Transactions of the Royal Society B: Biological Sciences, vol. 363, Is… [cited by applicant]
Laurinciukaite, et al., “Lithuanian Speech Corpus Liepa for Development of Human-Computer Interfaces Working in Voice Recognition and Synthesis Mode”, In Journal of Informatica vol. 29, Issue 3, Jan. 1, 2018, pp. 487-49… [cited by applicant]
Lewis, et al., “Ethnologue: Languages of the World”, Retrieved From: https://www.ethnologue.com/sites/default/files/Ethnologue-18-Honduras.pdf, 2015, 22 Pages. [cited by applicant]
Li, et al., “Neural Speech Synthesis with Transformer Network”, In Proceedings of Thirty-Third AAAI Conference on Artificial Intelligence, vol. 33, Issue 1, Jul. 17, 2019, pp. 6706-6713. [cited by applicant]
Liu, et al., “Completely Unsupervised Phoneme Recognition by Adversarially Learning Mapping Relationships from Audio Embeddings”, In Proceedings of 19th Annual Conference of the International Speech Communication Associ… [cited by applicant]
Liu, et al., “Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning”, In Repository of arXiv:1910.12729v1, Oct. 28, 2019, 5 Pages. [cited by applicant]
Luong, et al., “Effective Approaches to Attention-based Neural Machine Translation”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Sep. 17, 2015, pp. 1412-1421. [cited by applicant]
Panayotov, et al., “Librispeech: An ASR Corpus based on Public Domain Audio Books”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 19, 2015, pp. 5206-5210. [cited by applicant]
Park, et al., “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition”, In Proceedings of 20th Annual Conference of the International Speech Communication Association, Sep. 15, 2019, pp. 2613-26… [cited by applicant]
Ping, et al., “Deep Voice 3: Scaling Text-To-Speech with Convolutional Sequence Learning”, In Proceedings of Sixth International Conference on Learning Representations, Apr. 30, 2018, 16 Pages. [cited by applicant]
Press, et al., “Using the Output Embedding to Improve Language Models”, In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: vol. 2, Short Papers, Apr. 3, 2017,… [cited by applicant]
Ren, et al., “Almost Unsupervised Text to Speech and Automatic Speech Recognition”, In Proceedings of 36th International Conference on Machine Learning, Jun. 9, 2019, 10 Pages. [cited by applicant]
Ren, et al., “FastSpeech: Fast, Robust and Controllable Text to Speech”, In Proceedings of 33rd Conference on Neural Information Processing Systems, Dec. 8, 2019, 10 Pages. [cited by applicant]
Riviere, et al., “Unsupervised Pretraining Transfers Well Across Languages”, In Repository of arXiv:2002.02848v1, Feb. 7, 2020, 7 Pages. [cited by applicant]
Rosenberg, et al., “Speech Recognition with Augmented Synthesized Speech”, In Repository of arXiv:1909.11699v1, Sep. 25, 2019, 7 Pages. [cited by applicant]
Schneider, et al., “Wav2vec: Unsupervised Pre-training for Speech Recognition”, In Proceedings of 20th Annual Conference of the International Speech Communication Association, Sep. 15, 2019, pp. 3465-3469. [cited by applicant]
Sennrich, et al., “Improving Neural Machine Translation Models with Monolingual Data”, In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (vol. 1: Long Papers), Aug. 7, 2016, pp. … [cited by applicant]
Shen, et al., “Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 15, 2018, pp. 4779-4783. [cited by applicant]
Sun, et al., “Token-Level Ensemble Distillation for Grapheme-to-Phoneme Conversion”, In Proceedings of 20th Annual Conference of the International Speech Communication Association, Sep. 15, 2019, pp. 2115-2119. [cited by applicant]
Tan, et al., “Multilingual Neural Machine Translation with Knowledge Distillation”, In Proceedings of International Conference on Learning Representations, May 6, 2019, 15 Pages. [cited by applicant]
Thu, et al., “Comparison of Grapheme-to-Phoneme Conversion Methods on a Myanmar Pronunciation Dictionary”, In Proceedings of the 6th Workshop on South and Southeast Asian Natural Language Processing, Dec. 11, 2016, pp. … [cited by applicant]
Tjandra, et al., “Listening While Speaking: Speech Chain by Deep Learning”, In Proceedings of Automatic Speech Recognition and Understanding Workshop (ASRU), Dec. 16, 2017, pp. 301-308. [cited by applicant]
Vaswani, et al., “Attention Is All You Need”, In Proceedings of 31st Conference on Neural Information Processing Systems, Dec. 4, 2017, 11 Pages. [cited by applicant]
Wang, et al., “Tacotron: Towards End-to-End Speech Synthesis”, In Proceedings of 18th Annual Conference of the International Speech Communication Association, Aug. 20, 2017, pp. 4006-4010. [cited by applicant]
Xu, et al., “LRSpeech: Extremely Low-Resource Speech Synthesis and Recognition”, In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Aug. 23, 2020, pp. 2802-2812. [cited by applicant]
Yamagishi, et al., “Thousands of Voices for HMM-Based Speech Synthesis-Analysis and Application of TTS Systems Built on Various ASR Corpora”, In Journal of IEEE Transactions on Audio, Speech, and Language Processing vol… [cited by applicant]
Yamamoto, et al., “Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks With Multi-Resolution Spectrogram”, In Repository of arXiv:1910.11480v1, Oct. 25, 2019, 5 Pages. [cited by applicant]
Yeh, et al., “Unsupervised Speech Recognition via Segmental Empirical Output Distribution Matching”, In Proceedings of 7th International Conference on Learning Representations, May 6, 2019, 14 Pages. [cited by applicant]
Zhu, et al., “Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks”, In Proceedings of the IEEE International Conference on Computer Vision, Oct. 22, 2017, pp. 2242-2251. [cited by applicant]
Alarcon, Nefi, “Microsoft Leverages the Power of NVIDIA GPUs to Enhance Speech Recognition Algorithms”, Retrieved From: https://developer.nvidia.com/blog/microsoft-enhances-sra-tts-algorithms/, May 22, 2019, 2 Pages. [cited by applicant]
Bansal, et al., “Pre-Training on High-Resource Speech Recognition Improves Low-Resource Speech-to-Text Translation”, In Repository of arXiv:1809.01431v1, Sep. 5, 2018, 8 Pages. [cited by applicant]
Baskar, et al., “Self-supervised Sequence-to-sequence ASR using Unpaired Speech and Text”, In Repository of arXiv:1905.01152v1, Apr. 30, 2019, 6 Pages. [cited by applicant]
Jia, et al., “Leveraging Weakly Supervised Data to Improve End-To-End Speech-To-Text Translation”, In Repository of ofarXiv:1811.02050v1, Nov. 5, 2018, 5 Pages. [cited by applicant]
Lee, et al., “Learning Pronunciation from a Foreign Language in Speech Synthesis Networks”, In Repository of arXiv:1811.09364v1, Nov. 23, 2018, 10 Pages. [cited by applicant]
Luong, et al., “Bootstrapping Non-Parallel Voice Conversion from Speaker-Adaptive Text-to-Speech”, In Repository of arXiv:1909.06532v1, Sep. 14, 2019, 8 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US21/032128”, Mailed Date: Sep. 9, 2021, 9 Pages. [cited by applicant]
Toshniwal, et al., “Multilingual Speech Recognition With a Single End-To-End Model”, In Repository of arXiv:1711.01694v1, Nov. 6, 2017, 5 Pages. [cited by applicant]
Wind, Jan, “The Evolutionary History of the Human Speech Organs”, Published in Studies in Language Origins, vol. 1, Jan. 1989, pp. 173-197. [cited by applicant]
Nakayama, et al., “Speech Chain for Semi-supervised Learning of Japanese English Code-switching ASR and TTS”, IEEE Spoken Language Technology Workshop, 2018, pp. 182-189. [cited by applicant]
Office Action Received for Chinese Application No. 202010620533.5, mailed on Nov. 30, 2024, 19 pages. (English Translation Provided). [cited by applicant]
Wang et al., “Automatic Segmentation for TTS Units”, Microelectronics and computers, No. 12, Dec. 22, 2005, 4 pages. [cited by applicant]
Notice of Grant Received for Chinese Application No. 202010620533.5, mailed on Jul. 11, 2025, 4 pages. (English Translation Provided). [cited by applicant]