IP Library › Granted Patent US 12,205,596
Granted Patent B2
US 12,205,596 · App. 18/108,316 · Granted Jan 21, 2025

Generating and using text-to-speech data for speech recognition models

Inventors: Guoli Ye (Sammamish, WA); Yan Huang (Redmond, WA); Wenning Wei (Beijing, CN); Lei He (Beijing, CN); Eva Sharma (Vancouver, CA); Jian Wu (Bellevue, WA); Yao Tian (Beijing, CN); Edward C. Lin (Beijing, CN); Yifan Gong (Sammamish, WA); Rui Zhao (Bellevue, WA); Jinyu Li (Redmond, WA); William Maxwell Gale (Sunnyvale, CA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/26G10L13/08G10L15/063G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,596
App. No.
18/108,316
Granted
Jan 21, 2025
Kind
B2
Abstract

Systems, methods, and devices are provided for generating and using text-to-speech (TTS) data for improved speech recognition models. A main model is trained with keyword independent baseline training data. In some instances, acoustic and language model sub-components of the main model are modified with new TTS training data. In some instances, the new TTS training is obtained from a multi-speaker neural TTS system for a keyword that is underrepresented in the baseline training data. In some instances, the new TTS training data is used for pronunciation learning and normalization of keyword dependent confidence scores in keyword spotting (KWS) applications. In some instances, the new TTS training data is used for rapid speaker adaptation in speech recognition models.

Claims (53)

1. A computing system configured to modify a machine learning model with text-to-speech data, the machine learning model being used for speech recognition, wherein the computing system comprises:

one or more processors; and

one or more computer readable hardware storage devices that store computer-executable instructions that are structured to be executed by the one or more processors to cause the computing system to at least:

identify a main model trained with baseline training data, the main model including an acoustic model and a language model, the acoustic model and the language model being subcomponents of the main model;

receive an input comprising a keyword;

determining that the keyword is underrepresented in the baseline training data in response to determining that a computed confidence score for the keyword falls below a predefined confidence score threshold, the computed confidence score corresponding to an ability in which the main model is able to detect the keyword from spoken commands, wherein the keyword is one of a plurality of different keywords that the main model is trained to detect in spoken commands;

in response to determining that the keyword is underrepresented in the baseline training data, obtain new text-to-speech (TTS) training data that is specific to the keyword; and

modify at least the acoustic model and the language model of the main model with the new TTS training data to reduce a speech recognition error of the main model in performing speech recognition.

2. The computing system of claim 1 , wherein the each of the different keywords of the plurality of different keywords corresponding to a different computed confidence score and a different keyword-dependent confidence threshold.

3. The computing system of claim 1 , wherein the main model is structured as a recurrent neural network transducer (RNN-T) based model used for a keyword spotting system (KWS).

4. The computing system of claim 1 , wherein the new TTS training data is obtained from a multi-speaker, neural TTS system.

5. The computing system of claim 1 , wherein the execution of computer-executable instructions further causes the computing system to:

mix at least some new TTS training data with at least some baseline training data; and

modify at least the acoustic model and language model of the main model with the mixed training data to avoid overfitting during an adaptation of the main model.

6. The computing system of claim 5 , wherein the mixed training data is modified by applying a weighting to the at least some baseline training data to balance the ratio between new TTS training data and baseline training data.

7. The computing system of claim 1 , wherein the new TTS training data is further obtained by:

passing the synthesized audio data through at least one simulation process to generate simulated audio data;

collecting a second plurality of utterances of the particular keyword from the simulated audio data, wherein the second plurality of utterances is a plurality of simulated utterances; and

combining the first and second pluralities of utterances to form a combined set of utterances, wherein the new TTS training data comprises the combined set of utterances.

8. A computing system configured to modify a machine learning model with text-to-speech data, the machine learning model being used for speech recognition, wherein the computing system comprises:

one or more processors; and

one or more computer readable hardware storage devices that store computer-executable instructions that are structured to be executed by the one or more processors to cause the computing system to at least:

identify a main model trained with baseline training data, the main model including an acoustic model and a language model, the acoustic model and the language model being subcomponents of the main model;

receive an input comprising a keyword;

determining that the keyword is underrepresented in the baseline training data in response to determining that a computed confidence score for the keyword falls below a predefined confidence score threshold, the computed confidence score corresponding to a ratio of utterances that the main model correctly accepts as keyword utterances that contain the keyword relative to all utterances that the main model accepts as keyword utterances including utterances that do not contain the keyword;

in response to determining that the keyword is underrepresented in the baseline training data, obtain new text-to-speech (TTS) training data that is specific to the keyword; and

modify at least the acoustic model and the language model of the main model with the new TTS training data to reduce a speech recognition error of the main model in performing speech recognition.

9. The computing system of claim 8 , wherein the keyword is one of a plurality of different keywords that the main model is trained to detect in spoken commands, each of the different keywords of the plurality of different keywords corresponding to a different computed confidence score and a different keyword-dependent confidence threshold, the different computed confidence scores corresponding to an ability in which the main model is able to detect the different keywords from the spoken commands based on a determination that a corresponding computed confidence score is equal to or exceeds a value defined by the keyword-dependent confidence threshold.

10. The computing system of claim 8 , wherein the main model is structured as a recurrent neural network transducer (RNN-T) based model used for a keyword spotting system (KWS).

11. The computing system of claim 8 , wherein the new TTS training data is obtained from a multi-speaker, neural TTS system.

12. The computing system of claim 8 , wherein the execution of computer-executable instructions further causes the computing system to:

mix at least some new TTS training data with at least some baseline training data; and

modify at least the acoustic model and language model of the main model with the mixed training data to avoid overfitting during an adaptation of the main model.

13. The computing system of claim 12 , wherein the mixed training data is modified by applying a weighting to the at least some baseline training data to balance the ratio between new TTS training data and baseline training data.

14. The computing system of claim 1 , wherein the new TTS training data is further obtained by:

passing the synthesized audio data through at least one simulation process to generate simulated audio data;

collecting a second plurality of utterances of the particular keyword from the simulated audio data, wherein the second plurality of utterances is a plurality of simulated utterances; and

combining the first and second pluralities of utterances to form a combined set of utterances, wherein the new TTS training data comprises the combined set of utterances.

15. A computing system configured to modify a machine learning model with text-to-speech data, the machine learning model being used for speech recognition, wherein the computing system comprises:

one or more processors; and

one or more computer readable hardware storage devices that store computer-executable instructions that are structured to be executed by the one or more processors to cause the computing system to at least:

identify a main model trained with baseline training data, the main model including an acoustic model and a language model, the acoustic model and the language model being subcomponents of the main model;

receive an input comprising a keyword;

determining that the keyword is underrepresented in the baseline training data in response to determining that a computed confidence score for the keyword falls below a predefined confidence score threshold, the computed confidence score corresponding to a measure of utterances that the main model incorrectly accepts as keyword utterances;

in response to determining that the keyword is underrepresented in the baseline training data, obtain new text-to-speech (TTS) training data that is specific to the keyword; and

modify at least the acoustic model and the language model of the main model with the new TTS training data to reduce a speech recognition error of the main model in performing speech recognition.

16. The computing system of claim 15 , wherein the keyword is one of a plurality of different keywords that the main model is trained to detect in spoken commands, each of the different keywords of the plurality of different keywords corresponding to a different computed confidence score and a different keyword-dependent confidence threshold, the different computed confidence scores corresponding to an ability in which the main model is able to detect the different keywords from the spoken commands based on a determination that a corresponding computed confidence score is equal to or exceeds a value defined by the keyword-dependent confidence threshold.

17. The computing system of claim 15 , wherein the main model is structured as a recurrent neural network transducer (RNN-T) based model used for a keyword spotting system (KWS).

18. The computing system of claim 15 , wherein the new TTS training data is obtained from a multi-speaker, neural TTS system.

19. The computing system of claim 15 , wherein the execution of computer-executable instructions further causes the computing system to:

mix at least some new TTS training data with at least some baseline training data; and

modify at least the acoustic model and language model of the main model with the mixed training data to avoid overfitting during an adaptation of the main model.

20. The computing system of claim 19 , wherein the mixed training data is modified by applying a weighting to the at least some baseline training data to balance the ratio between new TTS training data and baseline training data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2023
From: YE, GUOLI; HUANG, YAN; WEI, WENNING; HE, LEI; SHARMA, EVA; WU, JIAN; TIAN, YAO; LIN, EDWARD C.; GONG, YIFAN; ZHAO, RUI; LI, JINYU; GALE, WILLIAM MAXWELL
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 062704/0529 →
Priority Claims (1)
CN 202010244661.4 · Mar 31, 2020 · national
Continuity (2)
Continuation 15931788 · May 14, 2020
Related Publication 20230186919A1 · Jun 15, 2023
References Cited (12)
US 10332508B1 · Hoffmeister · 2019 [cited by applicant]
US 20110288862A1 · Todic · 2011 [cited by applicant]
CN 107516511B · 2017 [cited by applicant]
CN 107680597B · 2019 [cited by applicant]
Junqua et al, Word Spotting and Rejection, 1996, Robustness in Automatic Speech Recognition: Kluwer Academic Publishers, pp. 1-21 (Year: 1996). [cited by examiner]
Rose, Richard C., Word Spotting From Continuous Speech Utterances, 1999, Automatic Speech and Speaker Recognition: Kluwer Academic Publishers, pp. 303-329 (Year: 1996). [cited by examiner]
Chai, Khe, et al., “Personalization of End-to-End Speech Recognition on Mobile Devices for Named Entities,” Dec. 14, 2019, pp. 1-8. [cited by applicant]
Murthy, Savitha, et al., “Effect of TTS Generated Audio on OOV Detection and Word Error Rate in ASR for Low-resource Languages,” Sep. 30, 2019, pp. 1-6. [cited by applicant]
Office Action Received for Chinese Application No. 202010244661.4, (MS #408172-CN01) mailed on Feb. 8, 2024, 14 pages (English Translation Provided). [cited by applicant]
Office Action Received for European Application No. 21709203.0, mailed on Jan. 4, 2024, 5 pages. [cited by applicant]
U.S. Appl. No. 15/931,788, filed May 14, 2020. [cited by applicant]
Office Action Received for Chinese Application No. 202010244661.4, (MS#408172-CN01), mailed on Jun. 6, 2024, 4 pages. (English Translation Provided). [cited by applicant]