IP Library Granted Patent US 12,307,213
Granted Patent B2
US 12,307,213 · App. 17/836,390 · Granted May 20, 2025

Automatic speech recognition systems and processes

Inventors: Kshitiz Kumar (Redmond, WA); Jian Wu (Bellevue, WA); Bo Ren (Bellevue, WA); Tianyu Wu (Suzhou, CN); Fahimeh Bahmaninezhad (San Mateo, CA); Edward C. Lin (Beijing, CN); Xiaoyang Chen (Suzhou, CN); Changliang Liu (Bellevue, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F40/58G10L15/005G10L15/063G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,213
App. No.
17/836,390
Granted
May 20, 2025
Kind
B2
Abstract

A data processing system is implemented for receiving speech data for a plurality of languages, and determining letters from the speech data. The data processing system also implements normalizing the speech data by applying linguistic based rules for Latin-based languages on the determined letters, building a computer model using the normalized speech data, fine-tuning the computer model using additional speech data, and recognizing words in a target language using the fine-tuned computer model.

Claims (42)

1. A data processing system comprising:

a processor; and

a machine-readable storage medium storing executable instructions that, when executed, cause the processor to perform operations comprising:

receiving speech data for a plurality of languages;

identifying and extracting graphemes from the speech data using a grapheme extraction engine;

normalizing the speech data using a normalizing engine that applies linguistic based rules for Latin-based languages to map the graphemes from the speech data to graphemes in a Latin-based language;

building a computer model using the normalized speech data;

fine-tuning the computer model using additional speech data; and

recognizing words in a target language using the fine-tuned computer model,

wherein the computer model is a Long Short-Term Memory model that has a top layer fine-tuned by the additional speech data.

2. The data processing system of claim 1 , wherein the plurality of languages include English, French, Italian, German, and Spanish languages and the speech data includes over 10,000 hours of data for each language.

3. The data processing system of claim 1 , wherein identifying and extracting the graphemes from the speech data using the grapheme extraction engine includes using natural language processing.

4. The data processing system of claim 1 , wherein the machine-readable storage medium includes instructions configured to cause the processor to perform an operation of:

receiving target speech data of the target language for the recognizing the words in the target language.

5. The data processing system of claim 1 , wherein the speech data includes data from video, broadcast news, and dictation sources for English, French, Italian, German, and Spanish languages.

6. The data processing system of claim 1 , wherein the machine-readable storage medium includes instructions configured to cause the processor to perform an operation of:

collecting the speech data from video, broadcast news, and dictation sources.

7. A method implemented in a data processing system, the method comprising:

receiving speech data for a plurality of languages;

identifying and extracting graphemes from the speech data using a grapheme extraction engine;

normalizing the speech data using a normalizing engine that applies linguistic based rules for Latin-based languages to map the graphemes from the speech data to graphemes in a Latin-based language;

building a computer model using the normalized speech data;

fine-tuning the computer model using additional speech data;

receiving target speech data of a target language; and

recognizing words of the target language in the target speech data using the fine-tuned computer model,

wherein the computer model is a transformer model that has a top layer fine-tuned by the additional speech data.

8. The method of claim 7 , further comprising:

collecting the speech data from video, broadcast news, and dictation sources for English, French, Italian, German, and Spanish languages.

9. The method of claim 7 , wherein identifying and extracting the graphemes from the speech data using the grapheme extraction engine includes using natural language processing.

10. The method of claim 7 , wherein the computer model is a Latency-Control Bidirectional Long Short-Term Memory model that has a top layer fine-tuned by the additional speech data.

11. The method of claim 7 , wherein the plurality of languages includes English, French, Italian, German, and Spanish languages and the speech data includes over 10,000 hours of data for each language.

12. A non-transitory machine-readable medium on which are stored instructions that, when executed, cause a processor of a programmable device to perform operations of:

receiving speech data for a plurality of different languages;

identifying and extracting graphemes from the speech data using a grapheme extraction engine;

normalizing the speech data using a normalizing engine that applies linguistic based rules for Latin-based languages to map the graphemes from the speech data to graphemes in a Latin-based language;

building a computer model using the normalized speech data;

fine-tuning the computer model using additional speech data;

receiving target speech data of a target language; and

recognizing target words of the target language in the target speech data using the fine-tuned computer model;

wherein the computer model is a Bidirectional Long Short-Term Memory model that has a top layer fine-tuned by the additional speech data.

13. The non-transitory machine-readable medium of claim 12 , wherein the plurality of different languages includes English, French, Italian, German, and Spanish languages and the speech data includes over 10,000 hours of data for each language.

14. The non-transitory machine-readable medium of claim 12 , wherein the computer model is a Latency-Control Bidirectional Long Short-Term Memory model that has a top layer fine-tuned by the additional speech data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2022
From: KUMAR, KSHITIZ; WU, JIAN; REN, BO; WU, TIANYU; BAHMANINEZHAD, FAHIMEH; LIN, EDWARD C.; CHEN, XIAOYANG; LIU, CHANGLIANG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 060150/0515 →
Continuity (1)
Related Publication 20230401392A1 · Dec 14, 2023
References Cited (58)
US 9483461B2 · Fleizach · 2016 [cited by examiner]
US 10593321B2 · Watanabe · 2020 [cited by examiner]
US 10657955B2 · Battenberg · 2020 [cited by examiner]
US 11048887B1 · Gupta et al. · 2021 [cited by applicant]
US 11158329B2 · Gopala et al. · 2021 [cited by applicant]
US 11437025B2 · Aleksic · 2022 [cited by examiner]
US 11475884B2 · Ghoshal · 2022 [cited by examiner]
US 11551668B1 · Baevski · 2023 [cited by examiner]
US 11809834B2 · Chen · 2023 [cited by examiner]
US 20200226327A1 · Matusov · 2020 [cited by examiner]
US 20200380215A1 · Kannan · 2020 [cited by examiner]
US 20210256961A1 · Garman et al. · 2021 [cited by applicant]
US 20210280202A1 · Wang et al. · 2021 [cited by applicant]
US 20220399006A1 · Jin · 2022 [cited by examiner]
CN 113345418A · 2021 [cited by applicant]
WO 2009151868A2 · 2009 [cited by applicant]
Vineel Pratap, Anuroop Sriram, Paden Tomasello, Awni Hannun, Vitaliy Liptchinsky, Gabriel Synnaeve, Ronan Collobert, “Massively Multilingual ASR: 50 Languages, 1 Model, 1 Billion Parameters”, arXiv: 2007.030016v2 [eess.… [cited by examiner]
Murat Akbacak, John H. L. Hansen, “Spoken Proper Name Retrieval for Limited Resource Languages Using Multilingual Hybrid Representations”, IEEE, vol. 18, No. 6, pp. 1486-1495, Aug. 2010 (Year: 2010). [cited by examiner]
Min Ma, “Adaptation and Augmentation: Towards Better Rescoring Strategies for Automatic Speech Recognition and Spoken Term Detection”, May 2018, The graduate center, City University of New York (Year: 2018). [cited by examiner]
Yuki Takashima, Shota Horiguchi, Shinji Watanabe, Paola Garcia, Yoheo Kawaguchi, “Updating Only Encoders Prevents Catastrophic Forgetting of End to End ASR models”, Jul. 2022, arXiv:2207.00216v1[eess.AS] (Year: 2022). [cited by examiner]
Anton Ragni, Edgar Dakin, Xie Chen, Mark J. F. Gales, Kate M. Knill, “Multi-Language Neural Network Language Models”, Sep. 2016, Interspeech, pp. 3042-3046 (Year: 2016). [cited by examiner]
Amrhein, et al., “On Romanization for Model Transfer Between Scripts in Neural Machine Translation”, In Proceedings of Conference on Empirical Methods in Natural Language Processing: Findings, Nov. 16, 2020, pp. 2461-24… [cited by applicant]
Chen, et al., “Training Deep Bidirectional LSTM Acoustic Model for LVCSR by a Context-Sensitive-Chunk BPTT Approach”, In Journal of IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 24, Issue 7, Jul.… [cited by applicant]
Chiu, et al., “State-of-the-Art Speech Recognition with Sequence-to Sequence Models”, In Proceeding of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 15, 2018, pp. 4774-4778. [cited by applicant]
Dahl, et al., “Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition”, In Journal of IEEE Transactions on Audio, Speech, and Language Processing, vol. 20, Issue 1, Jan. 2012, pp. 30-… [cited by applicant]
Datta, et al., “Language-Agnostic Multilingual Modeling”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 4, 2020, pp. 8239-8243. [cited by applicant]
Deng, et al., “Recent Advances in Deep Learning for Speech Research at Microsoft”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, May 26, 2013, pp. 8604-8608. [cited by applicant]
Eyben, et al., “From Speech to Letters-Using a Novel Neural Network Architecture for Grapheme Based ASR”, In Proceedings of IEEE Workshop on Automatic Speech Recognition & Understanding, Nov. 13, 2009, pp. 376-380. [cited by applicant]
Gales, et al., “Unicode-Based Graphemic Systems for Limited Resource Languages”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 19, 2015, pp. 5186-5190. [cited by applicant]
Ghahremani, et al., “Investigation of Transfer Learning for ASR Using LF-MMI Trained Neural Networks”, In Proceedings of IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Dec. 16, 2017, pp. 279-286. [cited by applicant]
Giollo, et al., “Bootstrap an End-to-End ASR System by Multilingual Training, Transfer Learning, Text-to-Text Mapping and Synthetic Audio”, In Repository of arXiv:2011.12696v1, Nov. 25, 2020, 5 Pages. [cited by applicant]
Heigold, et al., “Multilingual Acoustic Models Using Distributed Deep Neural Networks”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, May 26, 2013, pp. 8619-8623. [cited by applicant]
Hinton, et al., “Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups”, In Journal of IEEE Signal Processing Magazine, vol. 29, Issue 6, Nov. 2012, pp. 82-97. [cited by applicant]
Hochreiter, et al., “Long Short-Term Memory”, In Journal of Neural Computation, vol. 9, Issue 8, Nov. 15, 1997, 32 Pages. [cited by applicant]
Huang, et al., “Cross-Language Knowledge Transfer Using Multilingual Deep Neural Network with Shared Hidden Layers”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, May 26, 20… [cited by applicant]
Joshi, et al., “Transfer Learning Approaches for Streaming End-to-End Speech Recognition System”, In Repository of arXiv:2008.05086v2, Aug. 17, 2020, 5 Pages. [cited by applicant]
Kanthak, et al., “Context-Dependent Acoustic Modelling Using Graphemes for Large Vocabulary Speech Recognition”, In Proceedings of IEEE International Conference on Acoustics, Speech, and Signal Processing, May 13, 2002,… [cited by applicant]
Khare, et al., “Low Resource ASR: The Surprising Effectiveness of High Resource Transliteration”, In Proceedings of Interspeech, Aug. 30, 2021, pp. 1529-1533. [cited by applicant]
Kumar, et al., “Bandpass Noise Generation and Augmentation for Unified ASR”, In Proceedings of Interspeech, Oct. 25, 2020, pp. 1683-1687. [cited by applicant]
Le, et al., “From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition”, In Proceedings of IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Dec. 14, 2019, pp. 457-464. [cited by applicant]
Liu, et al., “Multilingual Graphemic Hybrid ASR with Massive Data Augmentation”, In Repository of arXiv:1909.06522v1, Sep. 14, 2019, 6 pages. [cited by applicant]
Sainath, et al., “Convolutional, Long Short-Term Memory, Fully Connected Deep Neural Networks”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 19, 2015, pp. 458… [cited by applicant]
Sak, et al., “Fast and Accurate Recurrent Neural Network Acoustic Models for Speech Recognition”, In Proceedings of Interspeech, Sep. 6, 2015, pp. 1468-1472. [cited by applicant]
Sak, et al., “Long Short-Term Memory Recurrent Neural Network Architectures for Large Scale Acoustic Modeling”, In Proceedings of Interspeech, Sep. 14, 2014, pp. 338-342. [cited by applicant]
Sak, et al., “Recurrent Neural Aligner: An Encoder-Decoder Neural Network Model for Sequence to Sequence Mapping”, In Proceedings of Interspeech, Aug. 20, 2017, pp. 1298-1302. [cited by applicant]
Schultz, et al., “Experiments on Cross-Language Acoustic Modeling”, In Proceedings of 7th European Conference on Speech Communication and Technology, Sep. 3, 2001, 4 Pages. [cited by applicant]
Simonyan, et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, In Repository of arxiv:1409.1556v1, Sep. 4, 2014, 10 Pages. [cited by applicant]
Stoian, et al., “Analyzing Asr Pretraining For Low-Resource Speech-To-Text Translation”, In Repository of arXiv:1910.10762v2, Feb. 9, 2020, 5 Pages. [cited by applicant]
Swietojanski, et al., “Unsupervised Cross-Lingual Knowledge Transfer in DNN-Based LVCSR”, In Proceedings of IEEE Spoken Language Technology Workshop (SLT), Dec. 2, 2012, pp. 246-251. [cited by applicant]
Vaswani, et al., “Attention is all You Need”, In Proceedings of 31st Conference on Neural Information Processing Systems, Dec. 4, 2017, 11 Pages. [cited by applicant]
Wang, et al., “Phonetic and Graphemic Systems for Multi-Genre Broadcast Transcription”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 15, 2018, pp. 5899-5903. [cited by applicant]
Wang, et al., “Transfer Learning for Speech and Language Processing”, In Proceedings of Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, Dec. 16, 2015, pp. 1225-1237. [cited by applicant]
Wang, et al., “Unispeech: Unified Speech Representation Learning with Labeled and Unlabeled Data”, In Repository of arXiv:2101.07597v1, Jan. 19, 2021, 10 Pages. [cited by applicant]
“Unidecode 1.3.4”, Retrieved from: https://pypi.org/project/Unidecode/, Retrieved Date: Mar. 17, 2022, 9 Pages. [cited by applicant]
Davis, et al., “Proposed Update Unicode Standard Annex #15 Unicode Normalization Forms”, In Technical Reports, Mar. 19, 2009, 30 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US23/019028”, Mailed Date: Jun. 29, 2023, 09 pages. [cited by applicant]
Pratap, et al., “Massively Multilingual ASR: 50 Languages, 1 Model, 1 Billion Parameters”, In Repository of arXiv:2007.03001v2, Jul. 8, 2020, 5 Pages. [cited by applicant]
Pratap, et al., “MLS: A Large-Scale Multilingual Dataset for Speech Research”, In Repository of arXiv:2012.03411v1, Dec. 7, 2020, 9 Pages. [cited by applicant]