IP Library Granted Patent US 12,254,875
Granted Patent B2
US 12,254,875 · App. 18/589,220 · Granted Mar 18, 2025

Multilingual re-scoring models for automatic speech recognition

Inventors: Neeraj Gaur (Mountain View, CA); Tongzhou Chen (Mountain View, CA); Ehsan Variani (Mountain View, CA); Bhuvana Ramabhadran (Mt. Kisco, NY); Parisa Haghani (Mountain View, CA); Pedro J. Moreno Mengibar (Jersey City, NJ)
Assignee: Google LLC
G10L15/197G10L15/005G10L15/16G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,875
App. No.
18/589,220
Granted
Mar 18, 2025
Kind
B2
Abstract

A method includes receiving a sequence of acoustic frames extracted from audio data corresponding to an utterance. During a first pass, the method includes processing the sequence of acoustic frames to generate N candidate hypotheses for the utterance. During a second pass, and for each candidate hypothesis, the method includes: generating a respective un-normalized likelihood score; generating a respective external language model score; generating a standalone score that models prior statistics of the corresponding candidate hypothesis; and generating a respective overall score for the candidate hypothesis based on the un-normalized likelihood score, the external language model score, and the standalone score. The method also includes selecting the candidate hypothesis having the highest respective overall score from among the N candidate hypotheses as a final transcription of the utterance.

Claims (34)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving transcribed audio training data comprising training audio data corresponding to an utterance paired with a ground-truth transcription of the utterance;

during a first pass, processing, using a speech recognition model, the training audio data to generate N candidate hypotheses for the utterance, each corresponding candidate hypothesis among the N candidate hypotheses having a respective first pass score;

during a second pass, for each corresponding candidate hypothesis of the N candidate hypotheses:

generating, using a neural network rescoring model, a respective second pass score based on the respective first pass score for the corresponding candidate hypothesis; and

applying a Softmax function to a respective negative edit-distance between the corresponding candidate hypothesis and the ground-truth transcription; and

optimizing model parameters of the neural network rescoring model based on the Softmax function applied to the respective negative edit-distance between the ground-truth transcription and each corresponding candidate hypothesis among the N candidate hypotheses.

2. The method of claim 1 , wherein the speech recognition model comprises a recurrent neural network-transducer (RNN-T) architecture.

3. The method of claim 1 , wherein each corresponding candidate hypothesis among the N candidate hypotheses comprises a respective sequence of word labels, each word label represented by a respective embedding vector.

4. The method of claim 1 , wherein each corresponding candidate hypothesis among the N candidate hypotheses comprises a respective sequence of sub-word labels, each sub-word label represented by a respective embedding vector.

5. The method of claim 1 , wherein the N candidate hypotheses generated using the speech recognition model during the first pass comprises an N-best list of candidate hypotheses.

6. The method of claim 1 , wherein the speech recognition model comprises an encoder-decoder architecture including a conformer encoder having a plurality of conformer layers.

7. The method of claim 1 , wherein the speech recognition model comprises an encoder-decoder architecture including a transformer encoder having a plurality of transformer layers.

8. The method of claim 1 , wherein the operations further comprise selecting one of the N candidate hypotheses as a final transcription of the utterance based on the respective second pass scores generated for the N candidate hypotheses.

9. The method of claim 1 , wherein the neural network rescoring model comprises a language-specific neural network rescoring model.

10. The method of claim 1 , wherein the neural network rescoring model comprises a multilingual neural network rescoring model.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving transcribed audio training data comprising training audio data corresponding to an utterance paired with a ground-truth transcription of the utterance;

during a first pass, processing, using a speech recognition model, the training audio data to generate N candidate hypotheses for the utterance, each corresponding candidate hypothesis among the N candidate hypotheses having a respective first pass score;

during a second pass, for each corresponding candidate hypothesis of the N candidate hypotheses:

generating, using a neural network rescoring model, a respective second pass score based on the respective first pass score for the corresponding candidate hypothesis; and

applying a Softmax function to a respective negative edit-distance between the corresponding candidate hypothesis and the ground-truth transcription; and

optimizing model parameters of the neural network rescoring model based on the Softmax function applied to the respective negative edit-distance between the ground-truth transcription and each corresponding candidate hypothesis among the N candidate hypotheses.

12. The system of claim 11 , wherein the speech recognition model comprises a recurrent neural network-transducer (RNN-T) architecture.

13. The system of claim 11 , wherein each corresponding candidate hypothesis among the N candidate hypotheses comprises a respective sequence of word labels, each word label represented by a respective embedding vector.

14. The system of claim 11 , wherein each corresponding candidate hypothesis among the N candidate hypotheses comprises a respective sequence of sub-word labels, each sub-word label represented by a respective embedding vector.

15. The system of claim 11 , wherein the N candidate hypotheses generated using the speech recognition model during the first pass comprises an N-best list of candidate hypotheses.

16. The system of claim 11 , wherein the speech recognition model comprises an encoder-decoder architecture including a conformer encoder having a plurality of conformer layers.

17. The system of claim 11 , wherein the speech recognition model comprises an encoder-decoder architecture including a transformer encoder having a plurality of transformer layers.

18. The system of claim 11 , wherein the operations further comprise selecting one of the N candidate hypotheses as a final transcription of the utterance based on the respective second pass scores generated for the N candidate hypotheses.

19. The system of claim 11 , wherein the neural network rescoring model comprises a language-specific neural network rescoring model.

20. The system of claim 11 , wherein the neural network rescoring model comprises a multilingual neural network rescoring model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2024
From: GAUR, NEERAJ; CHEN, TONGZHOU; VARIANI, EHSAN; RAMABHADRAN, BHUVANA; HAGHANI, PARISA; MENGIBAR, PEDRO J. MORENO
To: GOOGLE LLC
Reel/Frame 066580/0664 →
Continuity (3)
Continuation 17701635 · Mar 22, 2022
Provisional Application 63166916 · Mar 26, 2021
Related Publication 20240203409A1 · Jun 20, 2024
References Cited (29)
US 5677990A · Junqua · 1997 [cited by examiner]
US 5745649A · Lubensky · 1998 [cited by examiner]
US 11380308B1 · Pandey · 2022 [cited by examiner]
US 11741947B2 · Tripathi · 2023 [cited by examiner]
US 11749259B2 · Peyser · 2023 [cited by examiner]
US 11790899B2 · Rastogi · 2023 [cited by examiner]
US 12080283B2 · Gaur · 2024 [cited by examiner]
US 20100125458A1 · Franco · 2010 [cited by examiner]
US 20150243278A1 · Kibre · 2015 [cited by examiner]
US 20180068653A1 · Trawick · 2018 [cited by examiner]
US 20180082167A1 · Kurata · 2018 [cited by examiner]
US 20180330730A1 · Garg · 2018 [cited by examiner]
US 20190287519A1 · Ediz · 2019 [cited by examiner]
US 20200143806A1 · Sreedhara · 2020 [cited by examiner]
US 20200160838A1 · Lee · 2020 [cited by examiner]
US 20210343277A1 · Jaber · 2021 [cited by examiner]
US 20220115008A1 · Pust · 2022 [cited by examiner]
US 20220188361A1 · Botros · 2022 [cited by examiner]
US 20220246150A1 · Pust · 2022 [cited by examiner]
US 20220270597A1 · Qiu · 2022 [cited by examiner]
US 20220310080A1 · Qiu · 2022 [cited by examiner]
US 20220310081A1 · Gaur · 2022 [cited by examiner]
US 20230186907A1 · Hu · 2023 [cited by examiner]
US 20240203409A1 · Gaur · 2024 [cited by examiner]
Jul. 8, 2022 Written Opinion (WO) of the International Searching Authority (ISA) and International Search Report (ISR) issued in International Application No. PCT/US2022/021441. [cited by applicant]
Variani Ehsan et al: “Neural Oracle Search on N-Best Hypotheses”, ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 4, 2020 (May 4, 2020), pp. 7824-7828. [cited by applicant]
Ogawa Atsunori et al: “Rescoring N-Best Speech Recognition List Based on One-on-One Hypothesis Comparison Using Encoder-Classifier Model”, 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (I… [cited by applicant]
Ma Rao et al: “Neural Lattice Search for Speech Recognition”, ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 4, 2020 (May 4, 2020), pp. 7794-7798. [cited by applicant]
Hu He et al: “Transformer Based Deliberation for Two-Pass Speech Recognition”, 2021 IEEE Spoken Language Technology Workshop (SLT), IEEE, Jan. 19, 2021 (Jan. 19, 2021), pp. 68-74. [cited by applicant]