IP Library Granted Patent US 12,080,283
Granted Patent B2
US 12,080,283 · App. 17/701,635 · Granted Sep 3, 2024

Multilingual re-scoring models for automatic speech recognition

Inventors: Neeraj Gaur (Mountain View, CA); Tongzhou Chen (Mountain View, CA); Ehsan Variani (Mountain View, CA); Bhuvana Ramabhadran (Mt. Kisco, NY); Parisa Haghani (Mountain View, CA); Pedro J. Moreno Mengibar (Jersey City, NJ)
Assignee: Google LLC
G10L15/197G10L15/005G10L15/16G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,080,283
App. No.
17/701,635
Granted
Sep 3, 2024
Kind
B2
Abstract

A method includes receiving a sequence of acoustic frames extracted from audio data corresponding to an utterance. During a first pass, the method includes processing the sequence of acoustic frames to generate N candidate hypotheses for the utterance. During a second pass, and for each candidate hypothesis, the method includes: generating a respective un-normalized likelihood score; generating a respective external language model score; generating a standalone score that models prior statistics of the corresponding candidate hypothesis; and generating a respective overall score for the candidate hypothesis based on the un-normalized likelihood score, the external language model score, and the standalone score. The method also includes selecting the candidate hypothesis having the highest respective overall score from among the N candidate hypotheses as a final transcription of the utterance.

Claims (36)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving a sequence of acoustic frames extracted from audio data corresponding to an utterance;

receiving a language identifier indicating a language of the utterance;

during a first pass, processing, using a multilingual speech recognition model, the sequence of acoustic frames to generate N candidate hypotheses for the utterance;

during a second pass, for each candidate hypothesis of the N candidate hypotheses:

generating, using a neural oracle search (NOS) model, a respective un-normalized likelihood score based on the sequence of acoustic frames and the corresponding candidate hypothesis;

selecting a language-specific external language model from among a plurality of language-specific external language models each trained on a different respective language;

generating, using the language-specific external language model, a respective external language model score;

generating a standalone score that models prior statistics of the corresponding candidate hypothesis generated during the first pass; and

generating a respective overall score for the candidate hypothesis based on the un-normalized likelihood score, the external language model score, and the standalone score; and

selecting the candidate hypothesis having the highest respective overall score from among the N candidate hypotheses as a final transcription of the utterance.

2. The method of claim 1 , wherein each candidate hypothesis of the N candidate hypotheses comprises a respective sequence of word or sub-word labels, each word or sub-word label represented by a respective embedding vector.

3. The method of claim 1 , wherein the external language model is trained on text-only data.

4. The method of claim 1 , wherein the NOS model comprises a language-specific NOS model.

5. The method of claim 1 , wherein the NOS model comprises a multilingual NOS model.

6. The method of claim 1 , wherein the NOS model comprises two unidirectional long short-term memory (LSTM) layers.

7. The method of claim 1 , wherein the speech recognition model comprises an encoder-decoder architecture including a conformer encoder having a plurality of conformer layers and a LSTM decoder having two LSTM layers.

8. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a sequence of acoustic frames extracted from audio data corresponding to an utterance;

receiving a language identifier indicating a language of the utterance;

during a first pass, processing, using a multilingual speech recognition model, the sequence of acoustic frames to generate N candidate hypotheses for the utterance;

during a second pass, for each candidate hypothesis of the N candidate hypotheses:

generating, using a neural oracle search (NOS) model, a respective un-normalized likelihood score based on the sequence of acoustic frames and the corresponding candidate hypothesis;

selecting a language-specific external language model from among a plurality of language-specific external language models each trained on a different respective language;

generating, using the language-specific external language model, a respective external language model score;

generating a standalone score that models prior statistics of the corresponding candidate hypothesis generated during the first pass; and

generating a respective overall score for the candidate hypothesis based on the un-normalized likelihood score, the external language model score, and the standalone score; and

selecting the candidate hypothesis having the highest respective overall score from among the N candidate hypotheses as a final transcription of the utterance.

9. The system of claim 8 , wherein each candidate hypothesis of the N candidate hypotheses comprises a respective sequence of word or sub-word labels, each word or sub-word label represented by a respective embedding vector.

10. The system of claim 8 , wherein the external language model is trained on text-only data.

11. The system of claim 8 , wherein the NOS model comprises a language-specific NOS model.

12. The system of claim 8 , wherein the NOS model comprises a multilingual NOS model.

13. The system of claim 8 , wherein the NOS model comprises two unidirectional long short-term memory (LSTM) layers.

14. The system of claim 8 , wherein the speech recognition model comprises an encoder-decoder architecture including a conformer encoder having a plurality of conformer layers and a LSTM decoder having two LSTM layers.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE CORRECT THE THIRD NAMED INVENTOR FROM EHSAN VARIAN TO EHSAN VARIANI PREVIOUSLY RECORDED AT REEL: 63004 FRAME: 858. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jan 7, 2025
From: GAUR, NEERAJ; CHEN, TONGZHOU; VARIANI, EHSAN; RAMABHADRAN, BHUVANA; HAGHANI, PARISA; MORENO MENGIBAR, PEDRO J.
To: GOOGLE LLC
Reel/Frame 069828/0992 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2023
From: GAUR, NEERAJ; VARIAN, EHSAN; CHEN, TONGZHOU; MORENO MENGIBAR, PEDRO J.; RAMABHADRAN, BHUVANA; HAGHANI, PARISA
To: GOOGLE LLC
Reel/Frame 063004/0858 →
Continuity (2)
Provisional Application 63166916 · Mar 26, 2021
Related Publication 20220310081A1 · Sep 29, 2022
Cited By (1)
US 12,254,875