IP Library Granted Patent US 12,412,566
Granted Patent B2
US 12,412,566 · App. 17/650,566 · Granted Sep 9, 2025

Lookup-table recurrent language model

Inventors: Ronny Huang (Mountain View, CA); Tara N. Sainath (Jersey City, NJ); Trevor Strohman (Mountain View, CA); Shankar Kumar (Mountain View, CA)
Assignee: Google LLC
G10L15/083G06N3/04G10L15/16G10L15/187G10L15/26G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,566
App. No.
17/650,566
Granted
Sep 9, 2025
Kind
B2
Abstract

A computer-implemented method includes receiving audio data that corresponds to an utterance spoken by a user and captured by a user device. The method also includes processing the audio data to determine a candidate transcription that includes a sequence of tokens for the spoken utterance. Tor each token in the sequence of tokens, the method includes determining a token embedding for corresponding token, determining a n-gram token embedding for a previous sequence of n-gram tokens, and concatenating the token embedding and the n-gram token embedding to generate a concatenated output for the corresponding token. The method also includes rescoring the candidate transcription for the spoken utterance by processing the concatenated output generated for each corresponding token in the sequence of tokens.

Claims (38)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations, the operations comprising:

receiving audio data corresponding to an utterance spoken by a user and captured by a user device;

processing, using a speech recognizer, the audio data to determine a candidate transcription for the spoken utterance, the candidate transcription comprises a sequence of tokens;

for each corresponding token in the sequence of tokens subsequent to an initial token in the sequence of tokens:

determining, using a first embedding table, a token embedding for the corresponding token, the token embedding determined independent from each token in the sequence of tokens that precedes the corresponding token in the sequence of tokens;

determining, using a second embedding table, a n-gram token embedding for a previous sequence of n-gram tokens, based on each token in the sequence of tokens that precedes the corresponding token in the sequence of tokens; and

concatenating the token embedding and the n-gram token embedding to generate a concatenated output for the corresponding token; and

rescoring, using an external language model, the candidate transcription for the spoken utterance by processing the concatenated output generated for each corresponding token in the sequence of tokens.

2. The method of claim 1 , wherein the external language model comprises a recurrent neural network language model.

3. The method of claim 1 , wherein the external language model integrates with the speech recognizer by Hybrid Autoregressive Transducer (HAT) factorizing the speech recognizer.

4. The method of claim 1 , wherein the speech recognizer comprises a conformer audio encoder and a recurrent neural network-transducer decoder.

5. The method of claim 1 , wherein the speech recognizer comprises a transformer audio encoder and a recurrent neural network-transducer decoder.

6. The method of claim 1 , wherein each token in the sequence of tokens of the candidate transcription represents a word in the candidate transcription.

7. The method of claim 1 , wherein each token in the sequence of tokens of the candidate transcription represents a wordpiece in the candidate transcription.

8. The method of claim 1 , wherein each token in the sequence of tokens of the candidate transcription represents a n-gram, phoneme, or grapheme in the candidate transcription.

9. The method of claim 1 , wherein the first and second embedding tables are stored sparsely on memory hardware in communication with the data processing hardware.

10. The method of claim 1 , wherein determining the token embedding for the corresponding token comprises retrieving the token embedding from the first embedding table via a look-up without requiring access to any graphics processing units and/or tensor processing units.

11. The method of claim 1 , wherein the data processing hardware resides on the user device.

12. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed by the data processing hardware cause the data processing hardware to perform operations comprising:

receiving audio data corresponding to an utterance spoken by a user and captured by a user device;

processing, using a speech recognizer, the audio data to determine a candidate transcription for the spoken utterance, the candidate transcription comprises a sequence of tokens;

for each corresponding token in the sequence of tokens subsequent to an initial token in the sequence of tokens:

determining, using a first embedding table, a token embedding for the corresponding token, the token embedding determined independent from each token in the sequence of tokens that precedes the corresponding token in the sequence of tokens;

determining, using a second embedding table, a n-gram token embedding for a previous sequence of n-gram tokens based on each token in the sequence of tokens that precedes the corresponding token in the sequence of tokens; and

concatenating the token embedding and the n-gram token embedding to generate a concatenated output for the corresponding token; and

rescoring, using an external language model, the candidate transcription for the spoken utterance by processing the concatenated output generated for each corresponding token in the sequence of tokens.

13. The system of claim 12 , wherein the external language model comprises a recurrent neural network language model.

14. The system of claim 12 , wherein the external language model integrates with the speech recognizer by Hybrid Autoregressive Transducer (HAT) factorizing the speech recognizer.

15. The system of claim 12 , wherein the speech recognizer comprises a conformer audio encoder and a recurrent neural network-transducer decoder.

16. The system of claim 12 , wherein the speech recognizer comprises a transformer audio encoder and a recurrent neural network-transducer decoder.

17. The system of claim 12 , wherein each token in the sequence of tokens of the candidate transcription represents a word in the candidate transcription.

18. The system of claim 12 , wherein each token in the sequence of tokens of the candidate transcription represents a wordpiece in the candidate transcription.

19. The system of claim 12 , wherein each token in the sequence of tokens of the candidate transcription represents a n-gram, phoneme, or grapheme in the candidate transcription.

20. The system of claim 12 , wherein the first and second embedding tables are stored sparsely on memory hardware in communication with the data processing hardware.

21. The system of claim 12 , wherein determining the token embedding for the corresponding token comprises retrieving the token embedding from the first embedding table via a look-up without requiring access to any graphics processing units and/or tensor processing units.

22. The system of claim 12 , wherein the data processing hardware resides on the user device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2022
From: HUANG, RONNY; SAINATH, TARA; STROHMAN, TREVOR; KUMAR, SHANKAR
To: GOOGLE LLC
Reel/Frame 059045/0831 →
Continuity (2)
Provisional Application 63165725 · Mar 24, 2021
Related Publication 20220310067A1 · Sep 29, 2022
References Cited (22)
US 10176802B1 · Ladhak · 2019 [cited by examiner]
US 10431210B1 · Huang · 2019 [cited by examiner]
US 11256707B1 · Xiong · 2022 [cited by examiner]
US 11610586B2 · Qiu · 2023 [cited by examiner]
US 20170186432A1 · Aleksic et al. · 2017 [cited by applicant]
US 20170270100A1 · Audhkhasi · 2017 [cited by examiner]
US 20180329897A1 · Kalchbrenner · 2018 [cited by examiner]
US 20210232753A1 · He · 2021 [cited by examiner]
US 20210264220A1 · Wei · 2021 [cited by examiner]
US 20210279042A1 · Allamanis · 2021 [cited by examiner]
US 20220253502A1 · Alonichau · 2022 [cited by examiner]
CA 3039551A1 · 2018 [cited by examiner]
KR 20210154849A · 2021 [cited by examiner]
WO WO2019245916A1 · 2019 [cited by examiner]
Ehsan Variani, David Rybach, Cyril Allauzen, Michael Riley ““Hybrid Autoregressive Transducer (HAT)” (Mar. 12, 2020)” arXiv:2003.07705v1 [eess.AS] Mar. 12, 2020, (Year: 2020) (Year: 2020). [cited by examiner]
Ehsan Variani, David Rybach, Cyril Allauzen, Michael Riley ““Hybrid Autoregressive Transducer (HAT)” (Mar. 12, 2020)” arXiv:2003.07705v1 [eess.AS] Mar. 12, 2020, (Year: 2020). [cited by examiner]
Ogawa, A., Delcroix, M., Karita, S., & Nakatani, T. (2019). Improved Deep Duel Model for Rescoring N-Best Speech Recognition List Using Backward LSTMLM and Ensemble Encoders. In Interspeech (pp. 3900-3904). (Year: 2019). [cited by examiner]
Ronny Huang W et al.: “Lookup-Table Recurrent Language Models for Long Tail Speech Recognition”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, Ny 14853, Jun. 7, 2021 (Jun. 7, 2021). [cited by applicant]
Variani Ehsan et al.: “Hybrid Autoregressive Transducer (HAT)”, ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 4, 2020 (May 4, 2020), pp. 6139-6143. [cited by applicant]
Ofir Press et al.: “Using the Output Embedding to Improve Language Models”, Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: vol. 2, Short Papers, Feb. 21, 201… [cited by applicant]
May 18, 2022 Written Opinion (WO) of the International Searching Authority (ISA) and International Search Report (ISR) issued in International Application No. PCT/US2022/015956. [cited by applicant]
Indian Office Action for the related Application No. 202327062921 dated May 6, 2025. [cited by applicant]