IP Library Granted Patent US 12,586,586
Granted Patent B2
US 12,586,586 · App. 18/490,733 · Granted Mar 24, 2026

Speech recognition with selective use of dynamic language models

Inventors: Petar Aleksic (Jersey City, NJ); Pedro J. Moreno Mengibar (Jersey City, NJ)
Assignee: Google LLC
G10L15/26G10L15/07G10L15/1815G10L15/183G10L15/197G10L15/30G10L2015/0635G10L2015/088G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,586
App. No.
18/490,733
Granted
Mar 24, 2026
Kind
B2
Abstract

A computer-implemented method for transcribing an utterance includes receiving, at a computing system, speech data that characterizes an utterance of a user. A first set of candidate transcriptions of the utterance can be generated using a static class-based language model that includes a plurality of classes that are each populated with class-based terms selected independently of the utterance or the user. The computing system can then determine whether the first set of candidate transcriptions includes class-based terms. Based on whether the first set of candidate transcriptions includes class-based terms, the computing system can determine whether to generate a dynamic class-based language model that includes at least one class that is populated with class-based terms selected based on a context associated with at least one of the utterance and the user.

Claims (38)

1 . A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving an audio signal characterizing an utterance captured by a computing device associated with a user;

generating, using a class-based language model, a word lattice representing a transcription hypothesis path corresponding to a respective transcription for the utterance, the transcription hypothesis path including a class identifier;

obtaining, based on the word lattice, a list of user-specific class-based terms, each user-specific class-based term in the list of user-specific class-based terms belonging to a class associated with the class identifier included in the transcription hypothesis path;

modifying the word lattice to include the list of user-specific class-based terms;

re-training the class-based language model using the modified word lattice; and

providing, using the re-trained class-based language model, a speech recognition result for the utterance.

2 . The computer-implemented method of claim 1 , wherein the list of user-specific class-based terms comprise names derived from a contact list of the user.

3 . The computer-implemented method of claim 1 , wherein the word lattice is represented by a finite state transducer.

4 . The computer-implemented method of claim 1 , wherein the class-based language model comprises an n-gram language model.

5 . The computer-implemented method of claim 4 , wherein the class-based language model assigns each term in a sequence of terms in the word lattice a respective probability based on a statistical likelihood that the term would follow an immediately preceding n−1 terms.

6 . The computer-implemented method of claim 1 , wherein the data processing hardware resides on the computing device.

7 . The computer-implemented method of claim 1 , wherein the operations further comprise, prior to generating the word lattice, processing, using an acoustic model, the audio signal to generate a set of candidate phonemes or linguistic units of the utterance.

8 . The computer-implemented method of claim 1 , wherein the operations further comprise asynchronously obtaining context data associated with the user in parallel with generating the word lattice.

9 . The computer-implemented method of claim 1 , wherein the operations further comprise:

determining a sequence of terms traversed by the transcription hypothesis path includes the class identifier; and

in response to determining the sequence of terms includes the class identifier, obtaining context data associated with the user.

10 . The computer-implemented method of claim 1 , wherein the class identifier flanks a pre-defined class-based term that occurs in the word lattice.

11 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising

receiving an audio signal characterizing an utterance captured by a computing device associated with a user;

generating, using a class-based language model, a word lattice representing a transcription hypothesis path corresponding to a respective transcription for the utterance, the transcription hypothesis path including a class identifier;

obtaining, based on the word lattice, a list of user-specific class-based terms, each user-specific class-based term in the list of user-specific class-based terms belonging to a class associated with the class identifier included in the transcription hypothesis path;

modifying the word lattice to include the list of user-specific class-based terms;

re-training the class-based language model using the modified word lattice; and

providing, using the re-trained class-based language model, a speech recognition result for the utterance.

12 . The system of claim 11 , wherein the list of user-specific class-based terms comprise names derived from a contact list of the user.

13 . The system of claim 11 , wherein the word lattice is represented by a finite state transducer.

14 . The system of claim 11 , wherein the class-based language model comprises an n-gram language model.

15 . The system of claim 14 , wherein the class-based language model assigns each term in a sequence of terms in the word lattice a respective probability based on a statistical likelihood that the term would follow an immediately preceding n−1 terms.

16 . The system of claim 11 , wherein the data processing hardware resides on the computing device.

17 . The system of claim 11 , wherein the operations further comprise, prior to generating the word lattice, processing, using an acoustic model, the audio signal to generate a set of candidate phonemes or linguistic units of the utterance.

18 . The system of claim 11 , wherein the operations further comprise asynchronously obtaining context data associated with the user in parallel with generating the word lattice.

19 . The system of claim 11 , wherein the operations further comprise:

determining a sequence of terms traversed by the transcription hypothesis path includes the class identifier; and

in response to determining the sequence of terms in the word lattice includes the class identifier, obtaining context data associated with the user.

20 . The system of claim 11 , wherein the class identifier flanks a pre-defined class-based term that occurs in the word lattice.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2023
From: ALEKSIC, PETAR; MENGIBAR, PEDRO J. MORENO
To: GOOGLE INC.
Reel/Frame 065286/0701 →
CHANGE OF NAME Recorded Oct 19, 2023
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 065291/0621 →
Continuity (3)
Continuation 17118232 · Dec 10, 2020
Continuation 14982567 · Dec 29, 2015
Related Publication 20240046933A1 · Feb 8, 2024
References Cited (50)
US 6081779A · Besling et al. · 2000 [cited by applicant]
US 6138099A · Lewis et al. · 2000 [cited by applicant]
US 6157912A · Kneser et al. · 2000 [cited by applicant]
US 6901364B2 · Nguyen et al. · 2005 [cited by applicant]
US 7275033B1 · Zhao et al. · 2007 [cited by applicant]
US 8447608B1 · Chang et al. · 2013 [cited by applicant]
US 8489398B1 · Gruenstein · 2013 [cited by applicant]
US 8645138B1 · Weinstein et al. · 2014 [cited by applicant]
US 8775177B1 · Heigold et al. · 2014 [cited by applicant]
US 8812299B1 · Su · 2014 [cited by examiner]
US 9031839B2 · Thorsen et al. · 2015 [cited by applicant]
US 9165028B1 · Christensen et al. · 2015 [cited by applicant]
US 9324323B1 · Bikel et al. · 2016 [cited by applicant]
US 20020087309A1 · Lee et al. · 2002 [cited by applicant]
US 20020087315A1 · Lee et al. · 2002 [cited by applicant]
US 20030050778A1 · Nguyen et al. · 2003 [cited by applicant]
US 20040088162A1 · He et al. · 2004 [cited by applicant]
US 20050055210A1 · Venkataraman · 2005 [cited by examiner]
US 20050080632A1 · Endo et al. · 2005 [cited by applicant]
US 20050182628A1 · Choi · 2005 [cited by applicant]
US 20060074670A1 · Weng et al. · 2006 [cited by applicant]
US 20060100876A1 · Nishizaki et al. · 2006 [cited by applicant]
US 20070100618A1 · Lee et al. · 2007 [cited by applicant]
US 20090030698A1 · Cerra et al. · 2009 [cited by applicant]
US 20090313017A1 · Nakazawa et al. · 2009 [cited by applicant]
US 20100185448A1 · Meisel · 2010 [cited by applicant]
US 20100195806A1 · Zhang et al. · 2010 [cited by applicant]
US 20110029301A1 · Han et al. · 2011 [cited by applicant]
US 20110055256A1 · Phillips et al. · 2011 [cited by applicant]
US 20110144999A1 · Jang et al. · 2011 [cited by applicant]
US 20110153324A1 · Ballinger et al. · 2011 [cited by applicant]
US 20110208507A1 · Hughes · 2011 [cited by applicant]
US 20130018650A1 · Moore et al. · 2013 [cited by applicant]
US 20130317822A1 · Koshinaka · 2013 [cited by applicant]
US 20150025884A1 · White et al. · 2015 [cited by applicant]
US 20150058018A1 · Georges et al. · 2015 [cited by applicant]
US 20150134326A1 · Bell et al. · 2015 [cited by applicant]
US 20150279360A1 · Mengibar et al. · 2015 [cited by applicant]
US 20150340024A1 · Schogol et al. · 2015 [cited by applicant]
US 20160063994A1 · Skobeltsyn et al. · 2016 [cited by applicant]
US 20160104482A1 · Aleksic et al. · 2016 [cited by applicant]
US 20160224658A1 · Liu et al. · 2016 [cited by applicant]
US 20170365251A1 · Park et al. · 2017 [cited by applicant]
International Preliminary Report on Patentability issued in International Application No. PCT/US2016/056425, mailed on Jul. 12, 2018, 9 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2016/056425, mailed on Jan. 23, 2017, 13 pages. [cited by applicant]
Thadani, Kapil, Fadi Biadsy, and Daniel M. Bikel, “On-the-fly Topic Adaptation for YouTube Video Transcription.” Interspeech, Sep. 2012, pp. 1-4. [cited by applicant]
Mohri, “Speech Recognition-Lecture 12: Lattice Algorithms,” Nov. 25, 2012, Courant Institute of Mathematical Sciences, New York University, New York City, NY, 35 pages. [cited by applicant]
Odell, “The Use of Context in Large Vocabulary Speech Recognition,” Mar. 1995, Ph.D. Thesis, Queen's College, University of Cambridge, Cambridge, UK, 146 pages. [cited by applicant]
Povey et al. “Generating Exact Lattices in the WFST Framework,” Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) Mar. 25-30, 2012, 2012, pp. 4213-4216, IEEE, Kyoto, J… [cited by applicant]
USPTO. Office Action relating to U.S. Appl. No. 17/118,232, Dated Jan. 20, 2023. [cited by applicant]