IP Library Granted Patent US 12,437,756
Granted Patent B2
US 12,437,756 · App. 17/817,176 · Granted Oct 7, 2025

Cross-lingual speech recognition

Inventors: Petar Aleksic (Jersey City, NJ); Pedro J. Moreno Mengibar (Jersey City, NJ)
Assignee: Google LLC
G10L15/187G10L15/02G10L15/22G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,756
App. No.
17/817,176
Granted
Oct 7, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for cross-lingual speech recognition are disclosed. In one aspect, a method includes the actions of determining a context of a second computing device. The actions further include identifying, by a first computing device, an additional pronunciation for a term of multiple terms. The actions further include including the additional pronunciation for the term in the lexicon. The actions further include receiving audio data of an utterance. The actions further include generating a transcription of the utterance by using the lexicon that includes the multiple terms and the pronunciation for each of the multiple terms and the additional pronunciation for the term. The actions further include after generating the transcription of the utterance, removing the additional pronunciation for the term from the lexicon. The actions further include providing, for output, the transcription.

Claims (62)

1. A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising:

determining a context of a computing device, the computing device comprising a lexicon including multiple terms in a first language and a pronunciation for each of the multiple terms in the first language;

based on the context of the computing device:

identifying, for inclusion in the lexicon, one or more additional terms and a pronunciation for each of the one or more additional terms, the one or more additional terms in a second language different from the first language, wherein each respective term of the multiple terms in the first language and each respective term of the one or more additional terms in the second language comprises a corresponding likelihood score; and

for each respective term of the one or more additional terms in the second language, biasing the corresponding likelihood score based on the context of the computing device;

receiving audio data of an utterance comprising at least one word in the first language and at least one word in the second language;

based on the corresponding likelihood scores of the multiple terms in the first language and the biased corresponding likelihood scores of the one or more additional terms in the second language, generating, by performing speech recognition on the received audio data of the utterance using the lexicon, a transcription of the utterance, the transcription comprising the at least one word in the first language and the at least one word in the second language; and

providing, for output, the transcription of the utterance.

2. The method of claim 1 , wherein the operations further comprise, after generating the transcription of the utterance, removing the pronunciation for each of the one or more additional terms in the second language from the lexicon.

3. The method of claim 1 , wherein the operations further comprise:

receiving data indicating that the computing device is likely to receive the utterance,

wherein determining the context of the computing device is based on receiving the data indicating that the computing device is likely to receive the utterance.

4. The method of claim 3 , wherein the computing device is likely to receive the utterance based on the computing device receiving an initial utterance of a predefined hotword.

5. The method of claim 3 , wherein the computing device is likely to receive the utterance based on determining that a particular application is running in a foreground of the computing device.

6. The method of claim 1 , wherein the operations further comprise:

after providing the transcription of the utterance for output, determining an additional context of the computing device;

based on the additional context of the computing device, identifying, for inclusion in the lexicon, another term that is not included in the multiple terms and another pronunciation for the other term;

receiving additional audio data of an additional utterance;

generating, by performing speech recognition on the additional audio data of the additional utterance using the lexicon, an additional transcription of the additional utterance;

after generating the additional transcription of the additional utterance, removing the other term and the other pronunciation for the other term from the lexicon; and

providing, for output, the additional transcription of the additional utterance.

7. The method of claim 6 , wherein:

the pronunciation for each of the multiple terms in the first language includes phonemes of the first language;

the other term is in the second language; and

the other pronunciation for the other term includes phonemes of the second language.

8. The method of claim 6 , wherein the other term is not included in the lexicon.

9. The method of claim 1 , wherein:

the pronunciation for each of the multiple terms includes phonemes of the first language; and

an additional pronunciation for one of the terms of the multiple terms includes phonemes of the second language.

10. The method of claim 1 , wherein the operations further comprise adjusting probabilities of sequences of terms of the multiple terms.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

determining a context of a computing device, the computing device comprising a lexicon including multiple terms in a first language and a pronunciation for each of the multiple terms in the first language;

based on the context of the computing device:

identifying, for inclusion in the lexicon, one or more additional terms and a pronunciation for each of the one or more additional terms, the one or more additional terms in a second language different than the first language, wherein each respective term of the multiple terms in the first language and each respective term of the one or more additional terms in the second language comprises a corresponding likelihood score; and

for each respective term of the one or more additional terms in the second language, biasing the corresponding likelihood score based on the context of the computing device;

receiving audio data of an utterance comprising at least one word in the first language and at least one word in the second language;

based on the corresponding likelihood scores of the multiple terms in the first language and the biased corresponding likelihood scores of the one or more additional terms in the second language, generating, by performing speech recognition on the received audio data of the utterance using the lexicon, a transcription of the utterance, the transcription comprising the at least one word in the first language and the at least one word in the second language; and

providing, for output, the transcription of the utterance.

12. The system of claim 11 , wherein the operations further comprise, after generating the transcription of the utterance, removing the pronunciation for each of the one or more additional terms in the second language from the lexicon.

13. The system of claim 11 , wherein the operations further comprise:

receiving data indicating that the computing device is likely to receive the utterance,

wherein determining the context of the computing device is based on receiving the data indicating that the computing device is likely to receive the utterance.

14. The system of claim 13 , wherein the computing device is likely to receive the utterance based on the computing device receiving an initial utterance of a predefined hotword.

15. The system of claim 13 , wherein the computing device is likely to receive the utterance based on determining that a particular application is running in a foreground of the computing device.

16. The system of claim 11 , wherein the operations further comprise:

after providing the transcription of the utterance for output, determining an additional context of the computing device;

based on the additional context of the computing device, identifying, for inclusion in the lexicon, another term that is not included in the multiple terms and another pronunciation for the other term;

receiving additional audio data of an additional utterance;

generating an additional transcription of the additional utterance by performing speech recognition on the additional audio data of the additional utterance using the lexicon;

after generating the additional transcription of the additional utterance, removing the other term and the other pronunciation for the other term from the lexicon; and

providing, for output, the additional transcription of the additional utterance.

17. The system of claim 16 , wherein:

the pronunciation for each of the multiple terms in the first language includes phonemes of the first language;

the other term is in the second language; and

the other pronunciation for the other term includes phonemes of the second language.

18. The system of claim 16 , wherein the other term is not included in the lexicon.

19. The system of claim 11 , wherein:

the pronunciation for each of the multiple terms includes phonemes of the first language; and

an additional pronunciation for one of the terms of the multiple terms includes phonemes of the second language.

20. The system of claim 11 , wherein the operations further comprise adjusting probabilities of sequences of terms of the multiple terms.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 3, 2022
From: ALEKSIC, PETAR
To: GOOGLE LLC
Reel/Frame 060710/0492 →
Continuity (3)
Continuation 16593564 · Oct 4, 2019
Provisional Application 62741250 · Oct 4, 2018
Related Publication 20220383862A1 · Dec 1, 2022
References Cited (18)
US 8352246B1 · Lloyd · 2013 [cited by applicant]
US 8731929B2 · Kennewick · 2014 [cited by examiner]
US 10186256B2 · Li et al. · 2019 [cited by applicant]
US 20020095282A1 · Goronzy et al. · 2002 [cited by applicant]
US 20040230430A1 · Gupta et al. · 2004 [cited by applicant]
US 20090326945A1 · Tian · 2009 [cited by applicant]
US 20110307241A1 · Waibel · 2011 [cited by examiner]
US 20130317823A1 · Mengibar · 2013 [cited by applicant]
US 20150081270A1 · Kamatani · 2015 [cited by examiner]
US 20160103825A1 · Ehsani et al. · 2016 [cited by applicant]
US 20160336008A1 · Menezes · 2016 [cited by examiner]
US 20170069311A1 · Grost et al. · 2017 [cited by applicant]
US 20170221475A1 · Bruguier et al. · 2017 [cited by applicant]
US 20180053502A1 · Biadsy et al. · 2018 [cited by applicant]
US 20200302950A1 · Kawano et al. · 2020 [cited by applicant]
EP 2727103B1 · 2014 [cited by applicant]
JP 2012068345A · 2012 [cited by applicant]
JP 2019070799A · 2019 [cited by applicant]