IP Library › Granted Patent US 12,676,146
Granted Patent B2
US 12,676,146 · App. 18/382,886 · Granted Jul 7, 2026

Automatically determining language for speech recognition of spoken utterance received via an automated assistant interface

Inventors: Pu-sen Chao (Los Altos, CA); Diego Melendo Casado (Mountain View, CA); Ignacio Lopez Moreno (New York, NY)
Assignee: GOOGLE LLC
G10L15/197G10L13/00G10L15/005G10L15/08G10L15/14G10L15/1822G10L15/22G10L15/30G10L2015/088G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,676,146
App. No.
18/382,886
Filed
Oct 23, 2023
Granted
Jul 7, 2026
Kind
B2
Art Unit
2656
USPC
704/235
Abstract

Determining a language for speech recognition of a spoken utterance received via an automated assistant interface for interacting with an automated assistant. Implementations can enable multilingual interaction with the automated assistant, without necessitating a user explicitly designate a language to be utilized for each interaction. Implementations determine a user profile that corresponds to audio data that captures a spoken utterance, and utilize language(s), and optionally corresponding probabilities, assigned to the user profile in determining a language for speech recognition of the spoken utterance. Some implementations select only a subset of languages, assigned to the user profile, to utilize in speech recognition of a given spoken utterance of the user. Some implementations perform speech recognition in each of multiple languages assigned to the user profile, and utilize criteria to select only one of the speech recognitions as appropriate for generating and providing content that is responsive to the spoken utterance.

Claims (71)

1 . A method implemented by one or more processors, the method comprising:

processing audio data using an acoustic model corresponding to a given language to monitor for an occurrence of an invocation phrase configured to invoke an automated assistant, wherein the audio data is based on detection of spoken input of a user at a client device that includes an automated assistant interface for interacting with the automated assistant;

detecting, based on processing the audio data using the acoustic model, the occurrence of the invocation phrase in a portion of the audio data;

determining, based on processing of the audio data using the acoustic model, that the portion of the audio data that includes the invocation phrase corresponds to a user profile that is accessible to the automated assistant;

identifying a set of languages, including the given language, assigned to the user profile;

selecting a speech recognition model for the given language;

using the selected speech recognition model to process a subsequent portion of the audio data that follows the portion of the audio data;

causing the automated assistant to provide responsive content that is determined based on the processing of the subsequent portion using the selected speech recognition model;

subsequent to causing the automated assistant to provide the responsive content:

processing additional audio data based on detection of additional spoken input of the user;

using the selected speech recognition model to process the additional audio data to generate a first candidate text representation of the additional audio data in the given language;

identifying an additional language, in the set of languages assigned to the user profile;

selecting an additional speech recognition model, corresponding to the additional language;

using the selected additional speech recognition model to process the additional audio data to generate a second candidate text representation of the additional audio data in the additional language;

selecting the second candidate text representation in the additional language, in lieu of the first candidate text representation in the given language; and

causing the automated assistant to provide additional responsive content that is determined based on processing the second candidate text representation.

2 . The method of claim 1 , wherein selecting the second candidate text representation in the additional language, in lieu of the first candidate text representation in the given language comprises:

identifying one or more contextual parameters associated with the additional audio data; and

selecting the second candidate text representation in the additional language based on the one or more contextual parameters being more strongly associated, in the user profile, with the additional language than with the given language.

3 . The method of claim 2 , wherein the one or more contextual parameters comprise an identifier of the client device.

4 . The method of claim 2 , wherein the one or more contextual parameters comprise one or multiple of: a time of day, a day of the week, and a location of the client device.

5 . The method of claim 1 , wherein the automated assistant is configured to access multiple different user profiles that are: available at the client device, and associated with multiple different users of the client device.

6 . The method of claim 5 , wherein the multiple different user profiles each identify a set of corresponding languages and a corresponding language probability for each of the corresponding languages, the corresponding language probabilities each based on previous interactions between a corresponding one of the multiple different users and the automated assistant.

7 . A non-transitory computer readable storage medium configured to store instructions that, when executed by one or more processors, cause the one or more processors to perform operations that include:

processing audio data using an acoustic model corresponding to a given language to monitor for an occurrence of an invocation phrase configured to invoke an automated assistant, wherein the audio data is based on detection of spoken input of a user at a client device that includes an automated assistant interface for interacting with the automated assistant;

detecting, based on processing the audio data using the acoustic model, the occurrence of the invocation phrase in a portion of the audio data;

determining, based on processing of the audio data using the acoustic model, that the portion of the audio data that includes the invocation phrase corresponds to a user profile that is accessible to the automated assistant;

identifying a set of languages, including the given language, assigned to the user profile;

selecting a speech recognition model for the given language;

using the selected speech recognition model to process a subsequent portion of the audio data that follows the portion of the audio data;

causing the automated assistant to provide responsive content that is determined based on the processing of the subsequent portion using the selected speech recognition model;

subsequent to causing the automated assistant to provide the responsive content:

processing additional audio data based on detection of additional spoken input of the user;

using the selected speech recognition model to process the additional audio data to generate a first candidate text representation of the additional audio data in the given language;

identifying an additional language, in the set of languages assigned to the user profile;

selecting an additional speech recognition model, corresponding to the additional language;

using the selected additional speech recognition model to process the additional audio data to generate a second candidate text representation of the additional audio data in the additional language;

selecting the second candidate text representation in the additional language, in lieu of the first candidate text representation in the given language; and

causing the automated assistant to provide additional responsive content that is determined based on processing the second candidate text representation.

8 . The non-transitory computer readable storage medium of claim 7 , wherein selecting the second candidate text representation in the additional language, in lieu of the first candidate text representation in the given language comprises:

identifying one or more contextual parameters associated with the additional audio data; and

selecting the second candidate text representation in the additional language based on the one or more contextual parameters being more strongly associated, in the user profile, with the additional language than with the given language.

9 . The non-transitory computer readable storage medium of claim 8 , wherein the one or more contextual parameters comprise an identifier of the client device.

10 . The non-transitory computer readable storage medium of claim 8 , wherein the one or more contextual parameters comprise one or multiple of: a time of day, a day of the week, and a location of the client device.

11 . The non-transitory computer readable storage medium of claim 7 , wherein the automated assistant is configured to access multiple different user profiles that are: available at the client device, and associated with multiple different users of the client device.

12 . The non-transitory computer readable storage medium of claim 11 , wherein the multiple different user profiles each identify a set of corresponding languages and a corresponding language probability for each of the corresponding languages, the corresponding language probabilities each based on previous interactions between a corresponding one of the multiple different users and the automated assistant.

13 . A system comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations that include:

processing audio data using an acoustic model corresponding to a given language to monitor for an occurrence of an invocation phrase configured to invoke an automated assistant, wherein the audio data is based on detection of spoken input of a user at a client device that includes an automated assistant interface for interacting with the automated assistant;

detecting, based on processing the audio data using the acoustic model, the occurrence of the invocation phrase in a portion of the audio data;

determining, based on processing of the audio data using the acoustic model, that the portion of the audio data that includes the invocation phrase corresponds to a user profile that is accessible to the automated assistant;

identifying a set of languages, including the given language, assigned to the user profile;

selecting a speech recognition model for the given language;

using the selected speech recognition model to process a subsequent portion of the audio data that follows the portion of the audio data;

causing the automated assistant to provide responsive content that is determined based on the processing of the subsequent portion using the selected speech recognition model;

subsequent to causing the automated assistant to provide the responsive content:

processing additional audio data based on detection of additional spoken input of the user;

using the selected speech recognition model to process the additional audio data to generate a first candidate text representation of the additional audio data in the given language;

identifying an additional language, in the set of languages assigned to the user profile;

selecting an additional speech recognition model, corresponding to the additional language;

using the selected additional speech recognition model to process the additional audio data to generate a second candidate text representation of the additional audio data in the additional language;

selecting the second candidate text representation in the additional language, in lieu of the first candidate text representation in the given language; and

causing the automated assistant to provide additional responsive content that is determined based on processing the second candidate text representation.

14 . The system of claim 13 , wherein selecting the second candidate text representation in the additional language, in lieu of the first candidate text representation in the given language comprises:

identifying one or more contextual parameters associated with the additional audio data; and

selecting the second candidate text representation in the additional language based on the one or more contextual parameters being more strongly associated, in the user profile, with the additional language than with the given language.

15 . The system of claim 13 , wherein the one or more contextual parameters comprise an identifier of the client device.

16 . The system of claim 13 , wherein the one or more contextual parameters comprise one or multiple of: a time of day, a day of the week, and a location of the client device.

17 . The system of claim 13 , wherein the automated assistant is configured to access multiple different user profiles that are: available at the client device, and associated with multiple different users of the client device.

18 . The system of claim 17 , wherein the multiple different user profiles each identify a set of corresponding languages and a corresponding language probability for each of the corresponding languages, the corresponding language probabilities each based on previous interactions between a corresponding one of the multiple different users and the automated assistant.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2023
From: CHAO, PU-SEN; CASADO, DIEGO MELENDO; MORENO, IGNACIO LOPEZ
To: GOOGLE LLC
Reel/Frame 065439/0411 →
Continuity (3)
Continuation 17099367 · Nov 16, 2020
Continuation 15769013
Related Publication 20240054997A1 · Feb 15, 2024
References Cited (96)
US 117351A · Van Orden · 1871 [cited by applicant]
US 5515475A · Gupta et al. · 1996 [cited by applicant]
US 7873517B2 · Prieto et al. · 2011 [cited by applicant]
US 8909528B2 · Eide · 2014 [cited by applicant]
US 8935147B2 · Stern · 2015 [cited by applicant]
US 9031829B2 · Leydon et al. · 2015 [cited by applicant]
US 9418567B1 · Chen et al. · 2016 [cited by applicant]
US 9606767B2 · Corfield · 2017 [cited by applicant]
US 9786271B1 · Combs et al. · 2017 [cited by applicant]
US 9786281B1 · Adams et al. · 2017 [cited by applicant]
US 9953634B1 · Pearce et al. · 2018 [cited by applicant]
US 9953636B2 · Cohen et al. · 2018 [cited by applicant]
US 9971759B2 · Hobson · 2018 [cited by applicant]
US 10679615B2 · Chao et al. · 2020 [cited by applicant]
US 10839793B2 · Chao et al. · 2020 [cited by applicant]
US 11017766B2 · Chao et al. · 2021 [cited by applicant]
US 20030018475A1 · Basu et al. · 2003 [cited by applicant]
US 20050187770A1 · Kompe et al. · 2005 [cited by applicant]
US 20070294081A1 · Wang · 2007 [cited by applicant]
US 20080281598A1 · Eide · 2008 [cited by applicant]
US 20110055256A1 · Phillips et al. · 2011 [cited by applicant]
US 20120323557A1 · Koll et al. · 2012 [cited by applicant]
US 20130238336A1 · Sung et al. · 2013 [cited by applicant]
US 20130332147A1 · Corfield · 2013 [cited by applicant]
US 20140012577A1 · Freeman et al. · 2014 [cited by applicant]
US 20140012578A1 · Morioka · 2014 [cited by applicant]
US 20140272821A1 · Pitschel et al. · 2014 [cited by applicant]
US 20140280051A1 · Djugash · 2014 [cited by applicant]
US 20150006147A1 · Schmidt · 2015 [cited by applicant]
US 20150120288A1 · Thomson et al. · 2015 [cited by applicant]
US 20150142704A1 · London · 2015 [cited by applicant]
US 20150302855A1 · Kim et al. · 2015 [cited by applicant]
US 20150364129A1 · Gonzalez-Dominguez et al. · 2015 [cited by applicant]
US 20160035346A1 · Chengalvarayan · 2016 [cited by applicant]
US 20160125879A1 · Lovitt · 2016 [cited by examiner]
US 20160140218A1 · Moreno et al. · 2016 [cited by applicant]
US 20160162469A1 · Santos · 2016 [cited by applicant]
US 20160217788A1 · Stonehocker et al. · 2016 [cited by applicant]
US 20160217790A1 · Sharifi · 2016 [cited by applicant]
US 20160248768A1 · McLaren · 2016 [cited by examiner]
US 20160262017A1 · Lavee · 2016 [cited by examiner]
US 20160329048A1 · Li et al. · 2016 [cited by applicant]
US 20160350285A1 · Zhao et al. · 2016 [cited by applicant]
US 20160379638A1 · Basye et al. · 2016 [cited by applicant]
US 20170309271A1 · Chiang · 2017 [cited by applicant]
US 20170316305A1 · Liensberger · 2017 [cited by examiner]
US 20180018959A1 · Des Jardins et al. · 2018 [cited by applicant]
US 20180068653A1 · Trawick · 2018 [cited by applicant]
US 20180211650A1 · Knudson et al. · 2018 [cited by applicant]
US 20190102481A1 · Sreedhara · 2019 [cited by applicant]
US 20190318724A1 · Chao et al. · 2019 [cited by applicant]
US 20190318729A1 · Chao et al. · 2019 [cited by applicant]
US 20200104094A1 · White et al. · 2020 [cited by applicant]
US 20210074280A1 · Chao et al. · 2021 [cited by applicant]
CN 201332158 · 2009 [cited by applicant]
CN 101901599 · 2010 [cited by applicant]
CN 104282307 · 2015 [cited by applicant]
CN 104505091 · 2015 [cited by applicant]
CN 104575493 · 2015 [cited by applicant]
CN 104978015 · 2015 [cited by applicant]
CN 105190607 · 2015 [cited by applicant]
CN 105957516 · 2016 [cited by applicant]
CN 106710586 · 2017 [cited by applicant]
CN 106997762 · 2017 [cited by applicant]
CN 107623614 · 2018 [cited by applicant]
CN 107895578 · 2018 [cited by applicant]
WO 2015112149 · 2015 [cited by applicant]
WO 2015196063 · 2015 [cited by applicant]
Intellectual Property India; Extended Hearing Noticed issued in Application No. 201927050873; 4 pages; dated Jan. 30, 2024. [cited by applicant]
Intellectual Property India; Hearing Noticed issued in Application No. 201927050873; 3 pages; dated Dec. 19, 2023. [cited by applicant]
Levit, M. et al.; End-to-end speech recognition accuracy metric for voice-search tasks; IEEE International Conference on Acoustics; Speech and Signal Processing (ICASSP); Japan; pp. 5141-5144; dated 2012. [cited by applicant]
Sun, L. et al.; Generating language distance metrics by language recognition using acoustic features; 8th International Conference on Wireless communications & Signal Processing (WCSP); China; pp. 1-5; dated 2016. [cited by applicant]
China National Intellectual Property Administration; Notice of Allowance issued for Application No. 201880039581.6, 6 pages, dated Jul. 27, 2023. [cited by applicant]
European Patent Office; Intention to Grant issued in Application No. 20195508.5; 48 pages; dated Mar. 14, 2023. [cited by applicant]
China National Intellectual Property Administration; Notification of First Office Action issued for Application No. 201880039581.6, 19 pages, dated Mar. 1, 2023. [cited by applicant]
Intellectual Property Office of Singapore; Notice of Eligibility of Grant issued for Application No. 11201912061W, 4 pages, dated Dec. 13, 2022. [cited by applicant]
Intellectual Property India; Office Action issued in Application No. 201927050873; 6 pages; dated Mar. 12, 2021. [cited by applicant]
Eueropean Patent Office; Communication issued in Application No. 20195508.5; 11 pages; dated Mar. 11, 2021. [cited by applicant]
Eueropean Patent Office; Communication issue in Application No. 20195508.5; 13 pages; dated Dec. 7, 2020. [cited by applicant]
European Patent Office; Intention to Grant issue in Application No. 18722334.2; 48 pages; dated Jun. 8, 2020. [cited by applicant]
European Patent Office; International Search Report and Written Opinion of PCT Ser. No. PCT/US2018/027808 dated Nov. 26, 2018; 20 pages. [cited by applicant]
European Patent Office; Invitation to Pay Additional Fees in International Patent Application No. PCT Ser. No. PCT/US2018/027808 dated Oct. 2, 2018; 14 pages. [cited by applicant]
European Patent Office; International Search Report and Written Opinion of PCT Ser. No. PCT/US2018/027812; 14 pages; dated Oct. 1, 2018. [cited by applicant]
Gonzalez-Dominguez, J., et al. “A Real-Time End-to-End Multilingual Speech Recognition Architecture”. IEEE Journal of Selected Topics in Signal Processing, vol. 9, No. 4, Jun. 2015; pp. 749-759. [cited by applicant]
European Patent Office; Intention to Grant of EP Ser. No. 18722336.7; 44 pages; dated Dec. 20, 2019. [cited by applicant]
European Patent Office; Communication issue in Application No. 20177711.7; 9 pages; dated Aug. 25, 2020. [cited by applicant]
Intellectual Property India; Office Action issued in Application No. 201927051483; 6 pages; dated Mar. 18, 2021. [cited by applicant]
European Patent Office; Communication Pursuant to Article 94(3) EPC issue in Application No. 20177711.7; 5 pages; dated Sep. 28, 2021. [cited by applicant]
Intellectual Property Office of Singapore; Notice for Eligibility of Grant issued in Application No. 11201912053X; 4 pages; dated Dec. 13, 2022. [cited by applicant]
Intellectual Property Office of Singapore; Notice of Eligibility of Grant issued for Application No. 11201912053X, 4 pages, dated Dec. 13, 2022. [cited by applicant]
China National Intellectual Property Administration; Notification of First Office Action issued for Application No. 201880039579.9, 22 pages, dated Jan. 13, 2023. [cited by applicant]
China National Intellectual Property Administration; Notice of Allowance issued for Application No. 201880039579.9, 6 pages, dated May 29, 2023. [cited by applicant]
European Patent Office, Intention to Grant issue in Application No. 23191963.0; 49 pages; dated Jul. 22, 2024. [cited by applicant]
European Patent Office; Communication issued in Application No. 23191963.0; 5 pages; dated Nov. 21, 2023. [cited by applicant]
Intellectual Property India; Examination Report issued in Application No. 202428012271; 7 pages; dated Apr. 5, 2026. [cited by applicant]
Intellectual Property India; Hearing Notice issued in Application No. 202428012271; 2 pages; dated Apr. 15, 2026. [cited by applicant]