IP Library › Granted Patent US 11,798,541
Granted Patent B2
US 11,798,541 · App. 17/099,367 · Granted Oct 24, 2023

Automatically determining language for speech recognition of spoken utterance received via an automated assistant interface

Inventors: Pu-sen Chao (Los Altos, CA); Diego Melendo Casado (Mountain View, CA); Ignacio Lopez Moreno (New York, NY)
Assignee: GOOGLE LLC
G10L15/197G10L13/00G10L15/005G10L15/08G10L15/14G10L15/1822G10L15/22G10L15/30G10L2015/088G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,798,541
App. No.
17/099,367
Granted
Oct 24, 2023
Kind
B2
Abstract

Determining a language for speech recognition of a spoken utterance received via an automated assistant interface for interacting with an automated assistant. Implementations can enable multilingual interaction with the automated assistant, without necessitating a user explicitly designate a language to be utilized for each interaction. Implementations determine a user profile that corresponds to audio data that captures a spoken utterance, and utilize language(s), and optionally corresponding probabilities, assigned to the user profile in determining a language for speech recognition of the spoken utterance. Some implementations select only a subset of languages, assigned to the user profile, to utilize in speech recognition of a given spoken utterance of the user. Some implementations perform speech recognition in each of multiple languages assigned to the user profile, and utilize criteria to select only one of the speech recognitions as appropriate for generating and providing content that is responsive to the spoken utterance.

Claims (52)

1. A method implemented by one or more processors, the method comprising:

processing audio data, wherein the audio data is based on detection of spoken input of a user at a client device, the client device including an automated assistant interface for interacting with the automated assistant;

determining, based on processing of the audio data, that at least a portion of the audio data matches a user profile accessible to the automated assistant;

identifying at least one probabilistic metric assigned to the user profile and corresponding to a particular speech recognition model, for a particular language; and

based on the at least one probabilistic metric satisfying a threshold:

selecting the particular speech recognition model, for the particular language, for processing the audio data, and

processing the audio data, using the particular speech recognition model for to the particular language, to generate text, in the particular language, that corresponds to the spoken input; and

causing the automated assistant to provide responsive content that is determined based on the generated text.

2. The method of claim 1 , wherein the user profile further includes an additional probabilistic metric corresponding to at least one different speech recognition model, for a different language, and further comprising:

based on the additional probabilistic metric failing to satisfy the threshold:

refraining from processing the audio data using the different speech recognition model.

3. The method of claim 2 , further comprising:

identifying current contextual data associated with the audio data, wherein identifying the at least one probabilistic metric is based on a correspondence between the current contextual data and the at least one probabilistic metric.

4. The method of claim 3 , wherein the current contextual data identifies a location of the client device or an application that is being accessed via the client device when the spoken input is received.

5. The method of claim 3 , wherein the current contextual data identifies the client device.

6. The method of claim 1 , wherein the probabilistic metric is based on past interactions between the user and the automated assistant.

7. A method implemented by one or more processors, the method comprising:

receiving audio data, wherein the audio data is based on detection of spoken input of a user at a client device, the client device including an automated assistant interface for interacting with an automated assistant;

determining that the audio data corresponds to a user profile accessible to the automated assistant;

identifying a first language assigned to the user profile, and a first probability metric assigned to the first language in the user profile;

selecting a first speech recognition model for the first language, wherein selecting the first speech recognition model for the first language is based on identifying the first language as assigned to the user profile;

using the selected first speech recognition model to generate first text in the first language, and a first measure that indicates a likelihood the first text is an appropriate representation of the spoken input;

identifying a second language assigned to the user profile, and a second probability metric assigned to the second language in the user profile;

selecting a second speech recognition model for the second language, wherein selecting the second speech recognition model for the second language is based on identifying the second language as assigned to the user profile;

using the selected second speech recognition model to generate second text in the second language, and a second measure that indicates a likelihood the second text is an appropriate representation of the spoken input;

selecting the first text in the first language in lieu of the second text in the second language, wherein selecting the first text in the first language in lieu of the second text in the second language is based on: the first probability metric, the first measure, the second probability metric, and the second measure; and

in response to selecting the first text:

causing the automated assistant to provide responsive content that is determined based on the selected first text.

8. The method of claim 7 , further comprising:

identifying a current context associated with the audio data;

wherein identifying the first probability metric is based on the first probability metric corresponding to the current context; and

wherein identifying the second probability metric is based on the second probability metric corresponding to the current context.

9. The method of claim 7 , wherein determining that the audio data corresponds to the user profile is based on comparing features of the audio data to features of the user profile.

10. A system comprising:

one or more processors; and

memory configured to store instructions that, when executed by the one or more processors cause the one or more processors to perform operations that include:

processing audio data, wherein the audio data is based on detection of spoken input of a user at a client device, the client device including an automated assistant interface for interacting with the automated assistant;

determining, based on processing of the audio data, that at least a portion of the audio data matches a user profile accessible to the automated assistant;

identifying at least one probabilistic metric assigned to the user profile and corresponding to a particular speech recognition model, for a particular language; and

based on the at least one probabilistic metric satisfying a threshold:

selecting the particular speech recognition model, for the particular language, for processing the audio data, and

processing the audio data, using the particular speech recognition model for to the particular language, to generate text, in the particular language, that corresponds to the spoken input; and

causing the automated assistant to provide responsive content that is determined based on the generated text.

11. The system of claim 10 , wherein the user profile further includes an additional probabilistic metric corresponding to at least one different speech recognition model, for a different language, and wherein the operations further comprise:

based on the additional probabilistic metric failing to satisfy the threshold:

refraining from processing the audio data using the different speech recognition model.

12. The system of claim 11 , wherein the operations further comprise:

identifying current contextual data associated with the audio data, wherein

identifying the at least one probabilistic metric is based on a correspondence between the current contextual data and the at least one probabilistic metric.

13. The system of claim 11 , wherein the current contextual data identifies a location of the client device or an application that is being accessed via the client device when the spoken input is received.

14. The system of claim 11 , wherein the current contextual data identifies the client device.

15. The system of claim 10 , wherein the probabilistic metric is based on past interactions between the user and the automated assistant.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 29, 2020
From: CHAO, PU-SEN; CASADO, DIEGO MELENDO; MORENO, IGNACIO LOPEZ
To: GOOGLE LLC
Reel/Frame 054764/0769 →
Continuity (2)
Continuation 15769013
Related Publication 20210074280A1 · Mar 11, 2021