IP Library Granted Patent US 11,914,925
Granted Patent B2
US 11,914,925 · App. 17/812,320 · Granted Feb 27, 2024

Multi-modal input on an electronic device

Inventors: Brandon M. Ballinger (San Franciso, CA); Johan Schalkwyk (Scarsdale, NY); Michael H. Cohen (Portola Valley, CA); William J. Byrne (Davis, CA); Gudmundur Hafsteinsson (Los Gatos, CA); Michael J. Lebeau (New York, NY)
Assignee: Google LLC
G06F3/167G06F3/04886G06F40/284G06F40/58G10L15/005G10L15/18G10L15/183G10L15/22G10L15/26G10L15/30G10L15/197G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,914,925
App. No.
17/812,320
Granted
Feb 27, 2024
Kind
B2
Abstract

A computer-implemented input-method editor process includes receiving a request from a user for an application-independent input method editor having written and spoken input capabilities, identifying that the user is about to provide spoken input to the application-independent input method editor, and receiving a spoken input from the user. The spoken input corresponds to input to an application and is converted to text that represents the spoken input. The text is provided as input to the application.

Claims (42)

1. A computer-implemented method when executed on data processing hardware of an electronic device causes the data processing hardware to perform operations comprising:

receiving a speech utterance from a particular user, the speech utterance indicating that the particular user intends to provide a spoken query to an application executing on the electronic device;

receiving, from the particular user, the spoken query to the application;

obtaining a context-specific language model, the context-specific language model trained on training queries associated with a category associated with the application; and

converting, using a speech recognizer and the context-specific language model, the spoken query into corresponding text for processing by the application, the speech recognizer comprising a finite state transducer (FST) encoding a Hidden Markov Model (HMM).

2. The method of claim 1 , wherein the context-specific language model defines multiple language sub-models tailored to the application, each language sub-model defining a particular rule set for use in determining a likely intent of the particular user.

3. The method of claim 1 , wherein the speech utterance and the spoken query are captured by a microphone of the electronic device.

4. The method of claim 1 , wherein the application comprises a music player application.

5. The method of claim 1 , wherein the operations further comprise:

performing intermediate processing on the received spoken query,

wherein the speech recognizer uses the intermediate processing on the received spoken query to convert the spoken query into the corresponding text.

6. The method of claim 1 , wherein the operations further comprise:

executing an application-independent input method editor configured to receive spoken in put for a plurality of applications executable by the electronic device, the plurality of applications comprising the application,

wherein receiving the speech utterance comprises receiving the speech utterance at the application-independent input method editor.

7. The method of claim 6 , wherein the application-independent input method editor has written and spoken input capabilities.

8. The method of claim 1 , wherein the operations further comprise:

determining the category associated with the application,

wherein obtaining the context-specific language model is based on the category associated with the application.

9. The method of claim 1 , wherein the operations further comprise processing, by the application, the corresponding text converted from the spoken query.

10. The method of claim 1 , wherein the electronic device comprises an audio output device.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a speech utterance from a particular user, the speech utterance indicating that the particular user intends to provide a spoken query to an application executing on the electronic device;

receiving, from the particular user, the spoken query to the application;

obtaining a context-specific language model, the context-specific language model trained on training queries associated with a category associated with the application; and

converting, using a speech recognizer and the context-specific language model, the spoken query into corresponding text for processing by the application, the speech recognizer comprising a finite state transducer (FST) encoding a Hidden Markov Model (HMM).

12. The system of claim 11 , wherein the context-specific language model defines multiple language sub-models tailored to the application, each language sub-model defining a particular rule set for use in determining a likely intent of the particular user.

13. The system of claim 11 , wherein the speech utterance and the spoken query are captured by a microphone of the electronic device.

14. The system of claim 11 , wherein the application comprises a music player application.

15. The system of claim 11 , wherein the operations further comprise:

performing intermediate processing on the received spoken query,

wherein the speech recognizer uses the intermediate processing on the received spoken query to convert the spoken query into the corresponding text.

16. The system of claim 11 , wherein the operations further comprise:

executing an application-independent input method editor configured to receive spoken in put for a plurality of applications executable by the electronic device, the plurality of applications comprising the application,

wherein receiving the speech utterance comprises receiving the speech utterance at the application-independent input method editor.

17. The system of claim 16 , wherein the application-independent input method editor has written and spoken input capabilities.

18. The system of claim 11 , wherein the operations further comprise:

determining a category associated with the application,

wherein obtaining the context-specific language model is based on the category associated with the application.

19. The system of claim 11 , wherein the operations further comprise processing, by the application, the corresponding text converted from the spoken query n.

20. The system of claim 11 , wherein the electronic device comprises an audio output device.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2022
From: BALLINGER, BRANDON M.; SCHALKWYK, JOHAN; COHEN, MICHAEL H.; BYRNE, WILLIAM J.; HAFSTEINSSON, GUDMUNDUR; LEBEAU, MICHAEL J.
To: GOOGLE INC.
Reel/Frame 061093/0226 →
CHANGE OF NAME Recorded Sep 14, 2022
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 061433/0558 →
Continuity (9)
Continuation 16892749 · Jun 4, 2020
Continuation 16169279 · Oct 24, 2018
Continuation 14988408 · Jan 5, 2016
Continuation 14299837 · Jun 9, 2014
Continuation 13249172 · Sep 29, 2011
Continuation 12977003 · Dec 22, 2010
Provisional Application 61330219 · Apr 30, 2010
Provisional Application 61289968 · Dec 23, 2009
Related Publication 20220405046A1 · Dec 22, 2022
Cited By (1)
US 12,386,585