IP Library Granted Patent US 10,446,153
Granted Patent B2
US 10,446,153 · App. 16/216,752 · Granted Oct 15, 2019

Providing pre-computed hotword models

Inventor: Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L15/22G06F3/167G10L15/063G10L15/08G10L15/265G10L15/30G06F3/04842G10L15/18G10L2015/0631G10L2015/0638G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,446,153
App. No.
16/216,752
Granted
Oct 15, 2019
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for obtaining, for each of multiple words or sub-words, audio data corresponding to multiple users speaking the word or sub-word; training, for each of the multiple words or sub-words, a pre-computed hotword model for the word or sub-word based on the audio data for the word or sub-word; receiving a candidate hotword from a computing device; identifying one or more pre-computed hotword models that correspond to the candidate hotword; and providing the identified, pre-computed hotword models to the computing device.

Claims (40)

1. A method comprising:

displaying, by data processing hardware of a user device, a prompt in a graphical user interface executing on the data processing hardware, the prompt requesting a user to provide a personalized term for initiating the user device to perform a particular action;

receiving, at the data processing hardware, audio data corresponding to the user speaking the personalized term;

obtaining, by the data processing hardware using the audio data corresponding to the user speaking the personalized term, a detection model that corresponds to the personalized term;

after obtaining the detection model that corresponds to the personalized term, detecting, by the data processing hardware, the personalized term in an utterance spoken by the user using the obtained detection model; and

in response to detecting the personalized term in the utterance, initiating, by the data processing hardware, the user device to perform the particular action.

2. The method of claim 1 , further comprising, in response to receiving the audio data corresponding to the user speaking the personalized term, displaying, by the data processing hardware, a transcription of the personalized term spoken by the user in the graphical user interface.

3. The method of claim 2 , further comprising, when displaying the transcription of the personalized term in the graphical user interface, displaying, by the data processing hardware, a graphical button in the graphical user interface, the graphical button when selected by the user, causes the data processing hardware to prompt the user to speak the personalized term again.

4. The method of claim 3 , further comprising:

determining, by the data processing hardware, whether a user selection indication is received indicating selection of the graphical button displayed in the user interface; and

when the user selection indication is not received, obtaining the detection model that corresponds to the personalized term.

5. The method of claim 2 , further comprising, prior to obtaining the detection model that corresponds to the personalized term, receiving, at the data processing hardware, a user input indication indicating that the accepts the transcription of the personalized term displayed in the graphical user interface.

6. The method of claim 1 , wherein the personalized term comprises two or more words.

7. The method of claim 1 , further comprising providing the received audio data corresponding to the user speaking the predetermined term from the data processing hardware to a server-based configuration engine, the server-based configuration engine configured to dynamically create the detection model that corresponds to the personalized term based on the received audio data corresponding to the user speaking the predetermined term.

8. The method of claim 1 , further comprising providing the received audio data corresponding to the user speaking the predetermined term from the data processing hardware to a server-based configuration engine, the server-based configuration engine configured to dynamically create the detection model that corresponds to the personalized term based on a transcription of the received audio data corresponding to the user speaking the predetermined term.

9. The method of claim 1 , further comprising providing the received audio data corresponding to the user speaking the predetermined term from the data processing hardware to a server-based configuration engine, the server-based configuration engine configured to:

identify the detection model that corresponds to the personalized term in a vocabulary database, the detection model stored in the vocabulary database prior to the data processing hardware receiving the audio data corresponding to the user speaking the predetermined term; and

provide the detection model stored in the vocabulary database to the data processing hardware of the user device.

10. The method of claim 1 , further comprising, when initiating the user device to perform the particular action in response to detecting the personalized term in the utterance, displaying, by the data processing hardware, a description of the particular action to be performed by the user device.

11. A user device comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed by the data processing hardware cause the data processing hardware to perform operations comprising:

displaying a prompt in a graphical user interface executing on the data processing hardware, the prompt requesting a user to provide a personalized term for initiating the user device to perform a particular action;

receiving audio data corresponding to the user speaking the personalized term;

obtaining, using the audio data corresponding to the user speaking the personalized term, a detection model that corresponds to the personalized term;

after obtaining the detection model that corresponds to the personalized term, detecting the personalized term in an utterance spoken by the user using the obtained detection model; and

in response to detecting the personalized term in the utterance, initiating the user device to perform the particular action in the graphical user interface.

12. The user device of claim 11 , wherein the operations further comprise, in response to receiving the audio data corresponding to the user speaking the personalized term, displaying a transcription of the personalized term spoken by the user in the graphical user interface.

13. The user device of claim 12 , wherein the operations further comprise, when displaying the transcription of the personalized term in the graphical user interface, displaying a graphical button in the graphical user interface, the graphical button when selected by the user, causes the data processing hardware to prompt the user to speak the personalized term again.

14. The user device of claim 13 , wherein the operations further comprise:

determining whether a user selection indication is received indicating selection of the graphical button displayed in the user interface; and

when the user selection indication is not received, obtaining the detection model that corresponds to the personalized term.

15. The user device of claim 12 , wherein the operations further comprise, prior to obtaining the detection model that corresponds to the personalized term, receiving a user input indication indicating that the accepts the transcription of the personalized term displayed in the graphical user interface.

16. The user device of claim 11 , wherein the personalized term comprises two or more words.

17. The user device of claim 11 , wherein the operations further comprise providing the received audio data corresponding to the user speaking the predetermined term from the data processing hardware to a server-based configuration engine, the server-based configuration engine configured to dynamically create the detection model that corresponds to the personalized term based on the received audio data corresponding to the user speaking the predetermined term.

18. The user device of claim 11 , wherein the operations further comprise providing the received audio data corresponding to the user speaking the predetermined term from the data processing hardware to a server-based configuration engine, the server-based configuration engine configured to dynamically create the detection model that corresponds to the personalized term based on a transcription of the received audio data corresponding to the user speaking the predetermined term.

19. The user device of claim 11 , wherein the operations further comprise providing the received audio data corresponding to the user speaking the predetermined term from the data processing hardware to a server-based configuration engine, the server-based configuration engine configured to:

identify the detection model that corresponds to the personalized term in a vocabulary database, the detection model stored in the vocabulary database prior to the data processing hardware receiving the audio data corresponding to the user speaking the predetermined term; and

provide the detection model stored in the vocabulary database to the data processing hardware of the user device.

20. The user device of claim 11 , wherein the operations further comprise, when initiating the user device to perform the particular action in response to detecting the personalized term in the utterance, displaying a description of the particular action to be performed by the user device in the graphical user interface.

Continuity (6)
Continuation 15875996 · Jan 19, 2018
Continuation 15463786 · Mar 20, 2017
Continuation 15288241 · Oct 7, 2016
Continuation 15001894 · Jan 20, 2016
Continuation 14340833 · Jul 25, 2014
Related Publication 20190108840A1 · Apr 11, 2019