IP Library Granted Patent US 9,911,419
Granted Patent B2
US 9,911,419 · App. 15/463,786 · Granted Mar 6, 2018

Providing pre-computed hotword models

Inventor: Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L15/22G06F3/167G10L15/063G10L15/08G10L15/30G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,911,419
App. No.
15/463,786
Filed
Mar 20, 2017
Granted
Mar 6, 2018
Kind
B2
Art Unit
2667
USPC
704/235
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for obtaining, for each of multiple words or sub-words, audio data corresponding to multiple users speaking the word or sub-word; training, for each of the multiple words or sub-words, a pre-computed hotword model for the word or sub-word based on the audio data for the word or sub-word; receiving a candidate hotword from a computing device; identifying one or more pre-computed hotword models that correspond to the candidate hotword; and providing the identified, pre-computed hotword models to the computing device.

Claims (54)

1. A computer-implemented method comprising:

receiving audio data corresponding to a user's utterance of (i) a custom hotword that is unique to the user and that the user has previously associated with a particular action, wherein the custom hotword is a voice action initiation command that causes initiation of the particular action associated with the custom hotword, and (ii) one or more other terms, associated with the particular action, that include a question, command, declaration, or other request that can be speech recognized and acted on once the custom hotword causes initiation of the particular action; and

in response to receiving the audio data corresponding to the user's utterance of (i) the custom hotword that is a voice action initiation command, unique to the user, and that the user has previously associated with the particular action, and (ii) the one or more other terms, associated with the particular action, that include a question, command, declaration, or other request that can be speech recognized and acted on once the custom hotword causes initiation of the particular action, providing, for output, a user interface that includes (i) an indication that the particular action will be performed, and (ii) a transcription that an automated speech recognizer has generated for the one or more other terms.

2. The computer-implemented method of claim 1 , wherein the custom hotword comprises a sequence of two or more words that causes a mobile device to process the user's utterance beginning with the custom hotword as a voice action initiation command and a voice command.

3. The computer-implemented method of claim 2 , further comprising:

identifying, at a server, one or more pre-computed hotword models that correspond to a word of the two or more words of the custom hotword; and

receiving the identified, pre-computed hotword models that correspond to each word of the two or more words of the candidate hotword in response to receiving the audio data corresponding to the user's utterance.

4. The computer-implemented method of claim 1 , wherein receiving the audio data corresponding to the user's utterance of (i) the custom hotword that is a voice action initiation command, unique to the user and that the user has previously associated with the particular action, and (ii) one or more other terms, associated with the particular action, that include a question, command, declaration, or other request that can be speech recognized and acted on once the custom hotword causes initiation of the particular action comprises:

obtaining a locale of the user from a mobile device associated with the user when receiving the audio corresponding to the user's utterance.

5. The computer-implemented method of claim 1 , wherein providing, for output, a user interface that includes (i) the indication that the particular action will be performed, and (ii) the transcription that the automated speech recognizer has generated for the one or more other terms comprises:

providing, for output, the user interface in response to the user interacting with a plurality of other user interfaces provided for output, further comprising:

prompting a request that the user associates the custom hotword with the particular action;

displaying, in response to detecting the user's utterance in which the user associates the custom hotword with the particular action, a proposed transcription of the detected user's utterance; and

displaying, in response to a confirmation that the proposed transcription is correct, a confirmation of the transcription of the detected user's utterance.

6. The computer-implemented method of claim 3 , comprising:

generating, at the server, the one or more pre-computed hotword models based on i) a waveform of the user's utterance and ii) a transcription associated with the audio data corresponding to the user's utterance.

7. The computer-implemented method of claim 3 , wherein identifying the one or more pre-computed hotword models that correspond to the word of the two or more words of the custom hotword comprises:

training, by the server, the one or more pre-computed hotword models for each of the two or more words of the custom hotword.

8. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving audio data corresponding to a user's utterance of (i) a custom hotword that is unique to the user and that the user has previously associated with a particular action, wherein the custom hotword is a voice action initiation command that causes initiation of the particular action associated with the custom hotword, and (ii) one or more other terms, associated with the particular action, that include a question, command, declaration, or other request that can be speech recognized and acted on once the custom hotword causes initiation of the particular action; and

in response to receiving the audio data corresponding to the user's utterance of (i) the custom hotword that is a voice action initiation command, unique to the user, and that the user has previously associated with the particular action, and (ii) the one or more other terms, associated with the particular action, that include a question, command, declaration, or other request that can be speech recognized and acted on once the custom hotword causes initiation of the particular action, providing, for output, a user interface that includes (i) an indication that the particular action will be performed, and (ii) a transcription that an automated speech recognizer has generated for the one or more other terms.

9. The system of claim 8 , wherein the custom hotword comprises a sequence of two or more words that causes a mobile device to process the user's utterance beginning with the custom hotword as a voice action initiation command and a voice command.

10. The system of claim 9 , the operations further comprise:

identifying, at a server, one or more pre-computed hotword models that correspond to a word of the two or more words of the custom hotword; and

receiving the identified, pre-computed hotword models that correspond to each word of the two or more words of the candidate hotword in response to receiving the audio data corresponding to the user's utterance.

11. The system of claim 8 , wherein receiving the audio data corresponding to the user's utterance of (i) the custom hotword that is a voice action initiation command, unique to the user, and that the user has previously associated with the particular action, and (ii) one or more other terms, associated with the particular action, that include a question, command, declaration, or other request that can be speech recognized and acted on once the custom hotword causes initiation of the particular action, the operations further comprise:

obtaining a locale of the user from a mobile device associated with the user when receiving the audio corresponding to the user's utterance.

12. The system of claim 8 , wherein providing, for output, a user interface that includes (i) the indication that the particular action will be performed, and (ii) the transcription that the automated speech recognizer has generated for the one or more other terms the operations further comprise:

providing, for output, the user interface in response to the user interacting with a plurality of other user interfaces provided for output, further comprising:

prompting a request that the user associates the custom hotword with the particular action;

displaying, in response to detecting the user's utterance in which the user associates the custom hotword with the particular action, a proposed transcription of the detected user's utterance; and

displaying, in response to a confirmation that the proposed transcription is correct, a confirmation of the transcription of the detected user's utterance.

13. The system of claim 10 , the operations further comprise:

generating, at the server, the one or more pre-computed hotword models based on i) a waveform of the user's utterance and ii) a transcription associated with the audio data corresponding to the user's utterance.

14. The system of claim 10 , wherein identifying the one or more pre-computed hotword models that correspond to the word of the two or more words of the custom hotword the operations further comprise:

training, by the server, the one or more pre-computed hotword models for each of the two or more words of the custom hotword.

15. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving audio data corresponding to a user's utterance of (i) a custom hotword that is unique to the user and that the user has previously associated with a particular action, wherein the custom hotword is a voice action initiation command that causes initiation of the particular action associated with the custom hotword, and (ii) one or more other terms, associated with the particular action, that include a question, command, declaration, or other request that can be speech recognized and acted on once the custom hotword causes initiation of the particular action; and

in response to receiving the audio data corresponding to the user's utterance of (i) the custom hotword that is a voice action initiation command, unique to the user, and that the user has previously associated with the particular action, and (ii) the one or more other terms, associated with the particular action, that include a question, command, declaration, or other request that can be speech recognized and acted on once the custom hotword causes initiation of the particular action, providing, for output, a user interface that includes (i) an indication that the particular action will be performed, and (ii) a transcription that an automated speech recognizer has generated for the one or more other terms.

16. The computer readable-medium of claim 15 , wherein the custom hotword comprises a sequence of two or more words that causes a mobile device to process the user's utterance beginning with the custom hotword as a voice action initiation command and a voice command.

17. The computer readable-medium of claim 16 , the operations comprising:

identifying, at a server, one or more pre-computed hotword models that correspond to a word of the two or more words of the custom hotword; and

receiving the identified, pre-computed hotword models that correspond to each word of the two or more words of the candidate hotword in response to receiving the audio data corresponding to the user's utterance.

18. The computer readable-medium of claim 15 , wherein receiving the audio data corresponding to the user's utterance of (i) the custom hotword that is a voice action initiation command, unique to the user, and that the user has previously associated with the particular action, and (ii) one or more other terms, associated with the particular action, that include a question, command, declaration, or other request that can be speech recognized and acted on once the custom hotword causes initiation of the particular action, the operations comprising:

obtaining a locale of the user from a mobile device associated with the user when receiving the audio corresponding to the user's utterance.

19. The computer-readable medium of claim 15 , wherein providing, for output, a user interface that includes (i) the indication that the particular action will be performed, and (ii) the transcription that the automated speech recognizer has generated for the one or more other terms the operations comprising:

providing, for output, the user interface in response to the user interacting with a plurality of other user interfaces provided for output, further comprising:

prompting a request that the user associates the custom hotword with the particular action;

displaying, in response to detecting the user's utterance in which the user associates the custom hotword with the particular action, a proposed transcription of the detected user's utterance; and

displaying, in response to a confirmation that the proposed transcription is correct, a confirmation of the transcription of the detected user's utterance.

20. The computer-readable medium of claim 17 , the operations comprising:

generating, at the server, the one or more pre-computed hotword models based on i) a waveform of the user's utterance and ii) a transcription associated with the audio data corresponding to the user's utterance.

Assignments (2)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044129/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2017
From: SHARIFI, MATTHEW
To: GOOGLE INC.
Reel/Frame 041649/0079 →
Continuity (4)
Continuation 15288241 · Oct 7, 2016
Continuation 15001894 · Jan 20, 2016
Continuation 14340833 · Jul 25, 2014
Related Publication 20170193995A1 · Jul 6, 2017