IP Library Granted Patent US 10,186,268
Granted Patent B2
US 10,186,268 · App. 15/875,996 · Granted Jan 22, 2019

Providing pre-computed hotword models

Inventor: Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L15/22G06F3/167G10L15/063G10L15/08G10L15/265G10L15/30G06F3/04842G10L15/18G10L2015/0631G10L2015/0638G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,186,268
App. No.
15/875,996
Granted
Jan 22, 2019
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for obtaining, for each of multiple words or sub-words, audio data corresponding to multiple users speaking the word or sub-word; training, for each of the multiple words or sub-words, a pre-computed hotword model for the word or sub-word based on the audio data for the word or sub-word; receiving a candidate hotword from a computing device; identifying one or more pre-computed hotword models that correspond to the candidate hotword; and providing the identified, pre-computed hotword models to the computing device.

Claims (29)

1. A computer-implemented method comprising:

receiving, by a server-based hotword configuration engine, text that a user has typed or entered by touch input into a mobile computing device or digital assistant device for use as a custom hotword, wherein the custom hotword is a personalized term that, when spoken in advance of a voice command, initiates a wake up process on the mobile computing device or digital assistant device for processing the voice command;

obtaining, by the server-based hotword configuration engine, a hotword detection model that was trained by the server-based hotword configuration engine prior to receiving the text for use as the custom hotword from the mobile computing device or digital assistant device, the hotword detection model configured to detect only a word or sub-word of the custom hotword being spoken without using any audio samples of the user speaking the custom hotword; and

providing, by the server-based hotword configuration engine, the hotword detection model to the mobile computing device or the digital assistant device for use in detecting a spoken utterance of the custom hotword by the user without transcribing the custom hotword from the spoken utterance.

2. The method of claim 1 , further comprising receiving and processing a voice command that (i) was preceded by the custom hotword in an utterance that was received by the mobile computing device or the digital assistant device, and (ii) was received after the hotword detection model was provided to the mobile computing device or the digital assistant device.

3. The method of claim 1 , wherein the text is typed or entered into the mobile computing device or digital assistant device in response to a prompt.

4. The method of claim 1 , wherein the text comprises two or more words.

5. The method of claim 1 , wherein the hotword detection model is trained based on portions of audio samples of other users speaking other words that are not the custom hotword.

6. The method of claim 1 , wherein the hotword detection model comprises a pre-computed model that was already stored in a vocabulary database when the text was typed or entered.

7. A system comprising:

one or more computers of a server-based hotword configuration engine; and

a non-transitory computer-readable medium coupled to the one or more computers having instructions stored thereon which, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

receiving text that a user has typed or entered by touch input into a mobile computing device or digital assistant device for use as a custom hotword, wherein the custom hotword is a personalized term that, when spoken in advance of a voice command, initiates a wake up process on the mobile computing device or digital assistant device for processing the voice command;

obtaining a hotword detection model that was trained by the server-based hotword configuration engine prior to receiving the text for use as the custom hotword from the mobile computing device or digital assistant device, the hotword detection model configured to detect only a word or sub-word of the custom hotword being spoken without using any audio samples of the user speaking the custom hotword; and

providing the hotword detection model to the mobile computing device or the digital assistant device for use in detecting a spoken utterance of the custom hotword by the user without transcribing the custom hotword from the spoken utterance.

8. The system of claim 7 , wherein the operations further comprise receiving and processing a voice command that (i) was preceded by the custom hotword in an utterance that was received by the mobile computing device or the digital assistant device, and (ii) was received after the hotword detection model was provided to the mobile computing device or the digital assistant device.

9. The system of claim 7 , wherein the text is typed or entered into the mobile computing device or digital assistant device in response to a prompt.

10. The system of claim 7 , wherein the text comprises two or more words.

11. The system of claim 7 , wherein the hotword detection model is trained based on portions of audio samples of other users speaking other words that are not the custom hotword.

12. The system of claim 7 , wherein the hotword detection model comprises a pre-computed model that was already stored in a vocabulary database when the text was typed or entered.

13. A non-transitory computer storage medium encoded with a computer program, the computer program comprising instructions that when executed by one or more processors cause the one or more processors to perform operations comprising:

receiving, by a server-based hotword configuration engine, text that a user has typed or entered by touch input into a mobile computing device or digital assistant device for use as a custom hotword, wherein the custom hotword is a personalized term that, when spoken in advance of a voice command, initiates a wake up process on the mobile computing device or digital assistant device for processing the voice command;

obtaining, by the server-based hotword configuration engine, a hotword detection model that was trained by the server-based hotword configuration engine prior to receiving the text for use as the custom hotword from the mobile computing device or digital assistant device, the hotword detection model configured to detect only a word or sub-word the custom hotword being spoken without using any audio samples of the user speaking the custom hotword; and

providing, by the server-based hotword configuration engine, the hotword detection model to the mobile computing device or the digital assistant device for use in detecting the custom hotword by the user without transcribing the custom hotword from the spoken utterance.

14. The medium of claim 13 , further comprising receiving and processing a voice command that (i) was preceded by the custom hotword in an utterance that was received by the mobile computing device or the digital assistant device, and (ii) was received after the hotword detection model was provided to the mobile computing device or the digital assistant device.

15. The medium of claim 13 , wherein the text is typed or entered into the mobile computing device or digital assistant device in response to a prompt.

16. The medium of claim 13 , wherein the text comprises two or more words.

17. The medium of claim 13 , wherein the hotword detection model is trained based on portions of audio samples of other users speaking other words that are not the custom hotword.

18. The medium of claim 13 , wherein the hotword detection model comprises a pre-computed model that was already stored in a vocabulary database when the text was typed or entered.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2018
From: SHARIFI, MATTHEW
To: GOOGLE INC.
Reel/Frame 044680/0076 →
CHANGE OF NAME Recorded Jan 19, 2018
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 045102/0266 →
Continuity (5)
Continuation 15463786 · Mar 20, 2017
Continuation 15288241 · Oct 7, 2016
Continuation 15001894 · Jan 20, 2016
Continuation 14340833 · Jul 25, 2014
Related Publication 20180166078A1 · Jun 14, 2018
Cited By (1)
US 12,609,112