IP Library Granted Patent US 10,621,987
Granted Patent B2
US 10,621,987 · App. 16/669,503 · Granted Apr 14, 2020

Providing pre-computed hotword models

Inventor: Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L15/22G06F3/167G10L15/063G10L15/08G10L15/265G10L15/30G06F3/04842G10L15/18G10L2015/0631G10L2015/0638G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,621,987
App. No.
16/669,503
Granted
Apr 14, 2020
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for obtaining, for each of multiple words or sub-words, audio data corresponding to multiple users speaking the word or sub-word; training, for each of the multiple words or sub-words, a pre-computed hotword model for the word or sub-word based on the audio data for the word or sub-word; receiving a candidate hotword from a computing device; identifying one or more pre-computed hotword models that correspond to the candidate hotword; and providing the identified, pre-computed hotword models to the computing device.

Claims (32)

1. A method comprising:

receiving, at data processing hardware, from a user device, text corresponding to a personalized term for initiating the user device to perform a particular action, the user device configured to:

receive, in a graphical user interface executing on the user device, an input indication corresponding to a user entering the text corresponding to the personalized term; and

send the text corresponding to the personalized term to the data processing hardware;

receiving, at the data processing hardware, from the user device, audio data corresponding to an utterance spoken by the user; and

when the personalized term is detected in the utterance spoken by the user, initiating, by the data processing hardware, the user device to perform the particular action.

2. The method of claim 1 , wherein the user device is further configured to, prior to receiving the input indication in the graphical user interface, display a prompt in the graphical user interface, the prompt requesting the user to provide the personalized term.

3. The method of claim 1 , further comprising, prior to receiving the audio data corresponding to the utterance spoken by the user, obtaining, by the data processing hardware, using the text corresponding to the personalized term, a detection model that corresponds to the personalized term.

4. The method of claim 3 , further comprising detecting, by the data processing hardware, using the obtained detection model, the personalized term in the audio data corresponding to the utterance spoken by the user.

5. The method of claim 3 , wherein obtaining the detection model comprises dynamically creating the detection model based on the received text corresponding to the personalized term.

6. The method of claim 3 , wherein obtaining the detection model comprises retrieving the detection model that corresponds to the personalized term from a vocabulary database.

7. The method of claim 6 , further comprising storing, by the data processing hardware, the detection model in the vocabulary database prior to receiving the text corresponding to personalized term.

8. The method of claim 3 , wherein the detection model is trained based on portions of audio samples of other users speaking other words that are not the personalized term.

9. The method of claim 1 , wherein the personalized term comprises two or more words.

10. The method of claim 1 , wherein the personalized term comprises a single word.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed by the data processing hardware cause the data processing hardware to perform operations comprising:

receiving, from a user device, text corresponding to a personalized term for initiating the user device to perform a particular action, the user device configured to:

receive, in a graphical user interface executing on the user device, an input indication corresponding to a user entering the text corresponding to the personalized term; and

send the text corresponding to the personalized term to the data processing hardware;

receiving, from the user device, audio data corresponding to an utterance spoken by the user; and

when the personalized term is detected in the utterance spoken by the user, initiating the user device to perform the particular action.

12. The system of claim 11 , wherein the user device is further configured to, prior to receiving the input indication in the graphical user interface, display a prompt in the graphical user interface, the prompt requesting the user to provide the personalized term.

13. The system of claim 11 , wherein the operations further comprise, prior to receiving the audio data corresponding to the utterance spoken by the user, obtaining, using the text corresponding to the personalized term, a detection model that corresponds to the personalized term.

14. The system of claim 13 , wherein the operations further comprise detecting, using the obtained detection model, the personalized term in the audio data corresponding to the utterance spoken by the user.

15. The system of claim 13 , wherein obtaining the detection model comprises dynamically creating the detection model based on the received text corresponding to the personalized term.

16. The system of claim 13 , wherein obtaining the detection model comprises retrieving the detection model that corresponds to the personalized term from a vocabulary database.

17. The system of claim 16 , wherein the operations further comprise storing the detection model in the vocabulary database prior to receiving the text corresponding to personalized term.

18. The system of claim 13 , wherein the detection model is trained based on portions of audio samples of other users speaking other words that are not the personalized term.

19. The system of claim 11 , wherein the personalized term comprises two or more words.

20. The system of claim 11 , wherein the personalized term comprises a single word.

Continuity (8)
Continuation 16529300 · Aug 1, 2019
Continuation 16216752 · Dec 11, 2018
Continuation 15875996 · Jan 19, 2018
Continuation 15463786 · Mar 20, 2017
Continuation 15288241 · Oct 7, 2016
Continuation 15001894 · Jan 20, 2016
Continuation 14340833 · Jul 25, 2014
Related Publication 20200066275A1 · Feb 27, 2020