IP Library Granted Patent US 11,682,396
Granted Patent B2
US 11,682,396 · App. 17/304,459 · Granted Jun 20, 2023

Providing pre-computed hotword models

Inventor: Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L15/22G06F3/167G10L15/063G10L15/08G10L15/26G10L15/30G06F3/04842G10L15/18G10L2015/0631G10L2015/0638G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,682,396
App. No.
17/304,459
Granted
Jun 20, 2023
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for obtaining, for each of multiple words or sub-words, audio data corresponding to multiple users speaking the word or sub-word; training, for each of the multiple words or sub-words, a pre-computed hotword model for the word or sub-word based on the audio data for the word or sub-word; receiving a candidate hotword from a computing device; identifying one or more pre-computed hotword models that correspond to the candidate hotword; and providing the identified, pre-computed hotword models to the computing device.

Claims (36)

1. A computer-implemented method when executed on data processing hardware of a user device causes the data processing hardware to perform operations comprising:

displaying a prompt in a graphical user interface executing on the data processing hardware, the prompt requesting a user to:

provide a candidate term; and

associate the candidate term with a particular action for initiating the user device to perform when the user speaks the candidate term;

receiving, via the graphical user interface, a text-based input corresponding to the candidate term for initiating the user device to perform the particular action;

transmitting the text-based input to a server-based configuration engine, the text-based input when received by the server-based configuration engine causing the server-based configuration engine to create a detection model that corresponds to the candidate term;

receiving audio data corresponding to an utterance of the candidate term spoken by the user; and

when the candidate term is detected in the utterance spoken by the user using the detection model obtained by the server-based configuration engine, initiating, the user device to perform the particular action.

2. The computer-implemented method of claim 1 , wherein the operations further comprise displaying, in the graphical user interface, text corresponding to the candidate term.

3. The computer-implemented method of claim 2 , wherein the operations further comprise, after displaying the text corresponding to the candidate term, receiving a user input indication indicating that the user accepts the text corresponding to the candidate term displayed in the graphical user interface.

4. The computer-implemented method of claim 2 , wherein the operations further comprise displaying a graphical button in the graphical user interface, the graphical button when selected by the user causes the data processing hardware to prompt the user to speak the candidate term again.

5. The computer-implemented method of claim 2 , wherein the operations further comprise, when displaying the text corresponding to the candidate term in the graphical user interface, displaying, in the graphical user interface, a description of the particular action to be performed by the user device.

6. The computer-implemented method of claim 1 , wherein the candidate term comprises two or more words.

7. The computer-implemented method of claim 1 , wherein the server-based configuration engine is configured to obtain the detection model that corresponds to the candidate term by dynamically creating the detection model based on the text-based input corresponding to the candidate term.

8. The computer-implemented method of claim 1 , wherein the server-based configuration engine is configured to obtain the detection model that corresponds to the candidate term by dynamically creating the detection model based on the text-based input.

9. The computer-implemented method of claim 1 , wherein the server-based configuration engine is configured to obtain the detection model that corresponds to the candidate term by retrieving the detection model that corresponds to the candidate term in a vocabulary database, the detection model stored in the vocabulary database prior to the data processing hardware receiving text-based input corresponding to the candidate term.

10. The computer-implemented method of claim 1 , wherein the candidate term comprises a single word.

11. A user device comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed by the data processing hardware cause the data processing hardware to perform operations comprising:

displaying a prompt in a graphical user interface executing on the data processing hardware, the prompt requesting a user to:

provide a candidate term; and

associate the candidate term with a particular action for initiating the user device to perform when the user speaks the candidate term;

receiving, via the graphical user interface, a text-based input corresponding to the candidate term for initiating the user device to perform the particular action;

transmitting the text-based input to a server-based configuration engine, the text-based input when received by the server-based configuration engine causing the server-based configuration engine to create a detection model that corresponds to the candidate term;

receiving audio data corresponding to an utterance of the candidate term spoken by the user; and

when the candidate term is detected in the utterance spoken by the user using the detection model obtained by the server-based configuration engine, initiating, the user device to perform the particular action.

12. The user device of claim 11 , wherein the operations further comprise displaying, in the graphical user interface, text corresponding to the candidate term.

13. The user device of claim 12 , wherein the operations further comprise, after displaying the text corresponding to the candidate term, receiving a user input indication indicating that the user accepts the text corresponding to the candidate term displayed in the graphical user interface.

14. The user device of claim 12 , wherein the operations further comprise displaying a graphical button in the graphical user interface, the graphical button when selected by the user causes the data processing hardware to prompt the user to speak the candidate term again.

15. The user device of claim 12 , wherein the operations further comprise, when displaying the text corresponding to the candidate term in the graphical user interface, displaying, in the graphical user interface, a description of the particular action to be performed by the user device.

16. The user device of claim 11 , wherein the candidate term comprises two or more words.

17. The user device of claim 11 , wherein the server-based configuration engine is configured to obtain the detection model that corresponds to the candidate term by dynamically creating the detection model based on the text-based input corresponding to the candidate term.

18. The user device of claim 11 , wherein the server-based configuration engine is configured to obtain the detection model that corresponds to the candidate term by dynamically creating the detection model based on the text-based input.

19. The user device of claim 11 , wherein the server-based configuration engine is configured to obtain the detection model that corresponds to the candidate term by retrieving the detection model that corresponds to the candidate term in a vocabulary database, the detection model stored in the vocabulary database prior to the data processing hardware receiving text-based input corresponding to the candidate term.

20. The user device of claim 11 , wherein the candidate term comprises a single word.

Assignments (2)
CONVERSION Recorded Aug 7, 2023
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 064509/0754 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2023
From: SHARIFI, MATTHEW
To: GOOGLE INC.
Reel/Frame 063567/0614 →
Continuity (10)
Continuation 16806332 · Mar 2, 2020
Continuation 16669503 · Oct 30, 2019
Continuation 16529300 · Aug 1, 2019
Continuation 16216752 · Dec 11, 2018
Continuation 15875996 · Jan 19, 2018
Continuation 15463786 · Mar 20, 2017
Continuation 15288241 · Oct 7, 2016
Continuation 15001894 · Jan 20, 2016
Continuation 14340833 · Jul 25, 2014
Related Publication 20210312921A1 · Oct 7, 2021