Providing pre-computed hotword models
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for obtaining, for each of multiple words or sub-words, audio data corresponding to multiple users speaking the word or sub-word; training, for each of the multiple words or sub-words, a pre-computed hotword model for the word or sub-word based on the audio data for the word or sub-word; receiving a candidate hotword from a computing device; identifying one or more pre-computed hotword models that correspond to the candidate hotword; and providing the identified, pre-computed hotword models to the computing device.
1. A computer-implemented method when executed on data processing hardware of a user device causes the data processing hardware to perform operations comprising:
displaying a prompt in a graphical user interface executing on the data processing hardware, the prompt requesting a user to:
provide a candidate term; and
associate the candidate term with a particular action for initiating the user device to perform when the user speaks the candidate term;
receiving, via the graphical user interface, a text-based input corresponding to the candidate term for initiating the user device to perform the particular action;
transmitting the text-based input to a server-based configuration engine, the text-based input when received by the server-based configuration engine causing the server-based configuration engine to create a detection model that corresponds to the candidate term;
receiving audio data corresponding to an utterance of the candidate term spoken by the user; and
when the candidate term is detected in the utterance spoken by the user using the detection model obtained by the server-based configuration engine, initiating, the user device to perform the particular action.
2. The computer-implemented method of claim 1 , wherein the operations further comprise displaying, in the graphical user interface, text corresponding to the candidate term.
3. The computer-implemented method of claim 2 , wherein the operations further comprise, after displaying the text corresponding to the candidate term, receiving a user input indication indicating that the user accepts the text corresponding to the candidate term displayed in the graphical user interface.
4. The computer-implemented method of claim 2 , wherein the operations further comprise displaying a graphical button in the graphical user interface, the graphical button when selected by the user causes the data processing hardware to prompt the user to speak the candidate term again.
5. The computer-implemented method of claim 2 , wherein the operations further comprise, when displaying the text corresponding to the candidate term in the graphical user interface, displaying, in the graphical user interface, a description of the particular action to be performed by the user device.
6. The computer-implemented method of claim 1 , wherein the candidate term comprises two or more words.
7. The computer-implemented method of claim 1 , wherein the server-based configuration engine is configured to obtain the detection model that corresponds to the candidate term by dynamically creating the detection model based on the text-based input corresponding to the candidate term.
8. The computer-implemented method of claim 1 , wherein the server-based configuration engine is configured to obtain the detection model that corresponds to the candidate term by dynamically creating the detection model based on the text-based input.
9. The computer-implemented method of claim 1 , wherein the server-based configuration engine is configured to obtain the detection model that corresponds to the candidate term by retrieving the detection model that corresponds to the candidate term in a vocabulary database, the detection model stored in the vocabulary database prior to the data processing hardware receiving text-based input corresponding to the candidate term.
10. The computer-implemented method of claim 1 , wherein the candidate term comprises a single word.
11. A user device comprising:
data processing hardware; and
memory hardware in communication with the data processing hardware and storing instructions that when executed by the data processing hardware cause the data processing hardware to perform operations comprising:
displaying a prompt in a graphical user interface executing on the data processing hardware, the prompt requesting a user to:
provide a candidate term; and
associate the candidate term with a particular action for initiating the user device to perform when the user speaks the candidate term;
receiving, via the graphical user interface, a text-based input corresponding to the candidate term for initiating the user device to perform the particular action;
transmitting the text-based input to a server-based configuration engine, the text-based input when received by the server-based configuration engine causing the server-based configuration engine to create a detection model that corresponds to the candidate term;
receiving audio data corresponding to an utterance of the candidate term spoken by the user; and
when the candidate term is detected in the utterance spoken by the user using the detection model obtained by the server-based configuration engine, initiating, the user device to perform the particular action.
12. The user device of claim 11 , wherein the operations further comprise displaying, in the graphical user interface, text corresponding to the candidate term.
13. The user device of claim 12 , wherein the operations further comprise, after displaying the text corresponding to the candidate term, receiving a user input indication indicating that the user accepts the text corresponding to the candidate term displayed in the graphical user interface.
14. The user device of claim 12 , wherein the operations further comprise displaying a graphical button in the graphical user interface, the graphical button when selected by the user causes the data processing hardware to prompt the user to speak the candidate term again.
15. The user device of claim 12 , wherein the operations further comprise, when displaying the text corresponding to the candidate term in the graphical user interface, displaying, in the graphical user interface, a description of the particular action to be performed by the user device.
16. The user device of claim 11 , wherein the candidate term comprises two or more words.
17. The user device of claim 11 , wherein the server-based configuration engine is configured to obtain the detection model that corresponds to the candidate term by dynamically creating the detection model based on the text-based input corresponding to the candidate term.
18. The user device of claim 11 , wherein the server-based configuration engine is configured to obtain the detection model that corresponds to the candidate term by dynamically creating the detection model based on the text-based input.
19. The user device of claim 11 , wherein the server-based configuration engine is configured to obtain the detection model that corresponds to the candidate term by retrieving the detection model that corresponds to the candidate term in a vocabulary database, the detection model stored in the vocabulary database prior to the data processing hardware receiving text-based input corresponding to the candidate term.
20. The user device of claim 11 , wherein the candidate term comprises a single word.