Promoting voice actions to hotwords
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for designating certain voice commands as hotwords. The methods, systems, and apparatus include actions of receiving a hotword followed by a voice command. Additional actions include determining that the voice command satisfies one or more predetermined criteria associated with designating the voice command as a hotword, where a voice command that is designated as a hotword is treated as a voice input regardless of whether the voice command is preceded by another hotword. Further actions include, in response to determining that the voice command satisfies one or more predetermined criteria associated with designating the voice command as a hotword, designating the voice command as a hotword.
1. A computer-implemented method comprising:
receiving, by a computing device, a first speech utterance beginning with a hotword followed by a particular phrase, the particular phrase not currently designated by the computing device as a hotword;
in response to receiving the first speech utterance beginning with the hotword, triggering, by the computing device, semantic interpretation on the particular phrase following the hotword;
designating, by the computing device, the particular phrase as a new hotword based on the semantic interpretation determining that the particular phrase satisfies one or more predetermined criteria associated with designating voice commands as hotwords;
training, by the computing device, a hotword model to recognize the particular phrase now designated as the new hotword for subsequently received utterances; and
after training the hotword model to recognize the particular phrase and while the computing device is in a sleep state, receiving, by the computing device, a second speech utterance that begins with the particular phrase, the particular phrase when recognized in the second speech utterance by the hotword model causing the computing device to transition from the sleep state and process the second speech utterance as a voice command.
2. The method of claim 1 , further comprising, prior to designating the particular phrase as the new hotword, generating, by the computing device, a user interface prompt that prompts a user to designate the particular phrase as the new hotword.
3. The method of claim 1 , wherein a hotword comprises a sequence of one or more words that causes the computing device to exit the sleep state and process an utterance beginning with the hotword as a voice command.
4. The method of claim 1 , wherein a hotword manager of the computing device designates the particular phrase as the new hotword without requiring a user to explicitly designate the particular phrase as the new hotword.
5. The method of claim 1 , wherein determining that the particular phrase satisfies one or more predetermined criteria associated with designating voice commands as hotwords comprises determining that the particular phrase is not currently designated by the computing device as a hotword.
6. The method of claim 1 , further comprising, after receiving the first speech utterance beginning with the hotword followed by the particular phrase, generating, by an automated speech recognizer of the computing device, a transcription of the particular phrase.
7. The method of claim 1 , wherein training the hotword model to recognize the particular phrase comprises training the hotword model with acoustic features characteristic of the particular phrase.
8. The method of claim 1 wherein the one or more predetermined criteria includes at least one of a quantity of issuances of the voice command, a quality of acoustic features associated with the voice command, a score reflecting phonetic suitability of the voice command as a hotword, or an amount of acoustic features stored for the voice command.
9. A system comprising:
one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
receiving a first speech utterance beginning with a hotword followed by a particular phrase, the particular phrase not currently designated as a hotword;
in response to receiving the first speech utterance beginning with the hotword, triggering semantic interpretation on the particular phrase following the hotword;
designating the particular phrase as a new hotword based on the semantic interpretation determining that the particular phrase satisfies one or more predetermined criteria associated with designating voice commands as hotwords;
training a hotword model to recognize the particular phrase now designated as the new hotword for subsequently received utterances; and
after training the hotword model to recognize the particular phrase and while the one or more computers are in a sleep state, receiving a second speech utterance that begins with the particular phrase, the particular phrase when recognized in the second speech utterance by the hotword model causing the one or more computers to transition from the sleep state and process the second speech utterance as a voice command.
10. The system of claim 9 , wherein the operations further comprise, prior to designating the particular phrase as the new hotword, generating a user interface prompt that prompts a user to designate the particular phrase as the new hotword.
11. The system of claim 9 , wherein a hotword comprises a sequence of one or more words that causes the one or more computers to exit the sleep state and process an utterance beginning with the hotword as a voice command.
12. The system of claim 9 , wherein a hotword manager of the one or more computers designates the particular phrase as the new hotword without requiring a user to explicitly designate the particular phrase as the new hotword.
13. The system of claim 9 , wherein determining that the particular phrase satisfies one or more predetermined criteria associated with designating voice commands as hotwords comprises determining that the particular phrase is not currently designated by the computing device as a hotword.
14. The system of claim 9 , wherein the operations further comprise, after receiving the first speech utterance beginning with the hotword followed by the particular phrase, generating, by an automated speech recognizer of the one or more computers, a transcription of the particular phrase.
15. The system of claim 9 , wherein training the hotword model to recognize the particular phrase comprises training the hotword model with acoustic features characteristic of the particular phrase.
16. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:
receiving first speech utterance beginning with a hotword followed by a particular phrase, the particular phrase not currently designated as a hotword;
in response to receiving the first speech utterance beginning with the hotword, triggering semantic interpretation on the particular phrase following the hotword;
designating the particular phrase as a new hotword based on the semantic interpretation determining that the particular phrase satisfies one or more predetermined criteria associated with designating voice commands as hotwords;
training a hotword model to recognize the particular phrase now designated as the new hotword for subsequently received utterances; and
after training the hotword model to recognize the particular phrase and while the one or more computers are in a sleep state, receiving, a second speech utterance that begins with the particular phrase, the particular phrase when recognized in the second speech utterance by the hotword model causing the one or more computers to transition from the sleep state and process the second speech utterance as a voice command.
17. The non-transitory computer-readable medium of claim 16 , wherein the operations further comprise, prior to designating the particular phrase as the new hotword, generating a user interface prompt that prompts a user to designate the particular phrase as the new hotword.
18. The non-transitory computer-readable medium of claim 16 , wherein a hotword comprises a sequence of one or more words that causes the one or more computers to exit the sleep state and process an utterance beginning with the hotword as a voice command.
19. The non-transitory computer-readable medium of claim 16 , wherein a hotword manager of the one or more computers designates the particular phrase as the new hotword without requiring a user to explicitly designate the particular phrase as the new hotword.
20. The non-transitory computer-readable medium of claim 16 , wherein determining that the particular phrase satisfies one or more predetermined criteria associated with designating voice commands as hotwords comprises determining that the particular phrase is not currently designated by the computing device as a hotword.
21. The non-transitory computer-readable medium of claim 16 , wherein the operations further comprise, after receiving the first speech utterance beginning with the hotword followed by the particular phrase, generating, by an automated speech recognizer of the one or more computers, a transcription of the particular phrase.