IP Library › Granted Patent US 10,276,161
Granted Patent B2
US 10,276,161 · App. 15/391,358 · Granted Apr 30, 2019

Contextual hotwords

Inventors: Christopher Thaddeus Hughes (Redwood City, CA); Ignacio Lopez Moreno (New York, NY); Aleksandar Kracun (New York, NY)
Assignee: Google LLC
G10L15/22G10L15/02G10L15/08G10L15/20G10L2015/088G10L2015/223G10L2015/226G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,276,161
App. No.
15/391,358
Granted
Apr 30, 2019
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for contextual hotwords are disclosed. In one aspect, a method, during a boot process of a computing device, includes the actions of determining, by a computing device, a context associated with the computing device. The actions further include, based on the context associated with the computing device, determining a hotword. The actions further include, after determining the hotword, receiving audio data that corresponds to an utterance. The actions further include determining that the audio data includes the hotword. The actions further include, in response to determining that the audio data includes the hotword, performing an operation associated with the hotword.

Claims (78)

1. A computer-implemented method comprising:

determining, by a computing device, a context associated with the computing device;

based on the context associated with the computing device, identifying, by the computing device, a hotword that, upon receipt by the computing device, initiates execution of an operation that is available for execution by a user;

after identifying the hotword, receiving, by a microphone of the computing device, audio data that corresponds to an utterance;

determining, by the computing device, that the audio data does not include the hotword by providing the audio data as an input to a hotword model;

receiving, by the computing device and from a user, a non-audible request to perform the operation associated with the hotword;

in response to receiving the non-audible request to perform the operation associated with the hotword, executing, by the computing device, the operation associated with the hotword and updating the hotword model;

after updating the hotword model, receiving, by the microphone of the computing device, additional audio data that corresponds to an additional utterance;

determining, by the computing device, that the additional audio data includes the hotword by providing the additional audio data as an input to the updated hotword model; and

in response to determining that the additional audio data includes the hotword, executing, by the computing device, the operation associated with the hotword.

2. The method of claim 1 , wherein determining that the audio data does not include the hotword comprises determining that the audio data does not include the hotword without preforming speech recognition on the audio data.

3. The method of claim 1 , wherein determining that the audio data does not include the hotword comprises:

extracting audio features of the audio data that corresponds to the utterance;

generating a hotword confidence score by processing the audio features;

determining that the hotword confidence score does not satisfy a hotword confidence threshold; and

based on determining that the hotword confidence score does not satisfy a hotword confidence threshold, determining that the audio data that corresponds to the utterance does not include the hotword.

4. The method of claim 1 , comprising:

after identifying the hotword, receiving the hotword model that corresponds to the hotword.

5. The method of claim 1 , comprising:

identifying, by the computing device, an application that is running on the computing device,

wherein the context is based on the application that is running on the computing device, and

wherein the operation that is available for execution by the user is an operation of the application that is running on the computing device.

6. The method of claim 1 , comprising:

determining, by the computing device, that the context is no longer associated with the computing device; and

determining that subsequently received audio data that includes the hotword is not to trigger execution of the operation associated with the hotword.

7. The method of claim 1 , comprising:

identifying, by the computing device, movement of the computing device,

wherein the context is based on the movement of the computing device.

8. The method of claim 1 , comprising:

identifying, by the computing device, a location of the computing device,

wherein the context is based on the location of the computing device.

9. The method of claim 1 , wherein executing the operation associated with the hotword comprises:

performing speech recognition on a portion of the audio data that does not include the hotword,

wherein the operation is based on a transcription of the portion of the audio data that does not include the hotword.

10. The method of claim 1 , wherein the audio data only includes the hotword.

11. The method of claim 1 , wherein an initial portion of the audio data includes the hotword.

12. The method of claim 1 , comprising:

transmitting, by the computing device, a request for the hotword model that corresponds to the hotword and data identifying the context associated with the computing device; and

receiving, by the computing device, the hotword model that corresponds to the hotword and that was trained using speech samples that include a level of background noise associated with the context.

13. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

determining, by a computing device, a context associated with the computing device;

based on the context associated with the computing device, identifying, by the computing device, a hotword that, upon receipt by the computing device, initiates execution of an operation that is available for execution by a user;

after identifying the hotword, receiving, by a microphone of the computing device, audio data that corresponds to an utterance;

determining, by the computing device, that the audio data does not include the hotword by providing the audio data as an input to a hotword model;

receiving, by the computing device and from a user, a non-audible request to perform the operation associated with the hotword;

in response to receiving the non-audible request to perform the operation associated with the hotword, executing, by the computing device, the operation associated with the hotword and updating the hotword model;

after updating the hotword model, receiving, by the microphone of the computing device, additional audio data that corresponds to an additional utterance;

determining, by the computing device, that the additional audio data includes the hotword by providing the additional audio data as an input to the updated hotword model; and

in response to determining that the additional audio data includes the hotword, executing, by the computing device, the operation associated with the hotword.

14. The system of claim 13 , wherein determining that the audio data does not include the hotword comprises determining that the audio data does not include the hotword without preforming speech recognition on the audio data.

15. The system of claim 13 , wherein determining that the audio data does not include the hotword comprises:

extracting audio features of the audio data that corresponds to the utterance;

generating a hotword confidence score by processing the audio features;

determining that the hotword confidence score does not satisfy a hotword confidence threshold; and

based on determining that the hotword confidence score does not satisfy a hotword confidence threshold, determining that the audio data that corresponds to the utterance does not include the hotword.

16. The system of claim 13 , wherein the operations further comprise:

after identifying the hotword, receiving the hotword model that corresponds to the hotword.

17. The system of claim 13 , wherein the operations further comprise:

identifying, by the computing device, an application that is running on the computing device,

wherein the context is based on the application that is running on the computing device, and

wherein the operation that is available for execution by the user is an operation of the application that is running on the computing device.

18. The system of claim 13 , wherein the operations further comprise:

determining, by the computing device, that the context is no longer associated with the computing device; and

determining that subsequently received audio data that includes the hotword is not to trigger execution of the operation associated with the hotword.

19. The system of claim 13 , wherein executing the operation associated with the hotword comprises:

performing speech recognition on a portion of the audio data that does not include the hotword,

wherein the operation is based on a transcription of the portion of the audio data that does not include the hotword.

20. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

determining, by a computing device, a context associated with the computing device;

based on the context associated with the computing device, identifying, by the computing device, a hotword that, upon receipt by the computing device, initiates execution of an operation that is available for execution by a user;

after identifying the hotword, receiving, by a microphone of the computing device, audio data that corresponds to an utterance;

determining, by the computing device, that the audio data does not include the hotword by providing the audio data as an input to a hotword model;

receiving, by the computing device and from a user, a non-audible request to perform the operation associated with the hotword;

in response to receiving the non-audible request to perform the operation associated with the hotword, executing, by the computing device, the operation associated with the hotword and updating the hotword model;

after updating the hotword model, receiving, by the microphone of the computing device, additional audio data that corresponds to an additional utterance;

determining, by the computing device, that the additional audio data includes the hotword by providing the additional audio data as an input to the updated hotword model; and

in response to determining that the additional audio data includes the hotword, executing, by the computing device, the operation associated with the hotword.

Assignments (2)
CHANGE OF NAME Recorded Oct 20, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044567/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2016
From: HUGHES, CHRISTOPHER THADDEUS; MORENO, IGNACIO LOPEZ; KRACUN, ALEKSANDAR
To: GOOGLE INC.
Reel/Frame 040784/0658 →
Continuity (1)
Related Publication 20180182390A1 · Jun 28, 2018
Cited By (16)
US 12,211,490 US 12,217,748 US 12,230,291 US 12,236,932 US 12,283,269 US 12,327,549 US 12,327,556 US 12,360,734 US 12,387,716 US 12,424,220 US 12,505,832 US 12,513,479 US 12,518,756 US 12,699,543 US 12,711,962 US 12,732,547