IP Library Granted Patent US 10,134,398
Granted Patent B2
US 10,134,398 · App. 15/346,914 · Granted Nov 20, 2018

Hotword detection on multiple devices

Inventor: Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L15/285G10L15/01G10L15/08G10L15/22G10L15/32G10L17/22G06F3/167G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,134,398
App. No.
15/346,914
Filed
Nov 9, 2016
Granted
Nov 20, 2018
Kind
B2
Art Unit
2658
USPC
704/275
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for hotword detection on multiple devices are disclosed. In one aspect, a method includes the actions of receiving, by a first computing device, audio data that corresponds to an utterance. The actions further include determining a first value corresponding to a likelihood that the utterance includes a hotword. The actions further include receiving a second value corresponding to a likelihood that the utterance includes the hotword, the second value being determined by a second computing device. The actions further include comparing the first value and the second value. The actions further include based on comparing the first value to the second value, initiating speech recognition processing on the audio data.

Claims (70)

1. A computer-implemented method comprising:

receiving, by a computing device that is in a low power mode and that is configured to exit a low power mode upon detecting an utterance of a particular, predefined hotword using an on-device hotword detector, audio data that corresponds to an utterance of the particular, predefined hotword;

while the computing device remains in the low power mode, and in response to receiving the audio data that corresponds to the utterance of the particular, predefined hotword, transmitting, by the computing device and to another computing device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, an output of processing the audio data using the on-device hotword detector;

while the computing device remains in low power mode, receiving, by the computing device and from the other computing device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, an additional output of processing the audio data; and

after transmitting the output of processing the audio data using the on-device hotword detector and after receiving the additional output of processing the audio data from the other using device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, determining, by the computing device, to remain in the low power mode.

2. The method of claim 1 ,

wherein determining to remain in the low power mode is based at least in part on receiving the additional output of processing the audio from the other using device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword.

3. The method of claim 1 , comprising:

determining a hotword confidence score that is based on the audio data that corresponds to the utterance of the particular, predefined hotword and that reflects a likelihood that the audio data includes the particular, predefined hotword,

wherein the output of processing the audio data using the on-device hotword detector includes the hotword confidence score.

4. The method of claim 1 , comprising:

determining a hotword confidence score that is based on the audio data that corresponds to the utterance of the particular, predefined hotword and that reflects a likelihood that the audio data includes the particular, predefined hotword; and

determining that the hotword confidence score satisfies a threshold,

wherein transmitting the output of processing the audio data using the on-device hotword detector is based on determining that the hotword confidence score satisfies the threshold.

5. The method of claim 1 , wherein the computing device transmits the output of processing the audio data using the on-device hotword detector without performing speech recognition on the audio data that corresponds to an utterance of the particular, predefined hotword.

6. The method of claim 1 , wherein:

the computing device transmits the output of processing the audio data using the on-device hotword detector for a particular amount of time, and

the computing device determines to remain in the low power mode after transmitting the output of processing the audio data using the on-device hotword detector for the particular amount of time.

7. The method of claim 1 , comprising:

determining a hotword confidence score that is based on the audio data that corresponds to an utterance of the particular, predefined hotword and that reflects a likelihood that the audio data includes the particular, predefined hotword,

wherein receiving, by the computing device and from the other computing device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, an additional output of processing the audio data comprises receiving, by the computing device and from the other computing device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, a hotword confidence score that is based on the audio data that corresponds to the utterance of the particular, predefined hotword and that reflects an additional likelihood that the audio data includes the particular, predefined hotword,

wherein the method comprises comparing the hotword confidence score with the additional hotword confidence score, and

wherein determining to remain in the low power mode is based on comparing the hotword confidence score with the additional hotword confidence score.

8. The method of claim 1 , wherein:

the other computing device is in a vicinity of the computing device,

the computing device receives the utterance of the particular, predefined hotword through a microphone of the computing device,

the other computing device receives the utterance of the particular, predefined hotword through another microphone of the computing device, and

the additional output of processing the audio data is based on processing, by the other computing device, the utterance of the particular, predefined hotword received by the other computing device.

9. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, by a computing device that is in a low power mode and that is configured to exit a low power mode upon detecting an utterance of a particular, predefined hotword using an on-device hotword detector, audio data that corresponds to an utterance of a particular, predefined hotword;

while the computing device remains in the low power mode, and in response to receiving the audio data that corresponds to the utterance of the particular, predefined hotword, transmitting, by the computing device and to another computing device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, an output of processing the audio data using the on-device hotword detector;

while the computing device remains in low power mode, receiving, by the computing device and from the other computing device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, an additional output of processing the audio data; and

after transmitting the output of processing the audio data using the on-device hotword detector and after receiving the additional output of processing the audio data from the other using device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, determining, by the computing device, to remain in the low power mode.

10. The system of claim 9 ,

wherein determining to remain in the low power mode is based at least in part on receiving the additional output of processing the audio data from the other using device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword.

11. The system of claim 9 , wherein the operations further comprise:

determining a hotword confidence score that is based on the audio data that corresponds to the utterance of the particular, predefined hotword and that reflects a likelihood that the audio data includes the particular, predefined hotword,

wherein the output of processing the audio data using the on-device hotword detector includes the hotword confidence score.

12. The system of claim 9 , wherein the operations further comprise:

determining a hotword confidence score that is based on the audio data that corresponds to the utterance of the particular, predefined hotword and that reflects a likelihood that the audio data includes the particular, predefined hotword; and

determining that the hotword confidence score satisfies a threshold,

wherein transmitting the output of processing the audio data using the on-device hotword detector is based on determining that the hotword confidence score satisfies the threshold.

13. The system of claim 9 , wherein the computing device transmits the output of processing the audio data using the on-device hotword detector without performing speech recognition on the audio data that corresponds to an utterance of the particular, predefined hotword.

14. The system of claim 9 , wherein:

the computing device transmits the output of processing the audio data using the on-device hotword detector for a particular amount of time, and

the computing device determines to remain in the low power mode after transmitting the output of processing the audio data using the on-device hotword detector for the particular amount of time.

15. The system of claim 9 , wherein the operations further comprise:

determining a hotword confidence score that is based on the audio data that corresponds to an utterance of the particular, predefined hotword and that reflects a likelihood that the audio data includes the particular, predefined hotword,

wherein receiving, by the computing device and from the other computing device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, an additional output of processing the audio data comprises receiving, by the computing device and from the other computing device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, a hotword confidence score that is based on the audio data that corresponds to the utterance of the particular, predefined hotword and that reflects an additional likelihood that the audio data includes the particular, predefined hotword,

wherein the operations further comprise comparing the hotword confidence score with the additional hotword confidence score, and

wherein determining to remain in the low power mode is based on comparing the hotword confidence score with the additional hotword confidence score.

16. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, by a computing device that is in a low power mode and that is configured to exit a low power mode upon detecting an utterance of a particular, predefined hotword using an on-device hotword detector, audio data that corresponds to an utterance of a particular, predefined hotword;

while the computing device remains in the low power mode, and in response to receiving the audio data that corresponds to the utterance of the particular, predefined hotword, transmitting, by the computing device and to another computing device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, an output of processing the audio data using the on-device hotword detector;

while the computing device remains in low power mode, receiving, by the computing device and from the other computing device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, an additional output of processing the audio data; and

after transmitting the output of processing the audio data using the on-device hotword detector and after receiving the additional output of processing the audio data from the other using device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword, determining, by the computing device, to remain in the low power mode.

17. The medium of claim 16 ,

wherein determining to remain in the low power mode is based at least in part on receiving the additional output of processing the audio from the other using device that is configured to exit a low power mode upon detecting an utterance of the particular, predefined hotword.

18. The medium of claim 16 , wherein the operations further comprise:

determining a hotword confidence score that is based on the audio data that corresponds to the utterance of the particular, predefined hotword and that reflects a likelihood that the audio data includes the particular, predefined hotword,

wherein the output of processing the audio data using the on-device hotword detector includes the hotword confidence score.

19. The medium of claim 16 , wherein the operations further comprise:

determining a hotword confidence score that is based on the audio data that corresponds to the utterance of the particular, predefined hotword and that reflects a likelihood that the audio data includes the particular, predefined hotword; and

determining that the hotword confidence score satisfies a threshold,

wherein transmitting the output of processing the audio data using the on-device hotword detector is based on determining that the hotword confidence score satisfies the threshold.

20. The medium of claim 16 , wherein the computing device transmits the output of processing the audio data using the on-device hotword detector without performing speech recognition on the audio data that corresponds to an utterance of the particular, predefined hotword.

21. The medium of claim 16 , wherein:

the computing device transmits the output of processing the audio data using the on-device hotword detector for a particular amount of time, and

the computing device determines to remain in the low power mode after transmitting the output of processing the audio data using the on-device hotword detector for the particular amount of time.

Assignments (2)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044129/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2016
From: SHARIFI, MATTHEW
To: GOOGLE INC.
Reel/Frame 040268/0674 →
Continuity (4)
Continuation 15088477 · Apr 1, 2016
Continuation 14675932 · Apr 1, 2015
Provisional Application 62061830 · Oct 9, 2014
Related Publication 20170084277A1 · Mar 23, 2017
Cited By (29)
US 12,192,713 US 12,211,490 US 12,212,945 US 12,217,748 US 12,230,291 US 12,236,932 US 12,277,368 US 12,279,096 US 12,283,269 US 12,288,558 US 12,322,390 US 12,327,549 US 12,327,556 US 12,360,734 US 12,375,052 US 12,387,716 US 12,424,220 US 12,438,977 US 12,505,832 US 12,513,466 US 12,513,479 US 12,518,755 US 12,518,756 US 12,579,978 US 12,626,717 US 12,640,148 US 12,699,543 US 12,711,962 US 12,732,547