IP Library Granted Patent US 10,909,987
Granted Patent B2
US 10,909,987 · App. 16/553,883 · Granted Feb 2, 2021

Hotword detection on multiple devices

Inventor: Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L15/285G10L15/01G10L15/08G10L15/22G10L15/32G10L17/22G06F3/167G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,909,987
App. No.
16/553,883
Granted
Feb 2, 2021
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for hotword detection on multiple devices are disclosed. In one aspect, a method includes the actions of receiving, by a first computing device, audio data that corresponds to an utterance. The actions further include determining a first value corresponding to a likelihood that the utterance includes a hotword. The actions further include receiving a second value corresponding to a likelihood that the utterance includes the hotword, the second value being determined by a second computing device. The actions further include comparing the first value and the second value. The actions further include based on comparing the first value to the second value, initiating speech recognition processing on the audio data.

Claims (61)

1. A computer-implemented method comprising:

receiving, by a computing device, audio data;

determining, by the computing device, that the audio data likely includes an utterance of a particular, predefined hotword;

in response to determining that the audio data likely includes the utterance of the particular, predefined hotword:

transmitting, by the computing device, data to an additional computing device; and

performing, by the computing device, automated speech recognition processing on the audio data;

while performing the automated speech recognition processing on the audio data, receiving, by the computing device, additional data from the additional computing device, the additional data comprising a hotword score determined by the additional computing device that reflects a likelihood that the audio data includes the utterance of the particular, predefined hotword; and

based on the additional data, determining, by the computing device, whether to perform a command that is included in the utterance after the particular, predefined hotword.

2. The method of claim 1 , further comprising generating, by the computing device, the data based on the audio data.

3. The method of claim 1 , wherein determining that the audio data likely includes the utterance of the particular, predefined hotword comprises determining that the audio data likely includes the utterance of the particular, predefined hotword without performing automated speech recognition on the audio data.

4. The method of claim 1 , further comprising:

determining whether to perform the command that is included in the utterance after the particular, predefined hotword by determining to perform the command that is included in the utterance after the particular, predefined hotword;

performing, by the computing device, automated speech recognition on the audio data;

based on performing automated speech recognition on the audio data, identifying, by the computing device, the command that is included in the utterance; and

performing, by the computing device, the command.

5. The method of claim 1 , further comprising:

receiving the audio data by receiving the audio data while the computing device is in a low power mode;

determining whether to perform the command that is included in the utterance after the particular, predefined hotword by determining to bypass performing the command that is included in the utterance after the particular, predefined hotword; and

based on determining to bypass performing the command that is included in the utterance after the particular, predefined hotword, maintaining the computing device in the low power mode.

6. The method of claim 1 , further comprising generating, by the computing device, the data based on a portion of the audio data that includes the utterance of the particular, predefined hotword.

7. A system comprising:

one or more computers; and

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, by a computing device, audio data;

determining, by the computing device, that the audio data likely includes an utterance of a particular, predefined hotword;

in response to determining that the audio data likely includes the utterance of the particular, predefined hotword:

transmitting, by the computing device, data to an additional computing device; and

performing, by the computing device, automated speech recogntition processing on the audio data;

while performing the automated speech recognition processing on the audio data, receiving, by the computing device, additional data from the additional computing device, the additional data comprising a hotword score determined by the additional computing device that reflects a likelihood that the audio data includes the utterance of the particular, predefined hotword; and

based on the additional data, determining, by the computing device, whether to perform a command that is included in the utterance after the particular, predefined hotword.

8. The system of claim 7 , wherein the operations further comprise generating, by the computing device, the data based on the audio data.

9. The system of claim 7 , wherein determining that the audio data likely includes the utterance of the particular, predefined hotword comprises determining that the audio data likely includes the utterance of the particular, predefined hotword without performing automated speech recognition on the audio data.

10. The system of claim 7 , wherein the operations further comprise:

determining whether to perform the command that is included in the utterance after the particular, predefined hotword by determining to perform the command that is included in the utterance after the particular, predefined hotword;

performing, by the computing device, automated speech recognition on the audio data;

based on performing automated speech recognition on the audio data, identifying, by the computing device, the command that is included in the utterance; and

performing, by the computing device, the command.

11. The system of claim 7 , wherein the operations further comprise:

receiving the audio data by receiving the audio data while the computing device is in a low power mode;

determining whether to perform the command that is included in the utterance after the particular, predefined hotword by determining to bypass performing the command that is included in the utterance after the particular, predefined hotword; and

based on determining to bypass performing the command that is included in the utterance after the particular, predefined hotword, maintaining the computing device in the low power mode.

12. The system of claim 7 , wherein the operations further comprise generating, by the computing device, the data based on a portion of the audio data that includes the utterance of the particular, predefined hotword.

13. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, by a computing device, audio data;

determining, by the computing device, that the audio data likely includes an utterance of a particular, predefined hotword;

in response to determining that the audio data likely includes the utterance of the particular, predefined hotword:

transmitting, by the computing device, data to an additional computing device; and

performing, by the computing device, automated speech recognition processing on the audio data;

while performing the automated speech recognition processing on the audio data, receiving, by the computing device, additional data from the additional computing device, the additional data comprising a hotword score determined by the additional computing device that reflects a likelihood that the audio data includes the utterance of the particular, predefined hotword; and

based on the additional data, determining, by the computing device, whether to perform a command that is included in the utterance after the particular, predefined hotword.

14. The medium of claim 13 , wherein determining that the audio data likely includes the utterance of the particular, predefined hotword comprises determining that the audio data likely includes the utterance of the particular, predefined hotword without performing automated speech recognition on the audio data.

15. The medium of claim 13 , wherein the operations further comprise:

determining whether to perform the command that is included in the utterance after the particular, predefined hotword by determining to perform the command that is included in the utterance after the particular, predefined hotword;

performing, by the computing device, automated speech recognition on the audio data;

based on performing automated speech recognition on the audio data, identifying, by the computing device, the command that is included in the utterance; and

performing, by the computing device, the command.

16. The medium of claim 13 , wherein the operations further comprise:

receiving the audio data by receiving the audio data while the computing device is in a low power mode;

determining whether to perform the command that is included in the utterance after the particular, predefined hotword by determining to bypass performing the command that is included in the utterance after the particular, predefined hotword; and

based on determining to bypass performing the command that is included in the utterance after the particular, predefined hotword, maintaining the computing device in the low power mode.

17. The medium of claim 13 , wherein the operations further comprise generating, by the computing device, the data based on a portion of the audio data that includes the utterance of the particular, predefined hotword.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2019
From: SHARIFI, MATTHEW
To: GOOGLE INC.
Reel/Frame 050204/0595 →
ENTITY CONVERSION Recorded Aug 28, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 050206/0887 →
Continuity (6)
Continuation 16171495 · Oct 26, 2018
Continuation 15346914 · Nov 9, 2016
Continuation 15088477 · Apr 1, 2016
Continuation 14675932 · Apr 1, 2015
Provisional Application 62061830 · Oct 9, 2014
Related Publication 20200058306A1 · Feb 20, 2020