IP Library Granted Patent US 10,650,828
Granted Patent B2
US 10,650,828 · App. 16/362,990 · Granted May 12, 2020

Hotword recognition

Inventors: Matthew Sharifi (Kilchberg, CH); Jakob Nicolaus Foerster (San Francisco, CA)
Assignee: Google LLC
G10L17/02G06F3/167G06F21/32G07C9/25G10L15/1815G10L15/22G10L15/285G10L17/22G10L19/018G10L25/51G07C9/37G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,650,828
App. No.
16/362,990
Granted
May 12, 2020
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for receiving audio data corresponding to an utterance, determining that the audio data corresponds to a hotword, generating a hotword audio fingerprint of the audio data that is determined to correspond to the hotword, comparing the hotword audio fingerprint to one or more stored audio fingerprints of audio data that was previously determined to correspond to the hotword, detecting whether the hotword audio fingerprint matches a stored audio fingerprint of audio data that was previously determined to correspond to the hotword based on whether the comparison indicates a similarity between the hotword audio fingerprint and one of the one or more stored audio fingerprints that satisfies a predetermined threshold, and in response to detecting that the hotword audio fingerprint matches a stored audio fingerprint, disabling access to a computing device into which the utterance was spoken.

Claims (64)

1. A computer-implemented method comprising:

receiving, by a first computing device, first audio data of a first utterance;

determining, by the first computing device, that the first utterance includes a particular, predefined hotword;

based on determining that the first utterance includes the particular, predefined hotword, determining, by the first computing device, whether the first audio data includes an audio signal previously outputted by a second computing device while the second computing device received second audio data of a second utterance that included the particular, predefined hotword; and

based on determining whether the first audio data includes the audio signal, determining, by the first computing device, whether to perform a command that is included in the first utterance and that follows the particular, predefined hotword.

2. The method of claim 1 , wherein:

determining whether the first audio data includes the audio signal comprises:

determining that the first audio data includes the audio signal, and

determining whether to perform the command comprises:

determining to suppress performance of the command based on determining that the first audio data includes the audio signal.

3. The method of claim 2 , wherein:

the first computing device receives the first audio data while in a mode in which access to one or more resources is disabled,

the command is a command to exit the mode in which access to the one or more resources is disabled, and

determining to suppress performance of the command comprises maintaining the first computing device in the mode in which access to the one or more resources is disabled.

4. The method of claim 1 , wherein:

determining whether the first audio data includes the audio signal comprises:

determining that the first audio data does not include the audio signal, and

determining whether to perform the command comprises:

determining to perform the command based on determining that the first audio data does not include the audio signal.

5. The method of claim 1 , wherein the audio signal is an ultrasonic audio signal.

6. The method of claim 1 , wherein the audio signal is included in a portion of the audio data that includes the particular, predefined hotword.

7. The method of claim 1 , wherein the audio signal is included in a portion of the audio data that includes the command.

8. The method of claim 1 , comprising:

based on determining that the first utterance includes the particular, predefined hotword providing, for output, an additional audio signal.

9. The method of claim 1 , comprising:

based on determining whether the first audio data includes the audio signal, selecting performing, by the first computing device, a command that is included in the first utterance and that follows the particular, predefined hotword.

10. A system comprising:

one or more computers; and

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, by a first computing device, first audio data of a first utterance;

determining, by the first computing device, that the first utterance includes a particular, predefined hotword;

based on determining that the first utterance includes the particular, predefined hotword, determining, by the first computing device, whether the first audio data includes an audio signal previously outputted by a second computing device while the second computing device received second audio data of a second utterance that included the particular, predefined hotword; and

based on determining whether the first audio data includes the audio signal, determining, by the first computing device, whether to perform a command that is included in the first utterance and that follows the particular, predefined hotword.

11. The system of claim 10 , wherein:

determining whether the first audio data includes the audio signal comprises:

determining that the first audio data includes the audio signal, and

determining whether to perform the command comprises:

determining to suppress performance of the command based on determining that the first audio data includes the audio signal.

12. The system of claim 11 , wherein:

the first computing device receives the first audio data while in a mode in which access to one or more resources is disabled,

the command is a command to exit the mode in which access to the one or more resources is disabled, and

determining to suppress performance of the command comprises maintaining the first computing device in the mode in which access to the one or more resources is disabled.

13. The system of claim 10 , wherein:

determining whether the first audio data includes the audio signal comprises:

determining that the first audio data does not include the audio signal, and

determining whether to perform the command comprises:

determining to perform the command based on determining that the first audio data does not include the audio signal.

14. The system of claim 10 , wherein the audio signal is an ultrasonic audio signal.

15. The system of claim 10 , wherein the audio signal is included in a portion of the audio data that includes the particular, predefined hotword.

16. The system of claim 10 , wherein the audio signal is included in a portion of the audio data that includes the command.

17. The system of claim 10 , wherein the operations comprise:

based on determining that the first utterance includes the particular, predefined hotword providing, for output, an additional audio signal.

18. The system of claim 10 , wherein the operations comprise:

based on determining whether the first audio data includes the audio signal, selecting performing, by the first computing device, a command that is included in the first utterance and that follows the particular, predefined hotword.

19. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, by a first computing device, first audio data of a first utterance;

determining, by the first computing device, that the first utterance includes a particular, predefined hotword;

based on determining that the first utterance includes the particular, predefined hotword, determining, by the first computing device, whether the first audio data includes an audio signal previously outputted by a second computing device while the second computing device received second audio data of a second utterance that included the particular, predefined hotword; and

based on determining whether the first audio data includes the audio signal, determining, by the first computing device, whether to perform a command that is included in the first utterance and that follows the particular, predefined hotword.

20. The system of claim 19 , wherein:

determining whether the first audio data includes the audio signal comprises:

determining that the first audio data includes the audio signal, and

determining whether to perform the command comprises:

determining to suppress performance of the command based on determining that the first audio data includes the audio signal.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2019
From: SHARIFI, MATTHEW; FOERSTER, JAKOB NICOLAUS
To: GOOGLE INC.
Reel/Frame 048688/0816 →
ENTITY CONVERSION Recorded Mar 25, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 048690/0767 →
Continuity (6)
Continuation 15909519 · Mar 1, 2018
Continuation 15176830 · Jun 8, 2016
Continuation 15176482 · Jun 8, 2016
Continuation 14943287 · Nov 17, 2015
Provisional Application 62242650 · Oct 16, 2015
Related Publication 20190287536A1 · Sep 19, 2019