IP Library › Granted Patent US 10,262,659
Granted Patent B2
US 10,262,659 · App. 15/909,519 · Granted Apr 16, 2019

Hotword recognition

Inventors: Matthew Sharifi (Kilchberg, CH); Jakob Nicolaus Foerster (Oxford, GB)
Assignee: Google LLC
G10L17/02G06F21/32G07C9/00071G10L15/1815G10L15/22G10L15/285G10L17/22G10L19/018G10L25/51G07C9/00158G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,262,659
App. No.
15/909,519
Filed
Mar 1, 2018
Granted
Apr 16, 2019
Kind
B2
Art Unit
2659
USPC
704/251
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for receiving audio data corresponding to an utterance, determining that the audio data corresponds to a hotword, generating a hotword audio fingerprint of the audio data that is determined to correspond to the hotword, comparing the hotword audio fingerprint to one or more stored audio fingerprints of audio data that was previously determined to correspond to the hotword, detecting whether the hotword audio fingerprint matches a stored audio fingerprint of audio data that was previously determined to correspond to the hotword based on whether the comparison indicates a similarity between the hotword audio fingerprint and one of the one or more stored audio fingerprints that satisfies a predetermined threshold, and in response to detecting that the hotword audio fingerprint matches a stored audio fingerprint, disabling access to a computing device into which the utterance was spoken.

Claims (51)

1. A computer-implemented method comprising:

receiving, by a mobile computing device (i) that includes a replay attack engine, (ii) that is operating in a mode in which access to one or more resources is disabled, and (iii) that is configured to exit the mode in which access to the one or more resources is disabled based on receiving audio input corresponding to a particular utterance, an audio input corresponding to a recording of the particular utterance that was previously input to the same mobile computing device; and

in response to receiving, by the mobile computing device, the audio input corresponding to the recording of the particular utterance that was previously input to the same mobile computing device, preventing, by the replay attack engine of the mobile computing device, the mobile computing device from exiting the mode in which access to the one or more resources is disabled.

2. The method of claim 1 , wherein the particular utterance that was previously input is stored in a database.

3. The method of claim 1 , wherein an audio fingerprint of the audio input corresponding to the recording of the particular utterance that was previously input to the same mobile computing device corresponds to an audio fingerprint of additional audio input that was previously input to the same mobile computing device.

4. The method of claim 1 , comprising:

determining that the audio input corresponds to an utterance that was previously input based on a similarity between the audio input and one or more stored utterances of the hotword.

5. The method of claim 1 , wherein preventing the mobile computing device from exiting the mode in which access to the one or more resources is disabled comprises one or more of:

preventing the mobile computing device from being unlocked,

locking the mobile computing device,

initiating an authentication process, and

preventing the mobile computing device from waking.

6. The method of claim 1 , comprising:

receiving, by the mobile computing device, additional audio data corresponding to a voice command or query; and

determining a type of the voice command or query.

7. The method of claim 1 , comprising:

storing the audio input corresponding to the recording of an utterance that was previously input to the same mobile computing device in a database.

8. A system comprising:

one or more computers; and

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, by a mobile computing device (i) that includes a replay attack engine, (ii) that is operating in a mode in which access to one or more resources is disabled, and (iii) that is configured to exit the mode in which access to the one or more resources is disabled based on receiving audio input corresponding to a particular utterance, an audio input corresponding to a recording of the particular utterance that was previously input to the same mobile computing device; and

in response to receiving, by the mobile computing device, the audio input corresponding to the recording of the particular utterance that was previously input to the same mobile computing device, preventing, by the replay attack engine of the mobile computing device, the mobile computing device from exiting the mode in which access to the one or more resources is disabled.

9. The system of claim 8 , wherein the particular utterance that was previously input is stored in a database.

10. The system of claim 8 , wherein an audio fingerprint of the audio input corresponding to the recording of the particular utterance that was previously input to the same mobile computing device corresponds to an audio fingerprint of additional audio input that was previously input to the same mobile computing device.

11. The system of claim 8 , wherein the operations further comprise:

determining that the audio input corresponds to an utterance that was previously input based on a similarity between the audio input and one or more stored utterances of the hotword.

12. The system of claim 8 , wherein preventing the mobile computing device from exiting the mode in which access to the one or more resources is disabled comprises one or more of:

preventing the mobile computing device from being unlocked,

locking the mobile computing device,

initiating an authentication process, and

preventing the mobile computing device from waking.

13. The system of claim 8 , wherein the operations further comprise:

receiving, by the mobile computing device, additional audio data corresponding to a voice command or query; and

determining a type of the voice command or query.

14. The system of claim 8 , wherein the operations further comprise:

storing the audio input corresponding to the recording of an utterance that was previously input to the same mobile computing device in a database.

15. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, by a mobile computing device (i) that includes a replay attack engine, (ii) that is operating in a mode in which access to one or more resources is disabled, and (iii) that is configured to exit the mode in which access to the one or more resources is disabled based on receiving audio input corresponding to a particular utterance, an audio input corresponding to a recording of the particular utterance that was previously input to the same mobile computing device; and

in response to receiving, by the mobile computing device, the audio input corresponding to the recording of the particular utterance that was previously input to the same mobile computing device, preventing, by the replay attack engine of the mobile computing device, the mobile computing device from exiting the mode in which access to the one or more resources is disabled.

16. The medium of claim 15 , wherein the particular utterance that was previously input is stored in a database.

17. The medium of claim 15 , wherein an audio fingerprint of the audio input corresponding to the recording of the particular utterance that was previously input to the same mobile computing device corresponds to an audio fingerprint of additional audio input that was previously input to the same mobile computing device.

18. The medium of claim 15 , wherein the operations further comprise:

determining that the audio input corresponds to an utterance that was previously input based on a similarity between the audio input and one or more stored utterances of the hotword.

19. The medium of claim 15 , wherein preventing the mobile computing device from exiting the mode in which access to the one or more resources is disabled comprises one or more of:

preventing the mobile computing device from being unlocked,

locking the mobile computing device,

initiating an authentication process, and

preventing the mobile computing device from waking.

20. The medium of claim 15 , wherein the operations further comprise:

receiving, by the mobile computing device, additional audio data corresponding to a voice command or query; and

determining a type of the voice command or query.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2018
From: SHARIFI, MATTHEW; FOERSTER, JAKOB NICOLAUS
To: GOOGLE INC.
Reel/Frame 045128/0320 →
CHANGE OF NAME Recorded Mar 7, 2018
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 045513/0948 →
Continuity (5)
Continuation 15176830 · Jun 8, 2016
Continuation 15176482 · Jun 8, 2016
Continuation 14943287 · Nov 17, 2015
Provisional Application 62242650 · Oct 16, 2015
Related Publication 20180254045A1 · Sep 6, 2018