IP Library Granted Patent US 9,747,926
Granted Patent B2
US 9,747,926 · App. 14/943,287 · Granted Aug 29, 2017

Hotword recognition

Inventors: Matthew Sharifi (Kilchberg, CH); Jakob Nicolaus Foerster (Oxford, GB)
Assignee: Google Inc.
G10L25/51G10L15/02G10L15/1815G10L17/08G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,747,926
App. No.
14/943,287
Granted
Aug 29, 2017
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for receiving audio data corresponding to an utterance, determining that the audio data corresponds to a hotword, generating a hotword audio fingerprint of the audio data that is determined to correspond to the hotword, comparing the hotword audio fingerprint to one or more stored audio fingerprints of audio data that was previously determined to correspond to the hotword, detecting whether the hotword audio fingerprint matches a stored audio fingerprint of audio data that was previously determined to correspond to the hotword based on whether the comparison indicates a similarity between the hotword audio fingerprint and one of the one or more stored audio fingerprints that satisfies a predetermined threshold, and in response to detecting that the hotword audio fingerprint matches a stored audio fingerprint, disabling access to a computing device into which the utterance was spoken.

Claims (83)

1. A computer-implemented method comprising:

receiving audio data corresponding to an utterance that is received while a computing device is operating in a lock mode, the computing device being configured to exit the lock mode based on determining that the audio data corresponds to a hotword;

determining that the audio data corresponds to the hotword;

generating a hotword audio fingerprint of the audio data that is determined to correspond to the hotword;

determining a similarity between the hotword audio fingerprint and one or more stored audio fingerprints of audio data that was previously determined to correspond to the hotword;

detecting whether the hotword audio fingerprint matches a stored audio fingerprint of audio data that was previously determined to correspond to the hotword based on whether the similarity between the hotword audio fingerprint and one of the one or more stored audio fingerprints satisfies a predetermined threshold; and

in response to detecting that the hotword audio fingerprint matches a stored audio fingerprint, preventing the computing device into which the utterance was spoken from exiting the lock mode despite determining that the audio data corresponds to the hotword.

2. The computer-implemented method of claim 1 , wherein determining that the audio data corresponds to a hotword comprises:

identifying one or more acoustic features of the audio data;

comparing the one or more acoustic features of the audio data to one or more acoustic features associated with one or more hotwords stored in a database; and

determining that the audio data corresponds to one of the one or more hotwords stored in the database based on the comparison of the one or more acoustic features of the audio data to the one or more acoustic features associated with one or more hotwords stored in the database.

3. The computer-implemented method of claim 1 , comprising:

receiving additional audio data corresponding to an additional utterance;

identifying speaker-identification d-vectors using the additional audio data;

determining a similarity between the speaker-identification d-vectors from the additional audio data and hotword d-vectors from the audio data corresponding to the utterance;

detecting whether the audio data corresponding to the hotword matches the additional audio data based on whether the similarity between the speaker-identification d-vectors from the additional audio data and the hotword d-vectors from the audio data corresponding to the utterance satisfies a particular threshold; and

in response to detecting that the audio data corresponding to the hotword does not match the additional audio data, disabling access to the computing device.

4. The computer-implemented method of claim 1 , wherein the hotword is a particular term that triggers semantic interpretation of an additional term of one or more terms that follow the particular term.

5. The computer-implemented method of claim 1 , comprising:

receiving additional audio data corresponding to a voice command or query; and

determining a type of the voice command or query,

wherein the predetermined threshold is adjusted based on the type of the voice command or query.

6. The computer-implemented method of claim 1 , wherein determining that the audio data corresponds to a hotword comprises:

determining that an initial portion of the audio data corresponds to an initial portion of the hotword; and

in response to determining that the initial portion of the audio data corresponds to the initial portion of the hotword, causing one of a plurality of unique ultrasonic audio samples to be outputted after the initial portion of the audio data is received.

7. The computer-implemented method of claim 6 , comprising:

determining that the received audio data comprises audio data corresponding to one of the plurality of unique ultrasonic audio samples; and

in response to determining that the received audio data comprises audio data corresponding to one of the plurality of unique ultrasonic audio samples, disabling access to the computing device.

8. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving audio data corresponding to an utterance that is received while a computing device is operating in a lock mode, the computing device being configured to exit the lock mode based on determining that the audio data corresponds to a hotword;

determining that the audio data corresponds to the hotword;

generating a hotword audio fingerprint of the audio data that is determined to correspond to the hotword;

determining a similarity between the hotword audio fingerprint and one or more stored audio fingerprints of audio data that was previously determined to correspond to the hotword;

detecting whether the hotword audio fingerprint matches a stored audio fingerprint of audio data that was previously determined to correspond to the hotword based on whether the similarity between the hotword audio fingerprint and one of the one or more stored audio fingerprints satisfies a predetermined threshold; and

in response to detecting that the hotword audio fingerprint matches a stored audio fingerprint, preventing the computing device into which the utterance was spoken from exiting the lock mode despite determining that the audio data corresponds to the hotword.

9. The system of claim 8 , wherein determining that the audio data corresponds to a hotword comprises:

identifying one or more acoustic features of the audio data;

comparing the one or more acoustic features of the audio data to one or more acoustic features associated with one or more hotwords stored in a database; and

determining that the audio data corresponds to one of the one or more hotwords stored in the database based on the comparison of the one or more acoustic features of the audio data to the one or more acoustic features associated with one or more hotwords stored in the database.

10. The system of claim 8 , wherein the operations comprise:

receiving additional audio data corresponding to an additional utterance;

identifying speaker-identification d-vectors using the additional audio data;

determining a similarity between the speaker-identification d-vectors from the additional audio data and hotword d-vectors from the audio data corresponding to the utterance;

detecting whether the audio data corresponding to the hotword matches the additional audio data based on whether the similarity between the speaker-identification d-vectors from the additional audio data and the hotword d-vectors from the audio data corresponding to the utterance satisfies a particular threshold; and

in response to detecting that the audio data corresponding to the hotword does not match the additional audio data, disabling access to the computing device.

11. The system of claim 8 , wherein the hotword is a particular term that triggers semantic interpretation of an additional term of one or more terms that follow the particular term.

12. The system of claim 8 , wherein the operations comprise:

receiving additional audio data corresponding to a voice command or query; and

determining a type of the voice command or query,

wherein the predetermined threshold is weighted based on the type of the voice command or query.

13. The system of claim 8 , wherein determining that the audio data corresponds to a hotword comprises:

determining that an initial portion of the audio data corresponds to an initial portion of the hotword; and

in response to determining that the initial portion of the audio data corresponds to the initial portion of the hotword, causing one of a plurality of unique ultrasonic audio samples to be outputted after the initial portion of the audio data is received.

14. The system of claim 13 , wherein the operations comprise:

determining that the received audio data comprises audio data corresponding to one of the plurality of unique ultrasonic audio samples; and

in response to determining that the received audio data comprises audio data corresponding to one of the plurality of unique ultrasonic audio samples, disabling access to the computing device.

15. A computer-readable storage device storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving audio data corresponding to an utterance that is received while a computing device is operating in a lock mode, the computing device being configured to exit the lock mode based on determining that the audio data corresponds to a hotword;

determining that the audio data corresponds to the hotword;

generating a hotword audio fingerprint of the audio data that is determined to correspond to the hotword;

determining a similarity between the hotword audio fingerprint and one or more stored audio fingerprints of audio data that was previously determined to correspond to the hotword;

detecting whether the hotword audio fingerprint matches a stored audio fingerprint of audio data that was previously determined to correspond to the hotword based on whether the similarity between the hotword audio fingerprint and one of the one or more stored audio fingerprints satisfies a predetermined threshold; and

in response to detecting that the hotword audio fingerprint matches a stored audio fingerprint, preventing the computing device into which the utterance was spoken from exiting the lock mode despite determining that the audio data corresponds to the hotword.

16. The computer-readable storage device of claim 15 , wherein determining that the audio data corresponds to a hotword comprises:

identifying one or more acoustic features of the audio data;

comparing the one or more acoustic features of the audio data to one or more acoustic features associated with one or more hotwords stored in a database; and

determining that the audio data corresponds to one of the one or more hotwords stored in the database based on the comparison of the one or more acoustic features of the audio data to the one or more acoustic features associated with one or more hotwords stored in the database.

17. The computer-readable storage device of claim 15 , wherein the operations comprise:

receiving additional audio data corresponding to an additional utterance;

identifying speaker-identification d-vectors using the additional audio data;

determining a similarity between the speaker-identification d-vectors from the additional audio data and hotword d-vectors from the audio data corresponding to the utterance;

detecting whether the audio data corresponding to the hotword matches the additional audio data based on whether the similarity between the speaker-identification d-vectors from the additional audio data and the hotword d-vectors from the audio data corresponding to the utterance satisfies a particular threshold; and

in response to detecting that the audio data corresponding to the hotword does not match the additional audio data, disabling access to the computing device.

18. The computer-readable storage device of claim 15 , wherein the hotword is a particular term that triggers semantic interpretation of an additional term of one or more terms that follow the particular term.

19. The computer-readable storage device of claim 15 , wherein the operations comprise:

receiving additional audio data corresponding to a voice command or query;

determining a type of the voice command or query,

wherein the predetermined threshold is weighted based on the type of the voice command or query.

20. The computer-readable storage device of claim 15 , wherein determining that the audio data corresponds to a hotword comprises:

determining that an initial portion of the audio data corresponds to an initial portion of the hotword;

in response to determining that the initial portion of the audio data corresponds to the initial portion of the hotword, causing one of a plurality of unique ultrasonic audio samples to be outputted after the initial portion of the audio data is received;

determining that the received audio data comprises audio data corresponding to one of the plurality of unique ultrasonic audio samples; and

in response to determining that the received audio data comprises audio data corresponding to one of the plurality of unique ultrasonic audio samples, disabling access to the computing device.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044097/0658 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2015
From: SHARIFI, MATTHEW; FOERSTER, JAKOB NICOLAUS
To: GOOGLE INC.
Reel/Frame 037060/0664 →
Continuity (2)
Provisional Application 62242650 · Oct 16, 2015
Related Publication 20170110144A1 · Apr 20, 2017