IP Library Granted Patent US 10,867,600
Granted Patent B2
US 10,867,600 · App. 15/799,501 · Granted Dec 15, 2020

Recorded media hotword trigger suppression

Inventors: Alexander H. Gruenstein (Mountain View, CA); Johan Schalkwyk (Scarsdale, NY); Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L15/22G06F3/167G10L15/08G10L15/30G10L25/51G10L17/00G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,867,600
App. No.
15/799,501
Granted
Dec 15, 2020
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for hotword trigger suppression are disclosed. In one aspect, a method includes the actions of receiving, by a microphone of a computing device, audio corresponding to playback of an item of media content, the audio including an utterance of a predefined hotword that is associated with performing an operation on the computing device. The actions further include processing the audio. The actions further include in response to processing the audio, suppressing performance of the operation on the computing device.

Claims (59)

1. A computer-implemented method, the method comprising:

receiving, at a server, from a computing device, audio corresponding to playback of an item of media content captured by a microphone of the computing device, the audio including an utterance of a predefined hotword that is associated with performing an operation on the computing device;

while the computing device performs speech recognition on a first portion of the audio:

generating, by the server, an audio fingerprint of a second portion of the audio;

comparing, by the server, the audio fingerprint to one or more stored audio fingerprints that are each associated with a known audio recording that contains or is associated with the predefined hotword; and

based on comparing the audio fingerprint to the one or more stored audio fingerprints, determining, by the server, that the audio fingerprint corresponds to at least one of the one or more stored audio fingerprints; and

in response to determining that the audio fingerprint corresponds to at least one of the one or more stored audio fingerprints, transmitting, by the server, an instruction to the computing device, the instruction, when received by the computing device causing the computing device to:

cease performing the speech recognition on the first portion of the audio;

identify one or more nearby computing devices, each nearby computing device of the identified one or more nearby computing devices configured to respond to the predefined hotword; and

send a notification to each nearby computing device of the identified one or more nearby computing devices, the notification instructing each nearby computing device of the identified one or more nearby computing devices to not respond to the predefined hotword.

2. The method of claim 1 , wherein comparing the audio fingerprint to the one or more stored audio fingerprints comprises:

providing the audio fingerprint to a computing device that stores the one or more audio fingerprints;

providing, to the computing device that stores the one or more audio fingerprints, a request to compare the audio fingerprint to the one or more audio fingerprints; and

receiving, from the computing device that stores the one or more audio fingerprints, comparison data based on comparing the audio fingerprint to the one or more audio fingerprints.

3. The method of claim 1 , wherein the computing device remains in a low power/sleep/inactive state while (i) receiving the audio, (ii) generating the audio fingerprint, (iii) comparing the audio fingerprint, (iv) determining that the audio fingerprint corresponds to the at least one of the one or more stored audio fingerprints, and (v) ceasing to perform the speech recognition on the first portion of the audio.

4. The method of claim 1 , further comprising, while receiving the audio, providing, to a display of the computing device, data indicating that the computing device is receiving the audio.

5. The method of claim 1 , further comprising, while performing speech recognition on the first portion of the audio, providing, to a display of the computing device, data indicating that the computing device is processing the audio.

6. The method of claim 1 , wherein the instruction, when received by the computing device, further causes the computing device to at least one of:

deactivate a display of the computing device;

return the computing device to a low power/sleep/inactive state; or

provide, to the display of the computing device, data indicating that the computing device ceased to perform speech recognition on the first portion of the audio.

7. The method of claim 1 , further comprising, providing, for output, a selectable option that, upon selection by a user, provides an instruction to the computing device to perform the operation on the computing device.

8. The method of claim 7 , further comprising:

detecting a selection of the selectable option; and

adjusting a process for processing subsequently received audio that includes an utterance of the predefined hotword.

9. The method of claim 1 , wherein the second portion of the audio is received before the predefined hotword.

10. The method of claim 1 , wherein:

the first portion of the audio is received after the predefined hotword; and

the first portion of the audio and the second portion of the audio are a same portion of the audio.

11. The method of claim 1 , further comprising:

before performing speech recognition on the first portion of the audio, determining that the audio includes an utterance of the predefined hotword,

wherein generating the audio fingerprint of the second portion of the audio is based on determining that the audio includes an utterance of the predefined hotword.

12. The method of claim 11 , wherein determining that the audio includes an utterance of the predefined hotword comprises determining that the audio includes an utterance of the predefined hotword without performing speech recognition on the audio.

13. The method of claim 1 , wherein the first portion of the audio is different from the second portion of the audio.

14. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, from a computing device, audio corresponding to playback of an item of media content captured by a microphone of the computing device, the audio including an utterance of a predefined hotword that is associated with performing an operation on the computing device;

while the computing device performs speech recognition on a first portion of the audio:

generating an audio fingerprint of a second portion of the audio;

comparing the audio fingerprint to one or more stored audio fingerprints that are each associated with a known audio recording that contains or is associated with the predefined hotword; and

based on comparing the audio fingerprint to the one or more stored audio fingerprints, determining that the audio fingerprint corresponds to at least one of the one or more stored audio fingerprints; and

in response to determining that the audio fingerprint corresponds to at least one of the one or more stored audio fingerprints transmitting an instruction to the computing device, the instruction, when received by the computing device causing the computing device to:

cease performing the speech recognition on the first portion of the audio;

identify one or more nearby computing devices, each nearby computing device of the identified one or more nearby computing devices configured to respond to the predefined hotword; and

send a notification to each nearby computing device of the identified one or more nearby computing devices, the notification instructing each nearby computing device of the identified one or more nearby computing devices to not respond to the predefined hotword.

15. The system of claim 14 , wherein comparing the audio fingerprint to the one or more stored audio fingerprints comprises:

providing the audio fingerprint to a computing device that stores the one or more audio fingerprints;

providing, to the computing device that stores the one or more audio fingerprints, a request to compare the audio fingerprint to the one or more audio fingerprints; and

receiving, from the computing device that stores the one or more audio fingerprints, comparison data based on comparing the audio fingerprint to the one or more audio fingerprints.

16. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, at a server, from a computing device, audio corresponding to playback of an item of media content captured by a microphone of the computing device, the audio including an utterance of a predefined hotword that is associated with performing an operation on the computing device;

while the computing device performs speech recognition on a first portion of the audio:

generating an audio fingerprint of a second portion of the audio;

comparing the audio fingerprint to one or more stored audio fingerprints that are each associated with a known audio recording that contains or is associated with the predefined hotword; and

based on comparing the audio fingerprint to the one or more stored audio fingerprints, determining that the audio fingerprint corresponds to at least one of the one or more stored audio fingerprints; and

in response to determining that the audio fingerprint corresponds to at least one of the one or more stored audio fingerprints, transmitting an instruction to the computing device, the instruction, when received by the computing device causing the computing device to:

cease performing the speech recognition on the first portion of the audio;

identify one or more nearby computing devices, each nearby computing device of the identified one or more nearby computing devices configured to respond to the predefined hotword; and

send a notification to each nearby computing device of the identified one or more nearby computing devices, the notification instructing each nearby computing device of the identified one or more nearby computing devices to not respond to the predefined hotword.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2017
From: GRUENSTEIN, ALEXANDER H.; SCHALKWYK, JOHAN; SHARIFI, MATTHEW
To: GOOGLE LLC
Reel/Frame 044003/0354 →
Continuity (2)
Provisional Application 62497044 · Nov 7, 2016
Related Publication 20180130469A1 · May 10, 2018
Cited By (2)
US 12,374,353 US 12,744,040