IP Library › Granted Patent US 10,242,673
Granted Patent B2
US 10,242,673 · App. 15/371,907 · Granted Mar 26, 2019

Preventing of audio attacks using an input and an output hotword detection model

Inventors: Lee Campbell (Mountain View, CA); Samuel Kramer Beder (San Francisco, CA)
Assignee: Google LLC
G10L15/22G06F21/6218G10L15/16G10L15/30G10L15/20G10L15/222G10L17/24G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,242,673
App. No.
15/371,907
Filed
Dec 7, 2016
Granted
Mar 26, 2019
Kind
B2
Art Unit
2657
USPC
704/273
Abstract

In some implementations, a method includes receiving output audio data that is provided to a speaker of a device and that represents audio for output by the device, receiving, after the output audio data is provided to the speaker of the device, input audio data that represents audio detected by a microphone of the device, determining, by an output hotword detection model, that the output audio data that is provided to the speaker of the device includes a representation of a hotword, determining, by an input hotword detection model that is less accepting of hotwords than the output hotword detection model, that the input audio data that represents audio detected by a microphone of the device includes a representation of a hotword, and, in response, blocking use of the input audio data to initiate a command.

Claims (41)

1. A method comprising:

receiving, at a processing module of a device, output audio data that is provided to a speaker of the device and that represents audio for output by the device;

receiving, by the processing module and after the output audio data is provided to the speaker of the device, input audio data that represents audio detected by a microphone of the device;

determining, by an output hotword detection model of the processing module, that the output audio data that is provided to the speaker of the device includes a representation of a hotword, wherein the hotword is a word or phrase previously designated to precede a voice command;

determining, by an input hotword detection model that is less accepting of hotwords than the output hotword detection model, that the input audio data that represents audio detected by a microphone of the device includes a representation of a hotword; and

in response to determining, by the output hotword detection model, that the output audio data that is provided to the speaker of the device includes the representation of the hotword and, by the input hotword detection model that is less accepting of hotwords than the output hotword detection model, that the input audio data that represents input audio detected by the microphone of the device includes the representation of the hotword, blocking, by the processing module, use of the input audio data to initiate a command.

2. The method of claim 1 , wherein the determining that the output audio includes a representation of a hotword comprises:

generating by the output hotword detection model a hotword score for the output audio data,

comparing, by the output hotword detection model, the hotword score to a predetermined threshold; and

determining, by the output hotword detection model and based on the comparing, that the output audio includes a representation of a hotword.

3. The method of claim 2 , further comprising:

generating, by the input hotword detection model, a separate hotword score for the output audio data;

comparing by the input hotword detection model, the separate hotword score to a separate predetermined threshold;

confirming by the input hotword detection model and based on the comparing, that the output audio data includes a representation of a hotword; and

based on the confirming that the output audio data includes the presentation of the hotword, blocking, by the processing module, use of the input audio data to initiate a command.

4. The method of claim 3 , wherein the predetermined threshold is different from the separate predetermined threshold.

5. The method of claim 3 , wherein the output hotword detection model is a trained neural network, and wherein the input hotword detection model is a trained neural network.

6. The method of claim 5 , wherein the predetermined threshold is determined by the output hotword detection model during training, and wherein the separate predetermined threshold is determined by the output hotword detection model during training.

7. The method of claim 3 , wherein the input hotword detection model generates the separate hotword score after the determining that the output audio data includes the representation of the hotword.

8. The method of claim 2 , wherein blocking, by the processing module, use of the input audio data to initiate a command comprises blocking the command from being executed.

9. The method of claim 2 , further comprising outputting, by the processing module, data indicating that the device has been compromised.

10. The method of claim 1 , wherein the hotword is a predetermined word that has been designated to signal the beginning of a voice query or voice command that immediately follows the hotword.

11. The method of claim 1 , wherein the output hotword detection model and the input hotword detection model operate in parallel.

12. The method of claim 1 , wherein blocking, by the processing module, use of the input audio data to initiate a command comprises blocking use of the input audio data to initiate the command by preventing the device from transmitting the input audio data as a command to a remote server.

13. The method of claim 1 , wherein receiving, at a processing module of a device, output audio data that is provided to a speaker of the device and that represents audio for output by the device comprises:

receiving, at the processing module of the device, the output audio data before the audio is audibly output by the speaker.

14. The method of claim 1 , wherein determining, by the output hotword detection model of the processing model, that the output audio data that is provided to the speaker of the device includes the representation of the hotword occurs before the output audio data is audibly output by the speaker of the device.

15. A device comprising:

a processing module; and

one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the secure processing module to perform operations comprising:

receiving, at the processing module of the device, output audio data that is provided to a speaker of the device and that represents audio for output by the device;

receiving, by the processing module and after the output audio data is provided to the speaker of the device, input audio data that represents audio detected by a microphone of the device;

determining, by the processing module, that the output audio data that is provided to the speaker of the device includes a representation of a hotword, wherein the hotword is a word or phrase previously designated to precede a voice command;

determining, by an input hotword detection model that is less accepting of hotwords than an output hotword detection model, that the input audio data that represents audio detected by a microphone of the device includes a representation of a hotword; and

in response to determining, by the output hotword detection model, that the output audio data that is provided to the speaker of the device includes the representation of the hotword and, by the input hotword detection model that is less accepting of hotwords than the output hotword detection model, that the input audio data that represents input audio detected by the microphone of the device includes the representation of the hotword, blocking, by the processing module, use of the input audio data to initiate a command.

16. A computer-readable storage device storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, at a processing module of a device, output audio data that is provided to a speaker of the device and that represents audio for output by the device;

receiving, by the processing module and after the output audio data is provided to the speaker of the device, input audio data that represents audio detected by a microphone of the device;

determining, by an output hotword detection model of the processing module, that the output audio data that is provided to the speaker of the device includes a representation of a hotword, wherein the hotword is a word or phrase previously designated to precede a voice command;

determining, by an input hotword detection model that is less accepting of hotwords than the output hotword detection model, that the input audio data that represents audio detected by a microphone of the device includes a representation of a hotword; and

in response to determining, by the output hotword detection model, that the output audio data that is provided to the speaker of the device includes the representation of the hotword and, by the input hotword detection model that is less accepting of hotwords than the output hotword detection model, that the input audio data that represents input audio detected by the microphone of the device includes the representation of the hotword, blocking, by the processing module, use of the input audio data to initiate a command.

Assignments (2)
CHANGE OF NAME Recorded Oct 20, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044567/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2017
From: CAMPBELL, LEE; BEDER, SAMUEL KRAMER
To: GOOGLE INC.
Reel/Frame 040937/0687 →
Continuity (1)
Related Publication 20180158453A1 · Jun 7, 2018