IP Library Granted Patent US 10,692,496
Granted Patent B2
US 10,692,496 · App. 16/418,415 · Granted Jun 23, 2020

Hotword suppression

Inventors: Alexander H. Gruenstein (Mountain View, CA); Taral Pradeep Joglekar (Sunnyvale, CA); Vijayaditya Peddinti (San Jose, CA); Michiel A. U. Bacchiani (Summit, NJ)
Assignee: Google LLC
G10L15/22G10L15/063G10L15/08G10L15/30G10L17/005G10L17/22G10L25/51G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,692,496
App. No.
16/418,415
Granted
Jun 23, 2020
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for suppressing hotwords are disclosed. In one aspect, a method includes the actions of receiving audio data corresponding to playback of an utterance. The actions further include providing the audio data as an input to a model (i) that is configured to determine whether a given audio data sample includes an audio watermark and (ii) that was trained using watermarked audio data samples that each include an audio watermark sample and non-watermarked audio data samples that do not each include an audio watermark sample. The actions further include receiving, from the model, data indicating whether the audio data includes the audio watermark. The actions further include, based on the data indicating whether the audio data includes the audio watermark, determining to continue or cease processing of the audio data.

Claims (65)

1. A computer-implemented method comprising:

receiving, by a computing device, audio data corresponding to playback of an utterance;

providing, by the computing device, the audio data as an input to a model (i) that is configured to determine whether a given audio data sample includes an audio watermark and (ii) that was trained using watermarked audio data samples that each include an audio watermark sample and non-watermarked audio data samples that do not each include an audio watermark sample;

receiving, by the computing device and from the model (i) that is configured to determine whether the given audio data sample includes the audio watermark and (ii) that was trained using the watermarked audio data samples that include the audio watermark and the non-watermarked audio data samples that do not include the audio watermark, data indicating whether the audio data includes the audio watermark; and

based on the data indicating whether the audio data includes the audio watermark, determining, by the computing device, to continue or cease processing of the audio data.

2. The method of claim 1 , wherein:

receiving the data indicating whether the audio data includes the audio watermark comprises receiving the data indicating that the audio data includes the audio watermark,

determining to continue or cease processing of the audio data comprises determining to cease processing of the audio data based on receiving the data indicating that the audio data includes the audio watermark, and

the method further comprises, based on determining to cease processing of the audio data, ceasing, by the computing device, processing of the audio data.

3. The method of claim 1 , wherein:

receiving the data indicating whether the audio data includes the audio watermark comprises receiving the data indicating that the audio data does not include the audio watermark,

determining to continue or cease processing of the audio data comprises determining to continue processing of the audio data based on receiving the data indicating that the audio data does not include the audio watermark, and

the method further comprises, based on determining to continue processing of the audio data, continuing, by the computing device, processing of the audio data.

4. The method of claim 1 , wherein the processing of the audio data comprises:

generating a transcription of the utterance by performing speech recognition on the audio data.

5. The method of claim 1 , wherein the processing of the audio data comprises:

determining whether the audio data includes an utterance of a particular, predefined hotword.

6. The method of claim 1 , comprising:

before providing the audio data as an input to the model (i) that is configured to determine whether a given audio data sample includes an audio watermark and (ii) that was trained using watermarked audio data samples that each include an audio watermark sample and non-watermarked audio data samples that do not each include an audio watermark sample, determining, by the computing device, that the audio data includes an utterance of a particular, predefined hotword.

7. The method of claim 1 , comprising:

determining, by the computing device, that the audio data includes an utterance of a particular, predefined hotword,

wherein providing the audio data as an input to the model (i) that is configured to determine whether a given audio data sample includes an audio watermark and (ii) that was trained using watermarked audio data samples that each include an audio watermark sample and non-watermarked audio data samples that do not each include an audio watermark sample is in response to determining that the audio data includes an utterance of a particular, predefined hotword.

8. The method of claim 1 , comprising:

receiving, by the computing device, the watermarked audio data samples that each include an audio watermark, the non-watermarked audio data samples that do not each include an audio watermark, and data indicating whether each watermarked and non-watermarked audio sample includes an audio watermark; and

training, by the computing device and using machine learning, the model using the watermarked audio data samples that each include an audio watermark, the non-watermarked audio data samples that do not each include the audio watermark, and the data indicating whether each watermarked and non-watermarked audio sample includes an audio watermark.

9. The method of claim 8 , wherein at least a portion of the watermarked audio data samples each include an audio watermark at multiple, periodic locations.

10. The method of claim 8 , wherein audio watermarks in one of the watermarked audio data samples are different to audio watermark in another of the watermarked audio data samples.

11. The method of claim 1 , comprising:

determining, by the computing device, a first time of receipt of the audio data corresponding to playback of an utterance;

receiving, by the computing device, a second time that an additional computing device provided, for output, the audio data corresponding to playback of an utterance and data indicating whether the audio data included a watermark;

determining, by the computing device, that the first time matches the second time; and

based on determining that the first time matches the second time, updating, by the computing device, the model using the data indicating whether the audio data included a watermark.

12. A system comprising:

one or more computers; and

one or more non-transitory storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, by a computing device, audio data corresponding to playback of an utterance;

providing, by the computing device, the audio data as an input to a model (i) that is configured to determine whether a given audio data sample includes an audio watermark and (ii) that was trained using watermarked audio data samples that each include an audio watermark sample and non-watermarked audio data samples that do not each include an audio watermark sample;

receiving, by the computing device and from the model (i) that is configured to determine whether the given audio data sample includes the audio watermark and (ii) that was trained using the watermarked audio data samples that include the audio watermark and the non-watermarked audio data samples that do not include the audio watermark, data indicating whether the audio data includes the audio watermark; and

based on the data indicating whether the audio data includes the audio watermark, determining, by the computing device, to continue or cease processing of the audio data.

13. The system of claim 12 , wherein:

receiving the data indicating whether the audio data includes the audio watermark comprises receiving the data indicating that the audio data includes the audio watermark,

determining to continue or cease processing of the audio data comprises determining to cease processing of the audio data based on receiving the data indicating that the audio data includes the audio watermark, and

the method further comprises, based on determining to cease processing of the audio data, ceasing, by the computing device, processing of the audio data.

14. The system of claim 12 , wherein the processing of the audio data comprises:

generating a transcription of the utterance by performing speech recognition on the audio data.

15. The system of claim 12 , wherein the processing of the audio data comprises:

determining whether the audio data includes an utterance of a particular, predefined hotword.

16. The system of claim 12 , wherein the operations comprise:

before providing the audio data as an input to the model (i) that is configured to determine whether a given audio data sample includes an audio watermark and (ii) that was trained using watermarked audio data samples that each include an audio watermark sample and non-watermarked audio data samples that do not each include an audio watermark sample, determining, by the computing device, that the audio data includes an utterance of a particular, predefined hotword.

17. The system of claim 12 , wherein the operations comprise:

determining, by the computing device, that the audio data includes an utterance of a particular, predefined hotword,

wherein providing the audio data as an input to the model (i) that is configured to determine whether a given audio data sample includes an audio watermark and (ii) that was trained using watermarked audio data samples that each include an audio watermark sample and non-watermarked audio data samples that do not each include an audio watermark sample is in response to determining that the audio data includes an utterance of a particular, predefined hotword.

18. The system of claim 12 , wherein the operations comprise:

receiving, by the computing device, the watermarked audio data samples that each include an audio watermark, the non-watermarked audio data samples that do not each include an audio watermark, and data indicating whether each watermarked and non-watermarked audio sample includes an audio watermark; and

training, by the computing device and using machine learning, the model using the watermarked audio data samples that each include an audio watermark, the non-watermarked audio data samples that do not each include the audio watermark, and the data indicating whether each watermarked and non-watermarked audio sample includes an audio watermark.

19. The system of claim 12 , wherein the operations comprise:

determining, by the computing device, a first time of receipt of the audio data corresponding to playback of an utterance;

receiving, by the computing device, a second time that an additional computing device provided, for output, the audio data corresponding to playback of an utterance and data indicating whether the audio data included a watermark;

determining, by the computing device, that the first time matches the second time; and

based on determining that the first time matches the second time, updating, by the computing device, the model using the data indicating whether the audio data included a watermark.

20. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, by a computing device, audio data corresponding to playback of an utterance;

providing, by the computing device, the audio data as an input to a model (i) that is configured to determine whether a given audio data sample includes an audio watermark and (ii) that was trained using watermarked audio data samples that each include an audio watermark sample and non-watermarked audio data samples that do not each include an audio watermark sample;

receiving, by the computing device and from the model (i) that is configured to determine whether the given audio data sample includes the audio watermark and (ii) that was trained using the watermarked audio data samples that include the audio watermark and the non-watermarked audio data samples that do not include the audio watermark, data indicating whether the audio data includes the audio watermark; and

based on the data indicating whether the audio data includes the audio watermark, determining, by the computing device, to continue or cease processing of the audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 19, 2019
From: GRUENSTEIN, ALEXANDER H.; JOGLEKAR, TARAL PRADEEP; PEDDINTI, VIJAYADITYA; BACCHIANI, MICHIEL A.U.
To: GOOGLE LLC
Reel/Frame 049517/0341 →
Continuity (2)
Provisional Application 62674973 · May 22, 2018
Related Publication 20190362719A1 · Nov 28, 2019
Cited By (1)
US 12,573,400