IP Library › Granted Patent US 11,967,323
Granted Patent B2
US 11,967,323 · App. 17/849,253 · Granted Apr 23, 2024

Hotword suppression

Inventors: Alexander H. Gruenstein (Mountain View, CA); Taral Pradeep Joglekar (Sunnyvale, CA); Vijayaditya Peddinti (San Jose, CA); Michiel A. U. Bacchiani (Summit, NJ)
Assignee: GOOGLE LLC
G10L15/22G10L15/063G10L15/08G10L15/30G10L17/00G10L17/22G10L25/51G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,967,323
App. No.
17/849,253
Granted
Apr 23, 2024
Kind
B2
Abstract

A method includes adding, by a first computing device, a first audio watermark to first speech data corresponding to playback of a first utterance including a hotword used to invoke an attention of a second computing device. The method includes outputting, by the first computing device, the playback of the first utterance corresponding to the watermarked first speech data. The second computing device is configured to receive the watermarked first speech data and determine to cease processing of the watermarked first speech data.

Claims (44)

1. A computer-implemented method, comprising:

adding, by a first computing device, a first audio watermark to first speech data corresponding to playback of a first utterance including a hotword used to invoke an attention of a second computing device; and

outputting, by the first computing device, the playback of the first utterance corresponding to the watermarked first speech data, to the second computing device, which is configured to receive the watermarked first speech data and determine to cease processing of the watermarked first speech data.

2. The method of claim 1 , further comprising:

storing, by the first computing device, at least one of (i) data and time of outputting the playback of the first utterance, (ii) a location of the first computing device, (iii) a transcription of the playback of the first utterance, (iv) the watermarked first speech data, or (v) the first speech data without the first audio watermarks.

3. The method of claim 1 , wherein the adding comprises:

identifying, by the first computing device, a portion of the first speech data for including audio samples of the hotword; and

adding, by the first computing device, the first audio watermark to the portion of the first speech data.

4. The method of claim 3 , wherein the identifying comprises:

identifying, by the first computing device, the portion of the first speech data by performing a speech recognition on the first speech data.

5. The method of claim 3 , wherein the first audio watermark is based on a spread spectrum technique, and the identifying comprises:

identifying, by the first computing device, the portion of the first speech data based on a minimum energy criterion.

6. The method of claim 1 , wherein the adding comprises:

adding, by the first computing device, the first audio watermark at periodic intervals to the first speech data.

7. The method of claim 1 , wherein the adding comprises:

adding, by the first computing device, an audio fingerprint to the first speech data.

8. The method of claim 1 , further comprising:

adding, by the first computing device, a second audio watermark to second speech data corresponding to playback of a second utterance including the hotword, a query included in the second utterance being different from a query included in the first utterance, and the second audio watermark being different from the first audio watermark.

9. The method of claim 1 , wherein the second computing device is configured to determine the watermarked first speech data includes the first audio watermark based on a model (i) that is configured to determine whether a given audio data sample includes an audio watermark and (ii) that was trained using watermarked audio data samples that each include an audio watermark sample and non-watermarked audio data samples that do not each include an audio watermark sample.

10. An apparatus, comprising:

processing circuitry configured to

add a first audio watermark to first speech data corresponding to playback of a first utterance including a hotword used to invoke an attention of a client device, and

output the playback of the first utterance corresponding to the watermarked first speech data, to the client device, which is configured to receive the watermarked first speech data and determine to cease processing of the watermarked first speech data.

11. The apparatus of claim 10 , wherein the processing circuitry is configured to:

store at least one of (i) data and time of outputting the playback of the first utterance, (ii) a location of the apparatus, (iii) a transcription of the playback of the first utterance, (iv) the watermarked first speech data, or (v) the first speech data without the first audio watermarks.

12. The apparatus of claim 10 , wherein the processing circuitry is configured to add the first audio watermark by:

identifying a portion of the first speech data for including audio samples of the hotword, and

adding the first audio watermark to the portion of the first speech data.

13. The apparatus of claim 12 , wherein the identifying of the portion of the first speech data comprises:

identifying the portion of the first speech data by performing a speech recognition on the first speech data.

14. The apparatus of claim 12 , wherein the first audio watermark is based on a spread spectrum technique, and the identifying of the portion of the first speech data comprises:

identifying the portion of the first speech data based on a minimum energy criterion.

15. The apparatus of claim 10 , wherein the processing circuitry is configured to add the first audio watermark by:

adding the first audio watermark at periodic intervals to the first speech data.

16. The apparatus of claim 10 , wherein the processing circuitry is configured to add the first audio watermark by:

adding an audio fingerprint to the first speech data.

17. The apparatus of claim 10 , wherein the processing circuitry is configured to:

add a second audio watermark to second speech data corresponding to playback of a second utterance including the hotword, a query included in the second utterance being different from a query included in the first utterance, and the second audio watermark being different from the first audio watermark.

18. The apparatus of claim 10 , wherein the client device is configured to determine that the watermarked first speech data includes the first audio watermark based on a model (i) that is configured to determine whether a given audio data sample includes an audio watermark and (ii) that was trained using watermarked audio data samples that each include an audio watermark sample and non-watermarked audio data samples that do not each include an audio watermark sample.

19. A non-transitory computer-readable storage medium including computer executable instructions, wherein the instructions, when executed by a computer, cause the computer to perform a method, the method comprising:

adding a first audio watermark to first speech data corresponding to playback of a first utterance including a hotword used to invoke an attention of a client device; and

outputting the playback of the first utterance corresponding to the watermarked first speech data, to the client device, which is configured to receive the watermarked first speech data and determine to cease processing of the watermarked first speech data.

20. The non-transitory computer-readable storage medium of claim 19 , wherein the method comprises:

adding a second audio watermark to second speech data corresponding to playback of a second utterance including the hotword, a query included in the second utterance being different from a query included in the first utterance, and the second audio watermark being different from the first audio watermark.

Continuity (4)
Continuation 16874646 · May 14, 2020
Continuation 16418415 · May 21, 2019
Provisional Application 62674973 · May 22, 2018
Related Publication 20220319519A1 · Oct 6, 2022
Cited By (1)
US 12,573,400