IP Library › Granted Patent US 11,727,947
Granted Patent B2
US 11,727,947 · App. 17/457,820 · Granted Aug 15, 2023

Key phrase detection with audio watermarking

Inventor: Ricardo Antonio Garcia (Mountain View, CA)
Assignee: Google LLC
G10L19/018G06F3/165G06F21/31G10L15/08G10L15/22G10L21/00G10L2015/088G10L2015/223H04N21/233H04N21/8358
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,727,947
App. No.
17/457,820
Granted
Aug 15, 2023
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for using audio watermarks with key phrases. One of the methods includes receiving, by a playback device, an audio data stream; determining, before the audio data stream is output by the playback device, whether a portion of the audio data stream encodes a particular key phrase by analyzing the portion using an automated speech recognizer; in response to determining that the portion of the audio data stream encodes the particular key phrase, modifying the audio data stream to include an audio watermark; and providing the modified audio data stream for output.

Claims (48)

1. A computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations comprising:

determining whether an audio data stream to be output through a speaker encodes a key phrase, the audio data stream corresponding to one of music content or video content;

when the audio stream encodes the key phrase, creating a modified audio data stream by:

dynamically generating multiple audio watermarks encoding data that indicates the audio data stream originated from a content provider; and

inserting the dynamically generated multiple audio watermarks into the audio data stream to create the modified audio data stream; and

providing the modified audio data stream for output through the speaker,

wherein after providing the modified audio data stream for output through the speaker, a listening device, while in an awake mode responsive to detecting a key phrase:

captures the modified audio data stream; and

determines an action to perform using the multiple audio watermarks encoding the data that indicates the audio data stream originated from the content provider.

2. The computer-implemented method of claim 1 , wherein:

the data processing hardware resides on a playback device; and

prior to determining whether the audio data stream to be output through the speaker encodes the key phrase, the playback device receives the audio data stream from the content provider through a wireless input connection other than a microphone.

3. The computer-implemented method of claim 2 , wherein the playback device:

receives the audio data stream in a video stream from the content provider through the wireless input connection; and

connects to a display using a digital audio and video connection.

4. The computer-implemented method of claim 3 , wherein the operations further comprise, when providing the modified audio data stream for output through the speaker, providing, using the digital audio and video connection, a video portion of the video stream for presentation by the display.

5. The computer-implemented method of claim 4 , wherein the playback device synchronizes presentation of the video portion of the video stream by the display with the modified audio data stream for output through the speaker.

6. The computer-implemented method of claim 3 , where the playback device connects to a television using the digital audio and video connection, the television comprising the display and the speaker.

7. The computer-implemented method of claim 2 , wherein the playback device comprises the speaker.

8. The computer-implemented method of claim 1 , wherein the listening device is located in a same room as the speaker.

9. The computer-implemented method of claim 1 , wherein:

a portion of the multiple audio watermarks in the modified audio data stream encode different data than the other multiple audio watermarks; or

each of the multiple audio watermarks encode the same data.

10. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

determining whether an audio data stream to be output through a speaker encodes a key phrase, the audio data stream corresponding to one of music content or video content;

when the audio stream encodes the key phrase, creating a modified audio data stream by:

dynamically generating multiple audio watermarks encoding data that indicates the audio data stream originated from a content provider; and

inserting the dynamically generated multiple audio watermarks into the audio data stream to create the modified audio data stream; and

providing the modified audio data stream for output through the speaker,

wherein after providing the modified audio data stream for output through the speaker, a listening device, while in an awake mode responsive to detecting a key phrase:

captures the modified audio data stream; and

determines an action to perform using the multiple audio watermarks encoding the data that indicates the audio data stream originated from the content provider.

11. The system of claim 10 , wherein:

the data processing hardware and the memory hardware reside on a playback device; and

prior to determining whether the audio data stream to be output through the speaker encodes the key phrase, the playback device receives the audio data stream from the content provider through a wireless input connection other than a microphone.

12. The system of claim 11 , wherein the playback device:

receives the audio data stream in a video stream from the content provider through the wireless input connection; and

connects to a display using a digital audio and video connection.

13. The system of claim 12 , wherein the operations further comprise, when providing the modified audio data stream for output through the speaker, providing, using the digital audio and video connection, a video portion of the video stream for presentation by the display.

14. The system of claim 12 , where the playback device connects to a television using the digital audio and video connection, the television comprising the display and the speaker.

15. The system of claim 13 , wherein the playback device synchronizes presentation of the video portion of the video stream by the display with the modified audio data stream for output through the speaker.

16. The system of claim 11 , wherein the playback device comprises the speaker.

17. The system of claim 10 , wherein the listening device is located in a same room as the speaker.

18. The system of claim 10 , wherein:

a portion of the multiple audio watermarks in the modified audio data stream encode different data than the other multiple audio watermarks; or

each of the multiple audio watermarks encode the same data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2021
From: GARCIA, RICARDO ANTONIO
To: GOOGLE LLC
Reel/Frame 058316/0231 →
Continuity (4)
Continuation 16992647 · Aug 13, 2020
Continuation 16358109 · Mar 19, 2019
Continuation 15824183 · Nov 28, 2017
Related Publication 20220093114A1 · Mar 24, 2022
Cited By (2)
US 12,573,407 US 12,602,197