IP Library Granted Patent US 12,254,888
Granted Patent B2
US 12,254,888 · App. 18/373,244 · Granted Mar 18, 2025

Multi-factor audio watermarking

Inventors: Aleks Kracun (New York, NY); Matthew Sharifi (Kilchberg, CH)
Assignee: GOOGLE LLC
G10L17/22G10L17/02G10L17/06G10L25/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,888
App. No.
18/373,244
Granted
Mar 18, 2025
Kind
B2
Abstract

Techniques are described herein for multi-factor audio watermarking. A method includes: receiving audio data; processing the audio data to generate predicted output that indicates a probability of one or more hotwords being present in the audio data; determining that the predicted output satisfies a threshold that is indicative of the one or more hotwords being present in the audio data; in response to determining that the predicted output satisfies the threshold, processing the audio data using automatic speech recognition to generate a speech transcription feature; detecting a watermark that is embedded in the audio data; and in response to detecting the watermark: determining that the speech transcription feature corresponds to one of a plurality of stored speech transcription features; and in response to determining that the speech transcription feature corresponds to one of the plurality of stored speech transcription features, suppressing processing of a query included in the audio data.

Claims (48)

1. A method implemented by one or more processors, the method comprising:

inserting, into an audio track of media content, at a particular point in the audio track, a particular watermark, wherein the particular point in the audio track is a temporal location in the audio track that is prior to, concurrent with, or subsequent to a hotword or query in the audio track;

determining, based on the audio track of the media content, watermark registration data comprising a plurality of features, the plurality of features including (i) an indication of the particular watermark that was inserted into the audio track and (ii) a speech transcription feature or intermediate embedding that is associated with the particular watermark; and

providing, to a client device, the watermark registration data associated with the audio track of the media content, wherein providing the watermark registration data associated with the audio track of the media content to the client device causes the client device to:

store, locally at the client device, the watermark registration data;

detect audio data, via one or more microphones of the client device, corresponding to the audio track of the media content;

determine, based on processing the audio data, that the audio data includes the particular watermark included in the watermark registration data that is stored locally at the client device; and

in response to determining that the audio data includes the particular watermark included in the watermark registration data that is stored locally at the client device:

determine, based on processing the audio data, that the audio data includes the speech transcription feature or the intermediate embedding that is associated with the particular watermark included in the watermark registration data that is stored locally at the client device; and

in response to determining that the audio data includes the speech transcription feature or the intermediate embedding that is associated with the particular watermark:

suppress, based on detecting the audio data that includes the particular watermark and that includes the speech transcription feature or the intermediate embedding that is associated with the particular watermark, processing of the hotword or the query in the audio track.

2. The method according to claim 1 , wherein providing the watermark registration data associated with the audio track of the media content to the client device further causes the client device to modify a threshold that is indicative of one or more hotwords being present in audio data.

3. The method according to claim 1 , further comprising causing the watermark registration data to be stored on a cloud computing node,

wherein providing the watermark registration data to the client device comprises the cloud computing node providing the watermark registration data to the client device.

4. The method according to claim 1 , wherein the particular watermark is an audio watermark that is imperceptible to humans.

5. The method according to claim 1 , wherein the speech transcription feature or intermediate embedding that is associated with the particular watermark corresponds to the particular point in the audio track where the particular watermark was inserted.

6. The method according to claim 1 , wherein the speech transcription feature or intermediate embedding that is associated with the particular watermark includes erroneous transcriptions.

7. A computer program product comprising one or more non-transitory computer-readable storage media having program instructions collectively stored on the one or more non-transitory computer-readable storage media, the program instructions executable to:

insert, into an audio track of media content, at a particular point in the audio track, a particular watermark, wherein the particular point in the audio track is a temporal location in the audio track that is prior to, concurrent with, or subsequent to a hotword or query in the audio track;

determine, based on the audio track of the media content, watermark registration data comprising a plurality of features, the plurality of features including (i) an indication of the particular watermark that was inserted into the audio track and (ii) a speech transcription feature or intermediate embedding that is associated with the particular watermark; and

provide, to a client device, the watermark registration data associated with the audio track of the media content, wherein the instructions to provide the watermark registration data associated with the audio track of the media content to the client device cause the client device to:

store, locally at the client device, the watermark registration data;

detect audio data, via one or more microphones of the client device, corresponding to the audio track of the media content;

determine, based on processing the audio data, that the audio data includes the particular watermark included in the watermark registration data that is stored locally at the client device; and

in response to determining that the audio data includes the particular watermark included in the watermark registration data that is stored locally at the client device:

determine, based on processing the audio data, that the audio data includes the speech transcription feature or the intermediate embedding that is associated with the particular watermark included in the watermark registration data that is stored locally at the client device; and

in response to determining that the audio data includes the speech transcription feature or the intermediate embedding that is associated with the particular watermark:

suppress, based on detecting the audio data that includes the particular watermark and that includes the speech transcription feature or the intermediate embedding that is associated with the particular watermark, processing of the hotword or the query in the audio track.

8. The computer program product according to claim 7 , wherein the instructions to provide the watermark registration data associated with the audio track of the media content to the client device further causes the client device to modify a threshold that is indicative of one or more hotwords being present in audio data.

9. The computer program product according to claim 7 , wherein:

the program instructions are further executable to cause the watermark registration data to be stored on a cloud computing node; and

providing the watermark registration data to the client device comprises the cloud computing node providing the watermark registration data to the client device.

10. The computer program product according to claim 7 , wherein the particular watermark is an audio watermark that is imperceptible to humans.

11. The computer program product according to claim 7 , wherein the speech transcription feature or intermediate embedding that is associated with the particular watermark corresponds to the particular point in the audio track where the particular watermark was inserted.

12. The computer program product according to claim 7 , wherein the speech transcription feature or intermediate embedding that is associated with the particular watermark includes erroneous transcriptions.

13. A system comprising:

a processor, a computer-readable memory, one or more computer-readable storage media, and program instructions collectively stored on the one or more computer-readable storage media, the program instructions executable to:

insert, into an audio track of media content, at a particular point in the audio track, a particular watermark, wherein the particular point in the audio track is a temporal location in the audio track that is prior to, concurrent with, or subsequent to a hotword or query in the audio track;

determine, based on the audio track of the media content, watermark registration data comprising a plurality of features, the plurality of features including (i) an indication of the particular watermark that was inserted into the audio track and (ii) a speech transcription feature or intermediate embedding that is associated with the particular watermark; and

provide, to a client device, the watermark registration data associated with the audio track of the media content, wherein the instructions to provide the watermark registration data associated with the audio track of the media content to the client device cause the client device to:

store, locally at the client device, the watermark registration data;

detect audio data, via one or more microphones of the client device, corresponding to the audio track of the media content;

determine, based on processing the audio data, that the audio data includes the particular watermark included in the watermark registration data that is stored locally at the client device; and

in response to determining that the audio data includes the particular watermark included in the watermark registration data that is stored locally at the client device:

determine, based on processing the audio data, that the audio data includes the speech transcription feature or the intermediate embedding that is associated with the particular watermark included in the watermark registration data that is stored locally at the client device; and

in response to determining that the audio data includes the speech transcription feature or the intermediate embedding that is associated with the particular watermark:

suppress, based on detecting the audio data that includes the particular watermark and that includes the speech transcription feature or the intermediate embedding that is associated with the particular watermark, processing of the hotword or the query in the audio track.

14. The system according to claim 13 , wherein the instructions to provide the watermark registration data associated with the audio track of the media content to the client device further causes the client device to modify a threshold that is indicative of one or more hotwords being present in audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 5, 2023
From: KRACUN, ALEKS; SHARIFI, MATTHEW
To: GOOGLE LLC
Reel/Frame 065137/0985 →
Continuity (3)
Continuation 17114118 · Dec 7, 2020
Provisional Application 63110959 · Nov 6, 2020
Related Publication 20240021207A1 · Jan 18, 2024
References Cited (18)
US 9548053B1 · Basye · 2017 [cited by examiner]
US 10276175B1 · Garcia · 2019 [cited by examiner]
US 11100930B1 · Salem · 2021 [cited by examiner]
US 11244693B2 · Chauhan · 2022 [cited by examiner]
US 20070033026A1 · Bartosik · 2007 [cited by examiner]
US 20160077794A1 · Kim · 2016 [cited by examiner]
US 20180130469A1 · Gruenstein · 2018 [cited by examiner]
US 20180350356A1 · Garcia · 2018 [cited by applicant]
US 20190362719A1 · Gruenstein et al. · 2019 [cited by applicant]
US 20200098380A1 · Tai · 2020 [cited by examiner]
US 20210090575A1 · Mahmood et al. · 2021 [cited by applicant]
US 20220148601A1 · Kracun et al. · 2022 [cited by applicant]
WO 2020005202 · 2020 [cited by applicant]
Bar-Yossef, et al.; Approximating edit distance efficiently; 45th Annual IEEE Symposium on Foundations of Computer Science; pp. 1-10; dated 2004. [cited by applicant]
Intellectual Property India; Examination Report issued in Application No. 202227062338; 8 pages; dated Jul. 27, 2023. [cited by applicant]
European Patent Office; International Search Report and Written Opinion issued in Application No. PCT/US2021/058315; 13 pages; dated Mar. 10, 2022. [cited by applicant]
Tai, Y.Y. and Mansour, M.F. “Audio Watermarking Over the Air with Modulated Self-Correlation”. Amazon Inc., USA; arXiv:1903.08238v1 [cs.MM] Mar. 19, 2019; 5 pages. [cited by applicant]
European Patent Office, Intention to Grant issued in Application No. 21816268.3; 50 pages; dated Aug. 16, 2024. [cited by applicant]