IP Library Granted Patent US 10,764,230
Granted Patent B2
US 10,764,230 · App. 15/985,421 · Granted Sep 1, 2020

Low latency audio watermark embedding

Inventor: John D. Lord (West Linn, OR)
Assignee: Digimarc Corporation
H04L51/32G06T1/0021G10L19/018H04L51/10H04N21/2187H04N21/21805H04N21/233H04N21/23418H04N21/2665H04N21/2668H04N21/2743H04N21/4223H04N21/42203H04N21/4622H04N21/6581H04N21/8358G06T2201/005H04N21/4788
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,764,230
App. No.
15/985,421
Granted
Sep 1, 2020
Kind
B2
Abstract

A method for low latency audio watermark embedding buffers samples of an audio stream in a buffer, including previous blocks of audio samples in the audio stream. It computes a perceptual mask from the audio samples in the buffer, generates a watermark signal; and applies the perceptual mask to the watermark signal for the first block to produce a mask-applied watermark signal. It inserts the mask-applied watermark signal into the audio samples of the first block without waiting for a subsequent audio block of samples in the audio stream and outputs watermarked audio of the first block.

Claims (57)

1. A method for low latency audio watermark embedding, the method comprising:

buffering N most recently received samples of an audio stream in a buffer, the buffer comprising a first block of M most recently received audio samples and previous blocks of audio samples in the audio stream, where M is less than N;

computing a perceptual mask from the audio samples in the buffer;

generating a watermark signal;

applying the perceptual mask to the watermark signal for the first block to produce a mask-applied watermark signal; and

inserting the mask-applied watermark signal into the audio samples of the first block without waiting for a subsequent audio block of M samples in the audio stream and outputting watermarked audio of the first block.

2. The method of claim 1 wherein new segments of the audio stream are added to the buffer in a rolling manner so that the perceptual mask is computed for a longer audio block than a current audio block arriving from a live audio stream, and the perceptual mask is applied to the watermark signal for the current audio block.

3. The method of claim 1 further comprising:

converting the samples of audio in the buffer to a frequency domain representation;

computing the perceptual mask from the frequency domain representation of the audio samples in the buffer;

converting a generated watermark signal to a frequency domain representation; and

applying the perceptual mask to the frequency domain representation of the generated watermark signal, wherein only a short segment of the watermark signal corresponding to a most recent segment of the audio stream is inserted into the audio samples of the first block, and wherein the perceptual mask is updated from audio samples of the first block and the previous audio blocks.

4. The method of claim 3 wherein the perceptual mask is re-computed for each new arriving block of audio from the audio stream based on the new arriving audio block and plural previous audio blocks that arrived previously from the audio stream.

5. A method for low latency audio watermark embedding, the method comprising:

buffering samples of an audio stream in a buffer, the buffer comprising a first block of audio samples and previous blocks of audio samples in the audio stream;

computing a perceptual mask from the audio samples in the buffer;

generating a watermark signal for the first block;

applying the perceptual mask to the watermark signal for the first block; and

inserting the watermark signal for the first block into the audio samples of the first block; the method further comprising:

converting the samples of audio in the buffer to a frequency domain representation;

computing the perceptual mask from the frequency domain representation of the audio samples in the buffer;

converting a generated watermark signal to a frequency domain representation; and

applying the perceptual mask to the frequency domain representation of the generated watermark signal, wherein only a short segment of the watermark signal corresponding to a most recent segment of the audio stream is inserted into the audio signal of the first block; wherein the perceptual mask is re-computed for each new arriving block of audio from the audio stream based on the new arriving audio block and plural previous audio blocks that arrived previously from the audio stream; wherein the buffering comprises buffering over 1000 samples of most recent samples of the audio stream, wherein the audio stream is sampled at greater than 40,000 Hz and latency of the inserting of the watermarking signal in the audio stream is on the order of tens of microseconds.

6. The method of claim 5 wherein the method is executed in an audio processing system in which the audio stream is captured from a live event at a venue via a microphone, the watermark signal is inserted, and watermarked audio blocks are output at the venue.

7. An audio processing system for low latency insertion of data into an audio signal, the system comprising:

plural buffers, including a perceptual mask buffer configured to buffer N most recently received samples of an audio stream, the perceptual mask buffer comprising a first block of M most recently received audio samples and previous blocks of audio samples in the audio stream, where M is less than N;

a processor in communication with the perceptual mask buffer, the processor configured to compute a perceptual mask from the audio samples in the buffer;

the processor further configured to convert a variable digital payload into a watermark signal;

the processor configured to apply the perceptual mask to the watermark signal for the first block to produce a mask-applied watermark signal, insert the mask-applied watermark signal into the audio samples of the first block without waiting for a subsequent audio block of M samples in the audio stream, and output watermarked audio of the first block.

8. The system of claim 7 wherein new segments of the audio stream are added to the perceptual mask buffer in a rolling manner so that the perceptual mask is computed for a longer audio block than a current audio block arriving from a live audio stream, and the perceptual mask is applied to the watermark signal for the current audio block.

9. The system of claim 7 wherein the processor is configured to execute instructions to:

convert the samples of audio in the buffer to a frequency domain representation;

compute the perceptual mask from the frequency domain representation of the audio samples in the buffer;

convert a generated watermark signal to a frequency domain representation; and

apply the perceptual mask to the frequency domain representation of the generated watermark signal, wherein only a short segment of the watermark signal corresponding to a most recent segment of the audio stream is inserted into the audio samples of the first block, and wherein the perceptual mask is updated from audio samples of the first block and the previous audio blocks.

10. The system of claim 9 wherein the processor is configured to execute instructions to:

re-compute the perceptual mask for each new arriving block of audio from the audio stream based on the new arriving audio block and plural previous audio blocks that arrived previously from the audio stream.

11. An audio processing system for low latency insertion of data into an audio signal, the system comprising:

plural buffers, including a perceptual mask buffer configured to buffer samples of an audio stream, the perceptual mask buffer comprising a first block of audio samples and previous blocks of audio samples in the audio stream;

a processor in communication with the perceptual mask buffer, the processor configured to compute a perceptual mask from the audio samples in the buffer;

the processor further configured to convert a variable digital payload into a watermark signal;

the processor configured to apply the perceptual mask to the watermark signal for the first block and insert the watermark signal into the audio samples of the first block;

wherein the processor is configured to execute instructions to:

convert the samples of audio in the perceptual mask buffer to a frequency domain representation;

compute the perceptual mask from the frequency domain representation of the audio samples in the perceptual mask buffer;

convert a generated watermark signal to a frequency domain representation; and

apply the perceptual mask to the frequency domain representation of the generated watermark signal, wherein only a short segment of the watermark signal corresponding to a most recent segment of the audio stream is inserted into the audio samples of the first block;

wherein the processor is configured to execute instructions to:

re-compute the perceptual mask for each new arriving block of audio from the audio stream based on the new arriving audio block and plural previous audio blocks that arrived previously from the audio stream;

wherein the perceptual mask buffer is configured to buffer over 1000 samples of most recent samples of the audio stream, wherein the audio stream is sampled at greater than 40,000 Hz and latency of the inserting of the watermarking signal in the audio stream is on the order of tens of microseconds.

12. The system of claim 11 comprising a microphone, wherein the system is configured to capture the audio stream from a live event at a venue via a microphone, to insert the watermark signal and to output watermarked audio blocks at the venue.

13. An apparatus for low latency audio watermark embedding, the apparatus comprising:

means for capturing and storing N most recently received digital samples of an audio signal in a first buffer, wherein N is an integer number of audio samples, and the N audio samples comprising a first block of M most recently received audio samples and previous blocks of audio samples in the audio stream;

means for deriving a perceptual mask from the N audio samples in the first buffer;

means for generating a portion of digital watermark signal corresponding to M samples of the first block from a variable data payload, where M is less than N;

means for adapting the portion according to the perceptual mask without waiting for a subsequent audio block of M samples in the audio stream; and

means for combining the adapted portion with the audio signal of the first block without waiting for a subsequent audio block of M samples in the audio stream and outputting watermarked audio of the first block.

Assignments (3)
ARTICLES OF CONVERSION Recorded Jun 19, 2026
From: DIGIMARC CORPORATION
To: DIGIMARC LLC
Reel/Frame 075863/0211 →
ARTICLES OF AMENDMENT OFTHE ARTICLES OF ORGANIZATION OF DIGIMARC LLC Recorded Jun 19, 2026
From: DIGIMARC LLC
To: DMRC LLC
Reel/Frame 075863/0266 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2019
From: LORD, JOHN D.
To: DIGIMARC CORPORATION
Reel/Frame 050373/0552 →
Continuity (4)
Continuation 15276505 · Sep 26, 2016
Continuation 14270163 · May 5, 2014
Provisional Application 61819506 · May 3, 2013
Related Publication 20180343224A1 · Nov 29, 2018
Cited By (2)
US 12,520,403 US 12,537,803