IP Library Granted Patent US 8,626,504
Granted Patent B2
US 8,626,504 · App. 13/599,992 · Granted Jan 7, 2014

Extracting features of audio signal content to provide reliable identification of the signals

Inventors: Regunathan Radhakrishnan (San Bruno, CA); Claus Bauer (Beijing, CN); Kent Bennett Terry (Millbrae, CA); Brian David Link (San Jose, CA); Hyung-Suk Kim (Valencia, CA); Eric Gsell (Valencia, CA)
Assignee: Dolby Laboratories Licensing Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,626,504
App. No.
13/599,992
Granted
Jan 7, 2014
Kind
B2
Abstract

Signatures that can be used to identify video and audio content are generated from the content by generating measures of dissimilarity between features of corresponding groups of pixels in frames of video content and by generating low-resolution time-frequency representations of audio segments. The signatures are generated by applying a hash function to intermediate values derived from the measures of dissimilarity and to the low-resolution time-frequency representations. The generated signatures may be used in a variety of applications such as restoring synchronization between video and audio content streams and identifying copies of original video and audio content. The generated signatures can provide reliable identifications despite intentional and unintentional modifications to the content.

Claims (39)

1. A method for generating a signature that identifies content of an audio signal, wherein the method performed by a device comprises:

obtaining, by a device, a time-frequency representation of a set of blocks within a sequence of blocks of the audio signal, wherein the time-frequency representation comprises sets of spectral values, each set of spectral values representing all spectral components within at least a portion of the bandwidth of the audio signal in a respective block in the set of blocks;

deriving, by a device, intermediate values from intensities of all the spectral values arranged in groups of one or more spectral values within a respective set of spectral values; and

generating, by a device the signature that identifies content of the audio signal by projecting the intermediate values onto a set of random vectors, wherein the signature is represented by bits and each bit of the signature is derived from contributions from all of the intermediate values;

wherein each respective vector in the set of random vectors has vector elements with values that are obtained from a difference between uniformly distributed random variables within a range from zero to one and an average of the uniformly distributed random variables for all vector elements of the respective vector;

the projection of the intermediate values onto a respective random vector is obtained from an inner product of the intermediate values with the vector elements of the respective vector; and

each component of the signature has either a first value when the projection of the intermediate values onto a corresponding random vector is greater than a threshold or has a second value when the projection of the intermediate values onto the corresponding random vector is less than a threshold, wherein the threshold is equal to a median of the projections of intermediate values onto the set of random vectors.

2. The method of claim 1 , wherein:

the time-frequency representation is obtained by applying a time-to-frequency transform to each block of the audio signal in the set of blocks to obtain a respective set of spectral values; and

a respective intermediate value is derived by calculating an average intensity of the one or more spectral values in a group within the respective set of spectral values.

3. The method of claim 1 , wherein the groups of spectral values have numbers of spectral values that vary with frequency.

4. The method of claim 3 , wherein the groups of spectral values for higher frequencies have a greater number of spectral values.

5. The method of claim 1 , wherein each component of the signature is derived from the projection of the intermediate values onto a respective random vector.

6. An apparatus for generating a signature that identifies content of an audio signal, wherein the apparatus comprises:

means for obtaining a time-frequency representation of a set of blocks within a sequence of blocks of the audio signal, wherein the time-frequency representation comprises sets of spectral values, each set of spectral values representing all spectral components within at least a portion of the bandwidth of the audio signal in a respective block in the set of blocks;

means for deriving intermediate values from intensities of all the spectral values arranged in groups of one or more spectral values within a respective set of spectral values; and

means for generating the signature that identifies content of the audio signal by projecting the intermediate values onto a set of random vectors, wherein the signature is represented by bits and each bit of the signature is derived from contributions from all of the intermediate values;

wherein each respective vector in the set of random vectors has vector elements with values that are obtained from a difference between uniformly distributed random variables within a range from zero to one and an average of the uniformly distributed random variables for all vector elements of the respective vector;

the projection of the intermediate values onto a respective random vector is obtained from an inner product of the intermediate values with the vector elements of the respective vector; and

each component of the signature has either a first value when the projection of the intermediate values onto a corresponding random vector is greater than a threshold or has a second value when the projection of the intermediate values onto the corresponding random vector is less than a threshold, wherein the threshold is equal to a median of the projections of intermediate values onto the set of random vectors.

7. The apparatus of claim 6 , wherein:

the time-frequency representation is obtained by applying a time-to-frequency transform to each block of the audio signal in the set of blocks to obtain a respective set of spectral values; and

a respective intermediate value is derived by calculating an average intensity of the one or more spectral values in a group within the respective set of spectral values.

8. The apparatus of claim 6 , wherein the groups of spectral values have numbers of spectral values that vary with frequency.

9. The apparatus of claim 8 , wherein the groups of spectral values for higher frequencies have a greater number of spectral values.

10. The apparatus of claim 6 , wherein each component of the signature is derived from the projection of the intermediate values onto a respective random vector.

11. A non-transitory computer readable storage medium that records a program of instructions that is executable by a device to perform a method for generating a signature that identifies content of an audio signal, wherein the method comprises:

obtaining a time-frequency representation of a set of blocks within a sequence of blocks of the audio signal, wherein the time-frequency representation comprises sets of spectral values, each set of spectral values representing all spectral components within at least a portion of the bandwidth of the audio signal in a respective block in the set of blocks;

deriving intermediate values from intensities of all the spectral values arranged in groups of one or more spectral values within a respective set of spectral values; and

generating the signature that identifies content of the audio signal by projecting the intermediate values onto a set of random vectors, wherein the signature is represented by bits and each bit of the signature is derived from contributions from all of the intermediate values;

wherein each respective vector in the set of random vectors has vector elements with values that are obtained from a difference between uniformly distributed random variables within a range from zero to one and an average of the uniformly distributed random variables for all vector elements of the respective vector;

the projection of the intermediate values onto a respective random vector is obtained from an inner product of the intermediate values with the vector elements of the respective vector; and

each component of the signature has either a first value when the projection of the intermediate values onto a corresponding random vector is greater than a threshold or has a second value when the projection of the intermediate values onto the corresponding random vector is less than a threshold, wherein the threshold is equal to a median of the projections of intermediate values onto the set of random vectors.

12. The storage non-transitory computer readable storage medium of claim 11 , wherein:

the time-frequency representation is obtained by applying a time-to-frequency transform to each block of the audio signal in the set of blocks to obtain a respective set of spectral values; and

a respective intermediate value is derived by calculating an average intensity of the one or more spectral values in a group within the respective set of spectral values.

13. The non-transitory computer readable storage medium of claim 11 , wherein the groups of spectral values have numbers of spectral values that vary with frequency.

14. The non-transitory computer readable medium of claim 13 , wherein the groups of spectral values for higher frequencies have a greater number of spectral values.

15. The non-transitory computer readable storage medium of claim 11 , wherein each component of the signature is derived from the projection of the intermediate values onto a respective random vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2012
From: RADHAKRISHNAN, REGUNATHAN; BAUER, CLAUS; TERRY, KENT B; LINK, BRIAN; GSELL, ERIC; KIM, HYUNH-SUK
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 029092/0629 →
Continuity (3)
Division 12312840
Provisional Application 60872090 · Nov 30, 2006
Related Publication 20130064416A1 · Mar 14, 2013