IP Library Granted Patent US 9,299,364
Granted Patent B1
US 9,299,364 · App. 13/647,996 · Granted Mar 29, 2016

Audio content fingerprinting based on two-dimensional constant Q-factor transform representation and robust audio identification for time-aligned applications

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,299,364
App. No.
13/647,996
Granted
Mar 29, 2016
Kind
B1
Abstract

Content identification methods for consumer devices determine robust audio fingerprints that are resilient to audio distortions. One method generates signatures representing audio content based on a constant Q-factor transform (CQT). A 2D spectral representation of a 1D audio signal facilitates generation of region based signatures within frequency octaves and across the entire 2D signal representation. Also, points of interest are detected within the 2D audio signal representation and interest regions are determined around selected points of interest. Another method generates audio descriptors using an accumulating filter function on bands of the audio spectrum and generates audio transform coefficients. A response of each spectral band is computed and transform coefficients are determined by filtering, by accumulating derivatives with different lags, and computing second order derivatives. Additionally, time and frequency based onset detection determines audio descriptors at events and enhances descriptors with information related to an event.

Claims (99)

1. A method for robust fingerprinting of audio signals in a processor, the method comprising:

structuring a received one dimensional (1D) audio signal into overlapping audio frames;

applying a constant Q-factor transform (CQT) to the overlapping audio frames to generate a two dimensional (2D) CQT data structure representation of the audio frames;

processing octaves of the 2D CQT data structure to determine regions of interest within the 2D CQT data structure and peak interest points within selected interest regions;

generating multidimensional descriptors in windows around the peak interest points;

applying a quantizer threshold to the multidimensional descriptors to generate audio signatures representing the received 1D audio signal; and

applying a compacting discrete cosine transform (DCT) on a set of regularly spaced regions generated across one or more octaves in 2D CQT array frequency axes directions, wherein the multidimensional descriptors are generated with a length based on a PxQ descriptor box and on generated DCT coefficients.

2. The method of claim 1 , wherein to determine regions of interest within the 2D CQT data structure representation of the audio frames, the 2D CQT data structure is tiled into spatial regions localized within the extent of each octave.

3. The method of claim 2 , wherein the tiled spatial regions are processed to determine local maxima for groups of coefficients belonging to the same tiled spatial region.

4. The method of claim 3 , wherein a collection of local maxima in each tiled region is sorted according to their magnitudes, and a local maxima having the strongest magnitude is selected in each tiled region as a selected interest point with an associated descriptor describing its spatial position and peak intensity.

5. The method of claim 1 further comprising:

generating a rectangular descriptor box of predetermined size for each of the peak interest points within the selected interest regions at their associated spatial (x, y) position; and

generating the multidimensional descriptors in the selected interest regions based on filtering by a set of filtering rules.

6. A method for robust fingerprinting of audio signals in a processor, the method comprising:

structuring a received one dimensional (1D) audio signal into overlapping audio frames;

applying a constant Q-factor transform (CQT) to the overlapping audio frames to generate a two dimensional (2D) CQT data structure representation of the audio frames;

processing octaves of the 2D CQT data structure to determine regions of interest within 2D CQT data structure and peak interest points within selected interest regions;

generating multidimensional descriptors in windows around the peak interest points;

applying a quantizer threshold to the multidimensional descriptors to generate audio signatures representing the received 1D audio signal;

generating a rectangular descriptor box of predetermined size for each of the peak interest points within the selected interest regions at their associated spatial (x, y) position; and

generating the multidimensional descriptors in the selected interest regions based on filtering by a set of filtering rules, wherein a filtering rule for generating the multidimensional descriptors in the selected interest regions based on filtering is to apply a gradient filter in raster fashion on each rectangular descriptor box to generate gradient filter transformed CQT coefficients.

7. A method for robust fingerprinting of audio signals in a processor, the method comprising:

structuring a received one dimensional (1D) audio signal into overlapping audio frames;

applying a constant Q-factor transform (CQT) to the overlapping audio frames to generate a two dimensional (2D) CQT data structure representation of the audio frames;

processing octaves of the 2D CQT data structure to determine regions of interest within the 2D CQT data structure and peak interest points within selected interest regions;

generating multidimensional descriptors in windows around the peak interest points;

applying a quantizer threshold to the multidimensional descriptors to generate audio signatures representing the received 1D audio signal;

generating a rectangular descriptor box of predetermined size for each of the peak interest points within the selected interest regions at their associated spatial (x, y) position; and

generating the multidimensional descriptors in the selected interest regions based on filtering by a set of filtering rules, wherein a filtering rule for generating the multidimensional descriptors in the selected interest regions based on filtering is to apply a 2D spatial filter in raster fashion on each rectangular descriptor box to generate 2D spatial filter transformed CQT coefficients.

8. A method for robust fingerprinting of audio signals in a processor, the method comprising:

structuring a received one dimensional (1D) audio signal into overlapping audio frames;

applying a constant Q-factor transform (CQT) to the overlapping audio frames to generate a two dimensional (2D) CQT data structure representation of the audio frames;

processing octaves of the 2D CQT data structure to determine regions of interest within the 2D CQT data structure and peak interest points within selected interest regions;

generating multidimensional descriptors in windows around the peak interest points;

applying a quantizer threshold to the multidimensional descriptors to generate audio signatures representing the received 1D audio signal;

generating a rectangular descriptor box of predetermined size for each of the peak interest points within the selected interest regions at their associated spatial (x, y) position; and

generating the multidimensional descriptors in the selected interest regions based on filtering by a set of filtering rules, wherein a filtering rule for generating the multidimensional descriptors in the selected interest regions based on filtering is to apply a 2D discrete cosine transform (DCT) on each rectangular descriptor box to generate 2D discrete cosine transformed CQT coefficients.

9. A method for robust fingerprinting of audio signals in a processor, the method comprising:

structuring a received one dimensional (1D) audio signal into overlapping audio frames;

applying a constant Q-factor transform (CQT) to the overlapping audio frames to generate a two dimensional (2D) CQT data structure representation of the audio frames;

processing octaves of the 2D CQT data structure to determine regions of interest within the 2D CQT data structure and peak interest points within selected interest regions;

generating multidimensional descriptors in windows around the peak interest points; and

applying a quantizer threshold to the multidimensional descriptors to generate audio signatures representing the received 1D audio signal, wherein to determine regions of interest within the 2D CQT data structure representation of the audio frames, the 2D CQT data structure is tiled into spatial regions localized within the extent of each octave, wherein the tiled spatial regions are processed to determine local maxima for groups of coefficients belonging to the same tiled spatial region, wherein a collection of local maxima in each tiled region is sorted according to their magnitudes, and a local maxima having the strongest magnitude is selected in each tiled region as a selected interest point with an associated descriptor describing its spatial position and peak intensity, and wherein a DCT transform is applied to each of the selected interest regions around each of the selected interest points to provide a new set of coefficients for each of the selected interest regions to be used for multidimensional descriptor generation.

10. The method of claim 9 , wherein a selected PxQ rectangular region of interest is generated around a selected interest point for 2D DCT computation, wherein Q represents a number of coefficients within an octave in a y frequency direction and P is chosen to form a square area symmetric around the interest point in an x time direction.

11. The method of claim 10 , wherein transformed coefficients, computed for each of the selected PxQ rectangular region of interest, are used to generate a multidimensional descriptor of length according to the selected PxQ rectangular region of interest.

12. The method of claim 1 , wherein the 2D CQT data structure is extended by interpolation to form a scale-space image with an identical number of coefficient vectors in each octave.

13. The method of claim 12 , wherein a rectangular array extended by interpolation of CQT coefficients is treated as a digital image frame and then methods for image feature extraction are used to determine the regions of interest.

14. A method to generate audio signatures in a processor, the method comprising:

receiving a one dimensional (1D) audio signal that has been filtered to reduce noise and organized into weighted frames;

applying a constant Q-factor transform (CQT) to the weighted frames to generate a two dimensional (2D) scale-space representation of the 1D audio signal;

processing the 2D scale-space representation to determine interest regions and peak interest points within selected interest regions;

generating multidimensional descriptors for the selected interest regions;

applying a quantizer to the multidimensional descriptors to generate a first set of audio signatures;

combining side information with the signatures of the first set of audio signatures to generate a processed set of audio signatures representing the received 1D audio signal;

generating a rectangular descriptor box of predetermined size for each of the peak interest points within the selected interest regions at their associated spatial (x, y) position; and

generating the multidimensional descriptors in the selected interest regions based on filtering by a set of filtering rules, wherein a filtering rule for generating the multidimensional descriptors in the selected interest regions based on filtering is to apply a gradient filter in raster fashion on each rectangular descriptor box to generate gradient filter transformed CQT coefficients.

15. The method of claim 14 , wherein the 2D CQT scale-space representation is processed by filtering to form a scale-space image which is then processed by image feature extraction methods to determine the interest regions.

16. A method for audio identification in a processor that is robust to interfering noise, the method comprising:

applying a filtering function to spectral components of a one dimensional (1D) audio signal to accumulate energy of spectral bands of the 1D audio signal;

generating audio query signatures from measurements of the accumulated energy of spectral bands;

correlating the audio query signatures and unique features with reference audio signatures and features;

calculating a frequency band coefficient filter response by accumulating delta values for a predetermined time window applied to the 1D audio signal; and

applying a lag factor to the predetermined time window for difference calculation for frequency band coefficients.

17. The method of claim 16 further comprising:

detecting a strongest response of the filter function to the spectral bands and transform coefficients to determine strong features of the 1D audio signal.

18. The method of claim 17 further comprising:

performing a fuzzy search for query signature using various combinations of the strong features detected from the spectral band.

19. The method of claim 16 further comprising:

adding to each audio signature one or more fields comprising an index of one or more of strongest detected features, wherein the index is a hash key that is used to access the reference audio signatures during search and correlation for audio content identification.

20. The method of claim 16 further comprising:

adding indexes and signatures for likely content based on popularity and real time broadcast at the first stage of search, wherein the additional indexes and signatures for likely content improves the accuracy of matching of the reference content.

21. The method of claim 16 further comprising:

promoting a likely match based on user profile where a likely user content identification is compared against initial signature candidates and likely content per a user profile is promoted so that it has a better chance of surviving a first stage of search and can be evaluated further for actual matching.

22. A computer readable non-transitory medium encoded with computer readable program data and code, the program data and code when executed operable to:

structure a received one dimensional (1D) audio signal into overlapping audio frames;

apply a constant Q-factor transform (CQT) to the overlapping audio frames to generate a two dimensional (2D) CQT data structure of the audio frames;

process octaves of the 2D CQT data structure to determine regions of interest within the 2D CQT data structure and peak interest points within selected interest regions;

generate multidimensional descriptors in windows around the peak interest points; and

apply a quantizer threshold to the multidimensional descriptors to generate audio signatures representing the receive 1D audio signal; and

applying a compacting discrete cosine transform (DCT) on a set of regularly spaced regions generated across one or more octaves in 2D CQT array frequency axes directions, wherein the multidimensional descriptors are generated with a length based on a PxQ descriptor box and on generated DCT coefficients.

23. The computer readable medium of claim 22 , wherein the 2D CQT data structure is processed by filtering to form a scale-space image which is then processed by image feature extraction methods to determine the selected interest regions.

24. A method to generate audio signatures in a processor, the method comprising:

receiving a one dimensional (1D) audio signal that has been filtered to reduce noise and organized into weighted frames;

applying a constant Q-factor transform (CQT) to the weighted frames to generate a two dimensional (2D) scale-space representation of the 1D audio signal;

processing the 2D scale-space representation to determine interest regions and peak interest points within selected interest regions;

generating multidimensional descriptors for the selected interest regions;

applying a quantizer to the multidimensional descriptors to generate a first set of audio signatures;

combining side information with the signatures of the first set of audio signatures to generate a processed set of audio signatures representing the received 1D audio signal;

generating a rectangular descriptor box of predetermined size for each of the peak interest points within the selected interest regions at their associated spatial (x, y) position; and

generating the multidimensional descriptors in the selected interest regions based on filtering by a set of filtering rules, wherein a filtering rule for generating the multidimensional descriptors in the selected interest regions based on filtering is to apply a 2D spatial filter in raster fashion on each rectangular descriptor box to generate 2D spatial filter transformed CQT coefficients.

25. A method to generate audio signatures in a processor, the method comprising:

receiving a one dimensional (1D) audio signal that has been filtered to reduce noise and organized into weighted frames;

applying a constant Q-factor transform (CQT) to the weighted frames to generate a two dimensional (2D) scale-space representation of the 1D audio signal;

processing the 2D scale-space representation to determine interest regions and peak interest points within selected interest regions;

generating multidimensional descriptors for the selected interest regions;

applying a quantizer to the multidimensional descriptors to generate a first set of audio signatures;

combining side information with the signatures of the first set of audio signatures to generate a processed set of audio signatures representing the received 1D audio signal;

generating a rectangular descriptor box of predetermined size for each of the peak interest points within the selected interest regions at their associated spatial (x, y) position; and

generating the multidimensional descriptors in the selected interest regions based on filtering by a set of filtering rules, wherein a filtering rule for generating the multidimensional descriptors in the selected interest regions based on filtering is to apply a 2D discrete cosine transform (DCT) on each rectangular descriptor box to generate 2D discrete cosine transformed CQT coefficients.

Assignments (16)
SECURITY INTEREST Recorded Sep 18, 2024
From: ROKU, INC.
To: CITIBANK, N.A.
Reel/Frame 068982/0377 →
RELEASE (REEL 053473 / FRAME 0001) Recorded May 11, 2023
From: CITIBANK, N.A.
To: A. C. NIELSEN COMPANY, LLC; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE MEDIA SERVICES, LLC; THE NIELSEN COMPANY (US), LLC; NETRATINGS, LLC
Reel/Frame 063603/0001 →
RELEASE (REEL 054066 / FRAME 0064) Recorded May 11, 2023
From: CITIBANK, N.A.
To: A. C. NIELSEN COMPANY, LLC; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE MEDIA SERVICES, LLC; THE NIELSEN COMPANY (US), LLC; NETRATINGS, LLC
Reel/Frame 063605/0001 →
TERMINATION AND RELEASE OF INTELLECTUAL PROPERTY SECURITY AGREEMENT (REEL/FRAME 056982/0194) Recorded Feb 22, 2023
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROKU, INC.; ROKU DX HOLDINGS, INC.
Reel/Frame 062826/0664 →
RELEASE (REEL 042262 / FRAME 0601) Recorded Oct 13, 2022
From: CITIBANK, N.A.
To: GRACENOTE, INC.; GRACENOTE DIGITAL VENTURES, LLC
Reel/Frame 061748/0001 →
PATENT SECURITY AGREEMENT SUPPLEMENT Recorded Jun 29, 2021
From: ROKU, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 056982/0194 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2021
From: GRACENOTE, INC.
To: ROKU, INC.
Reel/Frame 056103/0786 →
PARTIAL RELEASE OF SECURITY INTEREST Recorded Apr 20, 2021
From: CITIBANK, N.A.
To: THE NIELSEN COMPANY (US), LLC; GRACENOTE, INC.
Reel/Frame 056973/0280 →
CORRECTIVE ASSIGNMENT TO CORRECT THE PATENTS LISTED ON SCHEDULE 1 RECORDED ON 6-9-2020 PREVIOUSLY RECORDED ON REEL 053473 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE SUPPLEMENTAL IP SECURITY AGREEMENT. Recorded Oct 7, 2020
From: A.C. NIELSEN (ARGENTINA) S.A.; A.C. NIELSEN COMPANY, LLC; ACN HOLDINGS INC.; ACNIELSEN CORPORATION; ACNIELSEN ERATINGS.COM; AFFINNOVA, INC.; ART HOLDING, L.L.C.; ATHENIAN LEASING CORPORATION; CZT/ACN TRADEMARKS, L.L.C.; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; NETRATINGS, LLC; NIELSEN AUDIO, INC.; NIELSEN CONSUMER INSIGHTS, INC.; NIELSEN CONSUMER NEUROSCIENCE, INC.; NIELSEN FINANCE CO.; NIELSEN FINANCE LLC; NIELSEN INTERNATIONAL HOLDINGS, INC.; NIELSEN MOBILE, LLC; NMR INVESTING I, INC.; TCG DIVESTITURE INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC; VIZU CORPORATION; VNU MARKETING INFORMATION, INC.; NMR LICENSING ASSOCIATES, L.P.; NIELSEN HOLDING AND FINANCE B.V.; THE NIELSEN COMPANY B.V.; VNU INTERNATIONAL B.V.
To: CITIBANK, N.A
Reel/Frame 054066/0064 →
SUPPLEMENTAL SECURITY AGREEMENT Recorded Jun 9, 2020
From: A. C. NIELSEN COMPANY, LLC; ACN HOLDINGS INC.; ACNIELSEN CORPORATION; ACNIELSEN ERATINGS.COM; AFFINNOVA, INC.; ART HOLDING, L.L.C.; ATHENIAN LEASING CORPORATION; CZT/ACN TRADEMARKS, L.L.C.; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; NETRATINGS, LLC; NIELSEN AUDIO, INC.; NIELSEN CONSUMER INSIGHTS, INC.; NIELSEN CONSUMER NEUROSCIENCE, INC.; NIELSEN FINANCE CO.; NIELSEN FINANCE LLC; NIELSEN INTERNATIONAL HOLDINGS, INC.; NIELSEN MOBILE, LLC; NIELSEN UK FINANCE I, LLC; NMR INVESTING I, INC.; TCG DIVESTITURE INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC; VIZU CORPORATION; VNU MARKETING INFORMATION, INC.; NMR LICENSING ASSOCIATES, L.P.; NIELSEN HOLDING AND FINANCE B.V.; THE NIELSEN COMPANY B.V.; VNU INTERNATIONAL B.V.
To: CITIBANK, N.A.
Reel/Frame 053473/0001 →
SUPPLEMENTAL SECURITY AGREEMENT Recorded Apr 13, 2017
From: GRACENOTE, INC.; GRACENOTE MEDIA SERVICES, LLC; GRACENOTE DIGITAL VENTURES, LLC
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 042262/0601 →
RELEASE OF SECURITY INTEREST IN PATENT RIGHTS Recorded Feb 8, 2017
From: JPMORGAN CHASE BANK, N.A.
To: GRACENOTE, INC.; CASTTV INC.; TRIBUNE MEDIA SERVICES, LLC; TRIBUNE DIGITAL VENTURES, LLC
Reel/Frame 041656/0804 →
NOTICE AND CONFIRMATION OF GRANT OF SECURITY INTEREST IN PATENTS Recorded May 11, 2016
From: GRACENOTE, INC.; TRIBUNE BROADCASTING COMPANY, LLC; TRIBUNE MEDIA COMPANY
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 038679/0458 →
SECURITY AGREEMENT Recorded Aug 14, 2015
From: GRACENOTE, INC.; TRIBUNE BROADCASTING COMPANY, LLC; CASTTV INC.
To: JPMORGAN CHASE BANK, N.A., AS AGENT
Reel/Frame 036354/0793 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2015
From: ZEITERA, LLC
To: GRACENOTE, INC.
Reel/Frame 036027/0392 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2012
From: PEREIRA, JOSE PIO; STOJANCIC, MIHAILO M.; WENDT, PETER
To: ZEITERA, LLC
Reel/Frame 029099/0143 →