IP Library Granted Patent US 9,093,120
Granted Patent B2
US 9,093,120 · App. 13/025,060 · Granted Jul 28, 2015

Audio fingerprint extraction by scaling in time and resampling

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,093,120
App. No.
13/025,060
Granted
Jul 28, 2015
Kind
B2
Abstract

An audio fingerprint is extracted from an audio sample, where the fingerprint contains information that is characteristic of the content in the sample. The fingerprint may be generated by computing an energy spectrum for the audio sample, resampling the energy spectrum, transforming the resampled energy spectrum to produce a series of feature vectors, and computing the fingerprint using differential coding of the feature vectors. The generated fingerprint can be compared to a set of reference fingerprints in a database to identify the original audio content.

Claims (38)

1. A method for extracting an audio fingerprint from an audio frame, the method comprising:

filtering, by a processor, the audio frame into a plurality of frequency bands by varying the frequency bands with time to produce a corresponding plurality of filtered audio signals;

scaling in time, by the processor, for fingerprinting, a size of each filtered audio signal based on a frequency of the corresponding frequency band, the scaling based on at least one of a mid-frequency and a frequency range of the corresponding frequency band;

resampling, by the processor, the scaled filtered audio signals by sampling from different corresponding frequency bands as time changes to produce resampled audio signals;

transforming, by the processor, the resampled audio signals to produce a feature vector for each resampled audio signal; and

computing, by the processor, the audio fingerprint based on the feature vectors.

2. The method of claim 1 , wherein transforming a resampled audio signal comprises performing a transform along a time axis.

3. The method of claim 1 , wherein transforming a resampled audio signal comprises:

transforming the resampled audio signal along a time axis; and

transforming the resampled audio signal along a frequency axis.

4. The method of claim 3 , further comprising, by the processor, scaling in magnitude, windowing, and normalizing the resampled audio signal.

5. The method of claim 1 , wherein transforming a resampled audio signal comprises performing a Two Dimensional Discrete Cosine Transform (2D DCT).

6. A computing device, comprising:

a processor;

a storage medium for tangibly storing thereon program logic for execution by the processor, the program logic comprising:

filtering logic executed by the processor for filtering an audio frame into a plurality of frequency bands by varying the frequency bands with time to produce a corresponding plurality of filtered audio signals;

scaling logic executed by the processor for scaling in time, for fingerprinting, a size of each filtered audio signal based on a frequency of the corresponding frequency band, the scaling based on at least one of a mid-frequency and a frequency range of the corresponding frequency band;

resampling logic executed by the processor for resampling the scaled filtered audio signals by sampling from different corresponding frequency bands as time changes to produce resampled audio signals;

transforming logic executed by the processor for transforming the resampled audio signals to produce a feature vector for each resampled audio signal; and

computing logic executed by the processor for computing an audio fingerprint based on the feature vectors.

7. The computing device of claim 6 , wherein the transforming logic for transforming a resampled audio signal comprises performing logic for performing a transform along a time axis.

8. The computing device of claim 6 , wherein the transforming logic for transforming a resampled audio signal comprises:

second transforming logic for transforming the resampled audio signal along a time axis; and

third transforming logic for transforming the resampled audio signal along a frequency axis.

9. The computing device of claim 8 , further comprising second scaling logic executed by the processor for scaling in magnitude, windowing, and normalizing the resampled audio signal.

10. The computing device of claim 6 , wherein the transforming logic for transforming a resampled audio signal comprises performing logic for performing a Two Dimensional Discrete Cosine Transform (2D DCT).

11. A non-transitory computer readable storage medium for tangibly storing computer program instructions capable of being executed by a computer processor, the computer program instructions defining the steps of:

filtering, by the computer processor, an audio frame into a plurality of frequency bands by varying the frequency bands with time to produce a corresponding plurality of filtered audio signals;

scaling, by the computer processor, in time, for fingerprinting, a size of each filtered audio signal based on a frequency of the corresponding frequency band, the scaling based on at least one of a mid-frequency and a frequency range of the corresponding frequency band;

resampling, by the computer processor, the scaled filtered audio signals by sampling from different corresponding frequency bands as time changes to produce resampled audio signals;

transforming, by the computer processor, the resampled audio signals to produce a feature vector for each resampled audio signal; and

computing, by the computer processor, an audio fingerprint based on the feature vectors.

12. The non-transitory computer readable storage medium of claim 11 , wherein transforming a resampled audio signal comprises performing a transform along a time axis.

13. The non-transitory computer readable storage medium of claim 11 , wherein transforming a resampled audio signal comprises:

transforming the resampled audio signal along a time axis; and

transforming the resampled audio signal along a frequency axis.

14. The non-transitory computer readable storage medium of claim 13 , further comprising, by the computer processor, scaling in magnitude, windowing, and normalizing the resampled audio signal.

15. The non-transitory computer readable storage medium of claim 11 , wherein transforming a resampled audio signal comprises performing, by the computer processor, a Two Dimensional Discrete Cosine Transform (2D DCT).

Assignments (12)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE ASSIGNOR NAME PREVIOUSLY RECORDED AT REEL: 052853 FRAME: 0153. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 29, 2021
From: R2 SOLUTIONS LLC
To: STARBOARD VALUE INTERMEDIATE FUND LP, AS COLLATERAL AGENT
Reel/Frame 056832/0001 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED ON REEL 053654 FRAME 0254. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST GRANTED PURSUANT TO THE PATENT SECURITY AGREEMENT PREVIOUSLY RECORDED. Recorded Dec 30, 2020
From: STARBOARD VALUE INTERMEDIATE FUND LP
To: R2 SOLUTIONS LLC
Reel/Frame 054981/0377 →
RELEASE OF SECURITY INTEREST IN PATENTS Recorded Jul 8, 2020
From: STARBOARD VALUE INTERMEDIATE FUND LP
To: ACACIA RESEARCH GROUP LLC; AMERICAN VEHICULAR SCIENCES LLC; BONUTTI SKELETAL INNOVATIONS LLC; CELLULAR COMMUNICATIONS EQUIPMENT LLC; INNOVATIVE DISPLAY TECHNOLOGIES LLC; LIFEPORT SCIENCES LLC; LIMESTONE MEMORY SYSTEMS LLC; MOBILE ENHANCEMENT SOLUTIONS LLC; MONARCH NETWORKING SOLUTIONS LLC; NEXUS DISPLAY TECHNOLOGIES LLC; PARTHENON UNIFIED MEMORY ARCHITECTURE LLC; R2 SOLUTIONS LLC; SAINT LAWRENCE COMMUNICATIONS LLC; STINGRAY IP SOLUTIONS LLC; SUPER INTERCONNECT TECHNOLOGIES LLC; TELECONFERENCE SYSTEMS LLC; UNIFICATION TECHNOLOGIES LLC
Reel/Frame 053654/0254 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 25, 2020
From: EXCALIBUR IP, LLC
To: R2 SOLUTIONS LLC
Reel/Frame 053459/0059 →
PATENT SECURITY AGREEMENT Recorded Jun 5, 2020
From: ACACIA RESEARCH GROUP LLC; AMERICAN VEHICULAR SCIENCES LLC; BONUTTI SKELETAL INNOVATIONS LLC; CELLULAR COMMUNICATIONS EQUIPMENT LLC; INNOVATIVE DISPLAY TECHNOLOGIES LLC; LIFEPORT SCIENCES LLC; LIMESTONE MEMORY SYSTEMS LLC; MERTON ACQUISITION HOLDCO LLC; MOBILE ENHANCEMENT SOLUTIONS LLC; MONARCH NETWORKING SOLUTIONS LLC; NEXUS DISPLAY TECHNOLOGIES LLC; PARTHENON UNIFIED MEMORY ARCHITECTURE LLC; R2 SOLUTIONS LLC; SAINT LAWRENCE COMMUNICATIONS LLC; STINGRAY IP SOLUTIONS LLC; SUPER INTERCONNECT TECHNOLOGIES LLC; TELECONFERENCE SYSTEMS LLC; UNIFICATION TECHNOLOGIES LLC
To: STARBOARD VALUE INTERMEDIATE FUND LP, AS COLLATERAL AGENT
Reel/Frame 052853/0153 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2016
From: YAHOO! INC.
To: EXCALIBUR IP, LLC
Reel/Frame 038950/0592 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2016
From: EXCALIBUR IP, LLC
To: YAHOO! INC.
Reel/Frame 038951/0295 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2016
From: YAHOO! INC.
To: EXCALIBUR IP, LLC
Reel/Frame 038383/0466 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADDRESS OF ASSIGNEE PREVIOUSLY RECORDED ON REEL 026448 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded Jun 22, 2011
From: INTONOW, INC.
To: YAHOO! INC.
Reel/Frame 026467/0423 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2011
From: INTONOW, INC.
To: YAHOO! INC.
Reel/Frame 026448/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 10, 2011
From: INTONOW, INC.
To: YAHOO! INC.
Reel/Frame 026424/0933 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2011
From: BILOBROV, SERGIY
To: INTONOW, INC.
Reel/Frame 025798/0934 →