IP Library › Granted Patent US 12,361,955
Granted Patent B2
US 12,361,955 · App. 18/500,764 · Granted Jul 15, 2025

Audio fingerprinting

Inventors: Jinyu Han (Emeryville, CA); Robert Coover (Orinda, CA)
Assignee: Gracenote, Inc.
G10L19/018
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,955
App. No.
18/500,764
Granted
Jul 15, 2025
Kind
B2
Abstract

A machine may be configured to generate one or more audio fingerprints of one or more segments of audio data. The machine may access audio data to be fingerprinted and divide the audio data into segments. For any given segment, the machine may generate a spectral representation from the segment; generate a vector from the spectral representation; generate an ordered set of permutations of the vector; generate an ordered set of numbers from the permutations of the vector; and generate a fingerprint of the segment of the audio data, which may be considered a sub-fingerprint of the audio data. In addition, the machine or a separate device may be configured to determine a likelihood that candidate audio data matches reference audio data.

Claims (40)

1. A non-transitory computer-readable medium having instructions stored thereon, wherein the instructions, when executed by one or more processors, cause the one or more processors to perform a set of operations comprising:

determining a first and second group of frequencies, wherein the first group of frequencies includes frequencies that are different than frequencies of the second group of frequencies;

identifying a first subgroup of frequencies in the first group of frequencies and a second subgroup of frequencies in the second group of frequencies, wherein the first subgroup is identified based on energy values of the first group and the frequencies of the first subgroup comprise energy values that are greater than energy values of other frequencies in the first group, and wherein the second subgroup is identified based on energy values of the first group or the second group and the frequencies of the second subgroup comprise energy values that are greater than energy values of other frequencies in the second group;

generating a vector that assigns a first value to frequencies in the first subgroup and a second value to frequencies in the second subgroup;

generating an ordered set of permutations of the vector;

generating a sequence that indicates an instance of the first value or of the second value within a permutation of the ordered set of permutations; and

generating a fingerprint of audio data based on the sequence.

2. The non-transitory computer readable medium of claim 1 , wherein at least one of the first and second group of frequencies is based on spectral data derived from the audio data.

3. The non-transitory computer readable medium of claim 1 , wherein the first group of frequencies includes frequencies that are higher than frequencies of the second group of frequencies.

4. The non-transitory computer readable medium of claim 1 , wherein at least one of the first subgroup of frequencies and the second subgroup of frequencies is identified based on ranked energy values for at least one of the first group of frequencies and the second subgroup of frequencies.

5. The non-transitory computer readable medium of claim 1 , wherein the first subgroup of frequencies is identified based on ranked energy values for the first group of frequencies, and wherein the second subgroup of frequencies is identified based on ranked energy values for the second group of frequencies.

6. The non-transitory computer readable medium of claim 1 , wherein the ordered set of permutations is based on arranged instances of the first and second values.

7. The non-transitory computer readable medium of claim 1 , wherein the sequence is generated based on generating numbers by calculating a remainder from an operation performed on a numerical representation of a lowest relative position occupied by at least one value in the ordered set of permutations.

8. The non-transitory computer readable medium of claim 1 , wherein the generated fingerprint comprises storing the sequence with a timestamp that indicates the audio data.

9. The non-transitory computer readable medium of claim 1 , wherein the generated fingerprint comprises storing at least one portion of the sequence in a hash table corresponding to a timestamp that indicates the audio data.

10. The non-transitory computer readable medium of claim 1 , wherein the ordered set of permutations is ordered based on a position of a lowest frequency value.

11. The non-transitory computer readable medium of claim 1 , wherein the ordered set of permutations is generated based on performing a modulo operation.

12. The non-transitory computer readable medium of claim 11 , wherein the modulo operation is performed based on a position of a lowest frequency with a non-zero value.

13. A computer-implemented method comprising:

determining a first and second group of frequencies, wherein the first group of frequencies includes frequencies that are different than frequencies of the second group of frequencies;

identifying a first subgroup of frequencies in the first group of frequencies and a second subgroup of frequencies in the second group of frequencies, wherein the first subgroup is identified based on energy values of the first group and the frequencies of the first subgroup comprise energy values that are greater than energy values of other frequencies in the first group, and wherein the second subgroup is identified based on energy values of the first group or the second group and the frequencies of the second subgroup comprise energy values that are greater than energy values of other frequencies in the second group;

generating a vector that assigns a first value to frequencies in the first subgroup and a second value to frequencies in the second subgroup;

generating an ordered set of permutations of the vector;

generating a sequence that indicates an instance of the first value or of the second value within a permutation of the ordered set of permutations; and

generating a fingerprint of audio data based on the sequence.

14. The computer-implemented method of claim 13 , wherein at least one of the first and second group of frequencies is based on spectral data derived from the audio data.

15. The computer-implemented method of claim 13 , wherein the first group of frequencies includes frequencies that are higher than frequencies of the second group of frequencies.

16. The computer-implemented method of claim 13 , wherein at least one of the first subgroup of frequencies and the second subgroup of frequencies is identified based on ranked energy values for at least one of the first group of frequencies and the second subgroup of frequencies.

17. The computer-implemented method of claim 13 , wherein the first subgroup of frequencies is identified based on ranked energy values for the first group of frequencies, and wherein the second subgroup of frequencies is identified based on ranked energy values for the second group of frequencies.

18. The computer-implemented method of claim 13 , wherein the sequence is generated based on generating numbers by calculating a remainder from an operation performed on a numerical representation of a lowest relative position occupied by at least one value in the ordered set of permutations.

19. The computer-implemented method of claim 13 , wherein the generated fingerprint comprises storing the sequence with a timestamp that indicates the audio data.

20. A computing device comprising:

one or more processors; and

a non-transitory, computer-readable medium having instructions stored thereon, wherein the instructions, when executed by one or more processors, cause the one or more processors to perform a set of operations comprising:

determining a first and second group of frequencies, wherein the first group of frequencies includes frequencies that are different than frequencies of the second group of frequencies;

identifying a first subgroup of frequencies in the first group of frequencies and a second subgroup of frequencies in the second group of frequencies, wherein the first subgroup is identified based on energy values of the first group and the frequencies of the first subgroup comprise energy values that are greater than energy values of other frequencies in the first group, and wherein the second subgroup is identified based on energy values of the first group or the second group and the frequencies of the second subgroup comprise energy values that are greater than energy values of other frequencies in the second group;

generating a vector that assigns a first value to frequencies in the first subgroup and a second value to frequencies in the second subgroup;

generating an ordered set of permutations of the vector;

generating a sequence that indicates an instance of the first value or of the second value within a permutation of the ordered set of permutations; and

generating a fingerprint of audio data based on the sequence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 27, 2023
From: HAN, JINYU; COOVER, ROBERT
To: GRACENOTE, INC.
Reel/Frame 065671/0620 →
Continuity (6)
Continuation 18049882 · Oct 26, 2022
Continuation 16926286 · Jul 10, 2020
Continuation 16270113 · Feb 7, 2019
Continuation 15008042 · Jan 27, 2016
Continuation 14107923 · Dec 16, 2013
Related Publication 20240071397A1 · Feb 29, 2024
References Cited (44)
US 7082106B2 · Sharma et al. · 2006 [cited by applicant]
US 8158870B2 · Lyon et al. · 2012 [cited by applicant]
US 8165414B1 · Yagnik · 2012 [cited by applicant]
US 8411977B1 · Baluja et al. · 2013 [cited by applicant]
US 8447032B1 · Covell et al. · 2013 [cited by applicant]
US 8660296B1 · Ioffe · 2014 [cited by applicant]
US 9143784B2 · Yagnik et al. · 2015 [cited by applicant]
US 9158842B1 · Yagnik et al. · 2015 [cited by applicant]
US 9286902B2 · Han et al. · 2016 [cited by applicant]
US 9589283B2 · Kim et al. · 2017 [cited by applicant]
US 9684715B1 · Ross · 2017 [cited by examiner]
US 10229689B2 · Han et al. · 2019 [cited by applicant]
US 10714105B2 · Han et al. · 2020 [cited by applicant]
US 20020083060A1 · Wang · 2002 [cited by examiner]
US 20100017195A1 · Villemoes · 2010 [cited by applicant]
US 20100067710A1 · Hendriks et al. · 2010 [cited by applicant]
US 20100257129A1 · Lyon et al. · 2010 [cited by applicant]
US 20110075851A1 · LeBoeuf et al. · 2011 [cited by applicant]
US 20110191577A1 · Tian et al. · 2011 [cited by applicant]
US 20120065966A1 · Wang · 2012 [cited by applicant]
US 20140280265A1 · Wang · 2014 [cited by applicant]
US 20140310006A1 · Anguera Miro et al. · 2014 [cited by applicant]
US 20150170660A1 · Han et al. · 2015 [cited by applicant]
US 20160217799A1 · Han et al. · 2016 [cited by applicant]
US 20190244624A1 · Han et al. · 2019 [cited by applicant]
WO 0211123 · 2002 [cited by applicant]
Kiao, “A Study on Music Retrieval based on Audio Fingerprinting,” Department of Information Science and Intelligent Systems: Graduate School of Advanced Technology and Science: The University of Tokushima, Sep. 2013, 79… [cited by applicant]
Haitsma et al., “A Highly Robust Audio Fingerprinting System”, Ismir (vol. 2002, pp. 107-115). Oct. 13, 2002, 9 pages. [cited by applicant]
Covell et al., “Waveprint: Efficient wavelet-based audio fingerprinting,” Pattern recognition 41.11. Nov. 30, k! 2008, 4 pages. [cited by applicant]
Zhang et al., “Heuristic Approach for Generic Audio Data Segmentation and Annotation,” Proceedings of the seventh ACM international conference on Multimedia {Part 1). ACM, 1999. Oct. 30, 1999, 10 pages. [cited by applicant]
Baluja et al. “Audio Fingerprinting: Combining Computer Vision & Data Stream Processing,” Acoustics, Speech and Signal Processing, 2007. ICASSP 2007, IEEE International Conference on vol. 2., Apr. 15, · 007, 5 pages. [cited by applicant]
Lu et al., “A Robust Audio Classification and Segmentation Method,” Proceedings of the ninth ACM international conference on Multimedia, ACM, Oct. 1, 2001, 9 pages. [cited by applicant]
Haitsma, “Philips Audio Fingerprinting Technology,” Philips Research Presentation—Philips Audio Fingerprinting Technology, Sep. 26, 2002, 16 pages. [cited by applicant]
Kim et al., “Audio Classification Based on MPEG-7 Spectral Basis Representations,” IEEE Transactions on Circuits and Systems for Video Technology 14 .5 (2004 ), May 1, 2004. 10 pages. [cited by applicant]
Lu et al., “Content-Based Audio Segmentation Using Support Vector Machines,” Proceedings of the IEEE International Conference on Multimedia and Expo (ICME 2001 ), Aug. 22, 2001, 4 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowabilily,” issued in connection with U.S. Appl. No. 15/008,042, on Feb. 15, 2019, 2 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 15/008,042, on Oct. 12, 2018, 8 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office action,” issued in connection with U.S. Appl. No. 15/008,042, on Feb. 22, 2018, 16 pages. [cited by applicant]
United States Patent and Trademark Office, “Restriction Requirement,” issued in connection with U.S. Appl. No. 15/008,042, on Sep. 8, 2017, 6 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 14/107,923, on Dec. 18, 2015, 18 pages. [cited by applicant]
United States Patent and Trademark Office, “Restriction Requirement,” issued in connection with U.S. Appl. No. 14/107,923, on Sep. 8, 2015, 6 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 16/270,113, mailed on Oct. 30, 2019, 50 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 16/270,113, mailed on Mar. 11, 2020, 54 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowability,” issued in connection with U.S. Appl. No. 16/270,113, mailed on May 6, 2020, 5 pages. [cited by applicant]