IP Library Granted Patent US 12,306,871
Granted Patent B2
US 12,306,871 · App. 18/165,107 · Granted May 20, 2025

Audio identification during performance

Inventors: Dale T. Roberts (San Anselmo, CA); Bob Coover (Orinda, CA); Nicola Marcantonio (San Francisco, CA); Markus K. Cremer (Orinda, CA)
Assignee: Gracenote, Inc.
G06F16/683G06F18/231G06Q50/01G10L25/18G10L25/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,306,871
App. No.
18/165,107
Granted
May 20, 2025
Kind
B2
Abstract

Methods and apparatus for audio identification during a performance are disclosed herein. An example apparatus includes at least one memory and at least one processor to transform a segment of audio into a log-frequency spectrogram based on a constant Q transform using a logarithmic frequency resolution, transform the log-frequency spectrogram into a binary image, each pixel of the binary image corresponding to a time frame and frequency channel pair, each frequency channel representing a corresponding quarter tone frequency channel in a range from C3-C8, generate a matrix product of the binary image and a plurality of reference fingerprints, normalize the matrix product to form a similarity matrix, select an alignment of a line in the similarity matrix that intersects one or more bins in the similarity matrix with the largest calculated Hamming similarities, and select a reference fingerprint based on the alignment.

Claims (38)

1. A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a computer system, cause the computer system to perform a set of operations comprising:

receiving, from a computing device at a live performance of an audio piece, a fingerprint of a segment of a live version of the audio piece, wherein the fingerprint contains a query for identification of the audio piece during the live performance of the live version of the audio piece;

computing a similarity matrix between at least one reference fingerprint and the fingerprint, wherein computing the similarity matrix comprises:

generating a binary image of a log-frequency spectrogram representing the fingerprint, wherein a plurality of pixels of the binary image correspond to a time frame and frequency channel pair, and wherein at least one frequency channel represents a corresponding quarter tone frequency channel in a range from musical note C3 to musical note C8; and

generating a matrix product of the binary image and a plurality of reference fingerprints including the at least one reference fingerprint; and

identifying the audio piece, wherein identifying the audio piece is based on a match between the at least one reference fingerprint and the fingerprint, wherein the match is based on determining a threshold similarity between the at least one reference fingerprint and the fingerprint, and wherein determining the threshold similarity between the at least one reference fingerprint and the fingerprint is based on the similarity matrix.

2. The non-transitory machine-readable storage medium of claim 1 , wherein the set of operations further comprises, during the live performance of the live version of the audio piece, determining an identifier of the audio piece and based on the determined identifier, comparing the fingerprint to a set of reference fingerprints, wherein the set of reference fingerprints include the at least one reference fingerprint, and wherein the set of reference fingerprints is accessed based on the determined identifier, and wherein the determined identifier of the audio pieces comprises a name of a performer that is performing the audio piece during the live performance of the live version of the audio piece.

3. The non-transitory machine-readable storage medium of claim 2 , wherein, during the live performance of the live version of the audio piece, determining the identifier of the audio piece comprises identifying a venue at which the live performance of the live version of the audio piece is being performed and retrieving information associated with the venue related to the live performance.

4. The non-transitory machine-readable storage medium of claim 2 , wherein, during the live performance of the live version of the audio piece, determining the identifier of the audio piece comprises receiving, from a plurality of attendees of the live performance of the live version of the audio piece, information associated with the audio piece during the live performance.

5. The non-transitory machine-readable storage medium of claim 4 , wherein the information comprises one or more of: (i) a performer of the audio piece; (ii) a name of the audio piece; and (iii) an album title associated with the audio piece.

6. The non-transitory machine-readable storage medium of claim 4 , wherein receiving, from a plurality of attendees of the live performance of the live version of the audio piece, information associated with the audio piece during the live performance comprises receiving votes related to the information associated with the audio piece during the live performance, and wherein identifying the audio piece further comprises providing, to the computing device, an identification of a result of the received votes.

7. The non-transitory machine-readable storage medium of claim 1 , wherein identifying the audio piece comprises identifying a name of the audio piece.

8. The non-transitory machine-readable storage medium of claim 1 , wherein identifying the audio piece comprises identifying an album title associated with the audio piece.

9. The non-transitory machine-readable storage medium of claim 2 , wherein the set of reference fingerprints comprises reference fingerprints include fingerprints generated from segments of reference versions of audio pieces associated with the determined identifier.

10. The non-transitory machine-readable storage medium of claim 1 , wherein determining a threshold similarity between the at least one reference fingerprint and the fingerprint further comprises determining a threshold similarity between one or more core characteristics of the at least one reference fingerprint and the fingerprint.

11. The non-transitory machine-readable storage medium of claim 10 , wherein the one or more core characteristics comprise one or more of: (i) notes; or (ii) rhythms.

12. The non-transitory machine-readable storage medium of claim 10 , wherein the one or more core characteristics comprise one or more of: (i) tempo; (ii) vocal timber; (iii) vocal strength; (iv) vibrato; (v) instrument tuning; (vi) ambient noise; (vii) reverberation; or (viii) distortions.

13. The non-transitory machine-readable storage medium of claim 1 , wherein computing the similarity matrix further comprises:

normalizing the matrix product to form the similarity matrix.

14. A method comprising:

receiving, from a computing device at a live performance of an audio piece, a fingerprint of a segment of a live version of the audio piece, wherein the fingerprint contains a query for identification of the audio piece during the live performance of the live version of the audio piece;

computing a similarity matrix between at least one reference fingerprint of a reference version of the audio piece and the fingerprint, wherein computing the similarity matrix comprises:

generating a binary image of a log-frequency spectrogram representing the fingerprint, wherein a plurality of pixels of the binary image correspond to a time frame and frequency channel pair, and wherein at least one frequency channel represents a corresponding quarter tone frequency channel in a range from musical note C3 to musical note C8; and

generating a matrix product of the binary image and a plurality of reference fingerprints including the at least one reference fingerprint; and

identifying the audio piece, wherein identifying the audio piece is based on a match between the at least one reference fingerprint and the fingerprint, wherein the match is based on determining a threshold similarity between the at least one reference fingerprint and the fingerprint, and wherein determining the threshold similarity between the at least one reference fingerprint and the fingerprint is based on the similarity matrix.

15. The method of claim 14 , wherein the method further comprises, during the live performance of the live version of the audio piece, determining an identifier of the audio piece and based on the determined identifier, comparing the fingerprint to a set of reference fingerprints, wherein the set of reference fingerprints include the at least one reference fingerprint, and wherein the set of reference fingerprints is accessed based on the determined identifier, and wherein the determined identifier of the audio pieces comprises a name of a performer that is performing the audio piece during the live performance of the live version of the audio piece.

16. The method of claim 15 , wherein, during the live performance of the live version of the audio piece, determining the identifier of the audio piece comprises identifying a venue at which the live performance of the live version of the audio piece is being performed and retrieving information associated with the venue related to the live performance.

17. The method of claim 15 , wherein determining a threshold similarity between the at least one reference fingerprint of the reference version of the audio piece and the fingerprint further comprises determining a threshold similarity between one or more core characteristics of the at least one reference fingerprint of the reference version of the audio piece and the fingerprint.

18. The method of claim 17 , wherein the one or more core characteristics comprise one or more of: (i) notes; or (ii) rhythms.

19. The method of claim 17 , wherein the one or more core characteristics comprise one or more of: (i) tempo; (ii) vocal timber; (iii) vocal strength; (iv) vibrato; (v) instrument tuning; (vi) ambient noise; (vii) reverberation; or (viii) distortions.

20. A computing device comprising:

one or more processors; and

a non-transitory, computer-readable medium storing instructions that, when executed by the one or more processors, cause the computing device to perform a set of acts comprising:

receiving, from a computing device at a live performance of an audio piece, a fingerprint of a segment of a live version of the audio piece, wherein the fingerprint contains a query for identification of the audio piece during the live performance of the live version of the audio piece;

computing a similarity matrix between at least one reference fingerprint of a reference version of the audio piece and the fingerprint, wherein computing the similarity matrix comprises:

generating a binary image of a log-frequency spectrogram representing the fingerprint, wherein a plurality of pixels of the binary image correspond to a time frame and frequency channel pair, and wherein at least one frequency channel represents a corresponding quarter tone frequency channel in a range from musical note C3 to musical note C8; and

generating a matrix product of the binary image and a plurality of reference fingerprints including the at least one reference fingerprint; and

identifying the audio piece, wherein identifying the audio piece is based on a match between the at least one reference fingerprint and the fingerprint, wherein the match is based on determining a threshold similarity between the at least one reference fingerprint and the fingerprint, and wherein determining the threshold similarity between the at least one reference fingerprint and the fingerprint is based on the similarity matrix.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2025
From: ROBERTS, DALE T.; COOVER, BOB; MARCANTONIO, NICOLA; CREMER, MARKUS K.
To: GRACENOTE, INC.
Reel/Frame 070782/0490 →
Continuity (4)
Continuation 17102012 · Nov 23, 2020
Continuation 15888998 · Feb 5, 2018
Continuation 14258263 · Apr 22, 2014
Related Publication 20230185847A1 · Jun 15, 2023
References Cited (67)
US 5040081A · McCutcllen · 1991 [cited by applicant]
US 6931134B1 · Waller, Jr. et al. · 2005 [cited by applicant]
US 7164076B2 · McHale et al. · 2007 [cited by applicant]
US 7302574B2 · Conwell et al. · 2007 [cited by applicant]
US 8180063B2 · Henderson · 2012 [cited by applicant]
US 9153239B1 · Postelnicu · 2015 [cited by examiner]
US 10846334B2 · Roberts et al. · 2020 [cited by applicant]
US 20040199387A1 · Wang et al. · 2004 [cited by applicant]
US 20060074679A1 · Pifer et al. · 2006 [cited by applicant]
US 20070188657A1 · Basson et al. · 2007 [cited by applicant]
US 20080013614A1 · Fiesel et al. · 2008 [cited by applicant]
US 20080065699A1 · Bioebaum et al. · 2008 [cited by applicant]
US 20080082510A1 · Wang et al. · 2008 [cited by applicant]
US 20080133556A1 · Conwell · 2008 [cited by examiner]
US 20080320078A1 · Feldman et al. · 2008 [cited by applicant]
US 20110085781A1 · Olson · 2011 [cited by applicant]
US 20110112913A1 · Murray · 2011 [cited by applicant]
US 20110202687A1 · Glitsch et al. · 2011 [cited by applicant]
US 20110273455A1 · Powar et al. · 2011 [cited by applicant]
US 20110276333A1 · Wang et al. · 2011 [cited by applicant]
US 20110289530A1 · Dureau et al. · 2011 [cited by applicant]
US 20120059826A1 · Mate et al. · 2012 [cited by applicant]
US 20120124638A1 · King et al. · 2012 [cited by applicant]
US 20120209612A1 · Bilobrov · 2012 [cited by applicant]
US 20120210233A1 · Davis et al. · 2012 [cited by applicant]
US 20130007201A1 · Jeffrey et al. · 2013 [cited by applicant]
US 20130160038A1 · Slaney et al. · 2013 [cited by applicant]
US 20130339877A1 · Skeen et al. · 2013 [cited by applicant]
US 20140082651A1 · Sharifi · 2014 [cited by applicant]
US 20140129571A1 · Scavo et al. · 2014 [cited by applicant]
US 20140169768A1 · Webb et al. · 2014 [cited by applicant]
US 20140254820A1 · Gardenfors et al. · 2014 [cited by applicant]
US 20140324616A1 · Proiettie et al. · 2014 [cited by applicant]
US 20150016661A1 · Lord · 2015 [cited by applicant]
US 20150104023A1 · Bilobrov · 2015 [cited by examiner]
US 20150193701A1 · Sohn et al. · 2015 [cited by applicant]
US 20150199974A1 · Bilobrov et al. · 2015 [cited by applicant]
US 20150206544A1 · Carter · 2015 [cited by applicant]
US 20150302086A1 · Roberts et al. · 2015 [cited by applicant]
WO 0227600 · 2002 [cited by applicant]
WO 03091899 · 2003 [cited by applicant]
Bisio, et al. “Fast audio fingerprint comparison for real-time TV-channel recognition applications”, published by Wireless Communications and Mobile Computing Conference (IWCMC), 9th International. IEE, 2013, 6 pages. [cited by applicant]
Camarena-Ibarrola, et al., “Identifying music by performances using an entropy based audio-fingerprint”, published by Mexican International Conference on Artificial Intelligence (MICAI), 2006, 11 pages. [cited by applicant]
Camarena-Ibarrola, et al., “Real time tracking of musical performances”, published by Advances in Soft Computing, 2010, 11 pages. [cited by applicant]
Collins, “Ubiquitous Electronics: Technology and Live Performance 1966-1996”, published by Leonardo Music Journal, 1998, 6 pages. [cited by applicant]
Dixon, “Live tracking of musical performances using on-line time warping”, proceedings of the 8th international Conference on Digital Audio Effects, 2005, 6 pages. [cited by applicant]
Fenet, et al., “A Scalable Audio Fingerprint Method with Robustness to Pitch-Shifting”, ISMIR, 2001, 6 pages. [cited by applicant]
Grosche, et al., “Audio Content-Based Music Retrieval”, Dagstuhl Follow-Ups, vol. 3, 2012, 18 pages. [cited by applicant]
Guzman, et al., “On the Use of Locality Sensitive Hashing for Audio Foliowing”, Iberoamerican Congress on Pattern Recognition, Springer International Publishing, 2014, 2 pages. [cited by applicant]
Kennedy, et al., “Less talk, more rock: automated organization of community-contributed collections of concert videos”, proceedings of the 18th International Conference on World Wide Web, ACM, 2009, 10 pages. [cited by applicant]
Miotto, et al., “Automatic: identification of music: works through audio matching”, Research and Advanced Technology for Digital Libraries, 2007, 12 pages. [cited by applicant]
Park, et al., “Frequency filtering for a highly robust audio fingerprinting scheme in a real-noise environment”, IE!CE transactions on information and systems 89.7, 2006, 4 pages. [cited by applicant]
Rafii, et al., “An audio fingerprinting :system for live version identification using image processing techniques”, Acoustics, Speech and Si?inal Processing (ICASSP), 2014, IEEE International Conference on. IEEE, 5 page… [cited by applicant]
Riley, et al., “A text retrieval approach to content-based audio retrieval”, Int. Symp. On Music Information Retrieval (ISMIR), 2008, 6 pages. [cited by applicant]
Wang, “An Industrial Strength Audio Searc:ll Algorithm”, ISMIR. vol. 2003, 7 pages. [cited by applicant]
Wang, et al., “Automatic Set List Identification and Song Segmentation for Full-Length Concert Videos”, !SMIR, 2014, 6 pages. [cited by applicant]
Tachibana, “Audio watermarking for live performance”, Electronic Imaging, 2003, International Society for Optics and Photonics, 2003, 12 pages. [cited by applicant]
“Cash-strapped music industry pins hope on festivals”, published by CBS news, on Aug. 21, 2012, from http://www.cbsnews.com/8301-505263_ 162-57 497067 /cash-strapped-music-industry-pins-hope-on-festivals/, 2 pages. [cited by applicant]
“Who Says the Music Industry Is Kaput?”, published by Bloomberg Businessweek Magazine, on May 27, 2010, from http://www.businessweek.com/magazine/content/10_23/b4181077568125.htm, 1 pages. [cited by applicant]
Grose, “Live, at a Field Near You: \!\i11y the Music Industry Is Singing a Happy Tune”, published by Time, on Nov. 14, 201 .1, from http://content.time.com/time/printout/0,8816, 2098639,00.html, 4 pages. [cited by applicant]
Sisario, “Strong Earnings for Live Nation in Concert Season”, published by The New York Times, on Aug. 6, 2013, from http://www.nytimes.com/2013/08/07/business/media/strong-earn ings-for-1 ive-nation-in-conce1i-season.h… [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action”, issued in connection with U.S. Appl. No. 14/258,263, on Jan. 15, 2016, 53 pages. [cited by applicant]
United States Patent and Trademark Office, “Final Office Action”, issued in connection with U.S. Appl. No. 14/258,263, on Aug. 12, 2016, 67 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action”, issued in connection with U.S. Appl. No. 14/258,263, on Feb. 23, 2017, 69 pa?ies. [cited by applicant]
United States Patent and Trademark Office, “Final Office Action”, issued in connection with U.S. Appl. No. 14/258,263, on Sep. 5, 2017, 87 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 15/888,998, dated Mar. 20, 2020, 13 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance,” issued in connection with U.S. Appl. No. 15/888,998, dated Jul. 23, 2020, 5 pages. [cited by applicant]