IP Library Granted Patent US 9,558,272
Granted Patent B2
US 9,558,272 · App. 15/103,994 · Granted Jan 31, 2017

Method of and a system for matching audio tracks using chromaprints with a fast candidate selection routine

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,558,272
App. No.
15/103,994
Granted
Jan 31, 2017
Kind
B2
Abstract

A computer-implemented method of matching of a first incoming audio track with an indexed audio track, the method executable at a server, the method comprising: selecting the indexed audio track as a candidate audio track from a plurality of indexed audio tracks; validating the candidate audio track against the first audio track.

Claims (47)

1. A computer-implemented method of matching of a first incoming audio track with an indexed audio track, the method executable at a server, the method comprising:

selecting the indexed audio track as a candidate audio track from a plurality of indexed audio tracks, the selecting by executing steps of:

determining a first short audio fingerprint, the first short audio fingerprint being an audio fingerprint of a first portion of the first incoming audio track, the first short audio fingerprint comprising a first chroma word, the first portion of the first audio track being of a first predetermined duration from a start of the first incoming audio track;

determining the candidate audio track from a set of indexed audio tracks, the candidate audio track having a second short audio fingerprint that contains a second chroma word, a beginning portion of the second chroma word being identical to a beginning portion of the first chroma word, the beginning portion of the second chroma word being a sub-plurality of bytes having a first byte and a following byte, the candidate audio track being indexed in a posting list within a set of posting lists amongst a plurality of sets of posting lists, each posting list within the set of posting lists being associated with respective chroma words having a same first byte and a different following byte, the different following byte being unique for each posting list, the second short audio fingerprint being an audio fingerprint of a first portion of the candidate audio track, the first portion of the candidate audio track being of said first predetermined duration from a start of the candidate audio track,

validating the candidate audio track against the first audio track by executing steps of:

determining a first long audio fingerprint, the first long audio fingerprint being an audio fingerprint of a second portion of the first incoming audio track;

retrieving a second long audio fingerprint, the second long audio fingerprint being an audio fingerprint of a second portion of the candidate audio track;

each of the second portion of the first audio track and the second portion of the candidate audio track are of a second predetermined duration from the start of the respective one of the first audio track and the candidate audio track;

each of the first portion of the respective one of the first audio track and the candidate audio track being fully contained within the second portion of the respective one of the first audio track and the candidate audio track

performing bit-by-bit comparing of the first long audio fingerprint with the second long audio fingerprint.

2. The method of claim 1 , wherein a beginning portion of the chroma word comprises a combination of

any one of a first byte and a first multi-byte sequence, the first multi-byte sequence being the sequence of bytes in the beginning of the respective chroma word, the first multi-byte sequence having a pre-determined number of bytes, and any one of a following byte and a second multi-byte sequence, the second multi-byte sequence being the sequence of bytes following any one of the first multi-byte sequence and the first byte of each respective chroma word, the second multi-byte sequence having the pre-determined number of bytes.

3. The method of claim 1 , wherein the first predetermined duration is lesser of:

a predetermined duration within a time range from 9 to 27 seconds, and

a respective audio track duration.

4. The method of claim 3 , wherein the first predetermined duration is lesser of:

21 seconds, and

a respective audio track duration.

5. The method of claim 1 , wherein the second predetermined duration is lesser one of:

a predetermined duration within a time range from 96 to 141 seconds, and

a respective audio track duration.

6. The method of claim 5 , wherein the second predetermined duration is lesser of:

120 seconds, and

a respective audio track duration.

7. The method of claim 1 , wherein each of said first chroma word and said second chroma word describes a portion of a respective audio track, duration of the portion of the audio track being within time range from ½ second to 8 seconds.

8. The method of claim 7 , further comprising generating said first chroma word and said second chroma word.

9. The method of claim 1 , wherein each of said first and second chroma words are associated with an indication of a track ID associated with a respective audio track.

10. The method of claim 9 , wherein the track ID is described with a third multi-byte sequence which is located after any one of

the following byte,

the second multi-byte sequence.

11. The method of claim 1 , wherein each of said first and second chroma words are associated with an indication of a track duration information associated with the respective audio track.

12. The method of claim 11 , wherein the indication of a track duration is described within one byte which immediately follows any one of

the following byte,

the second multi-byte sequence.

13. The method of claim 1 , said determining candidate audio track comprises comparing respective track duration of the first incoming audio track and the candidate audio track.

14. The method of claim 13 , further comprising determining that the candidate audio track is not a matched candidate to the first incoming audio track responsive to the track duration varying by more than a pre-set value.

15. The method of claim 13 , wherein the candidate audio track comprises a plurality of candidate audio tracks and wherein the method further comprises selecting a sub-set of the plurality of candidate audio tracks based on a pre-set candidate threshold number.

16. The method of claim 1 , wherein the bit-by-bit comparing of the first long audio fingerprint with the second long audio fingerprint comprises shifting the first long audio fingerprint relative to the second long audio fingerprint.

17. The method of claim 16 , wherein said shifting comprises an amplitude of a shift, and wherein said amplitude ranges between plus 20 seconds and minus 20 seconds.

18. The method of claim 1 , wherein said determining, that the beginning portion of the second chroma word is identical to the beginning portion of the first chroma word is executed by determining that an entire sequence of bytes in the beginning portion of the second chroma word matches an entire sequence of bytes in the beginning portion of the first chroma.

19. The method of claim 1 , wherein at least one of the short audio fingerprint and the long audio fingerprint contains an indication of a track ID associated with a respective audio track.

20. The method of claim 1 , further comprising, prior to said determining the first short audio fingerprint, receiving, by the server, at least a portion of the first incoming audio track.

21. The method of claim 1 , wherein retrieving of the second short audio fingerprint comprises retrieving using an index.

22. The method of claim 21 , said index is the audio track inverted index.

23. The method of claim 22 , the audio track inverted index is any one of

a pruning index, the pruning index being built for a plurality of short audio fingerprints, and

a validation index, the validation index being built for a plurality of long audio fingerprints.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2024
From: DIRECT CURSUS TECHNOLOGY L.L.C
To: Y.E. HUB ARMENIA LLC
Reel/Frame 068524/0184 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: YANDEX EUROPE AG
To: DIRECT CURSUS TECHNOLOGY L.L.C
Reel/Frame 064925/0808 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2016
From: YANDEX LLC
To: YANDEX EUROPE AG
Reel/Frame 040521/0363 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2016
From: KALININA, ELENA ANDREEVNA
To: YANDEX LLC
Reel/Frame 040816/0896 →