IP Library › Granted Patent US 8,880,545
Granted Patent B2
US 8,880,545 · App. 13/110,185 · Granted Nov 4, 2014

Query and matching for content recognition

Inventors: Kazuhito Koishida (Redmond, WA); David Nister (Bellevue, WA); Ian Simon (Seattle, WA); Tom Butcher (Seattle, WA)
Assignee: Microsoft Corporation
G06F17/30743
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,880,545
App. No.
13/110,185
Granted
Nov 4, 2014
Kind
B2
Abstract

Various embodiments enable audio data, such as music data, to be captured, by a device, from a background environment and processed to formulate a query that can then be transmitted to a content recognition service. In one or more embodiments, multiple queries are transmitted to the content recognition service. In at least some embodiments, subsequent queries can progressively incorporate previous queries plus additional data that is captured. In one or more embodiments, responsive to receiving the query, the content recognition service can employ a multi-stage matching technique to identify content items responding to the query. This matching technique can be employed as queries are progressively received.

Claims (51)

1. One or more computer-readable storage media comprising instructions that are executable to cause a device to perform operations comprising:

capturing, using a computing device, audio data, at least some of which is processable for provision to a content recognition service;

formulating, by applying a Hamming window to the audio data and further processing the audio data at the computing device, a query for submission to the content recognition service to identify displayable content information associated with the audio data;

submitting a first query to a content recognition service, the first query being formulated using one or more features extracted from a first portion of the audio data, each of the one or more features comprising at least spectral peak data for use in identifying the displayable content information associated with the audio data;

responsive to an indication that no displayable content information is received based on the first query, submitting one or more subsequent queries to the content recognition service, the one or more subsequent queries comprising at least one of the one or more features extracted from the first portion of the audio data and used to formulate the first query, along with additional features not included in the first query; and

terminating said submitting the one or more subsequent queries responsive to receiving the displayable content information from the content recognition service.

2. The one or more computer-readable storage media of claim 1 , wherein the operations further comprise receiving, from the content recognition service, the displayable content information associated with the audio data.

3. The one or more computer-readable storage media of claim 2 , wherein the displayable content information comprises one or more of a song title, an artist, an album title, a date an audio clip was recorded, a writer, a producer, or group members.

4. The one or more computer-readable storage media of claim 2 , wherein the operations further comprise displaying the displayable information associated with the audio data.

5. The one or more computer-readable storage media of claim 1 , wherein further processing the audio data comprises:

processing the audio data effective to extract the spectral peak of each of the one or more features from the audio data; and

accumulating the one or more features extracted from the audio data to formulate the query.

6. The one or more computer-readable storage media of claim 5 , wherein processing the audio data further comprises:

zero padding the audio data to which the Hamming window was applied;

transforming, using a fast Fourier transform algorithm, the zero-padded audio data;

producing a log-power time-frequency spectrum by applying a log power to the audio data to which the fast Fourier transform algorithm was applied; and

extracting the spectral peak of each of the one or more features from the log-power time-frequency spectrum.

7. The one or more computer-readable storage media of claim 1 , wherein the computing device is a mobile device.

8. A system comprising:

one or more processors; and

one or more memories storing instructions that are executable by the one or more processors to perform operations including:

receiving, from a device, a first query associated with one or more features extracted from audio data captured by the device, each of the one or more features comprising at least spectral peak data associated with the audio data;

processing the first query effective to attempt to identify a song associated with the audio data;

receiving, from the device and independent of a prompt for a query, at least one additional query associated with one or more additional features comprising at least one of the one or more features associated with the first query and additional spectral peak data associated with additional audio data captured by the device and not included in the first query;

processing the at least one additional query effective to attempt to identify the song associated with the additional audio data by:

scanning a content database across a first beam width corresponding to a frequency range to produce one or more content item candidates having peak information corresponding to the spectral peak data of the at least one additional query; and

scanning the one or more content item candidates across a second beam width to produce a content item candidate with peak information corresponding to the spectral peak data of the first query and the at least one additional query;

identifying the song as the content item candidate corresponding to the first query and the at least one additional query; and

responsive to identifying the song, returning content information associated with the song to the device.

9. The system of claim 8 , wherein the one or more features associated with the first query and the additional features of the at least one additional query comprise a time index and a frequency location corresponding to the spectral peak data.

10. The system of claim 9 , wherein the second beam width is wider than the first beam width.

11. The system claim 8 , wherein the computing device is a server.

12. A computer-implemented method comprising:

receiving, from a device, a query comprising a time index and a frequency location corresponding to an extracted audio peak;

scanning a content database across a first beam width to produce one or more time positions at the frequency location corresponding to the extracted audio peak;

assigning a content score to each content item corresponding to the one or more time positions, the content score corresponding to a difference between the one or more time positions of a content item and the time index of the extracted audio peak of the query, each content item corresponding to one of a plurality of candidates that are content items;

scanning, based on the content score assigned to each of the plurality of candidates, at least some of the plurality of candidates across a second beam width to produce one or more time positions at the frequency location corresponding to the extracted audio peak of the query, the one or more time positions corresponding to a candidate that is a content item responsive to the query; and

transmitting to the device displayable information regarding the candidate.

13. The computer-implemented method of claim 12 , wherein the second beam width is wider than the first beam width.

14. The computer-implemented method of claim 12 , wherein the displayable information comprises one or more of a song title, an artist, an album title, a date an audio clip was recorded, a writer, a producer, or group members.

15. The computer-implemented method of claim 12 , further comprising receiving, from the device, at least one additional query comprising:

a time index and a frequency location corresponding to one or more additional extracted audio peaks; and

the time index and the frequency location corresponding to the one or more extracted audio peaks of the first query.

16. The computer-implemented method of claim 12 , wherein the content item is a song.

17. One or more computer-readable storage media comprising instructions that are executable to cause a device to perform the method of claim 12 .

18. The one or more computer-readable storage media of claim 1 , wherein the spectral peak data comprises a time index and a frequency location corresponding to at least one spectral peak.

19. The system of claim 8 , wherein the spectral peak data comprises a time index and a frequency location corresponding to at least one spectral peak.

20. The one or more computer-readable storage media of claim 9 , wherein said scanning the content database across the first beam width comprises:

producing one or more time positions at the frequency location corresponding to an extracted spectral peak; and

assigning a content score to a content item corresponding to each of the one or more time positions, the content score corresponding to a difference between the one or more time positions of the content item and the time index of the extracted spectral peak of the query, the content item corresponding to one of the one or more candidates; and

wherein the scanning the one or more candidates is based on the content score.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034544/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2011
From: KOISHIDA, KAZUHITO; NISTER, DAVID; SIMON, IAN; BUTCHER, TOM
To: MICROSOFT CORPORATION
Reel/Frame 026299/0227 →
Continuity (1)
Related Publication 20120296938A1 · Nov 22, 2012