IP Library Granted Patent US 7,295,977
Granted Patent B2
US 7,295,977 · App. 09/939,954 · Granted Nov 13, 2007

Extracting classifying data in music from an audio bitstream

Assignee: NEC Laboratories America, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,295,977
App. No.
09/939,954
Granted
Nov 13, 2007
Kind
B2
Abstract

The method of the present invention utilizes machine-learning techniques, particularly Support Vector Machines in combination with a neural network, to process a unique machine-learning enabled representation of the audio bitstream. Using this method, a classifying machine is able to autonomously detect characteristics of a piece of music, such as the artist or genre, and classify it accordingly. The method includes transforming digital time-domain representation of music into a frequency-domain representation, then dividing that frequency data into time slices, and compressing it into frequency bands to form multiple learning representations of each song. The learning representations that result are processed by a group of Support Vector Machines, then by a neural network, both previously trained to distinguish among a given set of characteristics, to determine the classification.

Claims (24)

1. A method of extracting classifying data from an audio signal, the method comprising the steps of:

transforming a perceptual representation of the audio signal into a learning representation of the audio signal;

transmitting the learning representation to a multi-stage classifier, the multi-stage classifier comprising:

a first stage having a plurality of support vector machine classifiers, each support vector machine classifier trained to identify one out of a plurality of audio classification categories and generate a metalearner vector value reflecting how closely the audio signal conforms to the one out of the plurality of audio classification categories, and

a final stage having a metalearner classifier, the metalearner classifier using the generated metalearner vector to classify the audio signal into one out of the plurality of audio classification categories; and

generating classification category information for the audio signal based on results produced by the metalearner classifier.

2. The method of claim 1 wherein the final stage metalearner classifier is a neural network classifier.

3. The method of claim 1 wherein said audio classification categories comprises classifications by musical artist.

4. The method of claim 1 wherein the learning representation comprises dividing the perceptual representation of the audio signal into a plurality of time slices.

5. The method of claim 1 wherein the learning representation comprises dividing the perceptual representation of the audio signal into a plurality of frequency bands.

6. A computer readable storage medium, storing therein a program of instructions for causing a computer to execute a process of extracting classifying data from an audio signal, the process comprising the steps of:

processing a perceptual representation of the audio signal into a learning representation of the audio signal; and

inputting the learning representation into a multi-stage classifier, the multi-stage classifier comprising a first stage of support vector machine classifiers and a final stage metalearner classifier, each support vector machine classifier trained to identify one out of a plurality of audio classification categories and where the support vector machine classifiers are used to generate a metalearner vector that allows the final stage metalearner classifier to classify the audio signal into one out of the plurality of audio classification categories, each support vector machine classifier outputting a value reflecting how closely the audio signal conforms to the one out of the plurality of audio classification categories, each value then used in the metalearner vector.

7. The computer readable storage medium of claim 6 wherein the final stage metalearner classifier is a neural network classifier.

8. The computer readable storage medium of claim 6 wherein said audio classification categories comprises classifications by musical artist.

9. The computer readable storage medium of claim 6 wherein the learning representation comprises dividing the perceptual representation of the audio signal into a plurality of time slices.

10. The computer readable storage medium of claim 6 wherein the learning representation comprises dividing the perceptual representation of the audio signal into a plurality of frequency bands.

11. An apparatus for classifying an audio signal comprising:

means for processing a perceptual representation of the audio signal into a learning representation of the audio signal; and

a multi-stage classifier, the multi-stage classifier further comprising a first stage of support vector machine classifiers and a final stage metalearner classifier, each support vector machine classifier trained to identify one out of a plurality of audio classification categories from the learning representation of the audio signal and where the support vector machine classifiers are used to generate a metalearner vector that allows the final stage metalearner classifier to classify the audio signal into one out of the plurality of audio classification categories, each support vector machine classifier outputting a value reflecting how closely the audio signal conforms to the one out of the plurality of audio classification categories, each value then used in the metalearner vector.

12. The apparatus of claim 11 wherein the final stage metalearner classifier is a neural network classifier.

13. The apparatus of claim 11 wherein said audio classification categories comprises classifications by musical artist.

14. The apparatus of claim 11 wherein the learning representation comprises dividing the perceptual representation of the audio signal into a plurality of time slices.

15. The apparatus of claim 11 wherein the learning representation comprises dividing the perceptual representation of the audio signal into a plurality of frequency bands.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 12, 2008
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 020487/0759 →
CHANGE OF NAME Recorded Dec 31, 2002
From: NEC RESEARCH INSTITUTE, INC.
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 013599/0895 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2001
From: WHITMAN, BRIAN; FLAKE, GARY W.; LAWRENCE, STEPHEN R.
To: NEC RESEARCH INSTITUTE, INC.
Reel/Frame 012136/0670 →
Continuity (1)
Related Publication 20030040904A1 · Feb 27, 2003