IP Library Granted Patent US 8,112,418
Granted Patent B2
US 8,112,418 · App. 12/052,299 · Granted Feb 7, 2012

Generating audio annotations for search and retrieval

Assignee: The Regents of the University of California
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,112,418
App. No.
12/052,299
Granted
Feb 7, 2012
Kind
B2
Abstract

Embodiments of a computer system to determine one or more annotation items associated with an audio file are described. During operation, the computer system provides an interactive environment in which multiple users listen to the audio file within a time interval. Next, the computer system receives one or more annotation items associated with the audio file from the multiple users. Then, the computer system displays the received one or more annotation items from the multiple users in the interactive environment, thereby enabling the multiple users to provide feedback to a given user in the multiple users.

Claims (40)

1. A computer-implemented method for characterizing songs, wherein the method comprises:

providing interactive environments that include games, having predetermined durations, in which training songs are played to groups of users and associated annotation items are dynamically received from the groups of users, wherein the annotation items include words that characterize the training songs, wherein at least an annotation item received from a given user in a given group of users includes an extrema in characterizing a given training song;

wherein, in a given interactive environment, the given training song is played to users in the given group of users, the annotation items are dynamically received from the users, and the annotation items received from other users in the given group of users during a given game is provided to the given user; and

wherein a score of the given user during the given game is based on attempts of the given user to match his provided annotation items to those of a majority of the given group of users within a time limit;

using the computer, determining models that specify statistical relationships between the training songs and the received annotation items, wherein a given model specifies a probability that a given annotation item is associated with the given training song;

aggregating the models into compound models that specify the statistical relationships between the training songs and the annotation items, wherein a given compound model identifies subsets of the training songs that are associated with the given annotation item; and

determining the annotation items associated with additional songs based on the compound models, wherein the determining characterizes the additional songs, and wherein the annotation items associated with the additional songs were not previously determined in the interactive environments.

2. The method of claim 1 , further comprising providing a reward to the given user in the given group of users based on agreement between one or more of the annotation items received from the given user and those of the majority of the given group of users.

3. The method of claim 1 , further comprising providing feedback to the given user in the given group of users based one or more of the annotation items received from the given user and those of the majority of the given group of users.

4. The method of claim 1 , wherein the one or more annotation items include semantic labels.

5. The method of claim 1 , wherein, for the given user in the given group of users, receiving the one or more annotation items involves the given user selecting one or more annotation items from pre-determined annotation items.

6. The method of claim 5 , wherein the pre-determined annotation items are determined from a document associated with the audio file.

7. The method of claim 6 , wherein the document includes a review of the audio file.

8. The method of claim 5 , wherein the pre-determined annotation items include types of music, emotions, descriptions of vocals, types of musical instruments, descriptions of musical styles, and rhythms.

9. The method of claim 1 , further comprising repeating the given interactive environment for the given training song multiple times prior to determining the given model.

10. The method of claim 1 , wherein the given model specifies the probability that a given annotation item is associated with content of the given training song.

11. The method of claim 10 , wherein the content includes audio content.

12. The method of claim 1 , wherein the given compound model corresponds to a joint distribution.

13. A computer-program product for use in conjunction with a computer system, the computer-program product comprising a non-transitory computer-readable storage medium and a computer-program mechanism embedded therein to characterize songs, the computer-program mechanism including:

instructions for providing interactive environments that include games, having predetermined durations, in which training songs are played to groups of users and associated annotation items are dynamically received from the groups of users, wherein the annotation items include words that characterize the training songs, wherein at least an annotation item received from a given user in a given group of users includes an extrema in characterizing a given training song;

wherein, in a given interactive environment, the given training song is played to users in the given group of users, the annotation items are dynamically received from the users, and the annotation items received from other users in the given group of users during a given game is provided to the given user; and

wherein a score of the given user during the given game is based on attempts of the given user to match his provided annotation items to those of a majority of the given group of users within a time limit;

instructions for determining models that specify statistical relationships between the training songs and the received annotation items, wherein a given model specifies a probability that a given annotation item is associated with the given training song;

instructions for aggregating the models into compound models that specify the statistical relationships between the training songs and the annotation items, wherein a given compound model identifies subsets of the training songs that are associated with the given annotation item; and

instructions for determining the annotation items associated with additional songs based on the compound models, wherein the determining characterizes the additional songs, and wherein the annotation items associated with the additional songs were not previously determined in the interactive environments.

14. The computer-program product of claim 13 , wherein the computer-program mechanism includes instructions for providing feedback to the given user in the given group of users based one or more of the annotation items received from the given user and those of the majority of the given group of users.

15. The computer-program product of claim 13 , wherein the computer-program mechanism includes instructions for repeating the given interactive environment for the given training song multiple times prior to determining the given model.

16. The computer-program product of claim 13 , wherein the given model specifies the probability that a given annotation item is associated with content of the given training song.

17. The computer-program product of claim 16 , wherein the content includes audio content.

18. The computer-program product of claim 13 , wherein the given compound model corresponds to a joint distribution.

19. A computer system, comprising:

a processor;

a memory;

a program module, wherein the program module is stored in the memory and configured to be executed by the processor to characterize songs, the program module including:

instructions for providing interactive environments that include games, having predetermined durations, in which training songs are played to groups of users and associated annotation items are dynamically received from the groups of users, wherein the annotation items include words that characterize the training songs, wherein at least an annotation item received from a given user in a given group of users includes an extrema in characterizing a given training song;

wherein, in a given interactive environment, the given training song is played to users in the given group of users, the annotation items are dynamically received from the users, and the annotation items received from other users in the given group of users during a given game is provided to the given user; and

wherein a score of the given user during the given game is based on attempts of the given user to match his provided annotation items to those of a majority of the given group of users within a time limit;

instructions for determining models that specify statistical relationships between the training songs and the received annotation items, wherein a given model specifies a probability that a given annotation item is associated with the given training song;

instructions for aggregating the models into compound models that specify the statistical relationships between the training songs and the annotation items, wherein a given compound model identifies subsets of the training songs that are associated with the given annotation item; and

instructions for determining the annotation items associated with additional songs based on the compound models, wherein the determining characterizes the additional songs, and wherein the annotation items associated with the additional songs were not previously determined in the interactive environments.

Assignments (4)
CONFIRMATORY LICENSE Recorded May 24, 2012
From: UNIVERSITY OF CALIFORNIA SAN DIEGO
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 028274/0873 →
CONFIRMATORY LICENSE Recorded Aug 12, 2009
From: UNIVERSITY OF CALIFORNIA, SAN DIEGO
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 023088/0263 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2008
From: TURNBULL, DOUGLAS; BARRINGTON, LUKE; LIU, RUORAN; LANCKRIET, GERT
To: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
Reel/Frame 021831/0635 →
CONFIRMATORY LICENSE Recorded Jun 10, 2008
From: NATIONAL SCIENCE FOUNDATION
To: SAN DIEGO, UNIVERSITY OF CALIFORNIA AT
Reel/Frame 021071/0415 →
Continuity (2)
Provisional Application 60896216 · Mar 21, 2007
Related Publication 20080235283A1 · Sep 25, 2008