IP Library Granted Patent US 9,111,547
Granted Patent B2
US 9,111,547 · App. 13/591,489 · Granted Aug 18, 2015

Audio signal semantic concept classification method

Inventors: Alexander C. Loui (Penfield, NY); Wei Jiang (Fairport, NY); Kevin Michael Gobeyn (Rochester, NY); Charles Parker (Corvallis, OR)
Assignee: KODAK ALARIS INC.
G10L25/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,111,547
App. No.
13/591,489
Granted
Aug 18, 2015
Kind
B2
Abstract

A method for determining a semantic concept associated with an audio signal captured using an audio sensor. A data processor is used to automatically analyze the audio signal using a plurality of semantic concept detectors to determine corresponding preliminary semantic concept detection values, each semantic concept detector being adapted to detect a particular semantic concept. The preliminary semantic concept detection values are analyzed using a joint likelihood model based on predetermined pair-wise likelihoods that particular pairs of semantic concepts co-occur to determine updated semantic concept detection values. One or more semantic concepts are determined based on the updated semantic concept detection values. The semantic concept detectors and the joint likelihood model are trained together with a joint training process using training audio signals, at least some of which are known to be associated with a plurality of semantic concepts.

Claims (48)

1. A method for determining a semantic concept associated with an audio signal captured using an audio sensor, comprising:

receiving the audio signal from the audio sensor;

using a data processor to automatically analyze the audio signal using a plurality of semantic concept detectors to determine corresponding preliminary semantic concept detection values, the semantic concept detectors being associated with a corresponding plurality of semantic concepts, each semantic concept detector being adapted to detect a particular semantic concept;

using a data processor to automatically analyze the preliminary semantic concept detection values using a joint likelihood model to determine updated semantic concept detection values; wherein the joint likelihood model determines the updated semantic concept detection values based on predetermined pair-wise likelihoods that particular pairs of semantic concepts co-occur;

identifying one or more semantic concept associated with the audio signal based on the updated semantic concept detection values; and

storing an indication of the identified semantic concepts in a processor-accessible memory;

wherein the semantic concept detectors and the joint likelihood model are trained together with a joint training process using training audio signals, at least some of which are known to be associated with a plurality of semantic concepts, and

wherein each of the semantic concept detectors determines the preliminary semantic concept detection values responsive to an associated set of audio features, the audio features being determined by analyzing the audio signal.

2. The method of claim 1 wherein the particular audio features associated with each semantic concept detector are determined during the joint training process.

3. The method of claim 1 wherein the audio signal is subdivided into a set of audio frames, and wherein the audio frames are analyzed to determine frame-level audio features.

4. The method of claim 3 wherein the frame-level audio features from a plurality of audio frames are aggregated to determine clip-level features.

5. The method of claim 4 wherein the frame-level audio features are aggregated by computing frame-level preliminary semantic concept detection values responsive to the frame-level audio features and then determining clip-level preliminary semantic concept detection values by determining an average or a maximum of the frame-level preliminary semantic concept detection values.

6. The method of claim 1 wherein the semantic concept detectors are Nearest Neighbor classifiers, Support Vector Machine classifiers or decision tree classifiers.

7. The method of claim 1 wherein the joint likelihood model is a Markov Random Field model having a set of nodes connected by edges, wherein each node corresponds to a particular semantic concept, and the edge connecting a pair of nodes corresponds to a pair-wise potential function between the corresponding pair of semantic concepts providing an indication of the pair-wise likelihood that the pair of semantic concepts co-occur.

8. The method of claim 1 further including applying a filtering process to discard any semantic concept having a preliminary semantic concept detection value below a predefined threshold.

9. The method of claim 1 wherein the joint training process determines the semantic concept detectors and the joint likelihood model that maximize a predefined performance assessment function.

10. A method for determining a semantic concept associated with an audio signal captured using an audio sensor, comprising:

receiving the audio signal from the audio sensor;

using a data processor to automatically analyze the audio signal using a plurality of semantic concept detectors to determine corresponding preliminary semantic concept detection values, the semantic concept detectors being associated with a corresponding plurality of semantic concepts, each semantic concept detector being adapted to detect a particular semantic concept;

using a data processor to automatically analyze the preliminary semantic concept detection values using a joint likelihood model to determine updated semantic concept detection values; wherein the joint likelihood model determines the updated semantic concept detection values based on predetermined pair-wise likelihoods that particular pairs of semantic concepts co-occur;

identifying one or more semantic concept associated with the audio signal based on the updated semantic concept detection values; and

storing an indication of the identified semantic concepts in a processor-accessible memory;

wherein the semantic concept detectors and the joint likelihood model are trained together with a joint training process using training audio signals, at least some of which are known to be associated with a plurality of semantic concepts, and

wherein the semantic concept detectors are Nearest Neighbor classifiers, Support Vector Machine classifiers or decision tree classifiers.

11. A method for determining a semantic concept associated with an audio signal captured using an audio sensor, comprising:

receiving the audio signal from the audio sensor;

using a data processor to automatically analyze the audio signal using a plurality of semantic concept detectors to determine corresponding preliminary semantic concept detection values, the semantic concept detectors being associated with a corresponding plurality of semantic concepts, each semantic concept detector being adapted to detect a particular semantic concept;

using a data processor to automatically analyze the preliminary semantic concept detection values using a joint likelihood model to determine updated semantic concept detection values; wherein the joint likelihood model determines the updated semantic concept detection values based on predetermined pair-wise likelihoods that particular pairs of semantic concepts co-occur;

identifying one or more semantic concept associated with the audio signal based on the updated semantic concept detection values; and

storing an indication of the identified semantic concepts in a processor-accessible memory;

wherein the semantic concept detectors and the joint likelihood model are trained together with a joint training process using training audio signals, at least some of which are known to be associated with a plurality of semantic concepts, and

wherein the joint likelihood model is a Markov Random Field model having a set of nodes connected by edges, wherein each node corresponds to a particular semantic concept, and the edge connecting a pair of nodes corresponds to a pair-wise potential function between the corresponding pair of semantic concepts providing an indication of the pair-wise likelihood that the pair of semantic concepts co-occur.

12. A method for determining a semantic concept associated with an audio signal captured using an audio sensor, comprising:

receiving the audio signal from the audio sensor;

using a data processor to automatically analyze the audio signal using a plurality of semantic concept detectors to determine corresponding preliminary semantic concept detection values, the semantic concept detectors being associated with a corresponding plurality of semantic concepts, each semantic concept detector being adapted to detect a particular semantic concept;

using a data processor to automatically analyze the preliminary semantic concept detection values using a joint likelihood model to determine updated semantic concept detection values; wherein the joint likelihood model determines the updated semantic concept detection values based on predetermined pair-wise likelihoods that particular pairs of semantic concepts co-occur;

identifying one or more semantic concept associated with the audio signal based on the updated semantic concept detection values;

storing an indication of the identified semantic concepts in a processor-accessible memory; and

applying a filtering process to discard any semantic concept having a preliminary semantic concept detection value below a predefined threshold;

wherein the semantic concept detectors and the joint likelihood model are trained together with a joint training process using training audio signals, at least some of which are known to be associated with a plurality of semantic concepts.

13. A method for determining a semantic concept associated with an audio signal captured using an audio sensor, comprising:

receiving the audio signal from the audio sensor;

using a data processor to automatically analyze the audio signal using a plurality of semantic concept detectors to determine corresponding preliminary semantic concept detection values, the semantic concept detectors being associated with a corresponding plurality of semantic concepts, each semantic concept detector being adapted to detect a particular semantic concept;

using a data processor to automatically analyze the preliminary semantic concept detection values using a joint likelihood model to determine updated semantic concept detection values; wherein the joint likelihood model determines the updated semantic concept detection values based on predetermined pair-wise likelihoods that particular pairs of semantic concepts co-occur;

identifying one or more semantic concept associated with the audio signal based on the updated semantic concept detection values; and

storing an indication of the identified semantic concepts in a processor-accessible memory;

wherein the semantic concept detectors and the joint likelihood model are trained together with a joint training process using training audio signals, at least some of which are known to be associated with a plurality of semantic concepts, and

wherein the joint training process determines the semantic concept detectors and the joint likelihood model that maximize a predefined performance assessment function.

Assignments (12)
SHORT-FORM PATENTS SECURITY AGREEMENT Recorded Sep 5, 2025
From: KODAK ALARIS LLC
To: ENCINA PRIVATE CREDIT SPV 2, LLC, AS COLLATERAL AGENT
Reel/Frame 072818/0674 →
RELEASE OF SECURITY INTEREST Recorded Aug 29, 2025
From: FGI WORLDWIDE LLC
To: KODAK ALARIS LLC
Reel/Frame 072740/0681 →
CHANGE OF NAME Recorded Oct 31, 2024
From: KODAK ALARIS INC.
To: KODAK ALARIS LLC
Reel/Frame 069282/0866 →
RELEASE OF SECURITY INTEREST Recorded Aug 7, 2024
From: THE BOARD OF THE PENSION PROTECTION FUND
To: KODAK ALARIS INC.
Reel/Frame 068481/0300 →
SECURITY AGREEMENT Recorded Aug 2, 2024
From: KODAK ALARIS INC.
To: FGI WORLDWIDE LLC
Reel/Frame 068325/0938 →
ASSIGNMENT OF SECURITY INTEREST Recorded Nov 17, 2021
From: KPP (NO. 2) TRUSTEES LIMITED
To: THE BOARD OF THE PENSION PROTECTION FUND
Reel/Frame 058175/0651 →
SECURITY INTEREST Recorded Oct 5, 2020
From: KODAK ALARIS INC.
To: KPP (NO. 2) TRUSTEES LIMITED
Reel/Frame 053993/0454 →
CHANGE OF NAME Recorded Oct 9, 2013
From: 111616 OPCO (DELAWARE) INC.
To: KODAK ALARIS INC.
Reel/Frame 031394/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2013
From: EASTMAN KODAK COMPANY
To: 111616 OPCO (DELAWARE) INC.
Reel/Frame 031172/0025 →
RELEASE OF SECURITY INTEREST IN PATENTS Recorded Sep 5, 2013
From: CITICORP NORTH AMERICA, INC., AS SENIOR DIP AGENT; WILMINGTON TRUST, NATIONAL ASSOCIATION, AS JUNIOR DIP AGENT
To: EASTMAN KODAK COMPANY; PAKON, INC.
Reel/Frame 031157/0451 →
PATENT SECURITY AGREEMENT Recorded Apr 1, 2013
From: EASTMAN KODAK COMPANY; PAKON, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 030122/0235 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2012
From: LOUI, ALEXANDER C.; JIANG, WEI; GOBEYN, KEVIN MICHAEL; PARKER, CHARLES
To: EASTMAN KODAK
Reel/Frame 028975/0272 →
Continuity (1)
Related Publication 20140056432A1 · Feb 27, 2014