IP Library Granted Patent US 10,089,977
Granted Patent B2
US 10,089,977 · App. 14/793,089 · Granted Oct 2, 2018

Method for system combination in an audio analytics application

Inventors: Sriram Ganapathy (Hartford, CT); Mohamed K. Omar (Chappaqua, NY); Robert Ward (Westchester, NY)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G10L15/005G10L15/32G06F17/275G10L15/02G10L15/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,089,977
App. No.
14/793,089
Granted
Oct 2, 2018
Kind
B2
Abstract

Exemplary embodiments of the present invention provide a method of system combination in an audio analytics application including providing a plurality of language identification systems in which each of the language identification systems includes a plurality of probabilities. Each probability is associated with the system's ability to detect a particular language. The method of system combination in the audio analytics application includes receiving data at the language identification systems. The received data is different from data used to train the language identification systems. A confidence measure is determined for each of the language identification systems. The confidence measure identifies which language its system predicts for the received data and combining the language identification systems according to the confidence measures.

Claims (33)

1. A method of system combination in an audio analytics application, comprising:

providing a plurality of language identification systems, wherein each of the language identification systems includes a plurality of probabilities, wherein each probability is associated with the system's ability to detect a particular language;

receiving data at the language identification systems, wherein the received data is different from data used to train the language identification systems;

determining a confidence measure for each of the language identification systems, wherein the confidence measure identifies a level of accuracy for each of the language identification systems at predicting a presence of the particular language in the received data based on an inverse entropy of a posterior probability distribution of each of the language identification systems, wherein the inverse entropy of the posterior probability of each of the language identification systems is based on a number of languages each of the language identification systems can identify and the plurality of probabilities of each of the language identification systems;

rank ordering each of the language identification systems based on the confidence measure; and

combining a subset of the language identification systems having confidence measures above a predetermined threshold.

2. The method of claim 1 , wherein the language identification systems have different feature extraction methods from each other.

3. The method of claim 1 , wherein the language identification systems have different modeling schemes from each other.

4. The method of claim 1 , wherein the language identification systems have different noise removal schemes from each other.

5. The method of claim 1 , wherein the received data and the data used to train the language identification systems include speech.

6. The method of claim 1 , wherein the received data includes an utterance.

7. The method of claim 1 , wherein determining a confidence measure for each of the language identification systems includes normalizing the inverse entropy value for each of the language identification systems.

8. The method of claim 1 , wherein the steps of determining the confidence measure and combining the language identification systems are repeated for each utterance in the received data.

9. The method of claim 1 , wherein the received data has different characteristics than the data used to train the language identification systems.

10. The method of claim 1 , further comprising using the confidence measure to identify which language its system is best at detecting in the received data.

11. The method of claim 1 , further comprising applying the combination of confidence measures to the received data to increase a performance metric of the language identification systems.

12. The method of claim 11 , wherein the performance metric is indicative of accuracy in detecting language.

13. A method of system combination in an audio analytics application, comprising:

providing a plurality of language identification systems, wherein each of the language identification systems includes a plurality of probabilities, wherein each probability is associated with the system's ability to detect a particular language;

receiving data at the language identification systems, wherein the received data has different characteristics than data used to train the language identification systems;

determining a confidence measure for each of the language identification systems using a portion of the received data, wherein the confidence measure identifies a level of accuracy for each of the language identification systems at predicting a presence of the particular language in the received data based on an inverse entropy of a posterior probability distribution of each of the language identification systems, wherein the inverse entropy of the posterior probability of each of the language identification systems is based on a number of languages each of the language identification systems can identify and the plurality of probabilities of each of the language identification systems; and

combining at least two of the language identification systems having confidence measures above a predetermined threshold.

14. The method of claim 13 , wherein the portion of the received data used to determine the confidence measure is less than 10% of the received data.

15. The method of claim 13 , wherein less than all of the language identification systems are combined according to the confidence measures.

16. The method of claim 13 , further comprising, before combining the at least two language identification systems, pruning the language identification systems based on their confidence measures.

17. A method of system combination in an audio analytics application, comprising:

providing a plurality of language identification systems trained on first data;

receiving a second data at the language identification systems, wherein the second data is different from the first data;

determining a confidence measure for each of the language identification systems, wherein the confidence measure identifies a level of accuracy for each of the language identification systems at predicting a presence of a particular language in the received data based on an inverse entropy of a posterior probability distribution of each of the language identification systems, wherein the inverse entropy of the posterior probability of each of the language identification systems is based on a number of languages each of the language identification systems can identify and a plurality of probabilities of each of the language identification systems;

combining a subset of the language identification systems having confidence measures above a predetermined threshold;

inputting third data different from the first and second data to the combination of language identification systems; and

identifying a language of the third data.

18. The method of claim 17 , wherein each of the language identification systems includes the plurality of probabilities, wherein each probability is associated with the system's ability to detect the particular language.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2015
From: GANAPATHY, SRIRAM; OMAR, MOHAMED K.; WARD, ROBERT
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 036010/0236 →
Continuity (1)
Related Publication 20170011734A1 · Jan 12, 2017