IP Library › Granted Patent US 7,181,392
Granted Patent B2
US 7,181,392 · App. 10/196,014 · Granted Feb 20, 2007

Determining speech recognition accuracy

Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,181,392
App. No.
10/196,014
Granted
Feb 20, 2007
Kind
B2
Abstract

A method of determining the accuracy of a speech recognition system can include identifying from a log of the speech recognition system a text result and attributes associated with the text result. An audio representation from which the text result was derived can be accessed. The audio representation can be processed with a reference speech recognition engine to determine a second text result. The text result can be compared with the second text result to determine an accuracy of the speech recognition system.

Claims (48)

1. A method of determining the accuracy of a speech recognition system comprising:

identifying from a log of said speech recognition system a text result and attributes associated with said text result, wherein said log comprises a plurality of log entries, each log entry correlated with a different electronically stored audio segment defining an audio representation corresponding to a portion of text for which a text result and at least one attribute is provided, and wherein each attribute characterizes at least one of a configuration of the speech recognition system, a type of audio channel over which an audio segment is received, and whether the audio segment is received in response to a user prompt;

accessing an audio representation from which said text result was derived;

processing said audio representation with a reference speech recognition engine to determine a second text result; and

comparing said text result with said second text result to determine an accuracy of said speech recognition system.

2. The method of claim 1 , further comprising:

repeating each step of claim 1 for additional text results of said speech recognition system to determine an accuracy statistic of said speech recognition system.

3. The method of claim 2 , further comprising:

identifying at least one error condition specified in said log; and

determining said accuracy statistic for said at least one of said identified error conditions.

4. The method of claim 2 , wherein said attributes specify timing information for said text result, said accessing an audio representation step further comprising:

identifying said audio representation from a plurality of audio representations according to said attributes of said text result.

5. The method of claim 2 , further comprising:

manually determining text from said audio representation; and

comparing said first text result and said second text result with said manually determined text.

6. The method of claim 5 , further comprising:

receiving audio properties of said audio representation.

7. The method of claim 6 , further comprising:

adjusting the configuration of said reference speech recognition engine according to said audio properties of said audio representation.

8. The method of claim 6 , further comprising:

altering acoustic models of said reference speech recognition engine according to said audio properties of said audio representation.

9. The method of claim 6 , wherein said accuracy statistic is selected from the group consisting of a ratio of failed recognitions to total recognitions minus failed recognitions due to uncontrollable environmental elements, and a ratio of failed recognitions due to uncontrollable environmental elements to total recognitions.

10. The method of claim 6 , wherein said accuracy statistic is selected from the group consisting of a total number of occurrences of unique words in which there was an attempt at recognition, a number of said occurrences of unique words which were successfully recognized, a number of said occurrences of unique words which were unsuccessfully recognized, and a number of failed attempts for said occurrences of unique words due to uncontrollable environmental elements.

11. The method of claim 2 , wherein said accuracy statistic is selected from the group consisting of a ratio of the successful recognitions to total recognitions, a ratio of successful recognitions to total recognitions minus failed recognitions, and a ratio of failed recognitions to total recognitions.

12. A machine-readable storage, having stored thereon a computer program having a plurality of code sections executable by a machine for causing the machine to perform the steps of:

identifying from a log of said speech recognition system a text result and attributes associated with said text result, wherein said log comprises a plurality of log entries, each log entry correlated with a different electronically stored audio segment defining an audio representation corresponding to a portion of text for which a text result and at least one attribute is provided, and wherein each attribute characterizes at least one of a configuration of the speech recognition system, a type of audio channel over which an audio segment is received, and whether the audio segment is received in response to a user prompt;

accessing an audio representation from which said text result was derived;

processing said audio representation with a reference speech recognition engine to determine a second text result; and

comparing said text result with said second text result to determine an accuracy of said speech recognition system.

13. The machine-readable storage of claim 12 , further comprising:

repeating each step of claim 12 for additional text results of said speech recognition system to determine an accuracy statistic of said speech recognition system.

14. The machine-readable storage of claim 13 , further comprising:

identifying at least one error condition specified in said log; and

determining said accuracy statistic for said at least one of said identified error conditions.

15. The machine-readable storage of claim 13 , wherein said attributes specify timing information for said text result, said accessing an audio representation step further comprising:

identifying said audio representation from a plurality of audio representations according to said attributes of said text result.

16. The machine-readable storage of claim 13 , further comprising:

manually determining text from said audio representation; and

comparing said first text result and said second text result with said manually determined text.

17. The machine-readable storage of claim 16 , further comprising:

receiving audio properties of said audio representation.

18. The machine-readable storage of claim 17 , further comprising:

adjusting the configuration of said reference speech recognition engine according to said audio properties of said audio representation.

19. The machine-readable storage of claim 17 , further comprising:

altering acoustic models of said reference speech recognition engine according to said audio properties of said audio representation.

20. The machine-readable storage of claim 17 , wherein said accuracy statistic is selected from the group consisting of a ratio of failed recognitions to total recognitions minus failed recognitions due to uncontrollable environmental elements, and a ratio of failed recognitions due to uncontrollable environmental elements to total recognitions.

21. The machine-readable storage of claim 17 , wherein said accuracy statistic is selected from the group consisting of a total number of occurrences of unique words in which there was an attempt at recognition, a number of said occurrences of unique words which were successfully recognized, a number of said occurrences of unique words which were unsuccessfully recognized, and a number of failed attempts for said occurrences of unique words due to uncontrollable environmental elements.

22. The machine-readable storage of claim 13 , wherein said accuracy statistic is selected from the group consisting of a ratio of the successful recognitions to total recognitions, a ratio of successful recognitions to total recognitions minus failed recognitions, and a ratio of failed recognitions to total recognitions.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022354/0566 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 16, 2002
From: GANDHI, SHAILESH B.; JAISWAL, PEEYUSH; MOORE, VICTOR S.; TOON, GREGORY L.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 013108/0028 →
Continuity (1)
Related Publication 20040015350A1 · Jan 22, 2004