IP Library Granted Patent US 10,019,984
Granted Patent B2
US 10,019,984 · App. 14/634,714 · Granted Jul 10, 2018

Speech recognition error diagnosis

Inventors: Shiun-Zu Kuo (Bothell, WA); Thomas Reutter (Redmond, WA); Yifan Gong (Sammamish, WA); Mark T. Hanson (Woodinville, WA); Ye Tian (Kenmore, WA); Shuangyu Chang (Fremont, CA); Jonathan Hamaker (Issaquah, WA); Qi Miao (Kenmore, WA); Yuancheng Tu (Issaquah, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/01G10L15/183
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,019,984
App. No.
14/634,714
Granted
Jul 10, 2018
Kind
B2
Abstract

Techniques and technologies for diagnosing speech recognition errors are described. In an example implementation, a system for diagnosing speech recognition errors may include an error detection module configured to determine that a speech recognition result is least partially erroneous, and a recognition error diagnostics module. The recognition error diagnostics module may be configured to (a) perform a first error analysis of the at least partially erroneous speech recognition result to provide a first error analysis result; (b) perform a second error analysis of the at least partially erroneous speech recognition result to provide a second error analysis result; and (c) determine at least one category of recognition error associated with the at least partially erroneous speech recognition result based on a combination of the first error analysis result and the second error analysis result.

Claims (109)

1. A system for diagnosing speech recognition errors, comprising:

at least one processing component; and

one or more media operably coupled to the at least one processing component and bearing one or more instructions that, when executed by the at least one processing component, perform operations including at least:

determine that a speech recognition result is at least partially erroneous;

perform a first error analysis of the at least partially erroneous speech recognition result to provide a first error analysis result;

perform a second error analysis of the at least partially erroneous speech recognition result to provide a second error analysis result; and

determine at least one category of recognition error associated with the at least partially erroneous speech recognition result based on a combination of the first error analysis result and the second error analysis result, including determine that the at least one category of recognition error includes at least an acoustic model error when

(a) the first error analysis result indicates that a reference language model score associated with a reference speech is higher than a recognition language model score associated with the at least partially erroneous speech recognition result; and

(b) the second error analysis result indicates that a reference acoustic model score associated with the reference speech is lower than a recognition acoustic model score associated with the at least partially erroneous speech recognition result;

determine at least one corrective action to at least partially correct at least one aspect of a speech recognition component based at least partially on the at least one category of recognition error associated with the at least partially erroneous speech recognition result; and

at least one of:

provide an indication of the at least one corrective action; or

adjust at least one aspect of the speech recognition component based on the at least one corrective action.

2. The system of claim 1 , wherein:

the first error analysis includes at least one language model scoring operation; and

the second error analysis includes at least one acoustic model scoring operation.

3. The system of claim 1 , wherein:

the first error analysis includes at least one dictionary check operation; and

the second error analysis includes at least one transcription analysis operation.

4. The system of claim 1 , wherein:

the first error analysis includes at least one emulation operation; and

the second error analysis includes at least one grammar analysis operation.

5. The system of claim 1 , wherein:

the first error analysis of the at least partially erroneous speech recognition result includes a comparison of a language model score associated with the at least partially erroneous speech recognition result with a language model score associated with a reference speech recognition result; and

the second error analysis of the at least partially erroneous speech recognition result includes a comparison of an acoustic model score associated with the at least partially erroneous speech recognition result with an acoustic model score associated with the reference speech recognition result.

6. The system of claim 1 , wherein at least one of the first error analysis or the second error analysis comprises:

one or more emulation operations that assume an ideal operation of an acoustic model to assess an actual operation of a language model.

7. The system of claim 1 , wherein the operations further comprise:

perform a third error analysis of the at least partially erroneous speech recognition result to provide a third error analysis result; and

determine at least one category of recognition error associated with the at least partially erroneous speech recognition result based on a combination of at least the first error analysis result, the second error analysis result, and the third error analysis result.

8. The system of claim 7 , wherein:

the first error analysis includes at least one language model scoring operation;

the second error analysis includes at least one acoustic model scoring operation; and

the third error analysis includes at least one of an engine setting check operation, a penalty model setting check operation, a force alignment operation, a 1:1 alignment test operation, an emulation operation, or a dictionary check operation.

9. The system of claim 1 , wherein determine at least one corrective action to at least partially correct at least one aspect of a speech recognition component based at least partially on the at least one category of recognition error associated with the at least partially erroneous speech recognition result comprises:

determine at least one corrective action to at least partially correct at least one aspect of at least one of a language model, an acoustic model, a transcription model, a pruning model, a penalty model, or a grammar of a speech recognition component based at least partially on the at least one category of recognition error associated with the at least partially erroneous speech recognition result.

10. The system of claim 1 , wherein provide an indication of the at least one corrective action comprises:

provide at least one recommended action to at least partially correct at least one aspect of at least one of a language model, an acoustic model, a transcription model, a pruning model, a penalty model, or a grammar of the speech recognition component based at least partially on the at least one category of recognition error associated with the at least partially erroneous speech recognition result.

11. The system of claim 1 , wherein adjust at least one aspect of the speech recognition component based on the at least one corrective action comprises:

adjust at least one aspect of at least one of a language model, an acoustic model, a transcription model, a pruning model, a penalty model, or a grammar of a speech recognition component based at least partially on the at least one category of recognition error associated with the at least partially erroneous speech recognition result.

12. A system for diagnosing speech recognition errors, comprising:

at least one processing component; and

one or more media operably coupled to the at least one processing component and bearing one or more instructions that, when executed by the at least one processing component, perform operations including at least:

determine that a speech recognition result is at least partially erroneous;

perform a first error analysis of the at least partially erroneous speech recognition result to provide a first error analysis result;

perform a second error analysis of the at least partially erroneous speech recognition result to provide a second error analysis result;

determine at least one category of recognition error associated with the at least partially erroneous speech recognition result based on a combination of the first error analysis result and the second error analysis result, including determine that the at least one category of recognition error includes at least an acoustic model error and a language model error when

(a) the first error analysis result indicates that a reference language model score associated with a reference speech is lower than a recognition language model score associated with the at least partially erroneous speech recognition result; and

(b) the second error analysis result indicates that a reference acoustic model score associated with the reference speech is lower than a recognition acoustic model score associated with the at least partially erroneous speech recognition result;

determine at least one corrective action to at least partially correct at least one aspect of a speech recognition component based at least partially on the at least one category of recognition error associated with the at least partially erroneous speech recognition result; and

at least one of:

provide an indication of the at least one corrective action; or

adjust at least one aspect of the speech recognition component based on the at least one corrective action.

13. The system of claim 1 , wherein determine at least one corrective action to at least partially correct at least one aspect of a speech recognition component comprises determine at least one corrective action to at least partially correct at least one aspect of an acoustic model of the speech recognition component.

14. A system for diagnosing speech recognition errors, comprising:

at least one processing component; and

one or more media operably coupled to the at least one processing component and bearing one or more instructions that, when executed by the at least one processing component, perform operations including at least:

determine that a speech recognition result is at least partially erroneous;

perform a first error analysis of the at least partially erroneous speech recognition result to provide a first error analysis result;

perform a second error analysis of the at least partially erroneous speech recognition result to provide a second error analysis result;

determine at least one category of recognition error associated with the at least partially erroneous speech recognition result based on a combination of the first error analysis result and the second error analysis result, including determine that the at least one category of recognition error includes at least an language model error and a pruning model error when

(a) the first error analysis result indicates that a reference language model score associated with a reference speech is lower than a recognition language model score associated with the at least partially erroneous speech recognition result; and

(b) the second error analysis result indicates that a reference acoustic model score associated with the reference speech is higher than a recognition acoustic model score associated with the at least partially erroneous speech recognition result;

determine at least one corrective action to at least partially correct at least one aspect of a speech recognition component based at least partially on the at least one category of recognition error associated with the at least partially erroneous speech recognition result; and

at least one of:

provide an indication of the at least one corrective action; or

adjust at least one aspect of the speech recognition component based on the at least one corrective action.

15. A system for diagnosing speech recognition errors, comprising:

at least one processing component; and

one or more media operably coupled to the at least one processing component and bearing one or more instructions that, when executed by the at least one processing component, perform operations including at least:

determine that a speech recognition result is at least partially erroneous;

perform a first error analysis of the at least partially erroneous speech recognition result to provide a first error analysis result;

perform a second error analysis of the at least partially erroneous speech recognition result to provide a second error analysis result;

determine at least one category of recognition error associated with the at least partially erroneous speech recognition result based on a combination of the first error analysis result and the second error analysis result, including determine that the at least one category of recognition error includes at least a penalty model error when

(a) the first error analysis result indicates that a reference language model score associated with a reference speech is higher than a recognition language model score associated with the at least partially erroneous speech recognition result; and

(b) the second error analysis result indicates that a reference acoustic model score associated with the reference speech is higher than a recognition acoustic model score associated with the at least partially erroneous speech recognition result;

determine at least one corrective action to at least partially correct at least one aspect of a speech recognition component based at least partially on the at least one category of recognition error associated with the at least partially erroneous speech recognition result; and

at least one of:

provide an indication of the at least one corrective action; or

adjust at least one aspect of the speech recognition component based on the at least one corrective action.

16. A method for diagnosing speech recognition errors, comprising:

performing one or more speech recognition operations to provide a speech recognition result;

performing a first error analysis of the speech recognition result to provide a first error analysis result;

performing a second error analysis of the speech recognition result to provide a second error analysis result;

determining at least one corrective action to at least partially increase an operability of at least one of the one or more speech recognition operations based on a combination of at least the first error analysis result and the second error analysis result;

determining at least one corrective action to at least partially correct at least one aspect of a speech recognition component based at least partially on at least one category of recognition error associated with the at least partially erroneous speech recognition result, including determining that the at least one category of recognition error includes at least an acoustic model error when

(a) the first error analysis result indicates that a reference language model score associated with a reference speech is higher than a recognition language model score associated with the at least partially erroneous speech recognition result; and

(b) the second error analysis result indicates that a reference acoustic model score associated with the reference speech is lower than a recognition acoustic model score associated with the at least partially erroneous speech recognition result; and

at least one of:

providing an indication of the at least one corrective action; or

adjusting at least one aspect of the speech recognition component based on the at least one corrective action.

17. The method of claim 16 , wherein adjusting at least one aspect of the speech recognition component based on the at least one corrective action comprises:

adjusting at least one aspect of at least one of a language model, an acoustic model, a transcription model, a pruning model, a penalty model, or a grammar of a speech recognition component based at least partially on the determined at least one corrective action.

18. The method of claim 16 , wherein the one or more instructions are configured wherein:

performing a first error analysis includes at least performing at least one language model scoring operation; and

performing a second error analysis includes at least performing at least one acoustic model scoring operation.

19. The method of claim 16 , wherein determining at least one corrective action to at least partially increase an operability of at least one of the one or more speech recognition operations based on a combination of at least the first error analysis result and the second error analysis result comprises:

determining at least one corrective action to at least one of reduce a speech recognition error of at least one of the one or more speech recognition operations, increase a computational efficiency of at least one of the one or more speech recognition operations, or reduce a resource usage of at least one of the one or more speech recognition operations.

20. A system for diagnosing a speech recognition error, comprising:

one or more processing devices that, when configured by one or more executable instructions, are configured as:

circuitry for performing at least one first error analysis operation on a speech recognition result generated by a speech recognition component to provide at least one first error analysis result;

circuitry for performing at least one second error analysis operation on the speech recognition result to provide at least one second error analysis result;

circuitry for determining, based on a combination of at least the first error analysis result and the second error analysis result, that an acoustic model error is indicated when

(a) the first error analysis result indicates that a reference language model score associated with a reference speech is higher than a recognition language model score associated with the at least partially erroneous speech recognition result; and

(b) the second error analysis result indicates that a reference acoustic model score associated with the reference speech is lower than a recognition acoustic model score associated with the at least partially erroneous speech recognition result;

circuitry for determining at least one corrective action to at least partially increase an operability of at least one speech recognition operation of the speech recognition component; and

at least one of:

circuitry for providing an indication of the at least one corrective action; or

circuitry for adjusting at least one aspect of the speech recognition component based on the at least one corrective action.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE SIGNATURE DATE OF MARK HANSON IS LISTED AS "02/27/2015" AND SHOULD INSTEAD BE "02/23/2015" PREVIOUSLY RECORDED ON REEL 035074 FRAME 0276. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNOR HEREBY SELLS, ASSIGNS AND TRANSFERS TO THE ASSIGNEE, ASSIGNOR'S ENTIRE AND EXCLUSIVE RIGHTS, TITLE AND INTEREST. Recorded Jan 20, 2016
From: KUO, SHIUN-ZU; REUTTER, THOMAS C.; GONG, YIFAN; TIAN, YE; CHANG, SHUANGYU; HAMAKER, JONATHAN; HANSON, MARK; MIAO, QI; TU, YUANCHENG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 037751/0142 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2015
From: KUO, SHIUN-ZU; REUTTER, THOMAS C.; GONG, YIFAN; TIAN, YE; CHANG, SHUANGYU; HAMAKER, JONATHAN; HANSON, MARK; MIAO, QI; TU, YUANCHENG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 035084/0276 →
Continuity (1)
Related Publication 20160253989A1 · Sep 1, 2016
Cited By (4)
US 12,328,181 US 12,424,236 US 12,681,587 US 12,694,865