IP Library Granted Patent US 8,924,213
Granted Patent B2
US 8,924,213 · App. 13/544,331 · Granted Dec 30, 2014

Detecting potential significant errors in speech recognition results

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,924,213
App. No.
13/544,331
Granted
Dec 30, 2014
Kind
B2
Abstract

In some embodiments, the recognition results produced by a speech processing system (which may include two or more recognition results, including a top recognition result and one or more alternative recognition results) based on an analysis of a speech input, are evaluated for indications of potential significant errors. In some embodiments, the recognition results may be evaluated using one or more sets of words and/or phrases, such as pairs of words/phrases that may include words/phrases that are acoustically similar to one another and/or that, when included in a result, would change a meaning of the result in a manner that would be significant for a domain. The recognition results may be evaluated using the set(s) of words/phrases to determine, when the top result includes a word/phrase from a set of words/phrases, whether any of the alternative recognition results includes any of the other, corresponding words/phrases from the set.

Claims (92)

1. A method of processing results of a recognition by an automatic speech recognition (ASR) system on a speech input, the results comprising a first recognition result identified by the ASR system as most likely to be a correct recognition result for the speech input, the results further comprising at least one alternative recognition result identified by the ASR system, the method comprising:

determining whether the first recognition result includes a member of a set of words or phrases, each member of the set comprising a word or phrase and being associated with at least one other member of the set, and whether the at least one alternative recognition result includes any of the at least one other member associated with the member in the set; and

in response to determining that the first recognition result includes the member of the set of words or phrases and that the at least one alternative recognition result includes any of the at least one other member associated with the member in the set, triggering an alert,

wherein:

the at least one alternative recognition result comprises a second recognition result;

the member of the set is a first member of the set;

the at least one other member of the set associated with the member comprises a second member of the set;

the determining whether the first recognition result includes a member of the set and whether the at least one alternative recognition result includes any of the at least one other member comprises determining whether the ASR system recognized a segment of the speech input as the first member for the first recognition result and as the second member of the set for the second recognition result;

the results of the recognition are arranged as a lattice; and

the determining whether the ASR system recognized the segment of the speech input as the first member for the first recognition result and as the second member for the second recognition result comprises determining whether the lattice identifies the second member of the set as an alternative to the first member of the set.

2. A method of processing results of a recognition by an automatic speech recognition (ASR) system on a speech input, the results comprising a first recognition result identified by the ASR system as most likely to be a correct recognition result for the speech input, the results further comprising at least one alternative recognition result identified by the ASR system, the method comprising:

determining whether the first recognition result includes a member of a set of words or phrases, each member of the set comprising a word or phrase and being associated with at least one other member of the set, and whether the at least one alternative recognition result includes any of the at least one other member associated with the member in the set; and

in response to determining that the first recognition result includes the member of the set of words or phrases and that the at least one alternative recognition result includes any of the at least one other member associated with the member in the set, triggering an alert,

wherein:

the at least one alternative recognition result comprises a second recognition result;

the member of the set is a first member of the set;

the at least one other member of the set associated with the member comprises a second member of the set;

the determining whether the first recognition result includes a member of the set and whether the at least one alternative recognition result includes any of the at least one other member comprises determining whether the ASR system recognized a segment of the speech input as the first member for the first recognition result and as the second member of the set for the second recognition result; and

the determining whether the ASR system identified the segment of the speech input as the first member for the first recognition result and as the second member for the second recognition result comprises determining whether a time within the speech input associated with the first member for the first recognition result matches a time within the speech input associated with the second member of the set for the second recognition result.

3. The method of claim 2 , wherein:

a first member of the set is associated with a second member of the set with which the first member is acoustically-confusable and that, when substituted for the first member in a recognition result, changes a meaning of the recognition result; and

the determining whether the first recognition result includes a member of the set and whether the at least one alternative recognition result includes any of the at least one other member comprises determining whether the first recognition result includes the first member and whether the at least one alternative recognition result includes the second member.

4. The method of claim 3 , wherein the second member of the set, when substituted for the first member in a recognition result, changes a medical meaning of the recognition result.

5. A method of processing results of a recognition by an automatic speech recognition (ASR) system on a speech input, the results comprising a first recognition result identified by the ASR system as most likely to be a correct recognition result for the speech input, the results further comprising at least one alternative recognition result identified by the ASR system, the method comprising:

determining whether the first recognition result includes a member of a set of words or phrases, each member of the set comprising a word or phrase and being associated with at least one other member of the set, and whether the at least one alternative recognition result includes any of the at least one other member associated with the member in the set; and

in response to determining that the first recognition result includes the member of the set of words or phrases and that the at least one alternative recognition result includes any of the at least one other member associated with the member in the set, triggering an alert,

wherein:

at least one member of the set comprises a null word; and

the determining whether the first recognition result includes a member of the set and whether the at least one alternative recognition result includes any of the at least one other member comprises determining whether either of the first recognition result or the at least one alternative recognition result comprises the null word.

6. The method of claim 5 , wherein the determining whether either of the first recognition result or the at least one alternative recognition result comprises the null word comprises determining that the first recognition result or the at least one alternative recognition result includes the null word without evaluating words of the first recognition result or the at least one alternative recognition result.

7. At least one computer-readable storage medium having encoded thereon computer-executable instructions that, when executed by at least one computer, cause the at least one computer to carry out a method of processing results of a recognition by an automatic speech recognition (ASR) system on an utterance, the results comprising a first recognition result identified by the ASR system as most likely to be a correct recognition result for the utterance, the results further comprising at least one alternative recognition result identified by the ASR system, the method comprising:

determining whether the first recognition result includes a member of a set of words or phrases, each member of the set comprising a word or phrase and being associated with at least one other member of the set, and whether the at least one alternative recognition result includes any of the at least one other member associated with the member in the set; and

in response to determining that the first recognition result includes the member of the set of words or phrases and that the at least one alternative recognition result includes any of the at least one other member associated with the member in the set, triggering an alert,

wherein:

the at least one alternative recognition result comprises a second recognition result;

the member of the set is a first member of the set;

the at least one other member of the set associated with the member comprises a second member of the set;

the determining whether the first recognition result includes a member of the set and whether the at least one alternative recognition result includes any of the at least one other member comprises determining whether the ASR system recognized a segment of the speech input as the first member for the first recognition result and as the second member of the set for the second recognition result;

the results of the recognition are arranged as a lattice; and

the determining whether the ASR system recognized the segment of the speech input as the first member for the first recognition result and as the second member for the second recognition result comprises determining whether the lattice identifies the second member of the set as an alternative to the first member of the set.

8. The at least one computer-readable storage medium of claim 7 , wherein:

a first member of the set is associated with a second member of the set with which the first member is acoustically-confusable and that, when substituted for the first member in a recognition result, changes a meaning of the recognition result; and

the determining whether the first recognition result includes a member of the set and whether the at least one alternative recognition result includes any of the at least one other member comprises determining whether the first recognition result includes the first member and whether the at least one alternative recognition result includes the second member.

9. At least one computer-readable storage medium having encoded thereon computer-executable instructions that, when executed by at least one computer, cause the at least one computer to carry out a method of processing results of a recognition by an automatic speech recognition (ASR) system on an utterance, the results comprising a first recognition result identified by the ASR system as most likely to be a correct recognition result for the utterance, the results further comprising at least one alternative recognition result identified by the ASR system, the method comprising:

determining whether the first recognition result includes a member of a set of words or phrases, each member of the set comprising a word or phrase and being associated with at least one other member of the set, and whether the at least one alternative recognition result includes any of the at least one other member associated with the member in the set; and

in response to determining that the first recognition result includes the member of the set of words or phrases and that the at least one alternative recognition result includes any of the at least one other member associated with the member in the set, triggering an alert,

wherein:

the at least one alternative recognition result comprises a second recognition result;

the member of the set is a first member of the set;

the at least one other member of the set associated with the member comprises a second member of the set;

the determining whether the first recognition result includes a member of the set and whether the at least one alternative recognition result includes any of the at least one other member comprises determining whether the ASR system recognized a segment of the speech input as the first member for the first recognition result and as the second member of the set for the second recognition result; and

the determining whether the ASR system identified the segment of the speech input as the first member for the first recognition result and as the second member for the second recognition result comprises determining whether a time within the speech input associated with the first member for the first recognition result matches a time within the speech input associated with the second member of the set for the second recognition result.

10. At least one computer-readable storage medium having encoded thereon computer-executable instructions that, when executed by at least one computer, cause the at least one computer to carry out a method of processing results of a recognition by an automatic speech recognition (ASR) system on an utterance, the results comprising a first recognition result identified by the ASR system as most likely to be a correct recognition result for the utterance, the results further comprising at least one alternative recognition result identified by the ASR system, the method comprising:

determining whether the first recognition result includes a member of a set of words or phrases, each member of the set comprising a word or phrase and being associated with at least one other member of the set, and whether the at least one alternative recognition result includes any of the at least one other member associated with the member in the set; and

in response to determining that the first recognition result includes the member of the set of words or phrases and that the at least one alternative recognition result includes any of the at least one other member associated with the member in the set, triggering an alert,

wherein:

at least one member of the set comprises a null word; and

the determining whether the first recognition result includes a member of the set and whether the at least one alternative recognition result includes any of the at least one other member comprises determining that the first recognition result or the at least one alternative recognition result comprises the null word without evaluating words of the first recognition result or the at least one alternative recognition result.

11. An apparatus comprising:

at least one processor; and

at least one storage medium having encoded thereon processor-executable instructions that, when executed by the at least one processor, cause the at least one processor to carry out a method of processing results of a recognition by an automatic speech recognition (ASR) system on an utterance, the results comprising a first recognition result identified by the ASR system as most likely to be a correct recognition result for the utterance, the results further comprising at least one alternative recognition result identified by the ASR system, the method comprising:

determining whether the first recognition result includes a member of a set of words or phrases, each member of the set comprising a word or phrase and being associated with at least one other member of the set, and whether the at least one alternative recognition result includes any of the at least one other member associated with the member in the set; and

in response to determining that the first recognition result includes the member of the set of words or phrases and that the at least one alternative recognition result includes any of the at least one other member associated with the member in the set, triggering an alert,

wherein:

the at least one alternative recognition result comprises a second recognition result;

the member of the set is a first member of the set;

the at least one other member of the set associated with the member comprises a second member of the set;

the determining whether the first recognition result includes a member of the set and whether the at least one alternative recognition result includes any of the at least one other member comprises determining whether the ASR system recognized a segment of the speech input as the first member for the first recognition result and as the second member of the set for the second recognition result;

the results of the recognition are arranged as a lattice; and

the determining whether the ASR system recognized the segment of the speech input as the first member for the first recognition result and as the second member for the second recognition result comprises determining whether the lattice identifies the second member of the set as an alternative to the first member of the set.

12. An apparatus comprising:

at least one processor; and

at least one storage medium having encoded thereon processor-executable instructions that, when executed by the at least one processor, cause the at least one processor to carry out a method of processing results of a recognition by an automatic speech recognition (ASR) system on an utterance, the results comprising a first recognition result identified by the ASR system as most likely to be a correct recognition result for the utterance, the results further comprising at least one alternative recognition result identified by the ASR system, the method comprising:

determining whether the first recognition result includes a member of a set of words or phrases, each member of the set comprising a word or phrase and being associated with at least one other member of the set, and whether the at least one alternative recognition result includes any of the at least one other member associated with the member in the set; and

in response to determining that the first recognition result includes the member of the set of words or phrases and that the at least one alternative recognition result includes any of the at least one other member associated with the member in the set, triggering an alert, wherein:

the at least one alternative recognition result comprises a second recognition result;

the member of the set is a first member of the set;

the at least one other member of the set associated with the member comprises a second member of the set;

the determining whether the first recognition result includes a member of the set and whether the at least one alternative recognition result includes any of the at least one other member comprises determining whether the ASR system recognized a segment of the speech input as the first member for the first recognition result and as the second member of the set for the second recognition result; and

the determining whether the ASR system identified the segment of the speech input as the first member for the first recognition result and as the second member for the second recognition result comprises determining whether a time within the speech input associated with the first member for the first recognition result matches a time within the speech input associated with the second member of the set for the second recognition result.

13. The apparatus of claim 12 , wherein:

a first member of the set is associated with a second member of the set with which the first member is acoustically-confusable and that, when substituted for the first member in a recognition result, changes a meaning of the recognition result; and

the determining whether the first recognition result includes a member of the set and whether the at least one alternative recognition result includes any of the at least one other member comprises determining whether the first recognition result includes the first member and whether the at least one alternative recognition result includes the second member.

14. The apparatus of claim 13 , wherein the second member of the set, when substituted for the first member in a recognition result, changes a medical meaning of the recognition result.

15. An apparatus comprising:

at least one processor; and

at least one storage medium having encoded thereon processor-executable instructions that, when executed by the at least one processor, cause the at least one processor to carry out a method of processing results of a recognition by an automatic speech recognition (ASR) system on an utterance, the results comprising a first recognition result identified by the ASR system as most likely to be a correct recognition result for the utterance, the results further comprising at least one alternative recognition result identified by the ASR system, the method comprising:

determining whether the first recognition result includes a member of a set of words or phrases, each member of the set comprising a word or phrase and being associated with at least one other member of the set, and whether the at least one alternative recognition result includes any of the at least one other member associated with the member in the set; and

in response to determining that the first recognition result includes the member of the set of words or phrases and that the at least one alternative recognition result includes any of the at least one other member associated with the member in the set, triggering an alert,

wherein:

at least one member of the set comprises a null word; and

the determining whether the first recognition result includes a member of the set and whether the at least one alternative recognition result includes any of the at least one other member comprises determining that the first recognition result or the at least one alternative recognition result comprises the null word without evaluating words of the first recognition result or the at least one alternative recognition result.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2012
From: GANONG, WILLIAM F., III; FLEMING, ROBERT; VEMULA, RAGHU
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 028861/0667 →