IP Library Granted Patent US 8,818,813
Granted Patent B2
US 8,818,813 · App. 14/046,765 · Granted Aug 26, 2014

Methods and system for grammar fitness evaluation as speech recognition error predictor

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,818,813
App. No.
14/046,765
Granted
Aug 26, 2014
Kind
B2
Abstract

A plurality of statements are received from within a grammar structure. Each of the statements is formed by a number of word sets. A number of alignment regions across the statements are identified by aligning the statements on a word set basis. Each aligned word set represents an alignment region. A number of potential confusion zones are identified across the statements. Each potential confusion zone is defined by words from two or more of the statements at corresponding positions outside the alignment regions. For each of the identified potential confusion zones, phonetic pronunciations of the words within the potential confusion zone are analyzed to determine a measure of confusion probability between the words when audibly processed by a speech recognition system during the computing event. An identity of the potential confusion zones across the statements and their corresponding measure of confusion probability are reported to facilitate grammar structure improvement.

Claims (40)

1. A method, comprising:

performing a word-level alignment process between a pair of statements to generate a word-level alignment sequence for the pair of statements, each statement in the pair of statements including one or more words that articulate a message for a computer application upon recognition by a speech recognition system;

identifying each potential confusion zone within the word-level alignment sequence generated for the pair of statements; and

determining a probability of the speech recognition system confusing words within each identified potential confusion zone within the word-level alignment sequence generated for the pair of statements,

wherein the operations of the method are performed by a computer processor.

2. The method as recited in claim 1 , wherein the word-level alignment sequence is characterized in units of a matching element, a substitution element, an insertion element, and a deletion element.

3. The method as recited in claim 2 , wherein each matching element corresponds to one or more words in a first statement identical to one or more words in a second statement, the first and second statements composing the pair of statements,

wherein each substitution element corresponds to one or more words in the first statement substituted with one or more different words in the second statement,

wherein each insertion element corresponds to one or more words in the second statement not present in the first statement, and

wherein each deletion element corresponds to one or more words in the first statement not present in the second statement.

4. The method as recited in claim 3 , wherein performing the word-level alignment process between the pair of statements includes maximizing a number of the matching element and minimizing a combined number of the substitution, insertion, and deletion elements.

5. The method as recited in claim 3 , wherein identifying each potential confusion zone within the word-level alignment sequence includes identifying each substitution element, insertion element, and deletion element as a respective potential confusion zone.

6. The method as recited in claim 5 , wherein determining the probability of the speech recognition system confusing words within a given potential confusion zone includes performing a phoneme-level alignment across phonemes of the one or more words of each of the pair of statements within the given potential confusion zone, and computing a phonetic accuracy value for the given potential confusion zone based on the phoneme-level alignment.

7. The method as recited in claim 6 , wherein a phoneme is a minimal distinct unit of a sound system of a language.

8. The method as recited in claim 6 , wherein performing the phoneme-level alignment includes determining a best overall alignment of identical phonemes of the one or more words of the first statement within the given potential confusion zone with the one or more words of the second statement within the given potential confusion zone.

9. The method as recited in claim 8 , wherein the best overall alignment of identical phonemes corresponds to a maximum number of aligned identical phonemes between the one or more words of the first statement within the given potential confusion zone and the one or more words of the second statement within the given potential confusion zone.

10. The method as recited in claim 6 , wherein the phonetic accuracy value corresponds to a measure of confusion probability between the one or more words of the first statement within the given potential confusion zone and the one or more words of the second statement within the given potential confusion zone when audibly processed by the speech recognition system.

11. The method as recited in claim 10 , further comprising:

comparing the phonetic accuracy value for the given potential confusion zone to a confusion probability threshold value to determine whether or not the probability of the speech recognition system confusing words within the given potential confusion zone is of concern; and

reporting the given potential confusion zone as of concern when the phonetic accuracy value for the given potential confusion zone is greater than or equal to the confusion probability threshold value.

12. The method as recited in claim 11 , further comprising:

applying the method to a plurality of statements defined for speech recognition by the speech recognition system for the computer application, such that each different combination of two statements within the plurality of statements is processed as the pair of statements within the method.

13. The method as recited in claim 12 , wherein the method is performed without auditory input.

14. The method as recited in claim 1 , wherein the method is performed without auditory input.

15. A non-transitory data storage device having program instructions stored thereon for a system for grammar fitness evaluation, comprising:

program instructions for a word-level alignment module defined to perform a word-level alignment process between a pair of statements to generate a word-level alignment sequence for the pair of statements, each statement in the pair of statements including one or more words that articulate a message for a computer application upon recognition by a speech recognition system;

program instructions for a confusion zone identification module defined to identify each potential confusion zone within the word-level alignment sequence generated for the pair of statements; and

program instructions for a confusion probability analysis module defined to determine a probability of the speech recognition system confusing words within each identified potential confusion zone within the word-level alignment sequence generated for the pair of statements.

16. The non-transitory data storage device as recited in claim 15 , wherein the word-level alignment sequence is characterized in units of a matching element, a substitution element, an insertion element, and a deletion element,

each matching element corresponding to one or more words in a first statement identical to one or more words in a second statement, the first and second statements composing the pair of statements,

each substitution element corresponding to one or more words in the first statement substituted with one or more different words in the second statement,

each insertion element corresponding to one or more words in the second statement not present in the first statement, and

each deletion element corresponding to one or more words in the first statement not present in the second statement.

17. The non-transitory data storage device as recited in claim 16 , wherein the confusion zone identification module is defined to identify each substitution element, insertion element, and deletion element as a respective potential confusion zone within the word-level alignment sequence.

18. The non-transitory data storage device as recited in claim 17 , wherein the confusion probability analysis module is defined to perform a phoneme-level alignment across phonemes of the one or more words of each of the pair of statements within the given potential confusion zone, and based on the phoneme-level alignment determine a confusion probability between the one or more words of the first statement within the given potential confusion zone and the one or more words of the second statement within the given potential confusion zone.

19. The non-transitory data storage device as recited in claim 18 , wherein the confusion probability analysis module is defined to perform the phoneme-level alignment process by determining a best overall alignment of identical phonemes of the one or more words of the first statement within the given potential confusion zone with the one or more words of the second statement within the given potential confusion zone,

wherein the confusion probability analysis module is defined to compute a phonetic accuracy value for the given potential confusion zone based on the determined confusion probability, and

wherein the confusion probability analysis module is defined to compare the phonetic accuracy value for the given potential confusion zone to a confusion probability threshold value to determine whether or not the probability of the speech recognition system confusing words within the given potential confusion zone is of concern.

20. The non-transitory data storage device as recited in claim 19 , further comprising:

program instructions for an output module defined to report the given potential confusion zone as of concern when the phonetic accuracy value for the given potential confusion zone is greater than or equal to the confusion probability threshold value.