Method and apparatus for predicting word error rates from text
A method of modeling a speech recognition system includes decoding a speech signal produced from a training text to produce a sequence of predicted speech units. The training text comprises a sequence of actual speech units that is used with the sequence of predicted speech units to form a confusion model. In further embodiments, the confusion model is used to decode a text to identify an error rate that would be expected if the speech recognition system decoded speech based on the text.
1. A method of modeling a speech recognition system, the method comprising:
decoding a speech signal produced from a training text, the training text comprising a sequence of actual speech units to produce a sequence of predicted speech units;
constructing a confusion model based on the sequence of actual speech units and the sequence of predicted speech units; and
decoding a test text using the confusion model and a language model to generate at least one model-predicted sequence of speech units.
2. The method of claim 1 wherein constructing a confusion model comprises constructing a Hidden Markov Model.
3. The method of claim 2 wherein constructing a Hidden Markov Model comprises constructing a four state Hidden Markov Model.
4. The method of claim 3 further comprising generating a probability for each model-predicted sequence of speech units.
5. The method of claim 4 further comprising using the probability for a model-predicted sequence of speech units to identify word sequences in the test text that are likely to generate an erroneous model-predicted sequence of speech units.
6. The method of claim 1 further comprising:
decoding the test text using a first language model with the confusion model to form a first set of model-predicted speech units; and
decoding the test text using a second language model with the confusion model to form a second set of model-predicted speech units.
7. The method of claim 6 further comprising using the first and second sets of model-predicted speech units to compare the performance of the first language model to the performance of the second language model.
8. The method of claim 6 wherein the steps of decoding the test text using the first and second language model form part of a method of discriminative training of a language model.
9. The method of claim 1 wherein decoding the speech signal comprises using a training language model to decode the speech signal, wherein the training language model does not perform as well as the language model used to decode the test text.