HMM DECODING WITH ACOUSTIC MODEL COMPENSATION FOR PHONEME MODELING AND PRONUNCIATION MODELINGS
Described are techniques to recognize spoken wake word (WW) or command for human-machine interface using a speech recognition system that does not require any WW or command-matching speech data for training. The system uses the text or grapheme representation of the WW or commands for training before deployment. The technique includes generating a matrix of statistics characterizing phonetic modeling of an acoustic model that distinguishes speech signals according to a plurality of acoustic units of a language. It further includes constructing a speech recognition model based on the matrix of statistics to recognize a target phrase in speech. It further includes processing an audio signal based on the acoustic model and the speech recognition model to detect a presence of the target phrase.
1 . A method of speech recognition by a device, the method comprising:
generating a matrix of statistics characterizing phonetic modeling of an acoustic model that distinguishes speech signals according to a plurality of acoustic units of a language;
constructing a speech recognition model based on the matrix of statistics to recognize a target phrase in speech; and
processing an audio signal based on the acoustic model and the speech recognition model to detect a presence of the target phrase.
2 . The method of claim 1 , wherein generating the matrix of statistics characterizing phonetic modeling of the acoustic model comprises:
providing speech signals from an annotated speech database to the acoustic model, wherein the speech signals includes annotations indicating a sequence of true acoustic units of the speech signals and time boundaries of each of the true acoustic units;
extracting, by the acoustic model, feature vectors from the speech signals;
processing, by the acoustic model, the feature vectors to determine a probability estimate of a presence of each one of the plurality of acoustic units in each of a plurality of time windows of the speech signals based on the annotations; and
analyzing, by the acoustic model, probability estimates of the presence of the plurality of acoustic units over the plurality of time windows to generate the matrix of statistics.
3 . The method of claim 2 , wherein processing the feature vectors comprises:
determining, for a target true acoustic unit within a time boundary of the speech signals, a probability estimate of the presence of each one of the plurality of acoustic units in a time window whose center lies within the time boundary, wherein the probability estimate associated with the target true acoustic unit is a maximum value among the probability estimates associated with the plurality of acoustic units.
4 . The method of claim 3 , wherein analyzing probability estimates of the presence of the plurality of acoustic units comprises:
averaging corresponding probability estimates associated with the plurality of acoustic units over the plurality of time windows when the target true acoustic unit is at the maximum value to generate a vector of average probability estimates associated with the target true acoustic unit;
repeating the averaging associated with a target true acoustic unit within each time boundary of the speech signals to generate a plurality of vectors of average probability estimates associated with a plurality of true acoustic units, wherein each one of the plurality of true acoustic units corresponds to each one of the plurality of acoustic units of the language; and
generating the matrix of statistics based on the plurality of vectors of average probability estimates associated with the plurality of true acoustic units.
5 . The method of claim 4 , wherein the matrix of statistics comprises:
the plurality of true acoustic units arranged on a first axis; and
the vector of average probability estimates associated with each one of the plurality of true acoustic units arranged on a second axis, wherein a maximum value among the vector of average probability estimates associated with each of the true acoustic units is on a diagonal of the matrix of statistics.
6 . The method of claim 4 , wherein the vector of average probability estimates associated with a given true acoustic unit of the plurality of true acoustic units indicates an average probability estimate of the presence of each of the acoustic units when the given true acoustic unit is at a maximum value.
7 . The method of claim 1 wherein constructing the speech recognition model comprises:
constructing the speech recognition model to recognize acoustically similar acoustic units based on the matrix of statistics, wherein the matrix of statistics comprises a probability estimate of a presence of each of the acoustically similar acoustic units when the acoustic model identifies an acoustic unit associated with a highest probability estimate as a most likely acoustic unit.
8 . The method of claim 7 , wherein constructing the speech recognition model to recognize acoustically similar acoustic units comprises:
determining a ratio of the probability estimate associated with each of the acoustically similar acoustic units and the highest probability estimate associated with the most likely acoustic unit; and
providing each of the ratio to increase a confidence when the acoustic model identifies the most likely acoustic unit in the audio signal.
9 . The method of claim 7 , wherein processing the audio signal based on the acoustic model and the speech recognition model comprises:
determining a degree of correlation between the probability estimate associated with each of the acoustically similar acoustic units and an observed probability of each of the acoustically similar acoustic units from the acoustic model to increase a confidence when the acoustic model identifies the most likely acoustic unit in the audio signal.
10 . The method of claim 7 , wherein the speech recognition model comprises a sequence decoding model, where the sequence decoding model includes a plurality of states, wherein each of the states represents an acoustic unit.
11 . The method of claim 10 , wherein the sequence decoding model comprises a Hidden Markov Model (HMM), and wherein the plurality of states comprises:
a plurality of intermediate states, wherein each of the intermediate states decodes an acoustic unit of the target phrase.
12 . The method of claim 11 , wherein one of the intermediate states decodes a plurality of alternate acoustic units to support alternate pronunciations based on the probability estimate associated with each of the acoustically similar acoustic units when said one of the intermediate states decode the most likely acoustic unit.
13 . The method of claim 12 , wherein the plurality of acoustic units is decoded by said one of the intermediate states based on a ranking of a ratio of the probability estimate associated with each of the acoustically similar acoustic units and the probability estimate associated with the most likely acoustic unit.
14 . The method of claim 11 , wherein processing the audio signal based on the acoustic model and the speech recognition model comprises:
determining a probability of observing one of the intermediate states based on a degree of correlation between the probability estimate associated with each of the acoustically similar acoustic units to a most likely acoustic unit for said state and an observed probability of each of the acoustically similar acoustic units for said state; and
determining a most likely path through the states of the HMM based on a probability of observing a plurality of the intermediate states.
15 . The method of claim 14 , wherein said state decodes a plurality of alternate acoustic units to support alternate pronunciations, and wherein determining the probability of observing one of the intermediate states for said state comprises:
determining a maximum of probabilities of observing said intermediate state for the plurality of alternative acoustic units decoded by said state.
16 . The method of claim 1 , wherein the acoustic units comprise:
phonemes of the language; and
a silence class.
17 . An apparatus comprising:
an input terminal configured to receive an audio signal from one or more microphones; and
a processing system configured to:
generate a matrix of statistics characterizing phonetic modeling of an acoustic model that distinguishes speech signals according to a plurality of acoustic units of a language;
construct a speech recognition model based on the matrix of statistics to recognize a target phrase in speech; and
process the audio signal based on the acoustic model and the speech recognition model to detect a presence of the target phrase.
18 . The apparatus of claim 17 , where to generate the matrix of statistics, the processing system is configured to:
extract features vectors from speech signals provided by an annotated speech database, wherein the speech signals includes annotations indicating a sequence of true acoustic units of the speech signals and time boundaries of each of the true acoustic units;
process the feature vectors to determine a probability estimate of a presence of each one of the plurality of acoustic units in each of a plurality of time windows of the speech signals based on the annotations; and
analyze probability estimates of the presence of the plurality of acoustic units over the plurality of time windows to generate the matrix of statistics.
19 . The apparatus of claim 18 , wherein the matrix of statistics indicates a probability estimate of a presence of each of the acoustic units when a most likely acoustic unit is associated with a highest probability estimate.
20 . An apparatus of claim 17 , wherein to construct the speech recognition model, the processing system is configured to:
construct the speech recognition model to recognize acoustically similar acoustic units based on the matrix of statistics, wherein the matrix of statistics comprises a probability estimate of a presence of each of the acoustically similar acoustic units when the acoustic model identifies an acoustic unit associated with a highest probability estimate as a most likely acoustic unit.