System and method for providing a compensated speech recognition model for speech recognition
View Patent ↗An automatic speech recognition (ASR) system and method is provided for controlling the recognition of speech utterances generated by an end user operating a communications device. The ASR system and method can be used with a communications device that is used in a communications network. The ASR system can be used for ASR of speech utterances input into a mobile device, to perform compensating techniques using at least one characteristic and for updating an ASR speech recognizer associated with the ASR system by determined and using a background noise value and a distortion value that is based on the features of the mobile device. The ASR system can be used to augment a limited data input capability of a mobile device, for example, caused by limited input devices physically located on the mobile device.
1. An automatic speech recognition system, comprising:
a memory that stores a user profile having data related to user vocal information and value associated with a probability of the user being in a particular acoustic environment based on a time of day;
a controller coupled with the memory that receives the user profile and then compensates at least one speech recognition model based on the user profile;
a communication device that receives speech utterances from the user over a network; and
a speech recognizer that recognizes the speech utterances by using the at least one compensated speech recognition model.
2. The automatic speech recognition system according to claim 1 , wherein user profile further includes transducer data related to a distortion value related to a transducer of a mobile communications device.
3. The automatic speech recognition system according to claim 1 , wherein the particular acoustic environmental data includes a background noise value that corresponds to an operating environment of a mobile communications device.
4. The automatic speech recognition system according to claim 1 , wherein the vocal information includes a distortion value related to the user associated with a mobile communications, device.
5. The automatic speech recognition system according to claim 1 , wherein a personal computer is used provide the data of the particular acoustic environmental.
6. The automatic speech recognition system according to claim 1 , wherein a personal digital assistant is used to provide the data of the particular acoustic environmental.
7. The automatic speech recognition system according to claim 1 , wherein the data of the particular acoustic environmental is provided through a satellite communications system.
8. The automatic speech recognition system according to claim 1 , wherein the speech recognizer is a network server using a hidden Markov model.
9. The automatic speech recognition system according to claim 1 , wherein the controller is a network server that includes a pronunciation circuit, an environment-transducer-speaker circuit and a feature space circuit.
10. The automatic speech recognition system according to claim 8 , wherein the network server updates the at least one speech recognition model and a pronunciation model to reflect a specific type of communications device.
11. The automatic speech recognition system according to claim 1 , wherein the memory further stores personal account information that includes administrative information relating to the user.
12. The automatic speech recognition system according to claim 1 , wherein the communications device can be configured by the user to select a specific speech recognition network.
13. A controller used in an automatic speech recognition system, comprising:
a receiving section that receives speech utterances over a network from a user;
a first section that determines user profile data related to user vocal information and value associated with a probability of the user being in a particular acoustic environment based on a time of day; and
a second section that compensates a speech recognition model for recognizing the speech utterances based on the user profile data.
14. The controller according to claim 13 , wherein the controller identifies a user by a radio frequency identification tag.
15. The controller according to claim 13 , wherein the acoustic environmental data is determined using at least one microphone in the user's environment.
16. The controller according to claim 13 , wherein the acoustic environmental data is determined using a plurality of microphones that are selectively initiated as the user walks in between the plurality of microphones.
17. The controller according to claim 13 , where the user profile data further includes transducer data related to a distortion value based on a difference between an actual transducer in the mobile device and a response characteristic of a transducer used to train the speech recognition model.
18. The controller according to claim 13 , wherein the vocal information represents a variability that exists in vocal tract shapes among speakers of a group.
19. The controller according to claim 13 , wherein the controller communicates with a memory that stores various acoustic environmental models and various features of a specific type of mobile device.
20. The controller according to claim 19 , wherein a third section stores personal account information for each end user.
21. A method of using an automatic-speech recognition system, comprising:
speech utterances over a network;
determining user profile data related to user vocal information and a value associated with a probability of the user being in a particular acoustic environment based on a time of day;
compensating a speech recognition model based on the user profile data; and
recognizing the speech utterances using the compensated speech recognition model.
22. The method according to claim 21 , wherein the user profile further includes transducer data related to a distortion value related to a transducer used in a mobile device.
23. The method according to claim 22 , wherein the user profile further includes data related to the acoustic environmental data includes a background noise value that corresponds to an operating environment of a mobile communications device.
24. The method according to claim 21 , wherein the data of the particular acoustic environmental is received from a cellular telephone.
25. The method according to claim 21 , wherein the data of the particular acoustic environmental is received from a personal digital assistant.
26. The method according to claim 21 , wherein the data of the particular acoustic environmental is received via a satellite communications system.
27. The method according to claim 21 , wherein the speech recognition model is a hidden Markov model.
28. The method according to claim 23 , wherein determining the acoustic environmental data is performed using a network server.
29. The method according to claim 23 , wherein the acoustic environmental data is determined using at least one microphone in the user's environment.
30. The method according to claim 22 , wherein the user profile includes data related to a transducer and a distortion value is determined based on a difference between an actual transducer in the mobile device and a response characteristic of a transducer used to train the speech recognition model.
31. The method according to claim 21 , further comprising updating the speech recognition model and a pronunciation model to reflect a specific type of mobile communications device.
32. The method according to claim 21 , further comprising configuring the communications device to select a specific speech recognition network.