Multi-language speech recognition system
A speech recognition system includes distributed processing across a client and server for recognizing a spoken query by a user. A number of different speech models for different natural languages are used to support and detect a natural language spoken by a user. In some implementations an interactive electronic agent responds in the user's native language to facilitate an real-time, human like dialogue.
1 . A method of performing recognition of a speech utterance from a user with a distributed client-server system comprising the steps of:
(a) receiving user speech data from a client device in streaming packets through a network interface of a network server system, said speech data resulting from a first set of speech recognition operations being performed on the speech utterance by a client device;
(b) recognizing the speech utterance as well as a natural language used in said speech utterance using processing routines executing at said network server system which implement a second set of speech recognition operations;
(c) providing a response to the user in a same natural language as recognized in step (b).
2 . The method of claim 1 , wherein said response is an audible response from an electronic agent generated by a text to speech engine.
3 . The method of claim 1 , wherein an interactive electronic agent provides said response, which interactive electronic agent exhibits characteristics that are adjusted by said processing routines executing at said network server system based on a type of application interacting with the user.
4 . The method of claim 1 , wherein an interactive electronic agent provides said response, which interactive electronic agent exhibits characteristics that are adjusted by said processing routines executing at the network server system based on an identity of the user.
5 . The method of claim 1 , wherein an interactive electronic agent presented within a browser or graphical interface of the client system provides said response, which interactive electronic agent responds to user queries presented in speech form and assists the user to navigate and select items from an Internet web page.
6 . The method of claim 2 , wherein said interactive electronic agent further provides one or more specific suggested queries to the user.
7 . The method of claim 2 , wherein said interactive electronic agent further provides audible confirmation of selections made by the user.
8 . The method of claim 1 , wherein a plurality of Hidden Markov Models are used to recognize said speech utterance and said natural language of said speech utterance such that the system supports multiple natural languages.
9 . The method of claim 1 , further including a step: performing a natural language processing operation on said speech utterance to determine a meaning of a query presented by the user.
10 . The method of claim 9 , further including a step: forming a database query based on identifying said query presented by the user to retrieve a predetermined answer for said query.
11 . The method of claim 1 , wherein said processing routines at the network server system use one or more speech recognition models that are trained and optimized based on speech characteristics of a group of persons residing in geographical regions served by the distributed client—server system.
12 . The method of claim 1 , wherein said client device is a portable Internet based appliance.
13 . The method of claim 1 , further including a step: adjusting said second set of speech recognition operations based on an evaluation of resources available at the network server system and/or the client device.
14 . The method of claim 13 , further including a step: adjusting said first set of speech recognition operations based on an evaluation of resources available at the client device.
15 . The method of claim 1 , wherein said processing routines include software programs executable on said network server system.
16 . A method of performing recognition of a speech utterance from a user with a distributed client-server system comprising the steps of:
(a) receiving user speech data from a client device through a network interface of a network server system, said speech data constituting partially recognized speech derived by a client device from a speech utterance;
(b) completing recognition of the speech utterance and identifying a language therein using software routines executing at said network server system and a plurality of speech models associated with a plurality of languages;
(c) processing the speech utterance with one or more natural language operations to identify a meaning of the speech utterance;
(d) identifying a query presented by the user based on said meaning of the speech utterance;
(e) providing a response to the query in a same language as recognized in step (b).
17 . The method of claim 16 , further including a step: calibrating noise data present at the client device.
18 . The method of claim 16 , wherein said response is an audible response from an electronic agent generated by a text to speech engine.
19 . The method of claim 18 , further including a step: conducting an interactive dialog with the user using said electronic agent in response to further speech utterances.
20 . A system for recognizing a speech utterance from a user comprising:
(a) a first processing routine adapted to receive user speech data from a client device in streaming packets through a network interface of a network server system, said speech data resulting from a first set of speech recognition operations being performed on the speech utterance by the client device;
(b) a second processing routine adapted to recognize the speech utterance as well as a natural language used in said speech utterance by executing a second set of speech recognition operations;
(c) a third processing routine for providing a response to the user in a same natural language as recognized by said second software routine.
21 . The system of claim 20 , wherein a number of speech recognition operations performed on the network server system can be adjusted.
22 . The system of claim 20 , wherein said response is an audible response from an electronic agent generated by a text to speech engine.
23 . The system of claim 20 wherein an interactive electronic agent provides said response, which electronic agent exhibits characteristics that are adjusted by said third processing routine based on a type of application interacting with the user.
24 . The system of claim 20 , wherein an interactive electronic agent presented within a browser of the client system provides said response, which interactive electronic agent responds to user queries presented in speech form and assists the user to navigate and select items from an Internet web page.
25 . The system of claim 22 , wherein said electronic agent further provides one or more specific suggested queries to the user.
26 . The system of claim 22 , wherein said interactive electronic agent further provides audible confirmation of selections made by the user.
27 . The system of claim 20 wherein a plurality of Hidden Markov Models are used to recognize said speech utterance and said language of said speech utterance such that multiple natural languages are supported.
28 . The system of claim 20 , further including a natural language processing routine which operates on said speech utterance to determine a meaning of a query presented by the user.
29 . The system of claim 28 , further including a database interface which generates a database query based on identifying said query presented by the user to retrieve a predetermined answer for said query.
30 . The system of claim 20 , wherein said second processing routine uses one or more speech recognition models that are trained and optimized based on speech characteristics of a group of persons residing in geographical regions served by the distributed client-server system.
31 . The system of claim 20 , wherein said client device is a portable Internet based appliance.
32 . The system of claim 19 , further including a routine which adjusts said second set of speech recognition operations based on an evaluation of resources available at the network server system.
33 . The system of claim 32 , further including a routine which adjusts said first set of speech recognition operations based on an evaluation of resources available at the client device.
34 . The system of claim 20 , wherein speech recognition operations can be allocated between the client device and the network server system on a query-by-query basis.
35 . The system of claim 20 , wherein the network server system is a group of interlinked computing servers.
36 . The system of claim 20 , wherein said processing routines include software programs executable on said network server system.
37 . A system for recognizing a speech utterance from a user comprising:
(a) a first routine adapted to receive user speech data from a client device through a network interface of a network server system, said speech data constituting partially recognized speech derived by a client device from a speech utterance;
(b) a second routine executing at said network server system which is adapted to complete recognition of the speech utterance and to identify a natural language therein using a plurality of speech models associated with a plurality of natural languages;
(c) a third routine adapted to process the speech utterance with one or more natural language operations to identify a meaning of the speech utterance;
(d) a fourth routine adapted to identify a query presented by the user based on said meaning of the speech utterance;
(e) a fifth routine adapted to provide a response to the query in a same natural language as recognized in step (b).
38 . The system of claim 37 , wherein all of said routines are implemented as executable software programs on said network server system.
39 . The system of claim 37 , wherein said fifth routine interacts with a second user using a second natural language identified by said second routine.
40 . The system of claim 37 wherein said speech data is received continuously during a speech utterance.
41 . The system of claim 40 wherein said speech data is received continuously until silence is detected.