IP Library › Granted Patent US 7,016,849
Granted Patent B2
US 7,016,849 · App. 10/105,890 · Granted Mar 21, 2006

Method and apparatus for providing speech-driven routing between spoken language applications

Assignee: SRI International
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,016,849
App. No.
10/105,890
Granted
Mar 21, 2006
Kind
B2
Abstract

An apparatus and a concomitant method for speech recognition. In one embodiment, a distributed speech recognition system provides speech-driven control and remote service access. The distributed speech recognition system comprises a client device and a central server, where the client device is equipped with two speech recognition modules: a foreground speech recognizer and a background speech recognizer. The foreground speech recognizer is implementing a particular spoken language application (SLA) to handle a particular task, whereas the background speech recognizer is monitoring a change in the topic and/or a change in the intent of the user. Upon detection of a change in topic or intent of the user, the background speech recognizer will effect the routing to a new SLA to address the new topic or intent.

Claims (135)

1. Method for performing speech recognition, said method comprising the steps of:

(a) receiving a speech signal from a user;

(b) performing speech recognition on said speech signal in accordance with a first speech recognizer to produce a recognizable text signal, wherein said speech recognizer employs a first language model;

(c) performing speech recognition on said speech signal in parallel in accordance with a second speech recognizer for detecting a change of topic; and

(d) forwarding a second language model to said first speech recognizer in response to said detected change of topic by said second speech recognizer.

2. The method of claim 1 , wherein said speech signal is received locally from said user via a client device.

3. The method of claim 2 , further comprising the step of:

(b′) adapting said performance of speech recognition by said first speech recognizer based on at least one local parameter.

4. The method of claim 2 , further comprising the step of:

(c′) forwarding a signal to a central server in accordance with said detected change of topic, where said second language model is provided by said central server.

5. The method of claim 3 , wherein said at least one local parameter is representative of an environmental noise.

6. The method of claim 3 , wherein said at least one local parameter is representative of an acoustic environment.

7. The method of claim 3 , wherein said at least one local parameter is representative of a pronunciation of said user.

8. The method of claim 1 , further comprising the step of:

(e) storing at least a portion of said second language model in a cache of said client device.

9. The method of claim 1 , wherein said second speech recognizer performs said speech recognition step (c) by employing a language model that employs a catch all model.

10. The method of claim 9 , wherein said catch all model is employed to model at least one word that is not defined within a vocabulary of said language model of said second speech recognizer.

11. Method for performing speech recognition, said method comprising the steps of:

(a) receiving a speech signal from a user;

(b) performing speech recognition on said speech signal in accordance with a first speech recognizer to produce a recognizable text signal, wherein said speech recognizer employs a first language model;

(c) performing topic spotting from said recognizable text signal in accordance with a second speech recognizer for detecting a change of topic; and

(d) forwarding a second language model to said first speech recognizer in response to a detected change of topic by said second speech recognizer.

12. The method of claim 11 , wherein said speech signal is received locally from said user via a client device.

13. The method of claim 12 , further comprising the step of:

(b′) adapting said performance of speech recognition by said first speech recognizer based on at least one local parameter.

14. The method of claim 12 , further comprising the step of:

(c′) forwarding a signal to a central server in accordance with said detected change of topic, where said second language model is provided by said central server.

15. The method of claim 13 , wherein said at least one local parameter is representative of an environmental noise.

16. The method of claim 13 , wherein said at least one local parameter is representative of an acoustic environment.

17. The method of claim 13 , wherein said at least one local parameter is representative of a pronunciation of said user.

18. The method of claim 11 , further comprising the step of:

(e) storing at least a portion of said second language model in a cache of said client device.

19. The method of claim 11 , wherein said second speech recognizer performs said speech recognition step (c) by employing a language model that employs a catch all model.

20. The method of claim 19 , wherein said catch all model is employed to model at least one word that is not defined within a vocabulary of said language model of said second speech recognizer.

21. Method for performing speech recognition, said method comprising the steps of:

(a) receiving a speech signal from a user;

(b) performing speech recognition on said speech signal in accordance with a first speech recognizer to produce a recognizable text signal, wherein said speech recognizer employs a first language model;

(c) performing speech recognition on said speech signal in parallel in accordance with a second speech recognizer for detecting a change of topic; and

(d) updating said first language model with a second language model of said first speech recognizer in response to a detected change of topic by said second speech recognizer.

22. The method of claim 21 , wherein said speech signal is received locally from said user via a client device.

23. The method of claim 22 , further comprising the step of:

(b′) adapting said performance of speech recognition by said first speech recognizer based on at least one local parameter.

24. The method of claim 22 , further comprising the step of:

(c′) forwarding a signal to a central server in accordance with said detected change of topic, where said second language model is provided by said central server.

25. The method of claim 23 , wherein said at least one local parameter is representative of an environmental noise.

26. The method of claim 23 , wherein said at least one local parameter is representative of an acoustic environment.

27. The method of claim 23 , wherein said at least one local parameter is representative of a pronunciation of said user.

28. The method of claim 21 , further comprising the step of:

(e) storing at least a portion of said second language model in a cache of said client device.

29. The method of claim 21 , wherein said second speech recognizer performs said speech recognition step (c) by employing a language model that employs a catch all model.

30. The method of claim 29 , wherein said catch all model is employed to model at least one word that is not defined within a vocabulary of said language model of said second speech recognizer.

31. Method for performing speech recognition, said method comprising the steps of:

(a) receiving a speech signal from a user;

(b) performing speech recognition on said speech signal in accordance with a first speech recognizer to produce a recognizable text signal, wherein said speech recognizer employs a first language model;

(c) performing speech recognition on said speech signal in parallel in accordance with a second speech recognizer for detecting a change of intent; and

(d) forwarding a second language model to said first speech recognizer in response to said detected change of intent by said second speech recognizer.

32. The method of claim 31 , wherein said speech signal is received locally from said user via a client device.

33. The method of claim 32 , further comprising the step of:

(b′) adapting said performance of speech recognition by said first speech recognizer based on at least one local parameter.

34. The method of claim 32 , further comprising the step of:

(c′) forwarding a signal to a central server in accordance with said detected change of intent, where said second language model is provided by said central server.

35. The method of claim 33 , wherein said at least one local parameter is representative of an environmental noise.

36. The method of claim 33 , wherein said at least one local parameter is representative of an acoustic environment.

37. The method of claim 33 , wherein said at least one local parameter is representative of a pronunciation of said user.

38. The method of claim 31 , further comprising the step of:

(e) storing at least a portion of said second language model in a cache of said client device.

39. The method of claim 31 , wherein said second speech recognizer performs said speech recognition step (c) by employing a language model that employs a catch all model.

40. The method of claim 39 , wherein said catch all model is employed to model at least one word that is not defined within a vocabulary of said language model of said second speech recognizer.

41. Method for performing speech recognition, said method comprising the steps of:

(a) receiving a speech signal from a user;

(b) performing speech recognition on said speech signal in accordance with a first spoken language application to produce a recognizable text signal, wherein said first spoken language application employs a first language model;

(c) performing topic spotting from said recognizable text signal in accordance with a second spoken language application for detecting a change of topic; and

(d) forwarding a second language model to said first spoken language application in response to a detected change of topic by said second spoken language application.

42. The method of claim 41 , wherein said speech signal is received locally from said user via a client device.

43. The method of claim 42 , further comprising the step of:

(b′) adapting said performance of speech recognition by said first spoken language application based on at least one local parameter.

44. The method of claim 42 , further comprising the step of:

(c′) forwarding a signal to a central server in accordance with said detected change of topic, where said second language model is provided by said central server.

45. The method of claim 43 , wherein said at least one local parameter is representative of an environmental noise.

46. The method of claim 43 , wherein said at least one local parameter is representative of an acoustic environment.

47. The method of claim 43 , wherein said at least one local parameter is representative of a pronunciation of said user.

48. The method of claim 41 , further comprising the step of:

(e) storing at least a portion of said second language model in a cache of said client device.

49. The method of claim 41 , wherein said second spoken language application performs said topic spotting step (c) by employing a language model that employs a catch all model.

50. The method of claim 49 , wherein said catch all model is employed to model at least one word that is not defined within a vocabulary of said language model of said second spoken language application.

51. A client device for performing speech recognition, said client device comprising:

means for receiving a speech signal from a user;

means for performing speech recognition on said speech signal in accordance with a first speech recognizer to produce a recognizable text signal, wherein said speech recognizer employs a first language model;

means for performing a speech recognition on said speech signal in parallel in accordance with a second speech recognizer for detecting a change of topic; and

means for forwarding a second language model to said first speech recognizer in response to said detected change of topic by said second speech recognizer.

52. A client device for performing speech recognition, said client device comprising:

means for receiving a speech signal from a user;

means for performing speech recognition on said speech signal in accordance with a first speech recognizer to produce a recognizable text signal, wherein said speech recognizer employs a first language model;

means for performing topic spotting from said recognizable text signal in accordance with a second speech recognizer for detecting a change of topic; and

means for forwarding a second language model to said first speech recognizer in response to said detected change of topic by said second speech recognizer.

53. A client device for performing speech recognition, said client device comprising:

means for receiving a speech signal from a user;

means for performing speech recognition on said speech signal in accordance with a first speech recognizer to produce a recognizable text signal, wherein said speech recognizer employs a first language model;

means for performing speech recognition on said speech signal in parallel in accordance with a second speech recognizer for detecting a change of topic; and

means for updating said first language model with a second language model of said first speech recognizer in response to a detected change of topic by said second speech recognizer.

54. A client device for performing speech recognition, said client device comprising:

means for receiving a speech signal from a user;

means for performing speech recognition on said speech signal in accordance with a first speech recognizer to produce a recognizable text signal, wherein said speech recognizer employs a first language model;

means performing a speech recognition on said speech signal in parallel in accordance with a second speech recognizer for detecting a change of intent; and

means forwarding a second language model to said first speech recognizer in response to said detected change of intent by said second speech recognizer.

55. A client device for performing speech recognition, said client device comprising:

means for receiving a speech signal from a user;

means for performing speech recognition on said speech signal in accordance with a first spoken language application to produce a recognizable text signal, wherein said first spoken language application employs a first language model;

means for performing topic spotting from said recognizable text signal in accordance with a second spoken language application for detecting a change of topic; and

means for forwarding a second language model to said first spoken language application in response to a detected change of topic by said second spoken language application.

56. A computer-readable medium having stored thereon a plurality of instructions, the plurality of instructions including instructions which, when executed by a processor, cause the processor to perform the steps comprising of:

(a) receiving a speech signal from a user;

(b) performing speech recognition on said speech signal in accordance with a first speech recognizer to produce a recognizable text signal, wherein said speech recognizer employs a first language model;

(c) performing a speech recognition on said speech signal in parallel in accordance with a second speech recognizer for detecting a change of topic; and

(d) forwarding a second language model to said first speech recognizer in response to said detected change of topic by said second speech recognizer.

57. A computer-readable medium having stored thereon a plurality of instructions, the plurality of instructions including instructions which, when executed by a processor, cause the processor to perform the steps comprising of:

(a) receiving a speech signal from a user;

(b) performing speech recognition on said speech signal in accordance with a first speech recognizer to produce a recognizable text signal, wherein said speech recognizer employs a first language model;

(c) performing topic spotting from said recognizable text signal in accordance with a second speech recognizer for detecting a change of topic; and

(d) forwarding a second language model to said first speech recognizer in response to a detected change of topic by said second speech recognizer.

58. A computer-readable medium having stored thereon a plurality of instructions, the plurality of instructions including instructions which, when executed by a processor, cause the processor to perform the steps comprising of:

(a) receiving a speech signal from a user;

(b) performing speech recognition on said speech signal in accordance with a first speech recognizer to produce a recognizable text signal, wherein said speech recognizer employs a first language model;

(c) performing speech recognition on said speech signal in parallel in accordance with a second speech recognizer for detecting a change of topic; and

(d) updating said first language model with a second language model of said first speech recognizer in response to a detected change of topic by said second speech recognizer.

59. A computer-readable medium having stored thereon a plurality of instructions, the plurality of instructions including instructions which, when executed by a processor, cause the processor to perform the steps comprising of:

(a) receiving a speech signal from a user;

(b) performing speech recognition on said speech signal in accordance with a first speech recognizer to produce a recognizable text signal, wherein said speech recognizer employs a first language model;

(c) performing a speech recognition on said speech signal in parallel in accordance with a second speech recognizer for detecting a change of intent; and

(d) forwarding a second language model to said first speech recognizer in response to said detected change of intent by said second speech recognizer.

60. A computer-readable medium having stored thereon a plurality of instructions, the plurality of instructions including instructions which, when executed by a processor, cause the processor to perform the steps comprising of:

(a) receiving a speech signal from a user;

(b) performing speech recognition on said speech signal in accordance with a first spoken language application to produce a recognizable text signal, wherein said first spoken language application employs a first language model;

(c) performing topic spotting from said recognizable text signal in accordance with a second spoken language application for detecting a change of topic; and

(d) forwarding a second language model to said first spoken language application in response to a detected change of topic by said second spoken language application.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2002
From: ARNOLD, JAMES F.; FRANCO, HORACIO E.; ISRAEL, DAVID J.
To: SRI INTERNATIONAL
Reel/Frame 013044/0861 →
Continuity (1)
Related Publication 20030182131A1 · Sep 25, 2003