IP Library Granted Patent US 9,082,408
Granted Patent B2
US 9,082,408 · App. 13/491,856 · Granted Jul 14, 2015

Speech recognition using loosely coupled components

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,082,408
App. No.
13/491,856
Granted
Jul 14, 2015
Kind
B2
Abstract

An automatic speech recognition system includes an audio capture component, a speech recognition processing component, and a result processing component which are distributed among two or more logical devices and/or two or more physical devices. In particular, the audio capture component may be located on a different logical device and/or physical device from the result processing component. For example, the audio capture component may be on a computer connected to a microphone into which a user speaks, while the result processing component may be on a terminal server which receives speech recognition results from a speech recognition processing server.

Claims (147)

1. A system comprising:

a first device including an audio capture component, the audio capture component comprising means for capturing an audio signal representing speech of a user to produce a captured audio signal;

a speech recognition processing component comprising means for performing automatic speech recognition on the captured audio signal to produce speech recognition results;

a second device including a result processing component;

a context sharing component comprising:

means for determining that the result processing component is associated with a current context of the user, comprising:

means for identifying a list of at least one result processing component currently authorized for use on behalf of the user; and

means for determining that the at least one result processing component in the list is associated with the current context of the user; and

wherein the result processing component comprises means for processing the speech recognition results to produce result output.

2. The system of claim 1 :

wherein the system further comprises means for providing the speech recognition results to the result processing component in response to the determination that the result processing component is associated with the current context of the user.

3. The system of claim 2 , wherein the context sharing component comprises the means for providing the speech recognition results to the result processing component.

4. The system of claim 2 , wherein the speech recognition processing component comprises the means for providing the speech recognition results to the result processing component.

5. The system of claim 2 , wherein the means for determining comprises means for determining, at run-time, that the result processing component is associated with the user.

6. The system of claim 2 , wherein the means for providing the speech recognition results to the result processing component comprises means for providing the speech recognition results to the result processing component in real-time.

7. The system of claim 1 , wherein the system further comprises an audio capture device coupled to the first device, and wherein the audio capture device comprises means for capturing the speech of the user, means for producing the audio signal representing the speech of the user, and means for providing the audio signal to the audio capture component; and

wherein the first device does not include the speech recognition processing component.

8. The system of claim 7 , wherein the first device further comprises means for transmitting the captured audio signal to the speech recognition processing component over a network connection.

9. The system of claim 1 , wherein the second device further includes a terminal session manager.

10. The system of claim 9 , wherein the second device further includes the speech recognition processing component.

11. The system of claim 9 , wherein the first device further comprises a terminal services client, wherein the terminal services client comprises means for establishing a terminal services connection with the terminal session manager.

12. The system of claim 9 , further comprising a third device, wherein the third device includes the speech recognition processing component, and wherein the third device does not include a terminal session manager.

13. The system of claim 1 , wherein the second device further includes the speech recognition processing component.

14. The system of claim 1 , further comprising a third device, wherein the third device includes the speech recognition processing component.

15. The system of claim 1 , wherein the first device comprises a logical device.

16. The system of claim 1 , wherein the first device comprises a physical device.

17. The system of claim 1 , wherein the second device comprises a logical device.

18. The system of claim 1 , wherein the second device comprises a physical device.

19. The system of claim 1 :

wherein the first device further comprises the speech recognition processing component;

wherein the system further comprises a third device;

wherein the second device further includes means for providing the result output to the third device; and

wherein the third device comprises means for providing output representing the result output to the user.

20. The system of claim 19 :

wherein the third device comprises a terminal services client;

wherein the means for providing the result output to the third device comprises a terminal session manager in the second device; and

wherein the terminal services client comprises the means for providing output representing the result output to the user.

21. The system of claim 20 , further comprising:

an audio capture device comprising means for capturing the speech of the user, means for producing the audio signal representing the speech of the user, and means for transmitting the audio signal to the audio capture component over a network connection.

22. The system of claim 21 , wherein the audio capture device is not connected to the third device.

23. The system of claim 20 , wherein the second device further includes the speech recognition processing component.

24. The system of claim 20 , further comprising a third device, wherein the third device includes the speech recognition processing component.

25. The system of claim 1 , wherein the result processing component further comprises:

means for providing the result output to an application;

means for obtaining data representing a state of the application; and

means for providing the data representing the state of the application to the speech recognition processing component.

26. The system of claim 25 , wherein the speech recognition processing component further comprises:

means for receiving the data representing the state of the application; and

means for changing a speech recognition context of the speech recognition processing component based on the state of the application.

27. The system of claim 26 , wherein the means for changing the speech recognition context comprises means for changing a language model of the speech recognition processing component.

28. The system of claim 26 , wherein the means for changing the speech recognition context comprises means for changing an acoustic model of the speech recognition processing component.

29. The system of claim 1 , wherein the means for performing automatic speech recognition comprises means for performing automatic speech recognition on the captured audio signal to produce the speech recognition results in real-time.

30. A method, for use with a system, the method performed by at least one processor executing computer program instructions stored on a non-transitory computer-readable medium:

wherein the system comprises:

a first device including an audio capture component;

a speech recognition processing component; and

a second device including a result processing component;

wherein the method comprises:

(A) using the audio capture component to capture an audio signal representing speech of a user to produce a captured audio signal;

(B) using the speech recognition processing component to perform automatic speech recognition on the captured audio signal to produce speech recognition results;

(C) determining that the result processing component is associated with a current context of the user, comprising:

a. identifying a list of at least one result processing component currently authorized for use on behalf of the user; and

b. determining that the at least one result processing component in the list is associated with the current context of the user;

(D) in response to the determination that the result processing component is associated with the current context of the user, providing the speech recognition results to the result processing component; and

(E) using the result processing component to process the speech recognition results to produce result output.

31. The method of claim 30 , further comprising:

(F) providing the speech recognition results to the result processing component in response to the determination that the result processing component is associated with the current context of the user.

32. The method of claim 30 , wherein the system further comprises a context sharing component, and wherein the context sharing component performs (C) and (D).

33. The method of claim 30 , wherein the speech recognition processing component performs (D).

34. The method of claim 30 , wherein (C) comprises determining at run-time that the result processing component is associated with the user.

35. The method of claim 30 , wherein (D) comprises providing the speech recognition results to the result processing component in real-time.

36. The method of claim 30 :

wherein the system further comprises:

an audio capture device coupled to the first device; and

wherein the method further comprises using the audio capture device to:

(F) capture the speech of the user;

(G) produce the audio signal representing the speech of the user;

(H) providing the audio signal to the audio capture component; and

wherein the first device does not include the speech recognition processing component.

37. The method of claim 36 , further comprising:

(I) using the first device transmit the captured audio signal to the speech recognition processing component over a network connection.

38. The method of claim 30 , wherein the second device further includes a terminal session manager.

39. The method of claim 38 , wherein the second device further includes the speech recognition processing component.

40. The method of claim 38 , wherein the first device further comprises a terminal services client, wherein the method further comprises:

(F) using the terminal services client to establish a terminal services connection with the terminal session manager.

41. The method of claim 38 , further comprising a third device, wherein the third device includes the speech recognition processing component, and wherein the third device does not include a terminal session manager.

42. The method of claim 30 , wherein the second device further includes the speech recognition processing component.

43. The method of claim 30 , further comprising a third device, wherein the third device includes the speech recognition processing component.

44. The method of claim 30 , wherein the first device comprises a logical device.

45. The method of claim 30 , wherein the first device comprises a physical device.

46. The method of claim 30 , wherein the second device comprises a logical device.

47. The method of claim 30 , wherein the second device comprises a physical device.

48. The method of claim 30 :

wherein the first device further comprises the speech recognition processing component;

wherein the system further comprises a third device;

wherein the second device further includes means for providing the result output to the third device; and

wherein the method further comprises:

(F) using the means for providing the result output to provide the result output to the third device; and

(G) using the third device to provide output representing the result output to the user.

49. The method of claim 48 :

wherein the third device comprises a terminal services client;

wherein the means for providing the result output to the third device comprises a terminal session manager in the second device; and

wherein the terminal services client comprises the means for providing output representing the result output to the user.

50. The method of claim 49 :

wherein the system further comprises an audio capture device; and

wherein the method further comprises using the audio capture device to:

(H) capture the speech of the user;

(I) produce the audio signal representing the speech of the user; and

(J) transmit the audio signal to the audio capture component over a network connection.

51. The method of claim 50 , wherein the audio capture device is not connected to the third device.

52. The method of claim 49 , wherein the second device further includes the speech recognition processing component.

53. The method of claim 49 , wherein the system further comprises a third device, and wherein the third device includes the speech recognition processing component.

54. The method of claim 30 , further comprising using the result processing component to:

(F) provide the result output to an application;

(G) obtain data representing a state of the application; and

(H) provide the data representing the state of the application to the speech recognition processing component.

55. The method of claim 30 , further comprising using the speech recognition processing component to:

(F) receive the data representing the state of the application; and

(G) change a speech recognition context of the speech recognition processing component based on the state of the application.

56. The method of claim 55 , wherein (E) comprises changing a language model of the speech recognition processing component.

57. The method of claim 55 , wherein (E) comprises changing an acoustic model of the speech recognition processing component.

58. The method of claim 30 , wherein (B) comprises performing automatic speech recognition on the captured audio signal to produce the speech recognition results in real-time.

59. A system comprising:

a first machine comprising:

a target application; and

a first result processing component comprising:

means for processing first speech recognition results to produce result output;

means for providing the result output to the target application;

an audio capture device, wherein the first machine does not include the audio capture device; and

a context sharing component comprising means for logically coupling the result processing component to the audio capture device, comprising:

means for identifying a list of at least one result processing component currently authorized for use on behalf of the user, wherein the list includes the first result processing component; and

means for logically coupling each result processing component in the list, including the first result processing component, to the audio capture device.

60. The system of claim 59 , wherein the audio capture device comprises a telephone.

61. A computer-implemented method for use with a system:

the system comprising:

a first machine comprising:

a target application; and

a first result processing component comprising:

an audio capture device, wherein the first machine does not include the audio capture device; and

a context sharing component;

wherein the method comprises:

(A) using the result processing component to process first speech recognition results to produce result output;

(B) using the result processing component to provide the result output to the target application; and

(C) using the context sharing component to logically couple the result processing component to the audio capture device, comprising:

(C) (1) identifying a list of at least one result processing component currently authorized for use on behalf of the user, wherein the list includes the first result processing component; and

(C) (2) logically coupling each result processing component in the list, including the first result processing component, to the audio capture device.

62. The method of claim 61 , wherein the audio capture device comprises a telephone.

Assignments (10)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2021
From: MMODAL SERVICES, LTD
To: 3M HEALTH INFORMATION SYSTEMS, INC.
Reel/Frame 057567/0069 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Feb 22, 2019
From: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
To: MMODAL IP LLC; MULTIMODAL TECHNOLOGIES, LLC; MEDQUIST OF DELAWARE, INC.; MMODAL MQ INC.; MEDQUIST CM LLC
Reel/Frame 048411/0712 →
RELEASE OF SECURITY INTEREST Recorded Feb 1, 2019
From: CORTLAND CAPITAL MARKET SERVICES LLC, AS ADMINISTRATIVE AGENT
To: MMODAL IP LLC
Reel/Frame 048211/0799 →
NUNC PRO TUNC ASSIGNMENT Recorded Jan 5, 2019
From: MMODAL IP LLC
To: MMODAL SERVICES, LTD.
Reel/Frame 047910/0215 →
CHANGE OF ADDRESS Recorded Apr 14, 2017
From: MMODAL IP LLC
To: MMODAL IP LLC
Reel/Frame 042271/0858 →
PATENT SECURITY AGREEMENT Recorded Oct 10, 2014
From: MMODAL IP LLC
To: CORTLAND CAPITAL MARKET SERVICES LLC
Reel/Frame 033958/0729 →
SECURITY AGREEMENT Recorded Oct 8, 2014
From: MMODAL IP LLC
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 034047/0527 →
RELEASE OF SECURITY INTEREST Recorded Aug 1, 2014
From: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
To: MMODAL IP LLC
Reel/Frame 033459/0935 →
SECURITY AGREEMENT Recorded Aug 22, 2012
From: MMODAL IP LLC; MULTIMODAL TECHNOLOGIES, LLC; POIESIS INFOMATICS INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 028824/0459 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2012
From: KOLL, DETLEF; FINKE, MICHAEL
To: MMODAL IP LLC
Reel/Frame 028631/0583 →