IP Library Granted Patent US 8,019,608
Granted Patent B2
US 8,019,608 · App. 12/550,381 · Granted Sep 13, 2011

Distributed speech recognition using one way communication

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,019,608
App. No.
12/550,381
Granted
Sep 13, 2011
Kind
B2
Abstract

A speech recognition client sends a speech stream and control stream in parallel to a server-side speech recognizer over a network. The network may be an unreliable, low-latency network. The server-side speech recognizer recognizes the speech stream continuously. The speech recognition client receives recognition results from the server-side recognizer in response to requests from the client. The client may remotely reconfigure the state of the server-side recognizer during recognition if a first speech recognition result satisfies a predetermined criterion specified by the control stream.

Claims (75)

1. A computer-implemented method comprising:

(A) at a client, transmitting a speech stream and a control stream to a speech recognition server using a Hypertext Transfer Protocol (HTTP) having a first timeout period;

(B) at the speech recognition server, using an automatic speech recognition engine to initiate recognition of the speech stream;

(C) at the client, transmitting a first request for a speech recognition result to the server using HTTP; and

(D) at the server, transmitting a notification to the client indicating that no speech recognition results have become available within a second timeout period that differs from the first timeout period; and

(E) at the client, in response to receiving the notification, transmitting a second request for the speech recognition result to the server using HTTP.

2. The method of claim 1 , further comprising:

(F) at the server, recognizing a first portion of the speech stream to produce a first speech recognition result; and

(G) transmitting the first speech recognition result to the client using HTTP in response to the second request.

3. The method of claim 2 , wherein (G) comprises:

(G)( 1 ) determining whether any speech recognition results are available;

(G)( 2 ) if no speech recognition results are available, returning to (G)(1);

(G)(3) otherwise, transmitting the first speech recognition result to the client.

4. The method of claim 3 , wherein the server performs (F) and (G) in parallel.

5. The method of claim 1 , wherein (A) comprises transmitting the speech stream and the control stream using a Hypertext Transfer Protocol over Secure Socket Layer (HTTPS), and wherein (C) comprises transmitting the first request using HTTPS.

6. A system comprising a client device and a speech recognition server:

wherein the client device comprises:

means for transmitting a speech stream and a control stream to a speech recognition server using a Hypertext Transfer Protocol (HTTP) having a first timeout period;

means for transmitting a first request for a speech recognition result to the server using HTTP; and

wherein the speech recognition server comprises:

means for using an automatic speech recognition engine to initiate recognition of the speech stream;

means for transmitting a notification to the client indicating that no speech recognition results have become available within a second timeout period that differs from the first timeout period; and

wherein the client further comprises means, responsive to receipt of the notification, for transmitting a second request for the speech recognition result to the server using HTTP.

7. A computer-implemented method performed by a client device, the method comprising:

(A) transmitting a speech stream and a control stream to a speech recognition server using a Hypertext Transfer Protocol (HTTP) having a first timeout period;

(B) transmitting a first request for a speech recognition result to a server using HTTP at a first time;

(C) receiving, at a second time that differs from the first time by less than the first timeout period, a notification from the server indicating that no speech recognition results are available; and

(D) in response to receiving the notification, transmitting a second request for the speech recognition result to the server using HTTP.

8. An apparatus comprising:

means for transmitting a speech stream and a control stream to a speech recognition server using a Hypertext Transfer Protocol (HTTP) having a first timeout period;

means for transmitting a first request for a speech recognition result to a server using HTTP at a first time;

means for receiving, at a second time that differs from the first time by less than the first timeout period, a notification from the server indicating that no speech recognition results are available; and

means for transmitting a second request for the speech recognition result to the server using HTTP in response to receiving the notification.

9. A computer-implemented method performed by a server, the method comprising:

(A) receiving a speech stream and a control stream from a client using a Hypertext Transfer Protocol (HTTP) having a first timeout period;

(B) using an automatic speech recognition engine to initiate recognition of the speech stream;

(C) receiving a first request for a speech recognition result from the client using HTTP; and

(D) transmitting a notification to the client indicating that no speech recognition results have become available within a second timeout period that differs from the first timeout period.

10. The method of claim 9 , further comprising:

(E) receiving a second request for the speech recognition result from the client using HTTP;

(F) recognizing a first portion of the speech stream to produce a first speech recognition result; and

(G) transmitting the first speech recognition result to the client using HTTP in response to the second request.

11. An apparatus comprising:

means for receiving a speech stream and a control stream from a client using a Hypertext Transfer Protocol (HTTP) having a first timeout period;

means for using an automatic speech recognition engine to initiate recognition of the speech stream;

means for receiving a first request for a speech recognition result from the client using HTTP; and

means for transmitting a notification to the client indicating that no speech recognition results have become available within a second timeout period that differs from the first timeout period.

12. A computer-implemented method comprising:

(A) at a speech recognition server: (A)(1) receiving a speech stream and a control stream from a client; (A)(2) using an automatic speech recognition engine to recognize a first portion of the speech stream and thereby to produce a first speech recognition result;

(B) at the speech recognition server, if the first speech recognition result satisfies a first predetermined criterion specified by the control stream, then waiting until the speech recognition engine has been reconfigured before continuing to (C); and

(C) at the speech recognition server, using the automatic speech recognition engine to recognize a second portion of the speech stream and thereby to produce a second speech recognition result.

13. The method of claim 12 , further comprising:

(D) at a client, before (A), transmitting the speech stream and a control stream to the speech recognition server.

14. The method of claim 13 , further comprising:

(E) at the client, transmitting a request for a speech recognition result to the server; and

(F) at the server:

(F)(1) determining whether any speech recognition results are available;

(F)(2) if no speech recognition results are available, returning to (F)(1);

(F)(3) otherwise, transmitting at least one of the speech recognition results to the client.

15. The method of claim 14 , wherein the server performs (B) in parallel with (F).

16. The method of claim 14 , wherein (D) comprises transmitting the speech stream and the control stream using a Hypertext Transfer Protocol (HTTP), and wherein (E) comprises transmitting the request for the speech recognition result using HTTP.

17. The method of claim 14 , wherein (D) comprises transmitting the speech stream and the control stream using a Hypertext Transfer Protocol over Secure Socket Layer (HTTPS), and wherein (E) comprises transmitting the request for the speech recognition result using HTTPS.

18. The method of claim 13 , wherein (D) comprises:

(A)(1) transmitting a first control message in the control stream to the speech recognition server;

(A)(2) detecting a failure of the transmission of the first portion; and

(A)(3) in response to detection of the failure:

(A) (3) (a) creating a second control message specifying a combination of a first state change represented by the first control message and a second state change; and

(A)(3)(b) transmitting the second control message in the control stream to the speech recognition server.

19. The method of claim 12 , wherein (C) comprises waiting until the automatic speech recognition engine is in a predetermined configuration state before continuing to (D).

20. The method of claim 12 , wherein (C) further comprises executing one of the control messages to reconfigure the speech recognition engine after (B).

21. An apparatus comprising:

first reception means for receiving a speech stream and a control stream from a client;

first portion recognition means for using an automatic speech recognition engine to recognize a first portion of the speech stream and thereby to produce a first speech recognition result;

waiting means for waiting until the speech recognition engine has been reconfigured before activating the second portion recognition means if the first speech recognition result satisfies a first predetermined criterion specified by the control stream; and

second portion recognition means for using the automatic speech recognition engine to recognize a second portion of the speech stream and thereby to produce a second speech recognition result.

Assignments (12)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2024
From: 3M INNOVATIVE PROPERTIES COMPANY
To: SOLVENTUM INTELLECTUAL PROPERTIES COMPANY
Reel/Frame 066435/0347 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2021
From: MMODAL IP LLC
To: 3M INNOVATIVE PROPERTIES COMPANY
Reel/Frame 057883/0129 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Feb 22, 2019
From: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
To: MMODAL IP LLC; MULTIMODAL TECHNOLOGIES, LLC; MEDQUIST OF DELAWARE, INC.; MMODAL MQ INC.; MEDQUIST CM LLC
Reel/Frame 048411/0712 →
RELEASE OF SECURITY INTEREST Recorded Feb 1, 2019
From: CORTLAND CAPITAL MARKET SERVICES LLC, AS ADMINISTRATIVE AGENT
To: MMODAL IP LLC
Reel/Frame 048211/0799 →
CHANGE OF ADDRESS Recorded Apr 14, 2017
From: MMODAL IP LLC
To: MMODAL IP LLC
Reel/Frame 042271/0858 →
PATENT SECURITY AGREEMENT Recorded Oct 10, 2014
From: MMODAL IP LLC
To: CORTLAND CAPITAL MARKET SERVICES LLC
Reel/Frame 033958/0729 →
SECURITY AGREEMENT Recorded Oct 8, 2014
From: MMODAL IP LLC
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 034047/0527 →
RELEASE OF SECURITY INTEREST Recorded Aug 1, 2014
From: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
To: MULTIMODAL TECHNOLOGIES, LLC
Reel/Frame 033459/0987 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2013
From: MULTIMODAL TECHNOLOGIES, LLC
To: MMODAL IP LLC
Reel/Frame 030060/0627 →
SECURITY AGREEMENT Recorded Aug 22, 2012
From: MMODAL IP LLC; MULTIMODAL TECHNOLOGIES, LLC; POIESIS INFOMATICS INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 028824/0459 →
CHANGE OF NAME Recorded Oct 14, 2011
From: MULTIMODAL TECHNOLOGIES, INC.
To: MULTIMODAL TECHNOLOGIES, LLC
Reel/Frame 027061/0492 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2009
From: CARRAUX, ERIC; KOLL, DETLEFT
To: MULTIMODAL TECHNOLOGIES, INC.
Reel/Frame 023293/0892 →