IP Library Granted Patent US 8,249,878
Granted Patent B2
US 8,249,878 · App. 13/196,188 · Granted Aug 21, 2012

Distributed speech recognition using one way communication

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,249,878
App. No.
13/196,188
Granted
Aug 21, 2012
Kind
B2
Abstract

A speech recognition client sends a speech stream and control stream in parallel to a server-side speech recognizer over a network. The network may be an unreliable, low-latency network. The server-side speech recognizer recognizes a first portion of the speech stream and, if a predetermined criterion is satisfied by the speech recognition result, waits until the speech recognizer has been reconfigured before recognizing a second portion of the speech stream. The speech recognition client receives recognition results from the server-side recognizer in response to requests from the client. The client may remotely reconfigure the state of the server-side recognizer during recognition.

Claims (50)

1. A method performed by at least one computer processor executing computer program instructions stored on a non-transitory computer-readable medium, the method comprising:

(A) at a speech recognition server:

(A)(1) receiving a speech stream and a control stream from a client;

(A) (2) using an automatic speech recognition engine in a first configuration state to recognize a first portion of the speech stream and thereby to produce a first speech recognition result;

(B) at the speech recognition server, if the first speech recognition result satisfies a first predetermined criterion specified by the control stream, then waiting until the speech recognition engine has been reconfigured before continuing to (C); and

(C) at the speech recognition server, using the automatic speech recognition engine in a second configuration state to recognize a second portion of the speech stream and thereby to produce a second speech recognition result.

2. The method of claim 1 , wherein (C) further comprises executing a control message in the control stream to reconfigure the speech recognition engine after (B), the control message identifying the second configuration state.

3. The method of claim 2 , further comprising:

(D) at the client, before (A), transmitting the speech stream and the control stream to the speech recognition server.

4. The method of claim 3 , further comprising:

(E) at the client, transmitting a speech recognition result request to the speech recognition server; and

(F) at the speech recognition server:

(F)(1) determining whether any speech recognition results are available;

(F)(2) if no speech recognition results are available, returning to (F)(1);

(F)(3) otherwise, transmitting at least one of the first and second speech recognition results to the client.

5. The method of claim 4 , wherein the speech recognition server performs (B) in parallel with (F).

6. The method of claim 4 , wherein (D) comprises transmitting the speech stream and the control stream using a Hypertext Transfer Protocol (HTTP), and wherein (E) comprises transmitting the speech recognition result request using HTTP.

7. The method of claim 4 , wherein (D) comprises transmitting the speech stream and the control stream using a Hypertext Transfer Protocol over Secure Sockets Layer (HTTPS), and wherein (E) comprises transmitting the speech recognition result request using HTTPS.

8. The method of claim 3 , wherein (D) comprises:

(D)(1) transmitting a first control message in the control stream to the speech recognition server;

(D)(2) detecting a failure of the transmission of the first portion; and

(D)(3) in response to detection of the failure:

(D)(3)(a) creating a second control message specifying a combination of a first state change represented by the first control message and a second state change; and

(D)(3)(b) transmitting the second control message in the control stream to the speech recognition server.

9. The method of claim 2 , wherein (C) comprises waiting until the automatic speech recognition engine is in a predetermined configuration state before continuing to (D).

10. A non-transitory computer-readable medium comprising computer program instructions stored on the computer-readable medium, wherein the computer program instructions are executable by at least one computer processor to perform a method comprising:

(A) at a speech recognition server:

(A)(1) receiving a speech stream and a control stream from a client;

(A) (2) using an automatic speech recognition engine in a first configuration state to recognize a first portion of the speech stream and thereby to produce a first speech recognition result;

(B) at the speech recognition server, if the first speech recognition result satisfies a first predetermined criterion specified by the control stream, then waiting until the speech recognition engine has been reconfigured before continuing to (C); and

(C) at the speech recognition server, using the automatic speech recognition engine in a second configuration state to recognize a second portion of the speech stream and thereby to produce a second speech recognition result.

11. The computer-readable medium of claim 10 , wherein (C) further comprises executing a control message in the control stream to reconfigure the speech recognition engine after (B), the control message identifying the second configuration state.

12. The computer-readable medium of claim 11 , wherein the method further comprises:

(D) at the client, before (A), transmitting the speech stream and the control stream to the speech recognition server.

13. The computer-readable medium of claim 12 , wherein the method further comprises:

(E) at the client, transmitting a speech recognition result request to the speech recognition server; and

(F) at the speech recognition server:

(F)(1) determining whether any speech recognition results are available;

(F)(2) if no speech recognition results are available, returning to (F)(1);

(F)(3) otherwise, transmitting at least one of the first and second speech recognition results to the client.

14. The computer-readable medium of claim 13 , wherein the speech recognition server performs (B) in parallel with (F).

15. The computer-readable medium of claim 13 , wherein (D) comprises transmitting the speech stream and the control stream using a Hypertext Transfer Protocol (HTTP), and wherein (E) comprises transmitting the speech recognition result request using HTTP.

16. The computer-readable medium of claim 13 , wherein (D) comprises transmitting the speech stream and the control stream using a Hypertext Transfer Protocol over Secure Sockets Layer (HTTPS), and wherein (E) comprises transmitting the speech recognition result request using HTTPS.

17. The computer-readable medium of claim 12 , wherein (D) comprises:

(D)(1) transmitting a first control message in the control stream to the speech recognition server;

(D)(2) detecting a failure of the transmission of the first portion; and

(D)(3) in response to detection of the failure:

(D)(3)(a) creating a second control message specifying a combination of a first state change represented by the first control message and a second state change; and

(D)(3)(b) transmitting the second control message in the control stream to the speech recognition server.

18. The computer-readable medium of claim 11 , wherein (C) comprises waiting until the automatic speech recognition engine is in a predetermined configuration state before continuing to (D).

Assignments (12)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2024
From: 3M INNOVATIVE PROPERTIES COMPANY
To: SOLVENTUM INTELLECTUAL PROPERTIES COMPANY
Reel/Frame 066435/0347 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2021
From: MMODAL IP LLC
To: 3M INNOVATIVE PROPERTIES COMPANY
Reel/Frame 057883/0129 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Feb 22, 2019
From: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
To: MMODAL IP LLC; MULTIMODAL TECHNOLOGIES, LLC; MEDQUIST OF DELAWARE, INC.; MMODAL MQ INC.; MEDQUIST CM LLC
Reel/Frame 048411/0712 →
RELEASE OF SECURITY INTEREST Recorded Feb 1, 2019
From: CORTLAND CAPITAL MARKET SERVICES LLC, AS ADMINISTRATIVE AGENT
To: MMODAL IP LLC
Reel/Frame 048211/0799 →
CHANGE OF ADDRESS Recorded Apr 14, 2017
From: MMODAL IP LLC
To: MMODAL IP LLC
Reel/Frame 042271/0858 →
PATENT SECURITY AGREEMENT Recorded Oct 10, 2014
From: MMODAL IP LLC
To: CORTLAND CAPITAL MARKET SERVICES LLC
Reel/Frame 033958/0729 →
SECURITY AGREEMENT Recorded Oct 8, 2014
From: MMODAL IP LLC
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 034047/0527 →
RELEASE OF SECURITY INTEREST Recorded Aug 1, 2014
From: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
To: MULTIMODAL TECHNOLOGIES, LLC
Reel/Frame 033459/0987 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2013
From: MULTIMODAL TECHNOLOGIES, LLC
To: MMODAL IP LLC
Reel/Frame 030060/0649 →
SECURITY AGREEMENT Recorded Aug 22, 2012
From: MMODAL IP LLC; MULTIMODAL TECHNOLOGIES, LLC; POIESIS INFOMATICS INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 028824/0459 →
CHANGE OF NAME Recorded Oct 14, 2011
From: MULTIMODAL TECHNOLOGIES, INC.
To: MULTIMODAL TECHNOLOGIES, LLC
Reel/Frame 027061/0492 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2011
From: CARRAUX, ERIC; KOLL, DETLEF
To: MULTIMODAL TECHNOLOGIES, INC.
Reel/Frame 026686/0443 →