IP Library Granted Patent US 8,504,372
Granted Patent B2
US 8,504,372 · App. 13/563,998 · Granted Aug 6, 2013

Distributed speech recognition using one way communication

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,504,372
App. No.
13/563,998
Granted
Aug 6, 2013
Kind
B2
Abstract

A speech recognition client sends a speech stream and control stream in parallel to a server-side speech recognizer over a network. The network may be an unreliable, low-latency network. The server-side speech recognizer recognizes a first portion of the speech stream and, if a predetermined criterion is satisfied by the speech recognition result, waits until the speech recognizer has been reconfigured before recognizing a second portion of the speech stream. The speech recognition client receives recognition results from the server-side recognizer in response to requests from the client. The client may remotely reconfigure the state of the server-side recognizer during recognition.

Claims (38)

1. A computer-implemented method comprising:

(A) at a speech recognition server:

(A) (1) receiving a speech stream and a control stream from a client;

(A) (2) using an automatic speech recognition engine in a first configuration state to recognize a first portion of the speech stream and thereby to produce a first speech recognition result;

(B) at the speech recognition server,

(B) (1) analyzing a first control message in the control stream to determine whether the first speech recognition result satisfies a first predetermined criterion specified by the control stream;

(B) (2) waiting until the automation speech recognition engine has been reconfigured before continuing to (C); and

(C) at the speech recognition server, responsive to receiving a second control message, using the automatic speech recognition engine in a second configuration state to recognize a second portion of the speech stream and thereby to produce a second speech recognition result.

2. The method of claim 1 , wherein (A)(1) further comprises receiving the speech stream and the control stream in a multiplexed stream from the client.

3. The method of claim 2 , wherein (A)(1) further comprises demultiplexing the multiplexed stream received from the client to produce the speech stream and the control stream.

4. The method of claim 1 , wherein (A)(1) further comprises receiving an encrypted speech stream and an encrypted control stream from the client.

5. The method of claim 4 , wherein (A)(1) further comprises decrypting the encrypted speech stream and the encrypted control stream.

6. A system comprising at least one non-transitory computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions are executable by at least one computer processor to perform a method comprising:

(A) at a speech recognition server:

(A) (1) receiving a speech stream and a control stream from a client;

(A) (2) using an automatic speech recognition engine in a first configuration state to recognize a first portion of the speech stream and thereby to produce a first speech recognition result;

(B) at the speech recognition server,

(B) (1) analyzing a first control message in the control stream to determine whether the first speech recognition result satisfies a first predetermined criterion specified by the control stream;

(B) (2) waiting until the automation speech recognition engine has been reconfigured before continuing to (C); and

(C) at the speech recognition server, responsive to receiving a second control message, using the automatic speech recognition engine in a second configuration state to recognize a second portion of the speech stream and thereby to produce a second speech recognition result.

7. The system of claim 6 , wherein (A)(1) further comprises receiving the speech stream and the control stream in a multiplexed stream from the client.

8. The system of claim 7 , wherein (A)(1) further comprises demultiplexing the multiplexed stream received from the client to produce the speech stream and the control stream.

9. The system of claim 6 , wherein (A)(1) further comprises receiving an encrypted speech stream and an encrypted control stream from the client.

10. The system of claim 9 , wherein (A)(1) further comprises decrypting the encrypted speech stream and the encrypted control stream.

11. A computer-implemented method comprising:

(A) at a speech recognition server:

(A) (1) receiving a speech stream and a control stream from a client;

(A) (2) using an automatic speech recognition engine in a first configuration state to recognize a first portion of the speech stream and thereby to produce a first speech recognition result;

(B) at the speech recognition server, if the first speech recognition result satisfies a first predetermined criterion specified by the control stream, then waiting until the speech recognition engine has been reconfigured before continuing to (C); and

(C) at the speech recognition server, using the automatic speech recognition engine in a second configuration state to recognize a second portion of the speech stream and thereby to produce a second speech recognition result; and

(D) during (B), receiving by the speech recognition server, from the client, a third portion of the speech stream.

12. A system comprising at least one non-transitory computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions are executable by at least one computer processor to perform a method comprising:

(A) at a speech recognition server:

(A) (1) receiving a speech stream and a control stream from a client;

(A) (2) using an automatic speech recognition engine in a first configuration state to recognize a first portion of the speech stream and thereby to produce a first speech recognition result;

(B) at the speech recognition server, if the first speech recognition result satisfies a first predetermined criterion specified by the control stream, then waiting until the speech recognition engine has been reconfigured before continuing to (C); and

(C) at the speech recognition server, using the automatic speech recognition engine in a second configuration state to recognize a second portion of the speech stream and thereby to produce a second speech recognition result; and

(D) during (B), receiving by the speech recognition server, from the client, a third portion of the speech stream.

Assignments (10)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2024
From: 3M INNOVATIVE PROPERTIES COMPANY
To: SOLVENTUM INTELLECTUAL PROPERTIES COMPANY
Reel/Frame 066435/0347 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2021
From: MMODAL IP LLC
To: 3M INNOVATIVE PROPERTIES COMPANY
Reel/Frame 057883/0129 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Feb 22, 2019
From: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
To: MMODAL IP LLC; MULTIMODAL TECHNOLOGIES, LLC; MEDQUIST OF DELAWARE, INC.; MMODAL MQ INC.; MEDQUIST CM LLC
Reel/Frame 048411/0712 →
RELEASE OF SECURITY INTEREST Recorded Feb 1, 2019
From: CORTLAND CAPITAL MARKET SERVICES LLC, AS ADMINISTRATIVE AGENT
To: MMODAL IP LLC
Reel/Frame 048211/0799 →
CHANGE OF ADDRESS Recorded Apr 14, 2017
From: MMODAL IP LLC
To: MMODAL IP LLC
Reel/Frame 042271/0858 →
PATENT SECURITY AGREEMENT Recorded Oct 10, 2014
From: MMODAL IP LLC
To: CORTLAND CAPITAL MARKET SERVICES LLC
Reel/Frame 033958/0729 →
SECURITY AGREEMENT Recorded Oct 8, 2014
From: MMODAL IP LLC
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 034047/0527 →
RELEASE OF SECURITY INTEREST Recorded Aug 1, 2014
From: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
To: MMODAL IP LLC
Reel/Frame 033459/0935 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2012
From: MULTIMODAL TECHNOLOGIES, LLC
To: MMODAL IP LLC
Reel/Frame 029205/0149 →
SECURITY AGREEMENT Recorded Aug 22, 2012
From: MMODAL IP LLC; MULTIMODAL TECHNOLOGIES, LLC; POIESIS INFOMATICS INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 028824/0459 →