IP Library Granted Patent US 7,672,841
Granted Patent B2
US 7,672,841 · App. 12/123,336 · Granted Mar 2, 2010

Method for processing speech data for a distributed recognition system

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,672,841
App. No.
12/123,336
Granted
Mar 2, 2010
Kind
B2
Abstract

Speech signal information is formatted, processed and transported in accordance with a format adapted for TCP/IP protocols used on the Internet and other communications networks. NULL characters are used for indicating the end of a voice segment. The method is useful for distributed speech recognition systems such as a client-server system, typically implemented on an intranet or over the Internet based on user queries at his/her computer, a PDA, or a workstation using a speech input interface.

Claims (35)

1. A method of processing speech data from an utterance for a distributed speech query recognition system comprising the steps of:

establishing a network connection between a server computing system and a client device suitable for transporting a streaming communication;

receiving a continuous speech byte data stream containing speech data processed by a first component of the distributed speech query recognition system situated in the client device;

wherein said speech data is characterized by a form and data content representing only a partial recognition of an utterance;

further wherein said data stream includes NULL data used to identify a silence in speech data from said client device said NULL data being inserted at the client device after other NULL data is removed prior to transmission of the speech byte data stream; and

further processing said speech data at a second component of the distributed speech query recognition system situated at said server computing system to generate additional speech related content and complete recognition of words in said speech data.

2. The method of claim 1 , further including a step of processing said words using a natural language engine at said server computing system to determine a meaning of said utterance.

3. The method of claim 2 , wherein said meaning of said words is recognized in real time.

4. The method of claim 3 , wherein said meaning is determined before a speech utterance representing said sentence is completed.

5. The method of claim 3 wherein a confidence threshold can be specified for determining said meaning.

6. The method of claim 1 , wherein speech recognition tasks required by the distributed speech recognition system for recognizing words are allocated to said server computing device on a connection by connection basis.

7. The method of claim 1 , wherein speech recognition tasks used by the distributed speech recognition system for recognizing words are allocated to said client computing device on a connection by connection basis.

8. The method of claim 1 , further including a step: transmitting a spoken answer or response in the form of answer speech data from said server computing system to said client device in response to a spoken query presented at said client device.

9. The method of claim 1 , wherein said speech data only includes NULL data during periods of silence.

10. The method of claim 1 , wherein said NULL data information is appended to an end of speech data in said speech byte stream.

11. The method of claim 10 , wherein said NULL data information is a single NULL character.

12. The method of claim 1 , wherein said speech processing is initiated by depressing a dedicated button on said client device.

13. The method of claim 1 , wherein said client computing device is used to formulate a speech based query to an Internet based search engine.

14. The method of claim 1 , wherein an amount of said speech data is configured in response to a real-time performance requirement set for the distributed speech recognition system during a speech utterance session.

15. The method of claim 1 , wherein said words in said speech data are recognized in real time.

16. The method of claim 15 , further including performing a query operation based on said meaning to identify an answer to said words in real-time.

17. The method of claim 15 further including a step: specifying that a natural language engine should return multiple results for the speech data from different servers.

18. The method of claim 1 further including a step: calibrating speech and silence components of said speech data.

19. A method of processing speech data for a distributed speech query recognition system comprising the steps of:

establishing a network connection between a server computing system and a client device suitable for transporting a streaming communication;

receiving a data stream containing speech vector data from the client device, said speech vector data representing acoustic features of speech data and being characterized by a form and data content insufficient to recognize words;

wherein said data stream includes NULL data information used to identify a silence in speech data from said client device, said NULL data being inserted at the client device after other NULL data is removed prior to transmission of the data stream; and

further processing said speech vector data at said server computing system to generate additional speech feature related content and identify words in said speech data.

20. A method of processing speech data for a distributed speech query recognition system comprising the steps of:

establishing a network connection suitable for transporting a streaming communication between a server computing system and a client device;

configuring speech processing operations to be performed by said client device and server computing system respectively;

wherein said speech processing operations are automatically configured based on computing capabilities of said client device and server computing system respectively, and such that said server computing system supports a number of client devices having different computing capabilities;

receiving a data stream containing speech vector data from the client device, said speech vector data representing acoustic features of speech data and being characterized by a data content insufficient to recognize words;

wherein said data stream includes at least some NULL data used to identify a silence in speech data from said client device said NULL data being inserted at the client device after other NULL data is removed prior to transmission of the data stream; and

further processing said speech vector data at said server computing system to generate additional speech feature related content and identify words in said speech data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 6, 2013
From: PHOENIX SOLUTIONS, INC.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 030949/0249 →