IP Library Granted Patent US 9,076,448
Granted Patent B2
US 9,076,448 · App. 10/684,357 · Granted Jul 7, 2015

Distributed real time speech recognition system

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,076,448
App. No.
10/684,357
Granted
Jul 7, 2015
Kind
B2
Abstract

A real-time system incorporating speech recognition and linguistic processing for recognizing a spoken query by a user and distributed between client and server, is disclosed. The system accepts user's queries in the form of speech at the client where minimal processing extracts a sufficient number of acoustic speech vectors representing the utterance. These vectors are sent via a communications channel to the server where additional acoustic vectors are derived. Using Hidden Markov Models (HMMs), and appropriate grammars and dictionaries conditioned by the selections made by the user, the speech representing the user's query is fully decoded into text (or some other suitable form) at the server. This text corresponding to the user's query is then simultaneously sent to a natural language engine and a database processor where optimized SQL statements are constructed for a full-text search from a database for a recordset of several stored questions that best matches the user's query. Further processing in the natural language engine narrows the search to a single stored question. The answer corresponding to this single stored question is next retrieved from the file path and sent to the client in compressed form. At the client, the answer to the user's query is articulated to the user using a text-to-speech engine in his or her native natural language. The system requires no training and can operate in several natural languages.

Claims (10)

1. A method of performing distributed voice recognition, the method comprising:

(a) receiving speech utterance signals representing utterances of a user, the speech utterance signals comprising one or more words;

(b) generating, via a processing circuit of a client device, speech data values from the utterance signals during an utterance evaluation time frame corresponding to each utterance signal,

wherein the speech data values comprise compressed mel-frequency cepstral coefficient vectors (MFCC vectors) further comprising MFCC delta parameters and MFCC acceleration parameters automatically determined based on at least an amount of computational resources available on the client device and a speed of a transceiver used to transmit data between the client device and a server;

(c) encoding the speech data values into a transmission format suitable for transmission over a communications channel to the server; and

(d) communicating user context information over the communications channel to the server, wherein the server uses the context information to dynamically select a grammar to use for recognizing the speech data values.

2. The method of claim 1 wherein receiving speech utterance signals further comprises receiving speech utterance signals representative of a speech based query issued by the user.

3. The method of claim 2 , wherein the user utters the speech based query in response to a prompt issued to the user by the client device.

4. The method of claim 1 , wherein automatically determining the MFCC delta parameters and the MFCC acceleration parameters further comprises determining the MFCC delta parameters and the MFCC acceleration parameters based on an amount of bandwidth available to the server.

5. The method of claim 1 , wherein the MFCC vectors are generated at a rate corresponding to at least 100 frames per second, and such that the MFCC vectors include a separate cepstral coefficient value for a corresponding frequency component of the speech utterance signals.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 6, 2013
From: PHOENIX SOLUTIONS, INC.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 030949/0249 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2012
From: GROSS, J NICHOLAS
To: PHOENIX SOLUTIONS, INC
Reel/Frame 027879/0473 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2006
From: GURURAJ, PALLAKI
To: PHOENIX SOLUTIONS, INC.
Reel/Frame 018433/0739 →
SECURITY INTEREST Recorded Feb 6, 2004
From: PHOENIX SOLUTIONS, INC.
To: J. NICHOLAS GROSS, ATTORNEY AT LAW
Reel/Frame 014313/0379 →