IP Library Granted Patent US 7,941,317
Granted Patent B1
US 7,941,317 · App. 11/758,037 · Granted May 10, 2011

Low latency real-time speech transcription

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,941,317
App. No.
11/758,037
Granted
May 10, 2011
Kind
B1
Abstract

Systems and methods for low-latency real-time speech recognition/transcription. A discriminative feature extraction, such as a heteroscedastic discriminant analysis transform, in combination with a maximum likelihood linear transform is applied during front-end processing of a digital speech signal. The extracted features reduce the word error rate. A discriminative acoustic model is applied by generating state-level lattices using Maximum Mutual Information Estimation. Recognition networks of language models are replaced by their closure. Latency is reduced by eliminating segmentation such that a number of words/sentences can be recognized as a single utterance. Latency is further reduced by performing front-end normalization in a causal fashion.

Claims (22)

1. A method for transcribing speech in real-time, the method comprising:

analyzing via a processor speech received in real-time using a discriminative front-end processing module, wherein the discriminative front-end processing module includes a heteroscedastic discriminant analysis and a maximum likelihood linear transform;

applying discriminative acoustic models to feature vectors produced by the discriminative front-end processing module;

requesting Gaussians calculated as Gaussian likelihoods from memory based on a current frame for a batch of frames that includes the current frame associated with a first time and a plurality of K future frames, the plurality of K future frames being associated with future speech not yet received, the future speech to be uttered at a second time which is later than the first time; and

transcribing the speech recognized based on the discriminative acoustic models, wherein the future speech is recognized based at least in part on a Gaussian likelihood computed for the current frame.

2. A method as defined in claim 1 , wherein analyzing the speech using a discriminative front-end processing module further comprises extracting feature vectors from the speech.

3. A method as defined in claim 1 , further comprising recognizing speech from an output of the discriminative front-end processing module by generating additional state level lattices by iteratively applying maximum mutual information estimation to resulting state-level lattices.

4. A method as defined in claim 1 , wherein applying discriminative acoustic models to feature vectors produced by the discriminative front-end processing module further comprises applying a language model wherein recognition networks of the language model have been replaced by their closure.

5. A method as defined in claim 4 , wherein applying a language model wherein recognition networks of the language model have been replaced by their closure further comprises adding at least one of an epsilon transition, a silence transition, and a non-speech transition from a set of final states to a start state.

6. A method as defined in claim 4 , wherein transcribing the speech recognized from the discriminative acoustic models further comprises:

outputting a common prefix when the common prefix is non-empty; and

discarding memory used for the common prefix in order to avoid memory exhaustion.

7. A method as defined in claim 1 , wherein analyzing the speech using a discriminative front-end processing module further comprises performing front-end normalization in a low-latency fashion by reducing a lookahead of the discriminative front-end processing module.

8. A method as defined in claim 1 , wherein there are M Gaussians in a mixture, K frames are in a batch and wherein the N features have been transformed into N′ features where N′ equals 2N+1, further comprising determining likelihoods by computing a product of an M×N′ matrix with a N′×K matrix.

9. In a speech recognition system where computation time is divided between a traversal of a recognition search space and an acoustic model likelihood computation and wherein data for the acoustic model is retrieved from memory for each frame, a method for recognizing speech by reducing the computation time, the method comprising:

retrieving a batch of acoustic model data from main memory, wherein the batch corresponds to a current frame of real-time speech data associated with a first time and a plurality of K future frames of speech data, the K future frames being associated with future speech not yet received, the future speech to be uttered at a second time which is later than the first time;

computing a Gaussian likelihood for the current frame and Gaussian likelihoods for the plurality of K future frames using the batch of acoustic model data and based on the current frame;

storing the Gaussian likelihood for the current frame and the Gaussian likelihoods for the plurality of K future frames; and

recognizing and transcribing the future speech, when received, based at least in part on the Gaussian likelihood computed for the current frame.

10. A method as defined in claim 9 , further comprising discarding Gaussian likelihoods that are not needed and retrieving a new batch of acoustic model data from main memory.

11. A method as defined in claim 9 , further comprising pre-fetching the acoustic model data.

12. A method as defined in claim 9 , wherein computing a Gaussian likelihood for the current frame and Gaussian likelihoods for the plurality of future frames further comprises computing a product of an M×N′ matrix and an N′×K matrik, wherein N is a dimension of feature vectors, N′ is a transform of N to 2N+1 features, M is a number of Gaussians in a mixture, and K is a number of frames in a batch.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2017
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 041045/0179 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2017
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 041045/0207 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2016
From: GOFFIN, VINCENT; RILEY, MICHAEL DENNIS; SARACLAR, MURAT
To: AT&T CORP.
Reel/Frame 039525/0328 →