IP Library › Granted Patent US 9,009,039
Granted Patent B2
US 9,009,039 · App. 12/483,262 · Granted Apr 14, 2015

Noise adaptive training for speech recognition

Inventors: Michael Lewis Seltzer (Seattle, WA); James Garnet Droppo (Carnation, WA); Ozlem Kalinli (Los Angeles, CA); Alejandro Acero (Bellevue, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/063G10L15/20G10L15/144
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,009,039
App. No.
12/483,262
Granted
Apr 14, 2015
Kind
B2
Abstract

Technologies are described herein for noise adaptive training to achieve robust automatic speech recognition. Through the use of these technologies, a noise adaptive training (NAT) approach may use both clean and corrupted speech for training. The NAT approach may normalize the environmental distortion as part of the model training. A set of underlying “pseudo-clean” model parameters may be estimated directly. This may be done without point estimation of clean speech features as an intermediate step. The pseudo-clean model parameters learned from the NAT technique may be used with a Vector Taylor Series (VTS) adaptation. Such adaptation may support decoding noisy utterances during the operating phase of a automatic voice recognition system.

Claims (34)

1. A computer-implemented method for speech recognition, the method comprising computer-implemented operations for:

receiving a set of training speech utterances, wherein one or more of the training speech utterances comprises speech corrupted by noise or distortion;

initializing parameters of a pseudo-clean speech model;

initializing parameters of a distortion model by setting a channel mean to zero and setting a noise mean and a covariance according to one of a first or a last sample of an utterance where no speech is present;

iteratively applying noise adaptive training using the set of training speech utterances to generate a speech recognition model based upon identifying a set of distortion model parameters from the distortion model and identifying a set of pseudo-clean model parameters from the pseudo-clean speech model;

receiving a speech signal to be recognized;

adapting the speech recognition model to a current environment associated with the received speech signal; and

performing speech recognition on the received speech signal using the adapted speech recognition model.

2. The computer-implemented method of claim 1 , wherein a first one of the training speech utterances comprises a different noise or distortion than a second one of the training speech utterances.

3. The computer-implemented method of claim 1 , wherein the speech recognition model comprises a Hidden Markov Model.

4. The computer-implemented method of claim 1 , wherein generating the speech recognition model further comprises incorporating a Vector Taylor Series adaptation into training of the speech recognition model.

5. The computer-implemented method of claim 1 , wherein generating the speech recognition model further comprises iterating over the training speech utterances while accumulating statistics to update the speech recognition model.

6. The computer-implemented method of claim 1 , wherein generating the pseudo-clean model comprises approximating a model trained with clean speech.

7. A computer system comprising:

a processing unit;

a memory operatively coupled to the processing unit; and

a program module which executes in the processing unit from the memory and which, when executed by the processing unit, causes the computer system to recognize speech by

receiving a set of training speech utterances, wherein one or more of the training speech utterances comprises speech corrupted by noise or distortion;

initializing parameters of a pseudo-clean speech model;

initializing parameters of a distortion model by setting a channel mean to zero and setting a noise mean and a covariance according to one of a first or a last sample of an utterance where no speech is present;

iteratively applying noise adaptive training using the set of training speech utterances to generate a speech recognition model based upon identifying a set of distortion model parameters from the distortion model and identifying a set of pseudo-clean model parameters from the pseudo-clean model;

receiving a speech signal to be recognized;

adapting the speech recognition model to a current environment associated with the received speech signal; and

recognizing speech elements from the received speech signal using the adapted speech recognition model.

8. The computer system of claim 7 , wherein a first one of the training speech utterances comprises a different noise or distortion than a second one of the training speech utterances.

9. The computer system of claim 7 , wherein the speech recognition model comprises a Hidden Markov Model.

10. The computer system of claim 7 , wherein generating the speech recognition model comprises incorporating a Vector Taylor Series adaptation into training of the speech recognition model.

11. The computer system of claim 7 , wherein generating the speech recognition model comprises iterating over the training speech utterances while accumulating statistics to update the speech recognition model.

12. A computer-readable medium having computer-executable instructions stored thereon which, when executed by a computer, cause the computer to:

receive a set of training speech utterances, wherein at least one of the training speech utterances comprises speech corrupted by noise or distortion;

initializing parameters of a hidden Markov speech recognition model;

initializing parameters of a distortion model by setting a channel mean to zero and setting a noise mean and a covariance according to one of a first or a last sample of an utterance where no speech is present;

generate an improved hidden Markov speech recognition model using the set of training speech utterances by applying a Vector Taylor Series adaptation to training of the hidden Markov speech recognition model, and generating a set of pseudo-clean model parameters from the hidden Markov speech recognition model and a set of distortion model parameters from the distortion model; and

recognize speech elements using the improved hidden Markov speech recognition model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034564/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2009
From: SELTZER, MICHAEL LEWIS; DROPPO, JAMES GARNET; KALINLI, OZLEM; ACERO, ALEJANDRO
To: MICROSOFT CORPORATION
Reel/Frame 023099/0767 →
Continuity (1)
Related Publication 20100318354A1 · Dec 16, 2010