IP Library › Granted Patent US 8,972,253
Granted Patent B2
US 8,972,253 · App. 12/882,233 · Granted Mar 3, 2015

Deep belief network for large vocabulary continuous speech recognition

Inventors: Li Deng (Redmond, WA); Dong Yu (Kirkland, WA); George Edward Dahl (Toronto, CA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/14
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,972,253
App. No.
12/882,233
Granted
Mar 3, 2015
Kind
B2
Abstract

A method is disclosed herein that includes an act of causing a processor to receive a sample, wherein the sample is one of spoken utterance, an online handwriting sample, or a moving image sample. The method also comprises the act of causing the processor to decode the sample based at least in part upon an output of a combination of a deep structure and a context-dependent Hidden Markov Model (HMM), wherein the deep structure is configured to output a posterior probability of a context-dependent unit. The deep structure is a Deep Belief Network consisting of many layers of nonlinear units with connecting weights between layers trained by a pretraining step followed by a fine-tuning step.

Claims (40)

1. A method executed by a processor, the method comprising:

receiving a sample at a context-dependent combination of a Deep Belief Network (DBN) and a Hidden Markov Model (HMM), wherein the sample is a spoken utterance

outputting, at the DBN, a posterior probability distribution over labeled senones;

outputting, at the HMM, transition probabilities between the labeled senones, the transition probabilities based upon the posterior probability distribution over the labeled senones; and

decoding the sample based at least in part upon the posterior probability distribution over the labeled senones and the transition probabilities between the labeled senones.

2. The method of claim 1 , wherein the DBN is a probabilistic generative model that comprises multiple layers of stochastic hidden units above a single bottom layer of observed variables that represent a data vector.

3. The method of claim 2 , wherein the DBN is a feed-forward Artificial Neural Network (ANN).

4. The method of claim 1 , further comprising, during a training phase for the combination of the DBN and the HMM, deriving the combination of the DBN and the HMM from a Gaussian Mixture Model (GMM)-HMM system.

5. The method of claim 1 configured to execute in a mobile computing apparatus.

6. The method of claim 1 , further comprising:

during a training phase for the context-dependent combination of the DBN and the HMM, performing pretraining with respect to the DBN, the DBN comprises a plurality of hidden stochastic layers, and wherein pretraining comprises utilizing an unsupervised algorithm to initialize weights of connections between the hidden stochastic layers.

7. The method of claim 6 , further comprising utilizing back-propagation to further refine the weights of the connections between the hidden stochastic layers.

8. The method of claim 6 , wherein a Restricted Boltzmann Machine is used in connection with the pretraining.

9. The method of claim 1 , further comprising:

during a training phase, aligning output units of the DBN with senones in the HMM, such that the output units are assigned the senones in the HMM.

10. The method of claim 1 , further comprising:

decoding the sample based part upon prior probabilities assigned to the labeled senones.

11. The method of claim 1 , further comprising:

receiving a Gaussian Mixture Model (GMM)-Hidden Markov Model (HMM) system that is trained to undertake automatic speech recognition; and

converting the GMM-HMM to the DBN-HMM.

12. The method of claim 11 , further comprising:

utilizing an unsupervised training algorithm to initialize weights of the connections in the DBN; and

utilizing back-propagation to refine the weights of the connections in the DBN.

13. A computer-implemented speech recognition system comprising:

a processor; and

a plurality of components that are executable by the processor, the plurality of components comprising:

a computer-executable combination of a Deep Belief Network (DBN) and a Hidden Markov Model (HMM) that is configured to receive an input sample, wherein the input sample is based upon a spoken utterance, wherein the DBN is configured to output a posterior probability distribution over labeled senones, and wherein the HMM is configured to output transition probabilities between states, the states corresponding to the labeled senones; and

a decoder component that is configured to decode a word sequence from the input sample based at least in part upon the posterior probability distribution over the labeled senones and the transition probabilities between the states.

14. The system of claim 13 , wherein the DBN is a probabilistic generative model that comprises multiple layers of stochastic hidden units above a single bottom layer of observed variables that represent a data vector.

15. The system of claim 14 , wherein the components further comprise a converter/training component that is configured to generate the combination of the DBN and the HMM based at least in part upon a Gaussian Mixture Model (GMM)-HMM system.

16. The system of claim 13 comprised by a portable computing apparatus.

17. The system of claim 16 , wherein the portable computing apparatus is a mobile telephone.

18. The system of claim 13 , wherein the DBN is pretrained with unlabeled data and refined through back-propagation.

19. The system of claim 13 , wherein the decoder component that is configured to decode the word sequence from the input sample based part upon prior probabilities assigned to the labeled senones.

20. A computer-readable memory comprising instructions that, when executed by a processor, cause the processor to perform acts comprising:

receiving a Gaussian Mixture Model (GMM)-Hidden Markov Model (HMM) system that is trained to undertake automatic speech recognition;

converting the GMM-HMM to a Deep Belief Network (DBN)-HMM system, wherein the DBN comprises a plurality of layers of stochastic hidden units above a bottom layer of observed variables that represent a data vector, wherein the DBN comprises a plurality of undirected weighted connections between an uppermost two layers and directed weighted connections at other layers, wherein the DBN is configured to output posterior probabilities of senones pertaining to spoken utterances and the HMM is configured to output transition probabilities between the senones;

utilizing an unsupervised training algorithm to initialize weights of the connections in the DBN;

utilizing back-propagation to refine the weights of the connections in the DBN; and

deploying the DBN-HMM in an automatic speech recognition system.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034544/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 17, 2010
From: DENG, LI; YU, DONG; DAHL, GEORGE EDWARD
To: MICROSOFT CORPORATION
Reel/Frame 025002/0203 →
Continuity (1)
Related Publication 20120065976A1 · Mar 15, 2012