IP Library Granted Patent US 12,217,743
Granted Patent B1
US 12,217,743 · App. 18/227,769 · Granted Feb 4, 2025

System and method for speech recognition using deep recurrent neural networks

Inventor: Alexander B. Graves (London, GB)
Assignee: Google LLC
G10L15/16G06N3/044G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,743
App. No.
18/227,769
Granted
Feb 4, 2025
Kind
B1
Abstract

Deep recurrent neural networks applied to speech recognition. The deep recurrent neural networks (RNNs) are preferably implemented by stacked long short-term memory bidirectional RNNs. The RNNs are trained using end-to-end training with suitable regularisation.

Claims (42)

1. A method performed by one or more computers and for training a speech recognition neural network system comprising one or more neural networks, the method comprising:

pre-training one or more of the neural networks in the speech recognition neural network system on a corpus comprising text data, wherein pre-training the one or more of the neural networks in the speech recognition neural network system on a corpus comprising text data comprises:

pre-training one or more of the neural networks in the speech recognition neural network system on the corpus comprising text data to perform next-step prediction; and

after the pre-training, training the one or more neural networks in the speech recognition neural network system on training data that maps audio to text transcriptions.

2. The method of claim 1 , wherein the prediction neural network is a uni-directional long short-term memory (LSTM) neural network.

3. The method of claim 2 , wherein the vocabulary of output labels also includes a blank symbol that represents a non-output.

4. The method of claim 1 , wherein the one or more neural networks that are pre-trained on the corpus comprising text data comprise a language model.

5. The method of claim 1 , wherein the speech recognition neural network comprises:

a transcription neural network configured to generate a sequence of transcription hidden vectors that includes a respective transcription hidden vector for each audio observation in a particular sequence of audio observations;

a prediction neural network configured to, at position u in an output sequence generated from the particular sequence of audio observations, process a current output sequence that includes output labels at positions 1 through u−1 in the output sequence to generate a prediction hidden vector; and

an output neural network, configured to, at the position u in the output sequence generated from the particular sequence of audio observations, process the prediction hidden vector and a respective transcription hidden vector for an audio observation at a position t in the particular sequence of audio observations, to generate a probability distribution comprising a respective probability for each output label in a vocabulary of output labels.

6. The method of claim 5 , wherein pre-training one or more of the neural networks in the speech recognition neural network system on a corpus comprising text data comprises:

pre-training the prediction neural network in the speech recognition neural network system on the corpus comprising text data.

7. The method of claim 5 , further comprising:

initializing the transcription neural network from a CTC-trained neural network.

8. The method of claim 5 , wherein the transcription neural network is a bi-directional recurrent neural network.

9. The method of claim 5 , wherein the transcription neural network is a bi-directional long short-term memory (LSTM) neural network.

10. The method of claim 9 , wherein each transcription hidden vector is a concatenation of hidden vectors for the corresponding audio input generated by a set of uppermost bidirectional layers in the bi-directional LSTM neural network.

11. The method of claim 5 , wherein the prediction neural network is a uni-directional recurrent neural network.

12. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations for training a speech recognition neural network system comprising one or more neural networks, the operations comprising:

pre-training one or more of the neural networks in the speech recognition neural network system on a corpus comprising text data, wherein pre-training the one or more of the neural networks in the speech recognition neural network system on a corpus comprising text data comprises:

pre-training one or more of the neural networks in the speech recognition neural network system on the corpus comprising text data to perform next-step prediction; and

after the pre-training, training the one or more neural networks in the speech recognition neural network system on training data that maps audio to text transcriptions.

13. The system of claim 12 , wherein the one or more neural networks that are pre-trained on the corpus comprising text data comprise a language model.

14. The system of claim 12 , wherein the speech recognition neural network comprises:

a transcription neural network configured to generate a sequence of transcription hidden vectors that includes a respective transcription hidden vector for each audio observation in a particular sequence of audio observations;

a prediction neural network configured to, at position u in an output sequence generated from the particular sequence of audio observations, process a current output sequence that includes output labels at positions 1 through u−1 in the output sequence to generate a prediction hidden vector; and

an output neural network, configured to, at the position u in the output sequence generated from the particular sequence of audio observations, process the prediction hidden vector and a respective transcription hidden vector for an audio observation at a position/in the particular sequence of audio observations, to generate a probability distribution comprising a respective probability for each output label in a vocabulary of output labels.

15. The system of claim 14 , wherein pre-training one or more of the neural networks in the speech recognition neural network system on a corpus comprising text data comprises:

pre-training the prediction neural network in the speech recognition neural network system on the corpus comprising text data.

16. The system of claim 14 , the operations further comprising:

initializing the transcription neural network from a CTC-trained neural network.

17. The system of claim 14 , wherein the vocabulary of output labels also includes a blank symbol that represents a non-output.

18. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for training a speech recognition neural network system comprising one or more neural networks, the operations comprising:

pre-training one or more of the neural networks in the speech recognition neural network system on a corpus comprising text data, wherein pre-training the one or more of the neural networks in the speech recognition neural network system on a corpus comprising text data comprises:

pre-training one or more of the neural networks in the speech recognition neural network system on the corpus comprising text data to perform next-step prediction; and

after the pre-training, training the one or more neural networks in the speech recognition neural network system on training data that maps audio to text transcriptions.

19. The non-transitory computer storage media of claim 18 , wherein the one or more neural networks that are pre-trained on the corpus comprising text data comprise a language model.

20. The non-transitory computer storage media of claim 18 , wherein the speech recognition neural network comprises:

a transcription neural network configured to generate a sequence of transcription hidden vectors that includes a respective transcription hidden vector for each audio observation in a particular sequence of audio observations;

a prediction neural network configured to, at position u in an output sequence generated from the particular sequence of audio observations, process a current output sequence that includes output labels at positions 1 through u−1 in the output sequence to generate a prediction hidden vector; and

an output neural network, configured to, at the position u in the output sequence generated from the particular sequence of audio observations, process the prediction hidden vector and a respective transcription hidden vector for an audio observation at a position t in the particular sequence of audio observations, to generate a probability distribution comprising a respective probability for each output label in a vocabulary of output labels.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: GRAVES, ALEXANDER B.
To: GOOGLE, INC.
Reel/Frame 065507/0879 →
CHANGE OF NAME Recorded Nov 9, 2023
From: GOOGLE, INC.
To: GOOGLE LLC
Reel/Frame 065531/0881 →
Continuity (7)
Continuation 17700234 · Mar 21, 2022
Continuation 17013276 · Sep 4, 2020
Continuation 16658697 · Oct 21, 2019
Continuation 16267078 · Feb 4, 2019
Continuation 15043341 · Feb 12, 2016
Continuation 14090761 · Nov 26, 2013
Provisional Application 61731047 · Nov 29, 2012
References Cited (33)
US 6519563B1 · Lee · 2003 [cited by examiner]
US 9141916B1 · Corrado et al. · 2015 [cited by applicant]
US 10672388B2 · Hori et al. · 2020 [cited by applicant]
‘Daan Wierstra—Home Page’ [online]. “Daan Wierstra,” [retrieved on Jun. 10, 2013]. Retrieved from the internet: URL<http://www.idsia.ch/˜daan/>, 2 pages. [cited by applicant]
“Fibonacci Web Design,” 2010 [retrieved on Jun. 10, 2013]. Retrieved from the internet: URL<http://www.idsia.ch/˜juergen/fibonacciwebdesign.html>, 1 page. [cited by applicant]
“Matteo Gagliolo—CoMo Home Page,” retrieved on Jun 10, 2013. Retrieved from the internet: URL<http://www.idsia.ch/˜matteo/>, 4 pages. [cited by applicant]
Bayer et al., “Evolving memory cell structures for sequence learning,” ICANN 2009, Part II, LNCS 5769, 755-764, 2009. [cited by applicant]
Eyben et al., “From Speech to Letters—Using a Novel Neural Network Architecture for Grapheme Based ASR,” Proc. Automatic Speech Recognition and Understanding Workshop (ASRU), IEEE, pp. 376-380, Merano, Italy, 2009. [cited by applicant]
Fernandez et al., “An application of recurrent neural networks to discriminative keyword spotting,” ICANN'07 Proceedings of the 17th international conference on Artificial neural networks, 220-229, 2007. [cited by applicant]
Forster et al., “RNN-based Learning of Compact Maps for Efficient Robot Localization,” In proceeding of: ESANN 2007, 15th European Symposium on Artificial Neural Networks, Bruges, Belgium, Apr. 25-27, 2007, 6 pages. [cited by applicant]
Gomez et al., “Accelerated Neural Evolution through Cooperatively Coevolved Synapses,” Journal of Machine Learning Research, 9:937-965, 2008. [cited by applicant]
Graves et al., “A Novel Connectionist System for Unconstrained Handwriting Recognition,” Journal IEEE Transactions on Pattern Analysis and Machine Intelligence, 31(5):855-868, May 2009. [cited by applicant]
Graves et al., “Multi-Dimensional Recurrent Neural Networks,” Feb. 11, 2013, 10 pages. [cited by applicant]
Graves et al., “Offline Handwriting Recognition with Multidimensional Recurrent Neural Networks,” Advances in Neural Information Processing Systems, 2009, 8 pages. [cited by applicant]
Graves et al., “Unconstrained Online Handwriting Recognition with Recurrent Neural Networks,” Advances in Neural Information Processing Systems 20, 8 pages. [cited by applicant]
Graves, et al., “Connectionist Temporal Classifications: Labelling Unsegmented Sequence Data with Recurrent Neural Networks,” Proceeding ICML '06 Proceedings of the 23rd International Conference on Machine Learning, pp.… [cited by applicant]
Hochreiter et al., “Flat Minima,” Neural Computation 9(1):1-42 (1997), [retrieved on Jun. 10, 2013]. Retrieved from the internet: URL<http://www.idsia.ch/˜juergen/fm/>, 2 pages (abstract only). [cited by applicant]
Hochreiter et al., “LSTM Can Solve Hard Long Time Lag Problems,” Advances in Neural Information Processing Systems 9, NIPS'9, 473-479, MIT Press, Cambridge MA, 1997 [retrieved on Jun. 10, 2013]. Retrieved from the inter… [cited by applicant]
Koutnik et al., “Searching for Minimal Neural Networks in Fourier Space,” [retrieved Jun. 10, 2013]. Retrieved from the internet: URL<http://www.idsia.ch/˜juergen/agi10koutnik.pdf>, 6 pages. [cited by applicant]
Liwicki et al., “A Novel Approach to On-Line Handwriting Recognition Based on Bidirectional Long Short-Term Memory Networks,” Proceedings of the 9th International Conference on Document Analysis and Recognition, ICDAR 2… [cited by applicant]
Monner and Reggia, “A Generalized LSTM-like Training Algorithm for Second-order Recurrent Neural Networks,” Neural Networks, 2010, pp. 1-35. [cited by applicant]
Rückstieß et al., “State-Dependent Exploration for Policy Gradient Methods,” ECML PKDD 2008, Part II, LNAI 5212, 234-249, 2008. [cited by applicant]
Schaul et al., “A Scalable Neural Network Architecture for Board Games,” Computational Intelligence and Games, 2008. CIG '08. IEEE Symposium on, Dec. 15-18, 2008, 357-364. [cited by applicant]
Schmidhuber et al., “Book on Recurrent Neural Networks,” Jan. 2011 [retrieved on Jun. 10, 2013]. Retrieved from the internet: URL<http://www.idsia.ch/˜juergen/rnnbook.html>, 2 pages. [cited by applicant]
Schmidhuber, “Recurrent Neural Networks,” 2011 [retrieved on Jun. 10, 2013]. Retrieved from the internet: URL<http://www.idsia.ch/˜juergen/rnn.html>, 8 pages. [cited by applicant]
Schuster and Paliwal, “Bidirectional Recurrent Neural Networks,” IEEE Transactions on Signal Processing, 1997, 45(11):2673-2681. [cited by applicant]
Schuster, “Bi-directional Recurrent Neural Networks for Speech Recognition,” Technical report, 1996, 2 pages. [cited by applicant]
Sehnke et al., “Policy Gradients with Parameter-Based Exploration for Control,” ICANN 2008, Part I, LNCS 5163, 387-396, 2008. [cited by applicant]
Unkelbach et al., “An EM based training algorithm for recurrent neural networks,” ICANN 2009, Part I, LNCS 5768, 964-974, 2009. [cited by applicant]
Vinyals, et al., “Revisiting Recurrent Neural Networks for Robust ASR,” IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2012, pp. 4085-4088. [cited by applicant]
Wierstra et al., “Fitness Expectation Maximization,” PPSN X, LNCS 5199, 337-346, 2008. [cited by applicant]
Wierstra et al., “Natural Evolution Strategies,” IEEE, Jun. 1-6, 2008, 3381-3387. [cited by applicant]
Wierstra et al., “Policy Gradient Critics,” ECML 2007, LNAI 4701, 466-477, 2007. [cited by applicant]