IP Library Granted Patent US 11,756,535
Granted Patent B1
US 11,756,535 · App. 17/700,234 · Granted Sep 12, 2023

System and method for speech recognition using deep recurrent neural networks

Inventor: Alexander B. Graves (London, GB)
Assignee: Google LLC
G10L15/16G06N3/044G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,756,535
App. No.
17/700,234
Granted
Sep 12, 2023
Kind
B1
Abstract

Deep recurrent neural networks applied to speech recognition. The deep recurrent neural networks (RNNs) are preferably implemented by stacked long short-term memory bidirectional RNNs. The RNNs are trained using end-to-end training with suitable regularisation.

Claims (36)

1. A method performed by one or more computers and for training

a deep recurrent neural network (“deep RNN”) implemented on the one or more computers to perform speech recognition, the method comprising:

receiving training data for training the deep RNN, the training data comprising a plurality of sequences of audio observations and, for each sequence of audio observations, a corresponding sequence of text symbols that represents the sequence of audio observations comprising, and the deep RNN comprising:

a transcription neural network configured to generate a sequence of transcription hidden vectors that includes a respective transcription hidden vector for each audio observation in each sequence of audio observations;

a prediction neural network configured to, at position u in an output sequence generated from a particular sequence of audio observations, process a current output sequence that includes output labels at positions 1 through u−1 in the output sequence to generate a prediction hidden vector; and

an output neural network, configured to, at the position u in the output sequence generated from the particular sequence of audio observations, process the prediction hidden vector and a respective transcription hidden vector for an audio observation at a position t in the particular sequence of audio observations, to generate a probability distribution comprising a respective probability for each output label in a vocabulary of output labels; and

training the deep RNN to map the sequences of audio observations to the corresponding sequences of text symbols.

2. The method of claim 1 , wherein the transcription neural network is a bi-directional recurrent neural network.

3. The method of claim 1 , wherein the transcription neural network is a bi-directional long short-term memory (LSTM) neural network.

4. The method of claim 3 , wherein each transcription hidden vector is a concatenation of hidden vectors for the corresponding audio input generated by a set of uppermost bidirectional layers in the bi-directional LSTM neural network.

5. The method of claim 1 , wherein the prediction neural network is a uni-directional recurrent neural network.

6. The method of claim 1 , wherein the prediction neural network is a uni-directional long short-term memory (LSTM) neural network.

7. The method of claim 1 , wherein the vocabulary of output labels also includes a blank symbol that represents a non-output.

8. The method of claim 1 , wherein the transcription neural network is initialized from a CTC-trained neural network.

9. The method of claim 1 , wherein the prediction neural network is initialized from a next-step prediction network.

10. A system comprising one or more computers and one or more storage devices storing instructions for training a deep recurrent neural network (“deep RNN”) implemented on the one or more computers to perform speech recognition, the operations comprising:

receiving training data for training the deep RNN, the training data comprising a plurality of sequences of audio observations and, for each sequence of audio observations, a corresponding sequence of text symbols that represents the sequence of audio observations comprising, and the deep RNN comprising:

a transcription neural network configured to generate a sequence of transcription hidden vectors that includes a respective transcription hidden vector for each audio observation in each sequence of audio observations;

a prediction neural network configured to, at position u in an output sequence generated from a particular sequence of audio observations, process a current output sequence that includes output labels at positions 1 through u−1 in the output sequence to generate a prediction hidden vector; and

an output neural network, configured to, at the position u in the output sequence generated from the particular sequence of audio observations, process the prediction hidden vector and a respective transcription hidden vector for an audio observation at a position t in the particular sequence of audio observations, to generate a probability distribution comprising a respective probability for each output label in a vocabulary of output labels; and

training the deep RNN to map the sequences of audio observations to the corresponding sequences of text symbols.

11. The system of claim 10 , wherein the transcription neural network is a bi-directional recurrent neural network.

12. The system of claim 10 , wherein the transcription neural network is a bi-directional long short-term memory (LSTM) neural network.

13. The system of claim 12 , wherein each transcription hidden vector is a concatenation of hidden vectors for the corresponding audio input generated by a set of uppermost bidirectional layers in the bi-directional LSTM neural network.

14. The system of claim 10 , wherein the prediction neural network is a uni-directional recurrent neural network.

15. The system of claim 10 , wherein the prediction neural network is a uni-directional long short-term memory (LSTM) neural network.

16. The system of claim 10 , wherein the vocabulary of output labels also includes a blank symbol that represents a non-output.

17. The system of claim 10 , wherein the transcription neural network is initialized from a CTC-trained neural network.

18. The system of claim 10 , wherein the prediction neural network is initialized from a next-step prediction network.

19. One or more non-transitory computer-readable media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for training a deep recurrent neural network (“deep RNN”) implemented on the one or more computers to perform speech recognition, the operations comprising:

receiving training data for training the deep RNN, the training data comprising a plurality of sequences of audio observations and, for each sequence of audio observations, a corresponding sequence of text symbols that represents the sequence of audio observations comprising, and the deep RNN comprising:

a transcription neural network configured to generate a sequence of transcription hidden vectors that includes a respective transcription hidden vector for each audio observation in each sequence of audio observations;

a prediction neural network configured to, at position u in an output sequence generated from a particular sequence of audio observations, process a current output sequence that includes output labels at positions 1 through u−1 in the output sequence to generate a prediction hidden vector; and

an output neural network, configured to, at the position u in the output sequence generated from the particular sequence of audio observations, process the prediction hidden vector and a respective transcription hidden vector for an audio observation at a position t in the particular sequence of audio observations, to generate a probability distribution comprising a respective probability for each output label in a vocabulary of output labels; and

training the deep RNN to map the sequences of audio observations to the corresponding sequences of text symbols.

20. The one or more non-transitory computer-readable media of claim 19 , wherein the vocabulary of output labels also includes a blank symbol that represents a non-output.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE CONVEYING PARTY EXECUTION DATE PREVIOUSLY RECORDED AT REEL: 059511 FRAME: 0249. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 9, 2023
From: GOOGLE, INC.
To: GOOGLE LLC
Reel/Frame 065531/0726 →
CHANGE OF NAME Recorded Mar 25, 2022
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 059511/0249 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2022
From: GRAVES, ALEXANDER B.
To: GOOGLE INC.
Reel/Frame 059402/0656 →
Continuity (6)
Continuation 17013276 · Sep 4, 2020
Continuation 16658697 · Oct 21, 2019
Continuation 16267078 · Feb 4, 2019
Continuation 15043341 · Feb 12, 2016
Continuation 14090761 · Nov 26, 2013
Provisional Application 61731047 · Nov 29, 2012