IP Library Granted Patent US 11,715,486
Granted Patent B2
US 11,715,486 · App. 16/731,464 · Granted Aug 1, 2023

Convolutional, long short-term memory, fully connected deep neural networks

Inventors: Tara N. Sainath (Jersey City, NJ); Andrew W. Senior (London, GB); Oriol Vinyals (London, GB); Hasim Sak (New York, NY)
Assignee: Google LLC
G10L25/30G06N3/044G06N3/045G10L15/16G10L15/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,715,486
App. No.
16/731,464
Granted
Aug 1, 2023
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for identifying the language of a spoken utterance. One of the methods includes receiving input features of an utterance; and processing the input features using an acoustic model that comprises one or more convolutional neural network (CNN) layers, one or more long short-term memory network (LSTM) layers, and one or more fully connected neural network layers to generate a transcription for the utterance.

Claims (43)

1. A method comprising:

receiving input features, wherein the input features include respective segment features for each of a plurality of segments; and

processing the input features using a model comprising one or more convolutional neural network (CNN) layers and one or more long short-term memory (LSTM) network layers, wherein the processing comprises:

for each of the plurality of segments of the received input features:

providing the respective segment features for the respective segment to:

the one or more CNN layers, and

the one or more LSTM network layers;

generating first features for the respective segment by processing the respective segment features for the respective segment using the one or more convolutional neural network (CNN) layers, wherein the convolutional neural network (CNN) layers perform spatial modeling on the input feature;

generating second features for the respective segment by processing both the respective segment features for the respective segment and the first features generated for the respective segment using the one or more long short-term memory network (LSTM) layers to perform temporal modeling over the first features and the respective segment features, wherein a first layer of the one or more LSTM layers is configured to receive, as input, both the respective segment features for the respective segment and the first features generated for the respective segment; and

determining an output feature based on at least the second features for the plurality of segments.

2. The method of claim 1 , wherein processing the first features using the one or more LSTM layers to generate the second features comprises:

processing the first features using a linear layer to generate reduced features having a reduced dimension from a dimension of the first features; and

processing the reduced features using the one or more LSTM layers to generate the second features.

3. The method of claim 1 , wherein the one or more CNN layers, the one or more LSTM layers, and the one or more fully connected neural network layers have been jointly trained to determine trained values of parameters of the one or more CNN layers, the one or more LSTM layers, and the one or more fully connected neural network layers.

4. The method of claim 1 , wherein the input features include log-mel features having multiple dimensions.

5. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising:

receiving input features, wherein the input features include respective segment features for each of a plurality of segments; and

processing the input features using a model comprising one or more convolutional neural network (CNN) layers and one or more long short-term memory (LSTM) network layers, wherein the processing comprises:

for each of the plurality of segments of the received input features:

providing the respective segment features for the respective segment to:

the one or more CNN layers, and

the one or more LSTM network layers;

generating first features for the respective segment by processing the respective segment features for the respective segment using one or more convolutional neural network (CNN) layers, wherein the convolutional neural network (CNN) layers perform spatial modeling on the input feature;

generating second features for the respective segment by processing both the respective segment features for the respective segment and the first features generated for the respective segment using one or more long short-term memory network (LSTM) layers to perform temporal modeling over the first features and the respective segment features, wherein a first layer of the one or more LSTM layers is configured to receive, as input, both the respective segment features for the respective segment and the first features generated for the respective segment; and

determining an output feature based on at least the second features for the plurality of segments.

6. The system of claim 5 , wherein processing the first features using the one or more LSTM layers to generate the second features comprises:

processing the first features using a linear layer to generate reduced features having a reduced dimension from a dimension of the first features; and

processing the reduced features using the one or more LSTM layers to generate the second features.

7. The system of claim 5 , wherein the one or more CNN layers, the one or more LSTM layers, and the one or more fully connected neural network layers have been jointly trained to determine trained values of parameters of the one or more CNN layers, the one or more LSTM layers, and the one or more fully connected neural network layers.

8. A computer program product encoded on one or more non-transitory computer storage media, the computer program product comprising instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

receiving input features, wherein the input features include respective segment features for each of a plurality of segments; and

processing the input features using a model comprising one or more convolutional neural network (CNN) layers and one or more long short-term memory (LSTM) network layers, wherein the processing comprises:

for each of the plurality of segments of the received input features:

providing the respective segment features for the respective segment to:

the one or more CNN layers, and

the one or more LSTM network layers;

generating first features for the respective segment by processing the respective segment features for the respective segment using one or more convolutional neural network (CNN) layers, wherein the convolutional neural network (CNN) layers perform spatial modeling on the input feature;

generating second features for the respective segment by processing both the respective segment features for the respective segment and the first features generated for the respective segment using one or more long short-term memory network (LSTM) layers to perform temporal modeling over the first features and the respective segment features, wherein a first layer of the one or more LSTM layers is configured to receive, as input, both the respective segment features for the respective segment and the first features generated for the respective segment; and

determining an output feature based on at least the second features for the plurality of segments.

9. The computer program product of claim 8 , wherein processing the first features using the one or more LSTM layers to generate the second features comprises:

processing the first features using a linear layer to generate reduced features having a reduced dimension from a dimension of the first features; and

processing the reduced features using the one or more LSTM layers to generate the second features.

10. The computer program product of claim 8 , wherein the one or more CNN layers, the one or more LSTM layers, and the one or more fully connected neural network layers have been jointly trained to determine trained values of parameters of the one or more CNN layers, the one or more LSTM layers, and the one or more fully connected neural network layers.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2020
From: SAINATH, TARA N.; SENIOR, ANDREW W.; VINYALS, ORIOL; SAK, HASIM
To: GOOGLE INC.
Reel/Frame 051411/0790 →
CHANGE OF NAME Recorded Jan 3, 2020
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 051484/0141 →
Continuity (3)
Continuation 14847133 · Sep 8, 2015
Provisional Application 62059494 · Oct 3, 2014
Related Publication 20200135227A1 · Apr 30, 2020
Cited By (1)
US 12,555,356