IP Library Granted Patent US 10,438,581
Granted Patent B2
US 10,438,581 · App. 13/955,483 · Granted Oct 8, 2019

Speech recognition using neural networks

Inventors: Andrew W. Senior (New York, NY); Ignacio L. Moreno (New York, NY)
Assignee: Google LLC
G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,438,581
App. No.
13/955,483
Granted
Oct 8, 2019
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for speech recognition using neural networks. A feature vector that models audio characteristics of a portion of an utterance is received. Data indicative of latent variables of multivariate factor analysis is received. The feature vector and the data indicative of the latent variables is provided as input to a neural network. A candidate transcription for the utterance is determined based on at least an output of the neural network.

Claims (78)

1. A method performed by data processing apparatus, the method comprising:

receiving a feature vector that models audio characteristics of a portion of an utterance;

receiving data indicative of latent variables of multivariate factor analysis;

providing the feature vector and the data indicative of the latent variables as input to an input layer of a neural network comprising the input layer, multiple hidden layers, and an output layer; and

determining a candidate transcription for the utterance based on at least an output of the neural network that the neural network provides at the output layer in response to the feature vector and the data indicative of the latent variables being provided at the input layer.

2. The method of claim 1 , wherein receiving data indicative of latent variables of multivariate factor analysis comprises receiving data indicative of latent variables of multivariate factor analysis of an audio signal that includes the utterance; and

wherein providing, as input to a neural network, the feature vector and the data indicative of the latent variables comprises providing, as input to a neural network, the feature vector and the data indicative of the latent variables of multivariate factor analysis of the audio signal that includes the utterance.

3. The method of claim 1 , wherein the utterance is uttered by a speaker; and

wherein receiving data indicative of latent variables of multivariate factor analysis comprises receiving data indicative of latent variables of multivariate factor analysis of an audio signal that (i) does not include the utterance and (ii) includes other utterances uttered by the speaker.

4. The method of claim 1 , wherein receiving data indicative of latent variables of multivariate factor analysis comprises receiving an i-vector indicating time-independent audio characteristics; and

wherein providing, as input to a neural network, the feature vector and the data indicative of the latent variables comprises providing, as input to a neural network, the feature vector and the i-vector.

5. The method of claim 1 , wherein the utterance is uttered by a speaker, and the data indicative of latent variables of multivariate factor analysis is first data indicative of latent variables of multivariate factor analysis;

further comprising:

determining that an amount of audio data received satisfies a threshold;

in response to determining that the amount of audio data received satisfies the threshold, obtaining second data indicative of latent variables of multivariate factor analysis, the second data being different from the first data;

providing, as input to the neural network, a second feature vector that models audio characteristics of a portion of a second utterance and the second data; and

determining a candidate transcription for the second utterance based on at least a second output of the neural network.

6. The method of claim 1 , wherein providing, as input to a neural network, the feature vector and the data indicative of the latent variables comprises providing the feature vector and the data indicative of the latent variables to a neural network trained using audio data and data indicative of latent variables of multivariate factor analysis corresponding to the audio data;

wherein determining the candidate transcription for the utterance based on at least an output of the neural network comprises:

receiving, as an output of the neural network, data indicating a likelihood that the feature vector corresponds to a particular phonetic unit; and

determining the candidate transcription based on the data indicating the likelihood that the feature vector corresponds to the particular phonetic unit.

7. The method of claim 1 , further comprising:

recognizing a first portion of a speech sequence by providing a feature vector to an acoustic model that does not receive data indicative of latent variables of multivariate factor analysis as input; and

after recognizing the first portion of the speech sequence, determining that at least a minimum amount of audio of the speech sequence has been received;

wherein:

the utterance occurs in the speech sequence after the first portion;

receiving data indicative of latent variables of multivariate factor analysis comprises receiving data indicative of latent variables of multivariate factor analysis of received audio including the first portion of the speech sequence; and

providing, as input to a neural network, the feature vector and the data indicative of the latent variables comprises providing the feature vector and the data indicative of the latent variables as input to a neural network that is different from the acoustic model.

8. The method of claim 1 , further comprising:

receiving a second feature vector that models audio characteristics of a second portion of the utterance;

providing, as input to a neural network, the second feature vector and the data indicative of the latent variables; and

receiving a second output from the neural network;

wherein determining the candidate transcription for the utterance is based on at least the output of the neural network and the second output of the neural network.

9. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving a feature vector that models audio characteristics of a portion of an utterance;

receiving data indicative of latent variables of multivariate factor analysis;

providing the feature vector and the data indicative of the latent variables as input to an input layer of a neural network comprising the input layer, multiple hidden layers, and an output layer; and

determining a candidate transcription for the utterance based on at least an output of the neural network that the neural network provides at the output layer in response to the feature vector and the data indicative of the latent variables being provided at the input layer.

10. The system of claim 9 , wherein receiving data indicative of latent variables of multivariate factor analysis comprises receiving data indicative of latent variables of multivariate factor analysis of an audio signal that includes the utterance; and

wherein providing, as input to a neural network, the feature vector and the data indicative of the latent variables comprises providing, as input to a neural network, the feature vector and the data indicative of the latent variables of multivariate factor analysis of the audio signal that includes the utterance.

11. The system of claim 9 , wherein the utterance is uttered by a speaker; and

wherein receiving data indicative of latent variables of multivariate factor analysis comprises receiving data indicative of latent variables of multivariate factor analysis of an audio signal that (i) does not include the utterance and (ii) includes other utterances uttered by the speaker.

12. The system of claim 9 , wherein receiving data indicative of latent variables of multivariate factor analysis comprises receiving an i-vector indicating time-independent audio characteristics; and

wherein providing, as input to a neural network, the feature vector and the data indicative of the latent variables comprises providing, as input to a neural network, the feature vector and the i-vector.

13. The system of claim 9 , wherein the utterance is uttered by a speaker, and the data indicative of latent variables of multivariate factor analysis is first data indicative of latent variables of multivariate factor analysis;

wherein the operations further comprise:

determining that an amount of audio data received satisfies a threshold;

in response to determining that the amount of audio data received satisfies the threshold, obtaining second data indicative of latent variables of multivariate factor analysis, the second data being different from the first data;

providing, as input to the neural network, a second feature vector that models audio characteristics of a portion of a second utterance and the second data; and

determining a candidate transcription for the second utterance based on at least a second output of the neural network.

14. The system of claim 9 , wherein providing, as input to a neural network, the feature vector and the data indicative of the latent variables comprises providing the feature vector and the data indicative of the latent variables to a neural network trained using audio data and data indicative of latent variables of multivariate factor analysis corresponding to the audio data;

wherein determining the candidate transcription for the utterance based on at least an output of the neural network comprises:

receiving, as an output of the neural network, data indicating a likelihood that the feature vector corresponds to a particular phonetic unit; and

determining the candidate transcription based on the data indicating the likelihood that the feature vector corresponds to the particular phonetic unit.

15. The system of claim 9 , wherein the operations further comprise:

recognizing a first portion of a speech sequence by providing a feature vector to an acoustic model that does not receive data indicative of latent variables of multivariate factor analysis as input; and

after recognizing the first portion of the speech sequence, determining that at least a minimum amount of audio of the speech sequence has been received;

wherein:

the utterance occurs in the speech sequence after the first portion;

receiving data indicative of latent variables of multivariate factor analysis comprises receiving data indicative of latent variables of multivariate factor analysis of received audio including the first portion of the speech sequence; and

providing, as input to a neural network, the feature vector and the data indicative of the latent variables comprises providing the feature vector and the data indicative of the latent variables as input to a neural network that is different from the acoustic model.

16. The system of claim 9 , wherein the operations further comprise:

receiving a second feature vector that models audio characteristics of a second portion of the utterance;

providing, as input to a neural network, the second feature vector and the data indicative of the latent variables; and

receiving a second output from the neural network;

wherein determining the candidate transcription for the utterance is based on at least the output of the neural network and the second output of the neural network.

17. A non-transitory computer-readable storage device encoded with a computer program, the program comprising instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

receiving a feature vector that models audio characteristics of a portion of an utterance;

receiving data indicative of latent variables of multivariate factor analysis;

providing the feature vector and the data indicative of the latent variables as input to an input layer of a neural network comprising the input layer, multiple hidden layers, and an output layer; and

determining a candidate transcription for the utterance based on at least an output of the neural network that the neural network provides at the output layer in response to the feature vector and the data indicative of the latent variables being provided at the input layer.

18. The non-transitory computer-readable storage device of claim 17 , wherein receiving data indicative of latent variables of multivariate factor analysis comprises receiving data indicative of latent variables of multivariate factor analysis of an audio signal that includes the utterance; and

wherein providing, as input to a neural network, the feature vector and the data indicative of the latent variables comprises providing, as input to a neural network, the feature vector and the data indicative of the latent variables of multivariate factor analysis of the audio signal that includes the utterance.

19. The non-transitory computer-readable storage device of claim 17 , wherein the utterance is uttered by a speaker; and

wherein receiving data indicative of latent variables of multivariate factor analysis comprises receiving data indicative of latent variables of multivariate factor analysis of an audio signal that (i) does not include the utterance and (ii) includes other utterances uttered by the speaker.

20. The non-transitory computer-readable storage device of claim 17 , wherein receiving data indicative of latent variables of multivariate factor analysis comprises receiving an i-vector indicating time-independent audio characteristics; and

wherein providing, as input to a neural network, the feature vector and the data indicative of the latent variables comprises providing, as input to a neural network, the feature vector and the i-vector.

Assignments (2)
CHANGE OF NAME Recorded Dec 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044695/0115 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2013
From: SENIOR, ANDREW W.; MORENO, IGNACIO L.
To: GOOGLE INC.
Reel/Frame 031068/0902 →
Continuity (1)
Related Publication 20150039301A1 · Feb 5, 2015