IP Library Granted Patent US 10,672,384
Granted Patent B2
US 10,672,384 · App. 16/573,323 · Granted Jun 2, 2020

Asynchronous optimization for sequence training of neural networks

Inventors: Georg Heigold (Mountain View, CA); Erik McDermott (San Francisco, CA); Vincent O. Vanhoucke (San Francisco, CA); Andrew W. Senior (London, GB); Michiel A. U. Bacchiani (Summit, NJ)
Assignee: Google LLC
G10L15/063G06N3/0454G10L15/16G10L15/183
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,672,384
App. No.
16/573,323
Granted
Jun 2, 2020
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for obtaining, by a first sequence-training speech model, a first batch of training frames that represent speech features of first training utterances; obtaining, by the first sequence-training speech model, one or more first neural network parameters; determining, by the first sequence-training speech model, one or more optimized first neural network parameters based on (i) the first batch of training frames and (ii) the one or more first neural network parameters; obtaining, by a second sequence-training speech model, a second batch of training frames that represent speech features of second training utterances; obtaining one or more second neural network parameters; and determining, by the second sequence-training speech model, one or more optimized second neural network parameters based on (i) the second batch of training frames and (ii) the one or more second neural network parameters.

Claims (34)

1. A method performed by one or more computers, the method comprising:

obtaining, by the one or more computers, a replica of a neural network to be trained;

updating, by the one or more computers, parameters of the replica of the neural network based on data indicating model parameters determined based on training operations for one or more other replicas of the neural network;

after updating the parameters of the replica of the neural network, performing, by the one or more computers, one or more training operations for the replica of the neural network based on a subset of training data for training the neural network; and

sending, by the one or more computers, data indicating results of the one or more training operations.

2. The method of claim 1 , wherein the subset of training data comprises data indicating of one or more utterances.

3. The method of claim 1 , wherein performing the one or more training operations comprises performing the one or more training operations asynchronously with respect to training operations for the one or more other replicas of the neural network.

4. The method of claim 1 , wherein sending the data indicating results of the one or more training operations comprises sending the data to indicating results of the one or more training operations to a server.

5. The method of claim 1 , wherein sending the data indicating results of the one or more training operations comprises sending data indicating one or more model parameter updates.

6. The method of claim 1 , wherein sending the data indicating results of the one or more training operations comprises sending data indicating one or more gradients.

7. The method of claim 1 , wherein the neural network is a speech recognition model.

8. The method of claim 1 , wherein the neural network is an acoustic model.

9. The method of claim 1 , wherein the neural network is a speech model trained to indicate likelihoods that acoustic feature vectors represent different phonetic units.

10. The method of claim 1 , wherein the subset of training data is pseudo-randomly selected from the training data.

11. A system comprising:

one or more computers; and

one or more computer-readable media storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

obtaining, by the one or more computers, a replica of a neural network to be trained;

updating, by the one or more computers, parameters of the replica of the neural network based on data indicating model parameters determined based on training operations for one or more other replicas of the neural network;

after updating the parameters of the replica of the neural network, performing, by the one or more computers, one or more training operations for the replica of the neural network based on a subset of training data for training the neural network; and

sending, by the one or more computers, data indicating results of the one or more training operations.

12. The system of claim 11 , wherein the subset of training data comprises data indicating of one or more utterances.

13. The system of claim 11 , wherein performing the one or more training operations comprises performing the one or more training operations asynchronously with respect to training operations for the one or more other replicas of the neural network.

14. The system of claim 11 , wherein sending the data indicating results of the one or more training operations comprises sending the data to indicating results of the one or more training operations to a server.

15. The system of claim 11 , wherein sending the data indicating results of the one or more training operations comprises sending data indicating one or more model parameter updates.

16. The system of claim 11 , wherein sending the data indicating results of the one or more training operations comprises sending data indicating one or more gradients.

17. The system of claim 11 , wherein the neural network is a speech recognition model.

18. The system of claim 11 , wherein the neural network is an acoustic model.

19. The system of claim 11 , wherein the neural network is a speech model trained to indicate likelihoods that acoustic feature vectors represent different phonetic units.

20. One or more non-transitory computer-readable media storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

obtaining, by the one or more computers, a replica of a neural network to be trained;

updating, by the one or more computers, parameters of the replica of the neural network based on data indicating model parameters determined based on training operations for one or more other replicas of the neural network;

after updating the parameters of the replica of the neural network, performing, by the one or more computers, one or more training operations for the replica of the neural network based on a subset of training data for training the neural network; and

sending, by the one or more computers, data indicating results of the one or more training operations.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 17, 2019
From: HEIGOLD, GEORG; MCDERMOTT, ERIK; VANHOUCKE, VINCENT O.; SENIOR, ANDREW W.; BACCHIANI, MICHIEL A.U.
To: GOOGLE INC.
Reel/Frame 050408/0114 →
ENTITY CONVERSION Recorded Sep 17, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 050409/0719 →
Continuity (4)
Continuation 15910720 · Mar 2, 2018
Continuation 14258139 · Apr 22, 2014
Provisional Application 61899466 · Nov 4, 2013
Related Publication 20200118549A1 · Apr 16, 2020