IP Library Granted Patent US 9,811,775
Granted Patent B2
US 9,811,775 · App. 14/030,938 · Granted Nov 7, 2017

Parallelizing neural networks during training

Inventors: Alexander Krizhevsky (Toronto, CA); Ilya Sutskever (Mountain View, CA); Geoffrey E. Hinton (Toronto, CA)
Assignee: Google Inc.
G06N3/04G06K9/4628G06K9/6256G06K9/66G06N3/0454G06N3/063G06T1/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,811,775
App. No.
14/030,938
Filed
Sep 18, 2013
Granted
Nov 7, 2017
Kind
B2
Art Unit
2124
USPC
706/31
Abstract

A parallel convolutional neural network is provided. The CNN is implemented by a plurality of convolutional neural networks each on a respective processing node. Each CNN has a plurality of layers. A subset of the layers are interconnected between processing nodes such that activations are fed forward across nodes. The remaining subset is not so interconnected.

Claims (36)

1. A system for parallelizing a neural network having a plurality of parameters during training of the neural network on a training set to determine a final parameter setting for the parameters of the neural network that produces correct classifications for the training set, the system comprising:

a plurality of parallel neural networks, wherein each of the plurality of parallel neural networks is implemented on a respective computing node, and wherein the plurality of parallel neural networks each receive a same input image from the training set and collectively generate an output that classifies the input image, wherein each of the neural networks comprises a respective plurality of layers, wherein each plurality of layers comprises an interconnected layer and a non-interconnected layer, and wherein processing data through the layers of each of the plurality of parallel neural networks comprises:

providing output from the interconnected layer to at least one layer in each of the other parallel neural networks in the plurality of parallel neural networks; and

providing output from the non-interconnected layer only to a layer of the same parallel neural network, and wherein the system is configured to:

after the training, store the final parameter setting for the parameters of the neural network on one or more non-transitory computer storage media.

2. The system of claim 1 , wherein processing data through the layers of each of the plurality of parallel neural networks further comprises: providing output from the interconnected layer to at least one layer of the same parallel neural network.

3. The system of claim 1 , wherein each of the plurality of layers comprises a respective plurality of nodes and wherein each node generates a respective output activation based on an input activation received from one or more other layers.

4. The system of claim 3 , wherein providing output from the interconnected layer to at least one layer of at least one different parallel neural network of the plurality of parallel neural networks comprises providing output activations from each node of the interconnected layer to each node in at least one layer of a subset of the other parallel neural networks in the plurality of parallel neural networks.

5. The system of claim 3 , wherein providing output from the interconnected layer to at least one layer of at least one different parallel neural network of the plurality of parallel neural networks comprises providing output activations from only a subset of the nodes of the interconnected layer to a subset of the nodes in at least one layer of at least one different parallel neural network in the plurality of parallel neural networks.

6. The system of claim 1 , wherein the parallel neural networks are convolutional neural networks.

7. The system of claim 1 , wherein the non-interconnected layer is a convolutional layer.

8. The system of claim 1 , wherein the interconnected layer is a fully-connected layer.

9. A method for parallelizing a neural network having a plurality of parameters during training of the neural network on a training set to determine a final parameter setting for the parameters of the neural network that produces correct classifications for the training set, the method comprising:

processing data using each of a plurality of parallel neural networks, wherein each of the plurality of parallel neural networks is implemented on a respective computing node, wherein the plurality of parallel neural networks each receive a same input image from the training set and collectively generate an output that classifies the input image, wherein each of the neural networks comprises a respective plurality of layers, wherein each plurality of layers comprises an interconnected layer and a non-interconnected layer, wherein processing data using each of the plurality of parallel neural networks comprises processing the data through the layers of each of the plurality of parallel neural networks, and wherein processing the data through the layers of each of the plurality of parallel neural networks comprises:

providing output from the interconnected layer to at least one layer in each of the other parallel neural networks in the plurality of parallel neural networks;

providing output from the non-interconnected layer only to a layer of the same parallel neural network; and

after the training, storing the final parameter setting for the parameters of the neural network on one or more non-transitory computer storage media.

10. The method of claim 9 , wherein processing data through the layers of each of the plurality of parallel neural networks further comprises: providing output from the interconnected layer to at least one layer of the same parallel neural network.

11. The method of claim 9 , wherein each of the plurality of layers comprises a respective plurality of nodes and wherein each node generates a respective output activation based on an input activation received from one or more other layers.

12. The method of claim 11 , wherein providing output from the interconnected layer to at least one layer of at least one different parallel neural network of the plurality of parallel neural networks comprises providing output activations from each node of the interconnected layer to each node in at least one layer of a subset of the other parallel neural networks in the plurality of parallel neural networks.

13. The method of claim 11 , wherein providing output from the interconnected layer to at least one layer of at least one different parallel neural network of the plurality of parallel neural networks comprises providing output activations from only a subset of the nodes of the interconnected layer to a subset of the nodes in at least one layer of at least one different parallel neural network in the plurality of parallel neural networks.

14. The method of claim 9 , wherein the parallel neural networks are convolutional neural networks.

15. The method of claim 9 , wherein the non-interconnected layer is a convolutional layer.

16. The method of claim 9 , wherein the interconnected layer is a fully-connected layer.

17. A non-transitory computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations for parallelizing a neural network having a plurality of parameters during training of the neural network on a training set to determine a final parameter setting for the parameters of the neural network that produces correct classifications for the training set, the operations comprising:

processing data using each of a plurality of parallel neural networks, wherein each of the plurality of parallel neural networks is implemented on a respective computing node, wherein the plurality of parallel neural networks each receive a same input image from the training set and collectively generate an output that classifies the input image, wherein each of the neural networks comprises a respective plurality of layers, wherein each plurality of layers comprises an interconnected layer and a non-interconnected layer, wherein processing data using each of the plurality of parallel neural networks comprises processing the data through the layers of each of the plurality of parallel neural networks, and wherein processing the data through the layers of each of the plurality of parallel neural networks comprises:

providing output from the interconnected layer to at least one layer in each of the other parallel neural networks in the plurality of parallel neural networks;

providing output from the non-interconnected layer only to a layer of the same parallel neural network; and

after the training, storing the final parameter setting for the parameters of the neural network on one or more non-transitory computer storage media.

18. The non-transitory computer storage medium of claim 17 , wherein processing data through the layers of each of the plurality of parallel neural networks further comprises: providing output from the interconnected layer to at least one layer of the same parallel neural network.

19. The non-transitory computer storage medium of claim 17 , wherein each of the plurality of layers comprises a respective plurality of nodes and wherein each node generates a respective output activation based on an input activation received from one or more other layers.

20. The non-transitory computer storage medium of claim 19 , wherein providing output from the interconnected layer to at least one layer of at least one different parallel neural network of the plurality of parallel neural networks comprises providing output activations from each node of the interconnected layer to each node in at least one layer of a subset of the other parallel neural networks in the plurality of parallel neural networks.

21. The non-transitory computer storage medium of claim 19 , wherein providing output from the interconnected layer to at least one layer of at least one different parallel neural network of the plurality of parallel neural networks comprises providing output activations from only a subset of the nodes of the interconnected layer to a subset of the nodes in at least one layer of at least one different parallel neural network in the plurality of parallel neural networks.

22. The non-transitory computer storage medium of claim 17 , wherein the parallel neural networks are convolutional neural networks.

23. The non-transitory computer storage medium of claim 17 , wherein the non-interconnected layer is a convolutional layer.

24. The non-transitory computer storage medium of claim 17 , wherein the interconnected layer is a fully-connected layer.

Assignments (4)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044129/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2014
From: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
To: DNNRESEARCH INC.
Reel/Frame 033015/0793 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2014
From: DNNRESEARCH INC.
To: GOOGLE INC.
Reel/Frame 033016/0013 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2014
From: KRIZHEVSKY, ALEXANDER; HINTON, GEOFFREY; SUTSKEVER, ILYA
To: GOOGLE INC.
Reel/Frame 033016/0144 →
Continuity (2)
Provisional Application 61745717 · Dec 24, 2012
Related Publication 20140180989A1 · Jun 26, 2014