Method and apparatus for using a packet architecture to process neural networks in a neural processing unit
Artificial intelligence is an increasingly important sector of the computer industry. However, artificial intelligence is an extremely computationally intensive field such that it can be expensive, time consuming, and energy consuming field. Fortunately, many of the calculations required for artificial intelligence can be performed in parallel such that specialized processors can greatly increase computational performance for AI applications. Specifically, artificial intelligence generally requires large numbers of matrix operations such that specialized matrix processor circuits can greatly improve performance. To efficiently execute all these matrix operations, the matrix processor circuits must be quickly and efficiently supplied with a stream of data and instructions to process or else the matrix processor circuits end up idle. Thus, this document discloses packet architecture for efficiently creating and supplying neural network processors with work packets to process.
1 . A method of processing a multi-layer neural network, the method comprising the steps of:
dividing, with a packet compiler, each neural network layer of a set of neural network layers into a set of individual work packets for the neural network layer, each of the neural network layers comprising a set of input data elements and a set of output data elements, each individual work packet comprising (a) a subset of the set of output data elements for the neural network layer, (b) an associated weight matrix reference identifying weights specific to the subset of the set of output data elements, and (c) metadata comprising at least scheduling information for the individual work packet, the set of individual work packets for the neural network layer collectively covering all of the set of output data elements for the neural network layer;
generating, with the packet compiler, a dependency graph identifying data dependencies between the individual work packets of different neural network layers and creating an optimized packet stream order based on the data dependencies, the optimized packet stream order minimizing sequential processing of dependent work packets; and
processing, with a neural processing unit having multiple matrix processors, the individual work packets from the optimized packet stream order as processing tasks in parallel while maintaining the data dependencies identified in the dependency graph.
2 . The method of claim 1 , wherein the metadata for the individual work packet comprises a specification for a convolution operation.
3 . The method of claim 1 , wherein the output data elements are multidimensional tensors.
4 . The method of claim 1 , wherein the processing occurs on multiple neural processing units.
5 . The method of claim 4 , wherein the processing of each of the individual work packets comprises processing each of the individual work packets as native execution on a neural processing unit.
6 . The method of claim 1 , wherein each individual work packet comprises a priority value to schedule the individual work packet.
7 . The method of claim 1 , wherein the metadata for the individual work packet comprises metadata for synchronizing individual work packet execution.
8 . The method of claim 1 , wherein the metadata for each individual work packet comprises resource metadata for allocating hardware resources associated with a neural processing unit.
9 . The method of claim 1 , wherein the metadata for the individual work packet comprises data movement metadata for efficiently moving data.
10 . The method of claim 1 , wherein each individual work packet comprises an encapsulation of one or more of the individual work packets.
11 . An apparatus for processing a multi-layer neural network having a set of neural network layers, each of the neural network layers comprising a set of input data elements and a set of output data elements, the apparatus comprising:
a packet compiler configured to divide each of the neural network layers into a set of individual work packets for the neural network layer, each individual work packet comprising (a) a subset of the set of output data elements for the neural network layer, (b) an associated weight matrix reference identifying weights specific to the subset of the set of output data elements, and (c) metadata comprising at least scheduling information for the individual work packet, the set of individual work packets for the neural network layer collectively covering all of the set of output data elements for the neural network layer, the packet compiler further configured to generate a dependency graph identifying data dependencies between the individual work packets of different neural network layers and create an optimized packet stream order based on the data dependencies, the optimized packet stream order minimizing sequential processing of dependent work packets; and
a neural processing unit having multiple matrix processors, the neural processing unit configured to process the individual work packets from the optimized packet stream order as processing tasks in parallel while maintaining the data dependencies identified in the dependency graph.
12 . The apparatus of claim 11 , wherein the metadata for the individual work packet comprises a specification for a convolution operation.
13 . The apparatus of claim 11 , wherein the output data elements are multidimensional tensors.
14 . The apparatus of claim 11 , wherein the apparatus comprises multiple neural processing units.
15 . The apparatus of claim 14 , wherein each of the multiple matrix processors is configured to execute each individual work packet as a native instruction operation.
16 . The apparatus of claim 11 , wherein each individual work packet comprises a priority value to schedule the individual work packet.
17 . The apparatus of claim 11 , wherein the metadata for the individual work packet comprises metadata for synchronizing individual work packet execution.
18 . The apparatus of claim 14 , wherein the metadata for each individual work packet comprises resource metadata for allocating hardware resources associated with a neural processing unit.
19 . The apparatus of claim 11 , wherein the metadata for the individual work packet comprises data movement metadata for efficiently moving data.
20 . The apparatus of claim 11 , wherein each individual work packet comprises an encapsulation of one or more of the individual work packets.