Parallel method and device for convolution computation and data loading of neural network accelerator
Disclosed are a parallel method and device for convolution computation and data loading of a neural network accelerator. The method needs two input feature maps and two convolution kernel cache blocks, and sequentially stores the input feature maps and 64 convolution kernels into cache sub-blocks according to a loading length, so as to execute convolution computation and simultaneously load data of a next group of 64 convolution kernels.
1 . A parallel method for convolution computation and data loading of a neural network accelerator, comprising:
S1: storing a frame of input feature maps into an input feature map cache, and dispersedly storing the input feature maps into input feature map sub-caches according to channels of the input feature maps;
S2: sequentially loading a group of convolution kernels into corresponding convolution kernel cache sub-blocks in a first convolution kernel cache;
S3: loading the input feature map cache and the first convolution kernel cache to execute convolution computation, putting a result into an output feature map cache, and storing a next group of convolution kernels into corresponding convolution kernel cache sub-blocks in a second convolution kernel cache, which comprises:
S31: loading an input feature map instruction parameter latch, loading a convolution kernel instruction parameter latch, and under the condition that a current instruction is an input feature map loading instruction, latching an off-chip input feature map storage address and loading an input feature map length; and latching the number of currently loaded convolution kernels, lengths of the loaded convolution kernels, a convolution kernel cache starting address and an off-chip convolution kernel storage address under the condition that the current instruction is a convolution kernel loading instruction;
S32: comparing the number of the convolution kernels, and under the condition that the total number of the convolution kernels is greater than a latch value of the number of the loaded convolution kernels and is an integer multiple, greater than 1 multiple of the number of the loaded convolution kernels in a convolution computation instruction, computing convolution and synchronously loading convolution kernels, the number of channels of the convolution kernels being the number of the loaded convolution kernels; under the condition that the total number of the convolution kernels is greater than the latch value of the number of the loaded convolution kernels and is 1 multiple of the number of the loaded convolution kernels in the convolution computation instruction, computing convolution and synchronously loading convolution kernels, the number of channels of the convolution kernels being a difference value between the total number of the convolution kernels and the latch value of the number of the loaded convolution kernels; and under the condition that the total number of the convolution kernels is equal to the latch value of the number of the loaded convolution kernels in the convolution computation instruction, determining whether to load convolution kernels in the next layer according to setting in the convolution computation instruction; and
S33: setting a loading flag according to a comparison result of the number of the convolution kernels, the loading lag representing that data is specifically loaded into which convolution kernel cache, and setting a convolution computation starting flag;
S34: synchronously carrying out convolution computation and data loading which are independent of each other; and
S35: computing the number of remaining convolution kernels after convolution computation and data loading are completed, latching data loading parameters of the convolution kernels in S34, and returning to S32 to continue being executed until convolution computation is completed;
S4: after convolution computation of the layer is completed, interchanging the input feature map cache and the output feature map cache, and using a convolution kernel cache storing an effective weight as the first convolution kernel cache to execute S3, wherein the effective weight is stored into the convolution kernel cache and is directly given by means of fields in the convolution computation instruction in S4; and
S5: determining that all convolution computation is completed.
2 . The parallel method for convolution computation and data loading of a neural network accelerator according to claim 1 , wherein the number of the remaining convolution kernels is computed after a convolution computation completion flag and a data loading completion flag are valid at the same time in S35.
3 . The parallel method for convolution computation and data loading of a neural network accelerator according to claim 1 , wherein in S4, the effective weight is stored into the convolution kernel cache, which convolution kernel cache stores the effective weight is determined according to the number of the convolution kernels of each of convolution layers and initial storage positions of the convolution kernels, which comprises:
S41: under the condition that the initial storage positions of the convolution kernels in S2 are set as the first convolution kernel cache;
S42: using the second convolution kernel cache for next convolution computation under the condition that the total number of the convolution kernels is within the number of columns of a convolution computation array;
S43: using the first convolution kernel cache for next convolution computation under the condition that the total number of the convolution kernels is an even multiple of the number of columns of the convolution computation array;
S44: using the second convolution kernel cache for next convolution computation under the condition that the total number of the convolution kernels is an odd multiple of the number of columns of the convolution computation array; and
S45: dividing the total number of the convolution kernels by the number of columns of the convolution computation array to round up to an integer under the condition that the total number of the convolution kernels is not the even multiple or the odd multiple of the number of columns of the convolution computation array,
the initial storage positions are the second convolution kernel cache, the steps being the same, and a result being opposite.
4 . The parallel method for convolution computation and data loading of a neural network accelerator according to claim 1 , wherein the channels of the input feature maps are sequentially and circularly stored into sequences corresponding to sub-caches in order in S1.
5 . The parallel method for convolution computation and data loading of a neural network accelerator according to claim 1 , wherein the sub-cache is expanded to make the plurality of sub-caches store one channel when a single feature map sub-cache is not enough to store a feature map of a single channel in S1.
6 . The parallel method for convolution computation and data loading of a neural network accelerator according to claim 1 , wherein lengths of the convolution kernel cache sub-blocks are W×H×C in S2, W representing widths of the convolution kernels, H representing heights of the convolution kernels, and C representing the number of the channels of each of the input feature maps.
7 . The parallel method for convolution computation and data loading of a neural network accelerator according to claim 1 , wherein a next frame of input feature map is loaded or convolution kernels in the next layer are loaded after all convolution computation is completed in S5.
8 . A parallel device for convolution computation and data loading of a neural network accelerator, comprising:
a convolution computation array;
an input feature map cache;
an output feature map cache; and
a convolution kernel cache, which are each connected to the convolution computation array, wherein the convolution kernel cache comprises:
a first convolution kernel cache; and
a second convolution kernel cache,
wherein the input feature map cache and the output feature map cache are consistent in structure, and a frame of input feature maps is stored into the input feature map cache, and are dispersedly stored into input feature map sub-caches according to channels of the input feature maps;
the first convolution kernel cache and the second convolution kernel cache are consistent in structure, and a group of convolution kernels are sequentially loaded into corresponding convolution kernel cache sub-blocks in the first convolution kernel cache; and
the convolution computation array is composed of a two-dimensional array, the input feature map cache and the first convolution kernel cache are loaded to execute convolution computation, a result is put into the output feature map cache, and moreover, a next group of convolution kernels are stored into corresponding convolution kernel cache sub-blocks in the second convolution kernel cache; and after convolution computation of the layer is completed, the input feature map cache and the output feature map cache are interchanged, and a convolution kernel cache storing an effective weight is used as the first convolution kernel cache to continue convolution computation until all convolution computation is completed.