Method and apparatus for optimizing batch size for artificial neural network accelerator
View Patent ↗A method for optimizing a batch size for an artificial neural network accelerator that processes at least one batch in an apparatus for optimizing a batch size is provided. The method for optimizing a batch size includes: receiving information from an artificial neural network to determine a batch size; and determining the batch size for optimizing basic performance of the artificial neural network according to the artificial neural network.
1 . A method performed by an artificial neural network accelerator, the method comprising:
receiving information via an input interface from an artificial neural network to determine a batch size for batches within internal memory banks of the artificial neural network,
wherein the information from the artificial neural network includes information associated with a hardware configuration of the internal memory banks of the artificial neural network;
optimizing a throughput based performance metric of the artificial neural network to determine the batch size for the artificial neural network, by:
determining a plurality of candidate numbers of batches constrained by a total number of the internal memory banks;
evaluating the throughput based performance metric of the artificial neural network for each of the plurality of the candidate numbers of batches;
selecting, as the batch size, a number of batches corresponding to a maximum value of the throughput based performance metric among the plurality of candidate numbers of batches; and
performing operations within the artificial neural network accelerator based on the determined batch size in parallel using the internal memory banks.
2 . The method of claim 1 , wherein the determining of the batch size includes determining one batch size for the entire artificial neural network.
3 . The method of claim 1 , wherein the determining of the batch size includes determining a different batch size for each layer of the artificial neural network.
4 . The method of claim 3 , wherein the determining of a different batch size for each layer includes:
optimizing the throughput based performance metric of the artificial neural network based on each batch size for each layer;
determining a batch size having a best throughput maximum value of the throughput based performance metric for each layer as the batch size of the layer.
5 . An apparatus, comprising:
an input interface that receives information from an artificial neural network to determine a batch size,
wherein the information from the artificial neural network includes information associated with internal memory banks of the artificial neural network, including an internal memory bank structure and an internal memory bank size of the artificial neural network;
a processor optimizes a throughput based performance metric of the artificial neural network to determine the batch size for the artificial neural network, by:
determining a plurality of candidate numbers of batches constrained by a total number of internal memory banks;
evaluating the throughput based performance metric of the artificial neural network for each of the plurality of the candidate numbers of batches;
selecting, as the batch size, a number of batches corresponding to a maximum value of the throughput based performance metric among the plurality of candidate numbers of batches; and
performing operations within a artificial neural network accelerator based on the determined batch size in parallel using the internal memory banks; and
an output interface that transmits the determined batch size to the artificial neural network accelerator.
6 . The apparatus of claim 5 , wherein the processor determines a different batch size for each layer of the artificial neural network.
7 . The apparatus of claim 6 , wherein the processor optimizes a throughput based performance metric for the artificial neural network based on each batch size for each layer and determines a batch size having a maximum value of the throughput based performance metric for each layer as the batch size of the layer.
8 . The apparatus of claim 5 , wherein the processor buffers input data to match the determined batch size.
9 . The apparatus of claim 5 , wherein the processor optimizes the throughput based performance metric of the artificial neural network by changing a number of batches associated with the batch size to a total number of internal memory banks of the artificial neural network.