Accelerator system for training deep neural network model using NAND flash memory and operating method thereof
A DNN accelerator system includes a plurality of accelerator nodes each including a plurality of NAND flash memories, a flash memory system (FMS) controller for controlling the plurality of NAND flash memories, and a tensor buffer, and a processor configured to generate an operation sequence of the plurality of accelerator nodes, in which a DNN model is trained in a data parallel manner using the plurality of accelerator nodes.
1 . A deep neural network (DNN) accelerator system comprising:
a plurality of accelerator nodes each including a plurality of NAND flash memories, a flash memory system (FMS) controller for controlling the plurality of NAND flash memories, and a tensor buffer; and
a processor configured to generate an operation sequence of the plurality of accelerator nodes,
wherein a DNN model is trained in a data parallel manner using the plurality of accelerator nodes,
wherein the plurality of accelerator nodes are configured to participate in a multi-phase training process including generation of intermediate training data and updating of model parameters across multiple layers of the DNN model,
wherein, during training of the DNN model:
intermediate training data generated by execution of one or more layers of the DNN model is stored in the plurality of NAND flash memories and retained for reuse during subsequent training operations;
different classes of training data generated during training are stored and managed differently within the plurality of NAND flash memories based on a role of the training data in the training process; and
the FMS controller manages allocation and access to the plurality of NAND flash memories based on a training phase associated with the stored training data.
2 . The DNN accelerator system according to claim 1 , wherein the plurality of accelerator nodes includes one or more activation nodes for performing a series of operations for DNN training and includes a weight node for managing weights.
3 . The DNN accelerator system according to claim 2 , wherein:
the multi-phase training process includes a forward propagation step, a back propagation step, and a step of updating a final weight by the one or more activation nodes and the weight node; and
the forward propagation step and the back propagation step are performed for each layer included in the DNN model, and the step of updating the final weight is performed after operations for all layers included in the DNN model are completed.
4 . The DNN accelerator system according to claim 3 , wherein:
in the forward propagation step, when calculation of each layer is completed in each of the one or more activation nodes such that intermediate training data comprising activation data is generated in an SRAM, the activation data is copied to a tensor buffer of each of the activation nodes, and the activation data stored in the tensor buffer is stored in the plurality of NAND flash memories of each of the activation nodes for reuse in the back propagation step;
in the back propagation step, data necessary for weight calculation of a current operation layer among pieces of the activation data stored in the plurality of NAND flash memories is read into the tensor buffer at each of the activation nodes, the activation data is loaded into the SRAM to start calculation in response to completion of reading of the activation data, and intermediate training data comprising gradient data is stored in the tensor buffer of each of the activation nodes in response to completion of calculation; and
in the step of updating the final weight, final weight gradient data calculated in each of the one or more activation nodes is transmitted to the tensor buffer of the weight node, the weight data is updated with a training result, and the updated weight data is stored in the plurality of NAND flash memories of the weight node.
5 . The DNN accelerator system according to claim 1 , wherein the FMS controller allocates blocks of the plurality of NAND flash memories based on a round-robin policy for incremental sequential writing during training of the DNN model.
6 . The DNN accelerator system according to claim 1 , wherein the accelerator nodes position data in one or more independent storage areas on an FMS providing different storage functions according to data characteristics associated with different classes of training data.
7 . The DNN accelerator system according to claim 6 , wherein the storage areas on the FMS providing different storage functions include:
a first storage area which is non-volatile, is configured to retain data across multiple training operations, and exclusively allows data to be read therefrom or data to be sequentially written thereto according to a function of the accelerator nodes;
a second storage area which is volatile, is configured to retain data during execution of one or more training operations, and allows data to be accessed exclusively in a sequential write/read manner; and
a third storage area which is non-volatile, is configured to retain data after completion of training, and exclusively allows data to be written thereto or allows data to be accessed in a sequential write/read manner according to a function of the accelerator nodes.
8 . The DNN accelerator system according to claim 1 , wherein the tensor buffer serves as a staging area between a compute core for performing a DNN training operation and the plurality of NAND flash memories during execution of the training process.
9 . The DNN accelerator system according to claim 8 , wherein the tensor buffer includes a double data rate (DDR) DRAM.
10 . The DNN accelerator system according to claim 1 , wherein a data path of the FMS is implemented to correspond to a physical hardware configuration of the plurality of the NAND flash memories.
11 . A method of training a DNN model, the method comprising:
a forward propagation step in which, while iterative training is performed for one or more layers of the DNN model, one or more activation nodes and a weight node perform an operation in each of the one or more layers, wherein intermediate training data generated during the forward propagation step is stored in non-volatile memory and retained for reuse during subsequent training operations;
a back propagation step in which the one or more activation nodes and the weight node generate gradient data according to the operation in response to each forward propagation step, wherein gradient data generated during the back propagation step is stored separately from the intermediate training data in the non-volatile memory; and
a step of updating, by the weight node, a final weight based on final gradient data in response to completion of operations of all the layers, wherein storage and access to the non-volatile memory are managed based on a training phase associated with the stored training data.
12 . The method according to claim 11 , wherein:
each of the one or more activation nodes and the weight node comprises a plurality of NAND flash memories, an FMS controller for controlling the plurality of NAND flash memories, and a tensor buffer; and
execution of a step of training the DNN model is started according to an operation sequence for training the DNN model that controls execution of the forward propagation step, the back propagation step, and the step of updating the final weight.
13 . The method according to claim 12 , wherein the forward propagation step includes:
a step of generating activation data in an SRAM by completing calculation in each of the one or more activation nodes;
a step of copying the activation data in the SRAM to the tensor buffer of each of the activation nodes; and
a step of storing the activation data stored in the tensor buffer in the plurality of NAND flash memories of each of the activation nodes as intermediate training data for reuse in the back propagation step.
14 . The method according to claim 13 , wherein the back propagation step includes:
a step of reading data necessary for weight calculation of a current operation layer among pieces of the activation data stored as intermediate training data in the plurality of NAND flash memories into the tensor buffer at each of the activation nodes;
a step of loading the activation data into the SRAM to start calculation; and
a step of storing intermediate training data comprising gradient data generated in response to completion of the calculation in the tensor buffer of each of the activation nodes.
15 . The method according to claim 14 , wherein the step of updating the final weight includes:
a step of transmitting intermediate training data comprising final weight gradient data calculated in each of the one or more activation nodes to the tensor buffer of the weight node;
updating the weight data based on a training result at the weight node; and
a step of storing the updated weight data in the plurality of NAND flash memories of the weight node as training data distinct from the intermediate training data used during the forward propagation step and the back propagation step.
16 . The method according to claim 12 , wherein each FMS controller allocates blocks of the plurality of NAND flash memories based on a round-robin policy for incremental sequential writing during execution of the forward propagation step, the back propagation step, and the step of updating the final weight.
17 . The method according to claim 12 , wherein the one or more activation nodes and the weight node position data in one or more independent storage areas on an FMS providing different storage functions according to data characteristics associated with different classes of training data.
18 . The method according to claim 17 , wherein the storage area on the FMS providing different functions includes:
a first storage area which is non-volatile, is configured to retain data across multiple training operations, and exclusively allows data to be read therefrom or data to be sequentially written thereto according to a function of the activation nodes or the weight node;
a second storage area which is volatile, is configured to retain data during execution of one or more training operations, and allows data to be accessed exclusively in a sequential write/read manner; and
a third storage area which is non-volatile, is configured to retain data after completion of training, and exclusively allows data to be written thereto or allows data to be accessed in a sequential write/read manner according to a function of the activation nodes or the weight node.
19 . The method according to claim 12 , wherein the tensor buffer includes a double data rate (DDR) DRAM.
20 . A computer-readable non-transitory recording medium storing a computer program including at least one instruction configured to execute, by a processor, the method of training the DNN model according to any one of claims 11 to 19 .