Neural network optimization device for edge device meeting on-demand instruction and method using the same
Disclosed is a neural network optimizing method performed by a device including performing a computation on an input value with a basic resource of a first neural network based on a customer-requested instruction for the basic resource used for computation execution of the first neural network, checking computational performance including a computational processing speed, a power consumption amount, and a chip area according to the computation execution of the first neural network, adjusting at least one of the basic resource based on the checked computational performance, and re-performing the computation on the input value based on the customer-requested instruction with a resource of the second neural network after changing to an environment of a second neural network by adjusting the at least one of the basic resource.
1 . A neural network optimizing method performed by a device, the method comprising:
performing, based on a customer-requested instruction for basic resources used for computation execution of a first neural network, a first computation on an input value with the basic resources, wherein the performing of the first computation comprises inputting the customer-requested instruction into a particular resource having a greatest size, among the basic resources;
identifying a bottleneck portion caused in a process of data exchange between data blocks of the first neural network, at an intermediate branch point between a data line and a buffer;
checking a first computational performance of the first computation, wherein the first computational performance includes a computational processing speed and a power consumption amount of the bottleneck portion;
adjusting, based on the checked computational performance and a customer-requested performance according to the customer-requested instruction, at least one resource of the basic resources, adjusting a data path of the bottleneck portion, and configuring a second neural network with the adjusted at least one resource and the adjusted data path;
performing a second computation on the input value with the adjusted at least one resource and the adjusted data path of the second neural network;
checking a second computational performance of the second computation; and
when the second computational performance of the second computation minimizes consumption of the adjusted at least one resource and latency of the adjusted data path, outputting the second neural network as a final neural network,
wherein the customer-requested instruction is an instruction for generating the final neural network that meets a customer's request,
wherein the adjusting of the at least one resource of the basic resources comprises:
adjusting a first resource among the basic resources, wherein the first resource is configured to adjust internal memory capacity and a number of channels;
determining whether the customer-requested performance is satisfied with the adjusted, first resource;
when the customer-requested performance is satisfied with the adjusted first resource, further adjusting a second resource among the basic resources, wherein the second resource is configured to adjust a pipeline and a data bandwidth;
determining whether the customer-requested performance is satisfied with the adjusted second resource; and
when the customer-requested performance is satisfied with the adjusted second resource, further adjusting a third resource among the basic resources, wherein the third resource is configured to perform an additional function of the customer's request, and
wherein the method performs all of the performing of the first computation, the identifying of the bottleneck portion, the checking of the first computational performance, the adjusting of the at least one resource, the adjusting of the data path, the configuring of the second neural network, the performing of the second computation, and the outputting of the second neural network, again, for a different customer-requested instruction.
2 . A neural network optimization device comprising:
a computation execution block configured to perform, based on a customer-requested instruction for basic resources used for computation execution of a first neural network, a first computation on an input value with the basic resources, wherein the first computation is performed by inputting the customer-requested instruction into a particular resource having a greatest size, among the basic resources;
a computation performance check block configured to identify a bottleneck portion caused in a process of data exchange between data blocks of the first neural network, at an intermediate branch point between a data line and a buffer, and check a first computational performance of the first computation, wherein the first computational performance includes a computational processing speed and a power consumption amount of the bottleneck portion;
a resource adjustment block configured to adjust, based on the checked computational performance and a customer-requested performance according to the customer-requested instruction, at least one resource of the basic resources, adjust a data path of the bottleneck portion, and configure a second neural network with the adjusted at least one resource and the adjusted data path; and
an optimization execution block configured to perform a second computation on the input value with the adjusted at least one resource and the adjusted data path of the second neural network, check a second computational performance of the second computation, and output the second neural network as a final neural network when the second computational performance of the second computation minimizes consumption of the adjusted at least one resource and latency of the adjusted data path,
wherein the customer-requested instruction is an instruction for generating the final neural network that meets a customer's request,
wherein the resource adjustment block is configured to:
adjust a first resource among the basic resources, wherein the first resource is configured to adjust internal memory capacity and a number of channels,
determine whether the customer-requested performance is satisfied with the adjusted first resource;
when the customer-requested performance is satisfied with the adjusted first resource, further adjust a second resource among the basic resources, wherein the second resource is configured to adjust a pipeline and a data bandwidth;
determine whether the customer-requested performance is satisfied with the adjusted second resource; and
when the customer-requested performance is satisfied with the adjusted second resource, further adjust a third resource among the basic resources, wherein the third resource is configured to perform an additional function of the customer's request, and
wherein the neural network optimization device performs all of performing of the first computation, the identifying of the bottleneck portion, the checking of the first computational performance, the adjusting of the at least one resource, the adjusting of the data path, the configuring of the second neural network, the performing of the second computation, and the outputting of the second neural network, again, for a different customer-requested instruction.
3 . A non-transitory computer-readable recording medium that stores a computer program which, when executed by one or more processors, performs following operations for performing a neural network optimizing method, the operations including:
performing, based on a customer-requested instruction for basic resources used for computation execution of a first neural network, a first computation on an input value with the basic resources, wherein the performing of the first computation comprises inputting the customer-requested instruction into a particular resource having a greatest size, among the basic resources;
identifying a bottleneck portion caused in a process of data exchange between data blocks of the first neural network, at an intermediate branch point between a data line and a buffer;
checking a first computational performance of the first computation, wherein the first computational performance includes a computational processing speed and a power consumption amount of the bottleneck portion;
adjusting, based on the checked computational performance and a customer-requested performance according to the customer-requested instruction, at least one resource of the basic resources, adjusting a data path of the bottleneck portion, and configuring a second neural network with the adjusted at least one resource and the adjusted data path;
performing a second computation on the input value with the adjusted at least one resource and the adjusted data path of the second neural network;
checking a second computational performance of the second computation; and
when the second computational performance of the second computation minimizes consumption of the adjusted at least one resource and latency of the adjusted data path, outputting the second neural network as a final neural network,
wherein the customer-requested instruction is an instruction for generating the final neural network that meets a customer's request,
wherein the adjusting of the at least one resource of the basic resources comprises:
adjusting a first resource among the basic resources, wherein the first resource is configured to adjust internal memory capacity and a number of channels;
determining whether the customer-requested performance is satisfied with the adjusted first resource;
when the customer-requested performance is satisfied with the adjusted first resource, further adjusting a second resource among the basic resources, wherein the second resource is configured to adjust a pipeline and a data bandwidth;
determining whether the customer-requested performance is satisfied with the adjusted second resource; and
when the customer-requested performance is satisfied with the adjusted second resource, further adjusting a third resource among the basic resources, wherein the third resource is configured to perform an additional function of the customer's request, and
wherein the method performs all of the performing of the first computation, the identifying of the bottleneck portion, the checking of the first computational performance, the adjusting of the at least one resource, the adjusting of the data path, the configuring of the second neural network, the performing of the second computation, and the outputting of the second neural network, again, for a different customer-requested instruction.