Method and system for lightweighting artificial neural network model, and non-transitory computer-readable recording medium
A method for light-weighting an artificial neural network model, the method comprising is provided. The method includes the steps of: learning, on the basis of a channel length of a 1×1 pruning unit applied in a channel direction to each of a plurality of kernels included in an artificial neural network model, a pruning factor for each of the pruning units and a weight for each of the pruning units; and determining, on the basis of at least one of the learned pruning factors and weights, which of the pruning units is to be removed from the artificial neural network model, wherein the channel length of the pruning unit is shorter than a channel length of at least a part of the plurality of kernels.
1 . A method for a device connecting to and communicating with an
artificial neural network model light-weighting system, wherein the device includes an application for a user to receive distribution of a light-weighted artificial neural network model services, the method comprising the steps of:
downloading, from the artificial neural network model light-weighting system, the application;
learning, by the application, on the basis of a channel length of a 1×1 pruning unit applied in a channel direction to each of a plurality of kernels included in an artificial neural network model, a pruning factor for each of the pruning units and a weight for each of the pruning units; and
determining, on the basis of at least one of the learned pruning factors and weights, which of the pruning units is to be removed from the artificial neural network model, wherein the channel length of the pruning unit is shorter than a channel length of at least a part of the plurality of kernels, wherein
artificial neural network models of various structures are distributed to the device with only one-time learning and ensures
that artificial neural network models suitable for various edge computing environments.
2 . The method of claim 1 , wherein the channel length of the pruning unit is commonly applied to a plurality of convolution layers included in the artificial neural network model.
3 . The method of claim 2 , wherein the plurality of convolution layers include a first convolution layer and a second convolution layer having a different channel length from the first convolution layer, and
wherein in the determining step, the pruning unit to be removed is shared in the first convolution layer and the second convolution layer.
4 . The method of claim 1 , wherein the channel length of the pruning unit is a power of 2.
5 . A non-transitory computer-readable recording medium having stored thereon a computer program for executing the method of claim 1 .
6 . An artificial neural network model light-weighting system where a device connects to and communicating with
artificial neural network model light-weighting system, wherein the device includes an application for a user to receive distribution of a light-weighted artificial neural network model services, the system comprising:
storing the application for downloading to the device;
a pruning factor learning unit configured to learn using the application, on the basis of a channel length of a 1×1 pruning unit applied in a channel direction to each of a plurality of kernels, a pruning factor for each of the pruning units and a weight for each of the pruning units; and
an artificial neural network model management unit configured to determine, on the basis of at least one of the learned pruning factors and weights, which of the pruning units is to be removed from an artificial neural network model, wherein the channel length of the pruning unit is shorter than a channel length of at least a part of the plurality of kernels, wherein
artificial neural network models of various structures are distributed to the device with only one-time learning and ensures that artificial
neural network models suitable for various edge computing environments.