IP Library › Granted Patent US 11,397,894
Granted Patent B2
US 11,397,894 · App. 16/022,391 · Granted Jul 26, 2022

Method and device for pruning a neural network

Inventors: Seunghwan Cho (Seoul, KR); Sungjoo Yoo (Seoul, KR); Youngjae Jin (Seoul, KR)
Assignees: SK hynix Inc.; Seoul National University R&DB Foundation
G06N3/082G06F17/13G06N3/0635G06N5/046G06N3/0454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,397,894
App. No.
16/022,391
Granted
Jul 26, 2022
Kind
B2
Abstract

A method for pruning a neural network includes initializing a plurality of threshold values respectively corresponding to a plurality of layers included in the neural network; selecting one of the plurality of layers; adjusting the threshold value of the selected layer; and adjusting a plurality of weights respectively corresponding to a plurality of synapses included in the neural network.

Claims (38)

1. A method for pruning a neural network, the method comprising:

initializing a plurality of threshold values respectively corresponding to a plurality of layers included in the neural network;

selecting one of the plurality of layers;

adjusting the threshold value of the selected layer; and

adjusting a plurality of weights respectively corresponding to a plurality of synapses included in the neural network,

wherein selecting one of the plurality of layers comprises:

adjusting the plurality of threshold values respectively corresponding to the plurality of layers;

determining reduction rates of accuracy and reduction rates of computational complexity respectively for the plurality of layers based on the adjusted plurality of threshold values; and

selecting a layer among the plurality of layers according to the reduction rates of accuracy and the reduction rates of computational complexity corresponding to the plurality of layers, and,

wherein the selected layer has a highest enhancement ratio among the plurality of layers, the enhancement ratio of the selected layer being a reduction rate of computational complexity of the selected layer divided by a reduction rate of accuracy of the selected layer.

2. The method of claim 1 , wherein the neural network is a convolutional neural network, and

wherein the plurality of layers includes an input layer and one or more inner layers.

3. The method of claim 1 , wherein a value of a neuron in the selected layer that is less than the threshold value is set to be 0.

4. The method of claim 1 , further comprising:

determining whether a pruning ratio of the neural network is less than a target value after the plurality of weights corresponding to the plurality of synapses have been adjusted.

5. The method of claim 4 , the selected layer being a first layer, the method further comprising:

when the pruning ratio is less than the target value, selecting a second layer among the plurality of layers, adjusting the threshold value of the second layer, and adjusting the plurality of weights respectively corresponding to the plurality of synapses after the threshold value of the second layer has been adjusted.

6. The method of claim 1 , wherein adjusting the plurality of weights of the plurality of synapses comprises:

calculating a loss function based on an output vector corresponding to output data, the output data being output from the neural network when input data is provided to the neural network;

calculating a plurality of partial differential equations of the loss function corresponding to the plurality of weights of the plurality of synapses; and

adjusting the plurality of weights according to a plurality of current weights of the plurality of synapses and the plurality of partial differential equations.

7. The method of claim 6 , wherein the loss function includes a variable corresponding to a distance between the output vector and a truth vector, the truth vector being given as a truth corresponding to the input data.

8. The method of claim 6 , wherein each of the plurality of partial differential equations comprises a partial differential equation of a synapse connecting a start neuron with an end neuron, the start neuron being included in a lower level layer than the end neuron, the partial differential equation of the synapse connecting the start neuron with the end neuron being based on a value of the start neuron and a partial differential equation of a synapse that connects the end neuron with a neuron in a higher level layer than the end neuron.

9. The method of claim 6 , wherein adjusting the plurality of weights of the plurality of synapses includes:

decreasing a weight of a synapse when the partial differential equation corresponding to the weight is positive; and

increasing the weight of the synapse when the partial differential equation corresponding to the weight is negative.

10. A device for pruning a neural network, the device comprising:

a computing circuit configured to perform a convolution operation including an addition operation and a multiplication operation;

an input signal generator configured to generate input data, and to input the input data to the computing circuit;

a threshold adjusting circuit configured to select a layer from among a plurality of layers in the neural network and adjust a threshold value of the selected layer;

a weight adjusting circuit configured to adjust a plurality of weights of a plurality of synapses, respectively, which are included in the neural network; and

a controller configured to prune the neural network by controlling the threshold adjusting circuit and the weight adjusting circuit,

wherein the selected layer has a highest enhancement ratio among the plurality of layers, the enhancement ratio of the selected layer being a reduction rate of computational complexity of the selected layer divided by a reduction rate of accuracy of the selected layer.

11. The device of claim 10 , further comprising a memory device configured to store initial weights of the plurality of synapses, the input data, or both.

12. The device of claim 10 , wherein the controller controls the threshold adjusting circuit to adjust the threshold value of the selected layer among the plurality of layers, and then the controller controls the weight adjusting circuit to adjust the plurality of weights of the plurality of synapses in the neural network.

13. The device of claim 10 , wherein the controller controls the threshold adjusting circuit, the computing circuit, or both, to assume a value of a neuron in the selected layer that is less than the threshold value of the selected layer is 0.

14. The device of claim 13 , wherein the controller further controls the threshold adjusting circuit and the weight adjusting circuit to operate when a pruning ratio of the neural network is smaller than a target value, the pruning ratio being equal to a ratio of a number of neurons in the neural network having a value of 0 to a total number of neurons in the neural network.

15. The device of claim 10 , wherein the weight adjusting circuit calculates a loss function and adjusts the plurality of weights of the plurality of synapses with a plurality of partial differential equations of the loss function relative to the plurality of synapses, respectively.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2018
From: CHO, SEUNGHWAN; YOO, SUNGJOO; JIN, YOUNGJAE
To: SK HYNIX INC.; SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
Reel/Frame 046242/0693 →
Priority Claims (1)
KR 10-2017-0103569 · Aug 16, 2017 · national
Continuity (1)
Related Publication 20190057308A1 · Feb 21, 2019
Cited By (2)
US 12,579,406 US 12,608,593