IP Library Granted Patent US 12,412,097
Granted Patent B2
US 12,412,097 · App. 17/298,801 · Granted Sep 9, 2025

Neural network compression device

Inventors: Hiroaki Ito (Hitachinaka, JP); Goichi Ono (Tokyo, JP); Riu Hirai (Tokyo, JP)
Assignee: Hitachi Astemo, Ltd.
G06N3/082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,097
App. No.
17/298,801
Granted
Sep 9, 2025
Kind
B2
Abstract

Processing time of a neural network is shortened, and the number of operations of the neural network is reduced such that a plurality of calculators can be effectively used. A neural network reduction device ( 100 ) that reduces the number of operations of a neural network by an operation device ( 140 ) including a plurality of calculators by reducing the neural network, the neural network reduction device including: a calculator allocation unit ( 102 ) that sets the number of calculators allocated to calculation processing of the neural network; a number-of-operations setting unit ( 103 ) that sets the number of operations of a reduced neural network based on the number of allocated calculators; and a neural network reduction unit ( 104 ) that reduces the neural network such that the number of operations of the neural network by the operation device ( 140 ) is equal to the number of operations set by the number-of-operations setting unit ( 103 ).

Claims (24)

1. A neural network reduction device that reduces a number of operations of a neural network by an operation device including a plurality of calculators by reducing the neural network, the neural network reduction device comprising:

a calculator allocation unit that sets a number of the calculators allocated for calculation processing of the neural network, each of the calculators configured to execute a respective operation portion in each cycle of the calculation processing;

a number-of-operations setting unit that determines a number of operations of the neural network based on the number of the allocated calculators and sets the number of operations for the calculation processing;

a neural network reduction unit that reduces the neural network to generate a reduced neural network, such that the number of operations of the neural network is equal to the number of operations set by the number-of-operations setting unit;

an accuracy verification unit that

calculates accuracy of the reduced neural network, and

compares the accuracy with target accuracy,

wherein in response to the comparison, the calculator allocation unit decreases the number of the allocated calculators based on the accuracy being greater than or equal to the target accuracy or increases the number of the allocated calculators based on the accuracy being less than the target accuracy; and

a number-of-operations correction unit that updates the number of operations of the reduced neural network in response to increases or decreases in the number of the allocated calculators by the calculator allocation unit, wherein the updated number of operations is an integral multiple of the number of the allocated calculators.

2. The neural network reduction device according to claim 1 , wherein the number-of-operations setting unit sets the number of operations of the reduced neural network to be smaller than the number of operations of the neural network before reduction and be an integral multiple of the number of the allocated calculators set by the calculator allocation unit before the reduction.

3. The neural network reduction device according to claim 1 , wherein the number-of-operations setting unit sets the number of operations of the reduced neural network such that a remainder, obtained by dividing the number of operations of the neural network before reduction by the number of the allocated calculators set by the calculator allocation unit, is equal to or more than half of the number of the allocated calculators.

4. The neural network reduction device according to claim 1 , wherein

the neural network includes a plurality of layers,

the calculator allocation unit sets the number of the allocated calculators for each of the layers of the neural network, and

the number-of-operations setting unit sets the number of operations of the reduced neural network for each of the layers of the neural network.

5. The neural network reduction device according to claim 1 , wherein the neural network reduction unit reduces the neural network by pruning processing.

6. The neural network reduction device according to claim 1 ,

wherein the number-of-operations setting unit sets the number of operations of the reduced neural network to be small when the accuracy is equal to or higher than the target accuracy, and the number-of-operations setting unit sets the number of operations of the reduced neural network to be large when the accuracy is lower than the target accuracy.

7. A neural network reduction device that reduces a number of operations of a neural network by an operation device including a plurality of calculators by reducing the neural network, the neural network reduction device comprising:

a number-of-operations setting unit that sets a number of operations for calculation processing of the neural network;

a neural network reduction unit that reduces the neural network to generate a reduced neural network, such that the number of operations of the neural network is equal to the number of operations set by the number-of-operations setting unit;

a calculator allocation unit that sets a number of the calculators allocated for the calculation processing of the neural network, each of the calculators configured to execute a respective operation portion in each cycle of the calculation processing;

a number-of-operations correction unit that receives, from the calculator allocation unit, an indication of decreases in the number of the allocated calculators based on accuracy of the reduced neural network being greater than or equal to a target accuracy or increases the number of the allocated calculators based on the accuracy of the reduced neural network being less than the target accuracy; and

updates the number of operations of the reduced neural network in response to increases or decreases in the number of the allocated calculators set by the calculator allocation unit, wherein the updated number of operations is an integral multiple of the number of the allocated calculators.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2021
From: ITO, HIROAKI; ONO, GOICHI; HIRAI, RIU
To: HITACHI ASTEMO, LTD.
Reel/Frame 056402/0804 →
Priority Claims (1)
JP 2019-006660 · Jan 18, 2019 · national
Continuity (1)
Related Publication 20220036190A1 · Feb 3, 2022
References Cited (15)
US 20160379108A1 · Chung · 2016 [cited by examiner]
US 20180204110A1 · Kim · 2018 [cited by examiner]
US 20200184333A1 · Oh · 2020 [cited by examiner]
US 20200285950A1 · Baum · 2020 [cited by examiner]
US 20210295174A1 · Zhang · 2021 [cited by examiner]
JP H11215380A · 1999 [cited by applicant]
Baskin, C., Liss, N., Zheltonozhskii, E., Bronstein, A. M., & Mendelson, A. (May 2018). Streaming architecture for large-scale quantized neural networks on an FPGA-based dataflow platform. In 2018 IEEE International Par… [cited by examiner]
Salamat, S., Imani, M., Gupta, S., & Rosing, T. (Nov. 2018). Rnsnet: In-memory neural network acceleration using residue number system. In 2018 IEEE International Conference on Rebooting Computing (ICRC) (pp. 1-12). IEE… [cited by examiner]
Chakradhar, S., Sankaradas, M., Jakkula, V., & Cadambi, S. (Jun. 2010). A dynamically configurable coprocessor for convolutional neural networks. In Proceedings of the 37th annual international symposium on Computer arc… [cited by examiner]
Fujii, T., Sato, S., Nakahara, H., & Motomura, M. (Apr. 2017). An FPGA realization ofa deep convolutional neural network using a threshold neuron pruning. In Applied Reconfigurable Computing: 13th Int Symposium, ARC 201… [cited by examiner]
Posewsky, T., & Ziener, D. (Apr. 2018). Throughput optimizations for FPGA-based deep neural network inference. Microprocessors and microsystems, 60, 151-161. (Year: 2018). [cited by examiner]
International Search Report with English translation and Written Opinion issued in corresponding application No. PCT/JP2020/000231 dated Mar. 24, 2020. [cited by applicant]
Han et al., “Learning both Weights and Connections for Efficient Neural Networks”, [online], Oct. 30, 2015, [searched on Dec. 24, 2018], Internet <URL: https://arxiv.org/pdf/1506.02626.pdf>. [cited by applicant]
Kojima et al., Hitachi Review, Driving Forward with Future Vehicles, “Intelligent technology to support the advancement of automated driving, 4. 1 Technique for Neural Network Compression”, Hitachi Review, vol. 99, No. … [cited by applicant]
Yoshida, “Layer-unified Execution Control Scheme for Coarse Grain Task Parallel Processing, 2.2 Execution system for coarse grain task parallel processing; 6.1 Comparison of performance between layer-unified execution c… [cited by applicant]