IP Library Granted Patent US 11,710,026
Granted Patent B2
US 11,710,026 · App. 17/989,761 · Granted Jul 25, 2023

Optimization for artificial neural network model and neural processing unit

Inventor: Lok Won Kim (Seongnam-si, KR)
Assignee: DEEPX CO., LTD.
G06N3/02G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,710,026
App. No.
17/989,761
Granted
Jul 25, 2023
Kind
B2
Abstract

A computer-implemented apparatus installed and executed in a computer to search an optimal design of a neural processing unit (NPU), a hardware accelerator used for driving a computer-implemented artificial neural network (ANN) is disclosed. The NPU comprises a plurality of blocks connected in a form of pipeline, and the number of the plurality blocks and the number of the layers within each block of the plurality blocks are in need of optimization to reduce hardware resources demand and electricity power consumption of the ANN while maintaining the inference accuracy of the ANN at an acceptable level. The computer-implemented apparatus searches for and then outputs an optimal L value and an optimal C value when a first set of candidate values for a number of layers L and a second set of candidate values for a number of channels C per each layer of the ANN is provided.

Claims (54)

1. A computer-implemented apparatus configured to search an optimal design of a binarized neural network (BNN) to be operated by a neural processing unit (NPU), the computer-implemented apparatus comprising:

a first operation unit configured to search for an optimal L value and an optimal C value with respect to the BNN, which of all weight parameters are binarized, when a first set of candidate values for a number of layers L and a second set of candidate values for a number of channels C per each layer of the BNN are provided; and

a second operation unit configured to output the optimal L value and the optimal C value,

wherein when the NPU includes a plurality of blocks connected in a form of pipeline, a number of the plurality of blocks is related to the optimal L value.

2. The computer-implemented apparatus of claim 1 ,

wherein each of the plurality of blocks includes at least one of a line buffer, a XNOR logic gate, a pop-count instruction execution unit, and a batch-normalization unit.

3. The computer-implemented apparatus of claim 1 ,

wherein the first operation unit is configured to search for the optimal L value and the optimal C value based on at least one of a hardware cost of the NPU and a power consumption of the NPU.

4. The computer-implemented apparatus of claim 3 ,

wherein the hardware cost of the NPU is determined based on a number of flip-flops.

5. The computer-implemented apparatus of claim 1 ,

wherein the first operation unit is configured to search for the optimal L value and the optimal C value based on a required minimum inference accuracy.

6. The computer-implemented apparatus of claim 5 ,

wherein the optimal C value is increased by one or two steps when inference accuracy degradation occurs.

7. The computer-implemented apparatus of claim 1 ,

wherein the optimal C value is a constant value that is the same in all layers of the BNN.

8. A computing system comprising:

at least one processor; and

at least one memory, operably electrically connected to the at least one processor, configured to store an instruction to search an optimal design of a binarized neural network (BNN) to be operated by a neural processing unit (NPU),

wherein an operation, configured to be performed based on the instructions for the computer-implemented apparatus being executed by the at least one processor, comprising:

searching for an optimal L value and an optimal C value with respect to the BNN, which of all weight parameters are binarized, when a first set of candidate values for a number of layers L and a second set of candidate values for a number of channels C per each layer of the BNN are provided; and

outputting the optimal L value and the optimal C value,

wherein when the NPU includes a plurality of blocks connected in a form of pipeline, a number of the plurality of blocks is related to the optimal L value.

9. The computing system of claim 8 ,

wherein each of the plurality of blocks includes at least one of a line buffer, a XNOR logic gate, a pop-count instruction execution unit, and a batch-normalization unit.

10. The computing system of claim 8 ,

wherein the searching for the optimal L value and the optimal C value is conducted based on at least one of a hardware cost of the NPU and a power consumption of the NPU.

11. The computing system of claim 10 ,

wherein the hardware cost of the NPU is determined based on a number of flip-flops.

12. The computing system of claim 8 ,

wherein the searching for the optimal L value and the optimal C value is conducted-based on a required minimum inference accuracy.

13. The computing system of claim 12 ,

wherein the optimal C value is increased by one or two steps when inference accuracy degradation occurs.

14. The computing system of claim 8 ,

wherein the optimal C value is a constant value that is the same in all layers of the BNN.

15. A non-transitory computer-readable storage medium storing instructions for a computer-implemented apparatus that searches an optimal design of a binarized neural network (BNN) to be operated by a neural processing unit (NPU), the computer readable storage medium is configured to store the instructions,

wherein the instructions, when executed by at least one processor, cause the at least one processor to:

search for an optimal L value and an optimal C value with respect to the BNN, which of all weight parameters are binarized, when a first set of candidate values for a number of layers L and a second set of candidate values for a number of channels C per each layer of the BNN are provided; and

output the optimal L value and the optimal C value,

wherein when the NPU includes a plurality of blocks connected in a form of pipeline, a number of the plurality of blocks is related to the optimal L value.

16. The computer readable storage medium of claim 15 ,

wherein each of the plurality of blocks includes at least one of a line buffer, a XNOR logic gate, a pop-count instruction execution unit, and a batch-normalization unit.

17. The computer readable storage medium of claim 15 ,

wherein the searching for the optimal L value and the optimal C value is based on at least one of a hardware cost of the NPU and a power consumption of the NPU.

18. The computer readable storage medium of claim 15 ,

wherein the hardware cost of the NPU is determined based on a number of flip-flops.

19. The computer readable storage medium of claim 15 ,

wherein the searching for the optimal L value and the optimal C value is based on a required minimum inference accuracy.

20. A neural processing unit (NPU) configured to operate a binarized neural network (BNN), the NPU comprising:

a plurality of blocks,

wherein the plurality of blocks is connected in a form of pipeline,

wherein each of the plurality of blocks includes at least one of a line buffer, a XNOR logic gate, a pop-count instruction execution unit, and a batch-normalization unit,

wherein a number of the plurality of blocks is equal to a number L of a plurality of layers in the BNN, which of all weight parameters are binarized, and

wherein, the number L of the plurality of layers and a number C of channels per layer are determined based on a power consumption and a hardware implementation cost.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2022
From: KIM, LOK WON
To: DEEPX CO., LTD.
Reel/Frame 061819/0459 →
Priority Claims (2)
KR 10-2021-0166868 · Nov 29, 2021 · national
KR 10-2022-0133807 · Oct 18, 2022 · national
Continuity (1)
Related Publication 20230090720A1 · Mar 23, 2023