IP Library › Granted Patent US 11,861,486
Granted Patent B2
US 11,861,486 · App. 17/985,257 · Granted Jan 2, 2024

Neural processing unit for binarized neural network

Inventors: Lok Won Kim (Seongnam-si, KR); Quang Hieu Vo (Quang Binh province, VN)
Assignee: DEEPX CO., LTD.
G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,861,486
App. No.
17/985,257
Granted
Jan 2, 2024
Kind
B2
Abstract

A neural processing unit of a binarized neural network (BNN) as a hardware accelerator is provided, for the purpose of reducing hardware resource demand and electricity consumption while maintaining acceptable output precision. The neural processing unit may include: a first block configured to perform convolution by using a binarized feature map with a binarized weight; and a second block configured to perform batch-normalization on an output of the first block. A register having a particular size may be disposed between the first block and the second block. Each of the first block and the second block may include one or more processing engines. The one or more processing engines may be connected in a form of pipeline.

Claims (63)

1. A neural processing unit of a binarized neural network (BNN), the neural processing unit comprising:

a plurality of circuits, wherein the plurality of circuits are connected in a form of pipeline and comprise a first circuitry comprising:

a first sub-circuitry including a NOT logic gate and configured to perform a NOT logic gate operation by using a binarized feature map with a binarized weight,

a second sub-circuitry including a pop-count performing unit and configured to perform an accumulation, and

a third sub-circuitry configured to perform batch-normalization,

wherein the first circuitry is configured to:

determine whether the binarized weight is zero (0) or one (1),

select the first sub-circuitry including the NOT logic gate to which an input value is inputted when the binarized weight is zero (0), and

bypass the input value, thereby directly delivering the input value to the second sub-circuitry including the pop-count performing unit, when the binarized weight is one (1), and

wherein the NOT logic gate of the first sub-circuitry removes or replaces a XNOR logic gate thereby reducing a size of the neural processing unit.

2. The neural processing unit of claim 1 ,

wherein the first circuitry further comprises: a plurality of registers which are disposed between the first sub-circuitry, the second sub-circuitry and the third sub-circuitry.

3. The neural processing unit of claim 1 , further comprising:

a second circuitry configured to perform max-pooling on an output of the first circuitry.

4. The neural processing unit of claim 3 ,

wherein the first circuitry corresponds to a first layer of the BNN, and the second circuitry corresponds to a second layer of the BNN.

5. The neural processing unit of claim 1 , further comprising a line-buffer or a memory configured to store a binarized parameter corresponding to a layer of the BNN.

6. The neural processing unit of claim 5 ,

wherein a size of the line-buffer is determined based on the size of a corresponding binarized feature map and the size of a corresponding binarized weight.

7. The neural processing unit of claim 1 ,

wherein the third sub-circuitry is configured to perform the batch-normalization based on a pre-determined threshold value.

8. The neural processing unit of claim 1 ,

wherein the first sub-circuitry further includes a K-mean cluster unit.

9. The neural processing unit of claim 1 ,

wherein the second sub-circuitry further includes a compressor.

10. The neural processing unit of claim 1 ,

wherein the second sub-circuitry further includes a pop-count reuse unit.

11. A neural processing unit of an artificial neural network (ANN) having a plurality of layers, the neural processing unit comprising:

a plurality of circuits,

wherein the plurality of circuits are connected in a form of pipeline,

wherein the number of the plurality of circuits is identical to the number of the plurality of layers of the ANN,

wherein, a first circuitry among the plurality of circuits includes:

a first sub-circuitry including a NOT logic gate and configured to perform a NOT logic gate operation by using a binarized feature map with a binarized weight,

a second sub-circuitry including a pop-count performing unit and configured to perform an accumulation, and

a third sub-circuitry configured to perform batch-normalization,

wherein the first circuitry is configured to:

determine whether the binarized weight is zero (0) or one (1),

select the first sub-circuitry including the NOT logic gate to which an input value is inputted when the binarized weight is zero (0), and

bypass the input value, thereby directly delivering the input value to the second sub-circuitry including the pop-count performing unit, when the binarized weight is one (1), and

wherein the NOT logic gate of the first sub-circuitry removes or replaces a XNOR logic gate thereby reducing a size of the neural processing unit.

12. The neural processing unit of claim 11 ,

wherein the first circuitry further comprises: a plurality of registers which are disposed between the first sub-circuitry, the second sub-circuitry and the third sub-circuitry.

13. The neural processing unit of claim 11 , further comprising:

a second circuitry configured to perform max-pooling on an output of the first circuitry.

14. An electronic apparatus comprising:

a main memory; and

a neural processing unit (NPU) configured to perform a function of an artificial neural network (ANN) having a plurality of layers,

wherein the NPU includes a plurality of circuits,

wherein the plurality of blocks is circuits are connected in a form of pipeline,

wherein the number of the plurality of circuits is identical to the number of the plurality of layers of the ANN,

wherein a first circuitry among the plurality of circuits includes:

a first sub-circuitry including a NOT logic gate and configured to perform a NOT logic gate operation by using a binarized feature map with a binarized weight,

a second sub-circuitry including a pop-count performing unit and configured to perform an accumulation, and

a third sub-circuitry configured to perform batch-normalization,

wherein the first circuitry is configured to:

determine whether the binarized weight is zero (0) or one (1),

select the first sub-circuitry including the NOT logic gate to which an input value is inputted when the binarized weight is zero (0), and

bypass the input value, thereby directly delivering the input value to the second sub-circuitry including the pop-count performing unit, when the binarized weight is one (1), and

wherein the NOT logic gate of the first sub-circuitry removes or replaces a XNOR logic gate thereby reducing a size of the neural processing unit.

15. The electronic apparatus of claim 14 ,

wherein the first circuitry further comprises: a plurality of registers which are disposed between the first sub-circuitry, the second sub-circuitry and the third sub-circuitry.

16. The electronic apparatus of claim 14 , further comprising:

a second circuitry configured to perform max-pooling on an output of the first circuitry.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 11, 2022
From: KIM, LOK WON; VO, QUANG HIEU
To: DEEPX CO., LTD.
Reel/Frame 061733/0520 →
Priority Claims (2)
KR 10-2021-0166866 · Nov 29, 2021 · national
KR 10-2022-0132254 · Oct 14, 2022 · national
Continuity (1)
Related Publication 20230082952A1 · Mar 16, 2023
Cited By (1)
US 12,737,692