IP Library › Granted Patent US 11,960,986
Granted Patent B2
US 11,960,986 · App. 17/944,454 · Granted Apr 16, 2024

Neural network accelerator and operating method thereof

Inventors: Seokhyeong Kang (Pohang-si, KR); Yesung Kang (Ulsan, KR); Sunghoon Kim (Seoul, KR); Yoonho Park (Seoul, KR)
Assignee: Samsung Electronics Co., Ltd.
G06N3/063G06F9/30145G06F9/5027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,960,986
App. No.
17/944,454
Granted
Apr 16, 2024
Kind
B2
Abstract

A neural network accelerator includes an operator that calculates a first operation result based on a first tiled input feature map and first tiled filter data, a quantizer that generates a quantization result by quantizing the first operation result based on a second bit width extended compared with a first bit width of the first tiled input feature map, a compressor that generates a partial sum by compressing the quantization result, and a decompressor that generates a second operation result by decompressing the partial sum, the operator calculates a third operation result based on a second tiled input feature map, second tiled filter data, and the second operation result, and an output feature map is generated based on the third operation result.

Claims (44)

1. A neural network accelerator comprising:

an operator configured to calculate a first operation result based on a first tiled input feature map and first tiled filter data;

a quantizer configured to generate a quantization result by quantizing the first operation result based on a second bit width extended compared with a first bit width of the first tiled input feature map;

a compressor configured to generate a partial sum by compressing the quantization result; and

a decompressor configured to generate a second operation result by decompressing the partial sum.

2. The neural network accelerator of claim 1 , wherein the operator calculates a third operation result based on a second tiled input feature map, second tiled filter data, and the second operation result, and

wherein an output feature map is generated based on the third operation result.

3. The neural network accelerator of claim 2 , wherein the operator includes:

a multiplier configured to generate a multiplication result by multiplying the second tiled input feature map and the second tiled filter data; and

an accumulator configured to generate the third operation result by adding the multiplication result and the second operation result.

4. The neural network accelerator of claim 1 , wherein the quantizer quantizes the first operation result through a round-off.

5. The neural network accelerator of claim 1 , wherein the quantization result includes a sign bit and remaining bits composed of an integer part and a fractional part, and

wherein a bit width of at least one of the integer part and the fractional part is extended depending on the second bit width.

6. The neural network accelerator of claim 5 , wherein the compressor includes:

absolute value generation logic configured to generate an absolute value of partial bits of the remaining bits of the quantization result; and

a run-length encoder configured to generate compression bits by performing run-length encoding based on the generated absolute value.

7. The neural network accelerator of claim 6 , wherein the partial bits are selected in the order from an upper bit to a lower bit of the remaining bits based on a difference between the first bit width and the second bit width.

8. The neural network accelerator of claim 6 , wherein the partial sum includes the compression bits generated from the run-length encoder and remaining bits of the quantization result other than the partial bits.

9. The neural network accelerator of claim 1 , wherein the compressor stores the generated partial sum in an external memory, and

wherein the decompressor receives the stored partial sum from the external memory.

10. The neural network accelerator of claim 2 , wherein the quantizer generates the output feature map by quantizing the third operation result based on the first bit width.

11. An operation method of a neural network accelerator which operates based on channel loop tiling, the method comprising:

generating a first operation result based on a first tiled input feature map and first tiled filter data;

generating a first quantization result by quantizing the first operation result based on a second bit width extended compared with a first bit width of the first tiled input feature map;

generating a first partial sum by compressing the first quantization result; and

generating a second operation result by decompressing the first partial sum.

12. The method of claim 11 , wherein the generating of the second operation result includes generating the second operation result based on a second tiled input feature map, second tiled filter data, and the first partial sum.

13. The method of claim 12 , wherein the generating of the second operation result includes:

generating a multiplication result by multiplying the second tiled input feature map and the second tiled filter data; and

generating the second operation result by adding the multiplication result and the first operation result.

14. The method of claim 12 , further comprising:

when the second tiled input feature map and the second tiled filter data are not data lastly received,

generating a second quantization result by quantizing the second operation result based on the second bit width;

generating a second partial sum by compressing the second quantization result;

storing the generated second partial sum in the external memory; and

generating a third operation result based on a third tiled input feature map, third tiled filter data, and the second partial sum provided from the external memory.

15. The method of claim 12 , further comprising:

when the second tiled input feature map and the second tiled filter data are data lastly received,

generating a third quantization result by quantizing the second operation result based on the first bit width; and

generating an output feature map based on the third quantization result.

16. The method of claim 12 , wherein the first quantization result includes a sign bit and remaining bits composed of an integer part and a fractional part, and

wherein a bit width of at least one of the integer part and the fractional part is extended depending on the second bit width.

17. The method of claim 11 , wherein the generating the first partial sum comprises generating the first partial sum by compressing a bit group corresponding to the first bit width in the quantization result.

18. The neural network accelerator of claim 1 , wherein the compressor is configured to generate the partial sum by compressing a bit group corresponding to the first bit width in the quantization result.

Priority Claims (2)
KR 10-2019-0010293 · Jan 28, 2019 · national
KR 10-2019-0034583 · Mar 26, 2019 · national
Continuity (2)
Continuation 16751503 · Jan 24, 2020
Related Publication 20230004790A1 · Jan 5, 2023