IP Library › Granted Patent US 11,630,984
Granted Patent B2
US 11,630,984 · App. 16/663,487 · Granted Apr 18, 2023

Method and apparatus for accelerating data processing in neural network

Inventors: Sung Joo Yoo (Seoul, KR); Eun Hyeok Park (Seoul, KR)
Assignee: SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
G06N3/04G06K9/6267
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,630,984
App. No.
16/663,487
Granted
Apr 18, 2023
Kind
B2
Abstract

Proposed are a method and apparatus for accelerating data processing in a neural network. The apparatus for accelerating data processing in a neural network may include: a control unit configured to quantize data by at least one method according to a characteristic of data calculated at a node forming at least one layer constituting the neural network, and to separately perform calculation at the node according to the quantized data; and memory configured to store the quantized data.

Claims (23)

1. An apparatus for accelerating data processing on a neural network, the apparatus comprising:

a control unit configured to quantize data by at least one method according to a characteristic of data calculated at a node forming at least one layer constituting the neural network, and to separately perform calculation at the node according to the quantized data; and

memory configured to store the quantized data,

wherein the data comprises an activation adapted to be output data at the node and a weight adapted to be an intensity at which the activation is reflected into the node; and

the control unit classifies the activation as any one of an outlier activation larger than a preset number of bits and a normal activation equal to or smaller than the preset number of bits based on a number of bits representing a quantized activation.

2. The apparatus of claim 1 , wherein the control unit determines the quantization method, to be applied to the data, based on a value of the data on the neural network.

3. The apparatus of claim 2 , wherein the control unit varies a number of bits of data attributable to the quantization based on an error between the data and the quantized data.

4. The apparatus of claim 1 , wherein the control unit separates the outlier activation and the normal activation from each other, and separately performs calculation on them.

5. The apparatus of claim 1 , wherein the control unit stores at least one of the normal activation and the weight in a buffer.

6. The apparatus of claim 1 , wherein the control unit, when an outlier weight, which is a weight larger than a preset number of bits, is present in the buffer, separates the outlier weight into a preset number of upper bits and remaining bits and then performs calculation.

7. A method for accelerating data processing on a neural network, the method comprising:

quantizing data by at least one method according to a characteristic of data calculated at a node forming at least one layer constituting the neural network;

storing the quantized data; and

separately performing calculation at the node according to the quantized data,

wherein the data comprises an activation adapted to be output data at the node and a weight adapted to be an intensity at which the activation is reflected into the node; and

the method comprises classifying the activation as any one of an outlier activation larger than a preset number of bits and a normal activation equal to or smaller than the preset number of bits based on a number of bits representing a quantized activation.

8. The method of claim 7 , wherein quantizing the data comprises determining the quantization method, to be applied to the data, based on a value of the data on the neural network.

9. The method of claim 8 , wherein determining the quantization method comprises varying a number of bits of data attributable to the quantization based on an error between the data and the quantized data.

10. The method of claim 7 , wherein separately performing the calculation comprises separating the outlier activation and the normal activation from each other and then performing calculation on them.

11. The method of claim 7 , wherein performing the calculation at the node comprises storing at least one of the normal activation and the weight in a buffer.

12. The method of claim 7 , further comprising, when an outlier weight, which is a weight larger than a preset number of bits, is present in the buffer, separating the outlier weight into a preset number of upper bits and remaining bits and then performing calculation.

13. A non-transitory computer-readable storage medium having stored thereon a program which performs the method set forth in claim 7 .

14. A computer program stored in a non-transitory computer-readable storage medium such that a neural network acceleration apparatus performs the method set forth in claim 7 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2019
From: YOO, SUNG JOO; PARK, EUN HYEOK
To: SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
Reel/Frame 050823/0631 →
Priority Claims (2)
KR 10-2017-0055001 · Apr 28, 2017 · national
KR 10-2017-0178747 · Dec 22, 2017 · national
Continuity (2)
Continuation PCTKR2018005022 · Apr 30, 2018
Related Publication 20200057934A1 · Feb 20, 2020