IP Library Granted Patent US 11,507,429
Granted Patent B2
US 11,507,429 · App. 16/038,243 · Granted Nov 22, 2022

Neural network accelerator including bidirectional processing element array

Inventors: Chun-Gi Lyuh (Daejeon, KR); Young-Su Kwon (Daejeon, KR); Chan Kim (Daejeon, KR); Hyun Mi Kim (Daejeon, KR); Jeongmin Yang (Busan, KR); Jaehoon Chung (Daejeon, KR); Yong Cheol Peter Cho (Daejeon, KR)
Assignee: Electronics and Telecommunications Research Institute
G06F9/5044G06F9/30101G06F9/3877G06F15/8046G06N3/063G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,507,429
App. No.
16/038,243
Filed
Jul 18, 2018
Granted
Nov 22, 2022
Kind
B2
Art Unit
2183
USPC
706/41
Abstract

Provided is a neural network accelerator which performs a calculation of a neural network provided with layers, the neural network accelerator including a kernel memory configured to store kernel data related to a filter, a feature map memory configured to store feature map data which are outputs of the layers, and a Processing Element (PE) array including PEs arranged along first and second directions, wherein each of the PEs performs a calculation using the feature map data transmitted in the first direction from the feature map memory and the kernel data transmitted in the second direction from the kernel memory, and transmits a calculation result to the feature map memory in a third direction opposite to the first direction.

Claims (45)

1. A neural network accelerator which performs a calculation of a neural network comprising layers, the neural network accelerator comprising:

a kernel memory configured to store kernel data related to a filter;

a feature map memory configured to store feature map data which are outputs of the layers; and

a processing element (PE) array comprising PEs arranged along first and second directions,

wherein each of the PEs is configured to perform a calculation using the feature map data transmitted in the first direction from the feature map memory and the kernel data transmitted in the second direction from the kernel memory, and transmit a calculation result to the feature map memory in a third direction opposite to the first direction,

wherein the calculation comprises a multiplication, an addition, an activation calculation, a normalization calculation, and a pooling calculation,

wherein the PE array further comprises:

kernel load units configured to transmit the kernel data to the PEs in the second direction; and

feature map input/output (I/O) units configured to transmit the feature map data to the PEs in the first direction, receive the calculation result transmitted in the third direction, and transmit the calculation result to the feature map memory,

wherein each of the PEs comprises:

a control register configured to store a command transmitted in the first direction;

a kernel register configured to store the kernel data;

a feature map register configured to store the feature map data;

a multiplier configured to perform a multiplication on data stored in the kernel register and the feature map register;

an adder configured to perform an addition on a multiplication result of the multiplier and a previous calculation result;

an accumulator register configured to accumulate the previous calculation result or an addition result of the adder; and

an output register configured to store a calculation result transmitted in the third direction from another PE or the addition result,

wherein the feature map I/O unit is further configured to receive, from the PEs, new feature map data or a partial sum for generating the new feature map data based on an output command, and transmit the new feature map data or the partial sum for generating the new feature map data to the feature map memory, and

wherein the output register further is further configured to store a valid flag bit indicating whether the calculation result is valid based on the output command transmitted to the control register, and a last flag register bit indicating whether the PE is arranged in a most distant column from the feature map memory based on the first direction.

2. The neural network accelerator of claim 1 , wherein the PEs comprise a first PE and a second PE located beside the first PE,

wherein the first PE is configured to receive the calculation result, the valid flag bit, and a last flag bit of the second PE in the third direction based on the valid flag bit and the last flag bit of the second PE.

3. The neural network accelerator of claim 2 , wherein the first PE is further configured to repeatedly receive the calculation result, the valid flag bit and the last flag bit of the second PE in the third direction until the last flag bit of the second PE is activated.

4. A neural network accelerator which performs a calculation of a neural network comprising layers, the neural network accelerator comprising:

a kernel memory configured to store kernel data related to a filter;

a feature map memory configured to store feature map data which are outputs of the layers; and

a processing element (PE) array comprising PEs arranged along first and second directions,

wherein each of the PEs is configured to perform a calculation using the feature map data transmitted in the first direction from the feature map memory and the kernel data transmitted in the second direction from the kernel memory, and transmit a calculation result to the feature map memory in a third direction opposite to the first direction,

wherein the calculation comprises a multiplication, an addition, an activation calculation, a normalization calculation, and a pooling calculation,

wherein the PE array further comprises:

kernel load units configured to transmit the kernel data to the PEs in the second direction; and

feature map input/output (I/O) units configured to transmit the feature map data to the PEs in the first direction, receive the calculation result transmitted in the third direction, and transmit the calculation result to the feature map memory,

wherein each of the PEs comprises:

a control register configured to store a command transmitted in the first direction;

a kernel register configured to store the kernel data;

a feature map register configured to store the feature map data;

a multiplier configured to perform a multiplication on data stored in the kernel register and the feature map register;

an adder configured to perform an addition on a multiplication result of the multiplier and a previous calculation result;

an accumulator register configured to accumulate the previous calculation result or an addition result of the adder; and

an output register configured to store a calculation result transmitted in the third direction from another PE or the addition result,

wherein the feature map I/O unit is further configured to receive, from the PEs, new feature map data or a partial sum for generating the new feature map data based on an output command, and transmit the new feature map data or the partial sum for generating the new feature map data to the feature map memory, and

wherein the feature map I/O unit is further configured to receive the partial sum from the feature map memory in the first direction and transmit the partial sum to the PEs, based on a load partial sum command and a pass partial sum command.

5. The neural network accelerator of claim 4 , wherein the accumulator register of each the PEs is further configured to store the partial sum transmitted in the first direction in response to the load partial sum command transmitted to the control register.

6. The neural network accelerator of claim 5 , wherein the PEs comprise a first PE and a second PE located beside the first PE in the first direction,

wherein the first PE is configured to transmit the load partial sum command to the second PE instead of the pass partial sum command when receiving the pass partial sum command after receiving the load partial sum command.

7. The neural network accelerator of claim 6 , wherein the first PE is further configured to transmit, to the second PE, a partial sum transmitted together with the received pass partial sum command in the first direction.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE FIRST INVENTOR'S NAME PREVIOUSLY RECORDED ON REEL 046946 FRAME 0441. ASSIGNOR(S) HEREBY CONFIRMS THE CORRECT INVENTOR'S NAME IS CHUN-GI LYUH. Recorded Sep 25, 2018
From: LYUH, CHUN-GI; KWON, YOUNG-SU; KIM, CHAN; KIM, HYUN MI; YANG, JEONGMIN; CHUNG, JAEHOON; CHO, YONG CHEOL PETER
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 047142/0866 →
CORRECTIVE ASSIGNMENT TO CORRECT THE SIXTH INVENTOR'S NAME PREVIOUSLY RECORDED AT REEL: 046377 FRAME: 0693. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 23, 2018
From: LYUH, CHUN-GIL; KWON, YOUNG-SU; KIM, CHAN; KIM, HYUN MI; YANG, JEONGMIN; CHUNG, JAEHOON; CHO, YONG CHEOL PETER
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 046946/0441 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2018
From: LYUH, CHUN-GI; KWON, YOUNG-SU; KIM, CHAN; KIM, HYUN MI; YANG, JEONGMIN; JUNG, JAEHOON; CHO, YONG CHEOL PETER
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 046377/0693 →
Priority Claims (2)
KR 10-2017-0118068 · Sep 14, 2017 · national
KR 10-2018-0042395 · Apr 11, 2018 · national
Continuity (1)
Related Publication 20190079801A1 · Mar 14, 2019
Cited By (3)
US 12,488,228 US 12,693,990 US 12,717,629