IP Library › Granted Patent US 11,544,213
Granted Patent B2
US 11,544,213 · App. 17/369,298 · Granted Jan 3, 2023

Neural processor

Inventors: Dongyoung Kim (Suwon-si, KR); Jung Ho Ahn (Seoul, KR); Sunjung Lee (Seoul, KR); Jaewan Choi (Incheon, KR)
Assignees: Samsung Electronics Co., Ltd.; Seoul National University R&DB Foundation
G06F15/8046G06F9/3001G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,544,213
App. No.
17/369,298
Filed
Jul 7, 2021
Granted
Jan 3, 2023
Kind
B2
Art Unit
2183
USPC
712/19
Abstract

A neural processor is provided. The neural processor includes a matrix device which is configured to generate an output feature map by processing a standard convolution operation and which has a systolic array architecture, and accelerators with an adder-tree structure which are configured to process depth-wise convolution operations for each of elements of the output feature map corresponding to lanes of the matrix device.

Claims (33)

1. A neural processor, comprising:

a matrix device configured to generate an output feature map by processing a standard convolution operation, the matrix device configured to have a systolic array architecture; and

accelerators, configured to process depth-wise convolution operations for each of elements of the output feature map corresponding to lanes of the matrix device.

2. The neural processor of claim 1 , wherein the output feature map generated by the standard convolution operation is provided as an input to the accelerators in a pipelined manner to perform the depth-wise convolution operations.

3. The neural processor of claim 1 , wherein each of the accelerators is configured to perform the depth-wise convolution operations in parallel for each of the elements of the output feature map that are output for each of the lanes corresponding to columns of the matrix device.

4. The neural processor of claim 1 , wherein the accelerators are configured to operate by implementing a lockstep scheme of processing a same set of operations at a same time in parallel.

5. The neural processor of claim 1 , wherein the accelerators respectively correspond to the lanes of the matrix device.

6. The neural processor of claim 1 , wherein each of the accelerators comprises:

a plurality of multipliers;

a plurality of depth-wise input feature map buffers, configured to store the output feature map for the depth-wise convolution operations; and

a depth-wise weight buffer, configured to store weights for the depth-wise convolution operations.

7. The neural processor of claim 6 , wherein the plurality of multipliers are configured to:

perform a multiplication operation for the depth-wise convolution operations based on the output feature map received from the plurality of depth-wise input feature map buffers and the weights received from the depth-wise weight buffer; and

transmit a result of the multiplication operation to first adders included in the adder-tree structure.

8. The neural processor of claim 6 , wherein each of the accelerators further comprises:

a second adder, configured to collect a multiplication operation result of the plurality of multipliers; and

a latch, configured to store values collected in the second adder.

9. The neural processor of claim 6 , wherein the plurality of depth-wise input feature map buffers are configured to be one-to-one connected to the plurality of multipliers.

10. The neural processor of claim 6 , wherein the depth-wise weight buffer is configured to read the weights directly from a memory through a memory interface and to simultaneously store and process the weights by a double buffering scheme.

11. The neural processor of claim 6 , wherein:

the depth-wise weight buffer comprises a barrel shifter, configured to shift a position of the weights, and

the depth-wise weight buffer is configured to map an element of the output feature map and an element of the weights to the plurality of multipliers with the barrel shifter.

12. The neural processor of claim 1 , further comprising:

an accumulator, configured to store the elements of the output feature map generated by the matrix device.

13. The neural processor of claim 1 , wherein the matrix device further comprises a postprocessing module, configured to perform at least one postprocessing operation among an activation operation, a normalization operation, and a pooling operation on the elements of the output feature map.

14. The neural processor of claim 1 , wherein the accelerators are configured to have an adder-tree structure.

15. A processor-implemented neural network method, the method comprising:

generating, with a matrix device having a systolic array architecture, an output feature map by processing a standard convolution operation;

processing, with accelerators, depth-wise convolution operations for each of elements of the output feature map corresponding to lanes of the matrix device; and

providing the generated output feature map as an input to the accelerators to perform the depth-wise convolution operations.

16. The method of claim 15 , further comprising processing the depth-wise convolution operations with adders included in an adder tree structure of the accelerators.

17. The method of claim 15 , further comprising performing the depth-wise convolution operations in parallel for each of the elements of the output feature map that are output for each of the lanes corresponding to columns of the matrix device.

18. The method of claim 15 , further comprising performing at least one postprocessing operation among an activation operation, a normalization operation, and a pooling operation on the elements of the output feature map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2021
From: KIM, DONGYOUNG; AHN, JUNG HO; LEE, SUNJUNG; CHOI, JAEWAN
To: SAMSUNG ELECTRONICS CO., LTD.; SNU R&DB FOUNDATION
Reel/Frame 056777/0388 →
Priority Claims (2)
KR 10-2021-0028932 · Mar 4, 2021 · national
KR 10-2021-0036051 · Mar 19, 2021 · national
Continuity (1)
Related Publication 20220283984A1 · Sep 8, 2022