IP Library Granted Patent US 11,675,997
Granted Patent B2
US 11,675,997 · App. 16/163,772 · Granted Jun 13, 2023

Device and method for processing convolution operation using kernel

Inventors: Kyoung-hoon Kim (Suwon-si, KR); Young-hwan Park (Yongin-si, KR); Dong-kwan Suh (Yongin-si, KR); Keshava Prasad (Suwon-si, KR); Dae-hyun Kim (Seoul, KR); Suk-jin Kim (Seoul, KR); Han-su Cho (Suwon-si, KR); Hyun-jung Kim (Suwon-si, KR)
Assignee: SAMSUNG ELEOTRONICC CO., LTD.
G06N3/04G06F12/06G06F17/15G06F21/52G06N3/045G06N3/063G06N5/046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,675,997
App. No.
16/163,772
Granted
Jun 13, 2023
Kind
B2
Abstract

Provided are a method and apparatus for processing a convolution operation in a neural network. The apparatus may include a memory, and a processor configured to read, from the memory, one of divided blocks of input data stored in a memory; generate an output block by performing the convolution operation on the one of the divided blocks with a kernel; generate a feature map by using the output block, and write the feature map to the memory.

Claims (44)

1. An apparatus for processing a convolution operation in a neural network, the apparatus comprising:

a memory; and

a processor configured to:

read, from the memory, a divided block among a plurality of divided blocks of input data stored in the memory, wherein a plurality of different addresses in the memory are assigned to the plurality of divided blocks, respectively,

perform the convolution operation on the divided block by applying a kernel to the divided block to generate respective output values corresponding to inner pixels inside the divided block and outer pixels that are included in adjacent blocks of the divided block, without reading values of the outer pixels from the memory to provide a conflict-free access to the memory,

generate an output block by using the respective output values,

generate a feature map by using the output block, and

write the feature map to the memory.

2. The apparatus of claim 1 , wherein a size of the output block is larger than a size of the one of the divided blocks.

3. The apparatus of claim 1 , wherein a size of the output block varies according to a size of the kernel.

4. The apparatus of claim 1 ,

wherein addresses respectively corresponding to the divided blocks are assigned with respect to the divided blocks, and

wherein the divided blocks are respectively stored in a plurality of banks of the memory and are accessible by the addresses.

5. The apparatus of claim 4 , wherein the processor is further configured to perform the conflict-free access to one of the plurality of banks with reference to an address of the one of the divided blocks and read data of the one of the divided blocks from the one of the plurality of banks, based on the conflict-free access.

6. The apparatus of claim 1 , wherein the processor is further configured to execute a code temporarily storing kernel information to prevent a stack overflow when performing the convolution operation.

7. The apparatus of claim 1 , further comprising a buffer,

wherein the processor is further configured to accumulate the output block and other outputs previously stored in the buffer by writing the output block to the buffer and generate the feature map based on results accumulated in the buffer.

8. The apparatus of claim 7 , wherein the processor is further configured to convert data of a vertical form of the output block into data of a horizontal form and write the converted data of the horizontal form to the buffer.

9. The apparatus of claim 7 , wherein the processor is further configured to perform accumulation using address information of data stored in the buffer and tag information indicating block type information.

10. A method of processing a convolution operation in a neural network, the method comprising:

reading, from a memory, a divided block among a plurality of divided blocks of input data stored in the memory, wherein a plurality of different addresses in the memory are assigned to the plurality of divided blocks, respectively;

performing the convolution operation on the divided block by applying a kernel to the divided block to generate respective output values corresponding to inner pixels inside the divided block and outer pixels that are included in adjacent blocks of the divided block, without reading values of the outer pixels from the memory to provide a conflict-free access to the memory;

generating an output block by using the respective output values;

generating, via a processor, a feature map by using the output block; and

writing the feature map to the memory.

11. The method of claim 10 , wherein a size of the output block is larger than a size of the one of the divided blocks.

12. The method of claim 10 , wherein a size of the output block varies according to a size of the kernel.

13. The method of claim 10 ,

wherein addresses respectively corresponding to the divided blocks are assigned with respect to the divided blocks, and

wherein the divided blocks are respectively stored in a plurality of banks of the memory and are accessible by the addresses.

14. The method of claim 13 , wherein the reading comprises:

performing the conflict-free access to one of the plurality of banks with reference to an address of the one of the divided blocks; and

reading data of the one of the divided blocks from the one of the plurality of banks, based on the conflict-free access.

15. The method of claim 10 , further comprising: executing a code temporarily storing kernel information to prevent a stack overflow when performing the convolution operation.

16. The method of claim 10 , wherein the generating of the feature map comprises:

accumulating the output block and other outputs previously stored in a buffer by writing the output block to the buffer;

generating the feature map based on results accumulated in the buffer.

17. The method of claim 16 , wherein the accumulating comprises converting data of a vertical form of the output block into data of a horizontal form and writing the converted data of the horizontal form to the buffer.

18. A non-transitory computer-readable recording medium having recorded thereon a program for performing, via a processor, operations comprising:

reading, from a memory, a divided block among a plurality of divided blocks of input data stored in the memory, wherein a plurality of different addresses in the memory are assigned to the plurality of divided blocks, respectively;

performing the convolution operation on the divided block by applying a kernel to the divided block to generate respective output values corresponding to inner pixels inside the divided block and outer pixels that are included in adjacent blocks of the divided block, without reading values of the outer pixels from the memory to provide a conflict-free access to the memory;

generating an output block by using the respective output values;

generating, via a processor, a feature map by using the output block; and

writing the feature map to the memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2018
From: KIM, KYOUNG-HOON; PARK, YOUNG-HWAN; SUH, DONG-KWAN; PRASAD, KESHAVA; KIM, DAE-HYUN; KIM, SUK-JIN; CHO, HAN-SU; KIM, HYUN-JUNG
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 047690/0143 →
Priority Claims (1)
KR 10-2017-0151722 · Nov 14, 2017 · national
Continuity (1)
Related Publication 20190147319A1 · May 16, 2019