IP Library › Granted Patent US 12,430,181
Granted Patent B2
US 12,430,181 · App. 17/467,890 · Granted Sep 30, 2025

Method and apparatus for partitioning neural network data

Inventors: Hanwoong Jung (Seoul, KR); Joonho Song (Hwaseong-si, KR); Seungwon Lee (Hwaseong-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06F9/5061G06F9/4881G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,181
App. No.
17/467,890
Granted
Sep 30, 2025
Kind
B2
Abstract

A method and apparatus for scheduling a neural network operation. The method includes receiving data on a layer of a neural network, generating partitions to be assigned to cores by dividing the data, generating tiles by dividing the partitions, and scheduling an operation order of the tiles based on whether the data are shared between the cores.

Claims (51)

1. A processor-implemented method, comprising:

dividing an input feature map or weight, respectively of a layer of a neural network, into partitions to be assigned to cores in different ways depending on whether an output of a previous layer is forwardable;

generating tiles by dividing each one of the partitions; and

scheduling an operation order of the tiles based on whether values of the input feature map or the weight are shared between the cores.

2. The method of claim 1 , wherein dividing of the input feature map or weigh comprises generating the partitions based on a partitioning policy of a layer previous to the layer.

3. The method of claim 1 , wherein dividing of the input feature map or weigh comprises:

generating the partitions based on a partitioning policy of the previous layer in response to the output of the previous layer being forwardable; and

generating the partitions by comparing a size of the input feature map of the layer to a size of the weight of the layer in response to the output of the previous layer being unforwardable.

4. The method of claim 3 , wherein the generating of the partitions by comparing the size of the input feature map of the layer to the size of the weight of the layer in response to the output of the previous layer being unforwardable comprises generating the partitions by comparing a loss caused by a memory size to a loss caused by unbalance in response to the input feature map or the weight being not uniformly divided and assigned to the cores.

5. The method of claim 1 , wherein the generating of the tiles comprises generating the tiles based on a partitioning policy of a layer previous to the layer, a partitioning policy of the layer, or a tile division policy of the previous layer.

6. The method of claim 5 , wherein the generating of the tiles comprises:

generating the tiles based on a partitioning policy of a previous layer, a partitioning policy of the layer, and a tile division policy of the previous layer in response to an output of the previous layer being forwardable; and

generating the tiles by comparing a size of the input feature map of the layer to a size of the weight of the layer in response to the output of the previous layer being unforwardable.

7. The method of claim 1 , wherein the generating of the partitions comprises dividing the data in a height or width direction of the data.

8. The method of claim 1 , wherein the scheduling comprises changing the operation order of the tiles based on whether operation results of the tiles are shared between the cores.

9. The method of claim 8 , wherein the changing of the operation order of the tiles comprises changing the operation order so as to operate one tile included in the tiles with priority in response to an operation result of the one tile being shared between a core to which the one tile is assigned and another core.

10. An apparatus, comprising:

a processor configured to:

dividing an input feature map or a weight, respectively of a layer of a neural network, into partitions to be assigned to cores in different ways depending on whether an output of a previous layer is forwardable;

generate tiles by dividing each one of the partitions; and

schedule an operation order of the tiles based on whether values of the input feature map or the weight are shared between the cores.

11. The apparatus of claim 10 , wherein, for the generation of the partitions, the processor is further configured to generate the partitions based on a partitioning policy of a layer previous to the layer.

12. The apparatus of claim 10 , wherein, for the generation of the partitions, the processor is further configured to generate the tiles based on a partitioning policy of a layer previous to the layer, a partitioning policy of the layer, or a tile division policy of the previous layer.

13. The apparatus of claim 10 , wherein the processor is further configured to divide the data in a height or width direction of the data.

14. The apparatus of claim 10 , wherein the processor is further configured to change the operation order of the tiles based on whether operation results of the tiles are shared between the cores.

15. The apparatus of claim 14 , wherein the processor is further configured to change the operation order so as to operate one tile included in the tiles with priority in response to an operation result of the one tile being shared between a core to which the one tile is assigned and another core.

16. The apparatus of claim 11 , wherein, for the generation of the partitions, the processor is further configured to:

generate the partitions based on the partitioning policy of the previous layer in response to an output of the previous layer being forwardable; and

generate the partitions by comparing a size of the input feature map of the layer to a size of the weight of the layer in response to the output of the previous layer being unforwardable.

17. The apparatus of claim 12 , wherein, for the generation of the partitions, the processor is further configured to:

generate the tiles based on the partitioning policy of the previous layer, the partitioning policy of the layer, and the tile division policy of the previous layer in response to an output of the previous layer being forwardable; and

generate the tiles by comparing a size of the input feature map of the layer to a size of the weight of the layer in response to the output of the previous layer being unforwardable.

18. An apparatus, comprising:

a receiver configured to receive data on a layer of a neural network; and

a processor configured to:

generate partitions to be assigned to cores by dividing the data;

generate tiles by dividing the partitions; and

schedule an operation order of the tiles based on whether the data are shared between the cores,

wherein, for the generation of the partitions, the processor is further configured to:

generate the partitions based on a partitioning policy of a previous layer in response to an output of the previous layer being forwardable; and

generate the partitions by comparing a size of an input feature map of the layer to a size of a weight of the layer in response to the output of the previous layer being unforwardable.

19. The apparatus of claim 18 , wherein, for the generation of the partitions, the processor is further configured to generate the partitions by comparing a loss caused by a memory size to a loss caused by unbalance in response to the input feature map or the weight being not uniformly divided and assigned to the cores.

20. An apparatus, comprising:

a receiver configured to receive data on a layer of a neural network; and

a processor configured to:

generate partitions to be assigned to cores by dividing the data;

generate tiles by dividing the partitions; and

schedule an operation order of the tiles based on whether the data are shared between the cores,

wherein, for the generation of the partitions, the processor is further configured to:

generate the tiles based on a partitioning policy of a previous layer, a partitioning policy of the layer, and a tile division policy of the previous layer in response to an output of the previous layer being forwardable; and

generate the tiles by comparing a size of an input feature map of the layer to a size of a weight of the layer in response to the output of the previous layer being unforwardable.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2021
From: JUNG, HANWOONG; SONG, JOONHO; LEE, SEUNGWON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 057400/0010 →
Priority Claims (1)
KR 10-2020-0167681 · Dec 3, 2020 · national
Continuity (1)
Related Publication 20220179714A1 · Jun 9, 2022
References Cited (14)
US 10223333B2 · Chetlur et al. · 2019 [cited by applicant]
US 10417555B2 · Brothers et al. · 2019 [cited by applicant]
US 20180322606A1 · Das · 2018 [cited by examiner]
US 20180373976A1 · Woo · 2018 [cited by examiner]
US 20190340491A1 · Norden et al. · 2019 [cited by applicant]
US 20200104690A1 · Bai et al. · 2020 [cited by applicant]
US 20200293866A1 · Guo · 2020 [cited by examiner]
US 20200371835A1 · Chuang · 2020 [cited by examiner]
US 20210191765A1 · Bokam · 2021 [cited by examiner]
US 20220147791A1 · Yao · 2022 [cited by examiner]
KR 1020090094256A · 2009 [cited by applicant]
KR 1020180062912A · 2018 [cited by applicant]
KR 1020190018888A · 2019 [cited by applicant]
Li, Jiajun, et al., “SmartShuttle: Optimizing Off-Chip Memory Accesses for Deep Learning Accelerators”, [cited by applicant]