IP Library › Granted Patent US 12,032,931
Granted Patent B2
US 12,032,931 · App. 17/215,712 · Granted Jul 9, 2024

Compiling method and apparatus for neural networks

Inventor: Keunmo Park (Seoul, KR)
Assignee: Samsung Electronics Co., Ltd.
G06F8/41G06N3/04G06N3/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,032,931
App. No.
17/215,712
Granted
Jul 9, 2024
Kind
B2
Abstract

Disclosed are compiling methods and apparatuses, where a compiling method includes receiving a single-core-based code and input data for an operation to be performed based on the single-core-based code, generating kernel clusters by performing graph clustering based on one or more operation kernels in the single-core-based code and the input data, and generating a multi-core-based code based on the kernel clusters.

Claims (40)

1. A compiling method, comprising:

receiving, a single-core neural network based code of an operation and a plurality of input data, each respective input data comprises an operand for the operation included in a neural network and a parameter of the neural network;

generating a graph of the single-core neural network based code, wherein the generating of the graph comprises:

generating a plurality of nodes, wherein each respective node corresponds to a respective operation kernel of a plurality of operation kernels of the single-core neural network based code; and

generating a plurality of edges connecting the nodes based on a similarity between respective input data received by operation kernels, wherein the similarity is based on an amount of similar input data shared between respective operation kernels and a distance between storage locations at which the respective input data used by the operation kernels are stored:

generating kernel clusters, based on the generated graph, wherein the kernel clusters comprise operation kernels using similar input data;

generating, based on the kernel clusters, multi-core neural network based code for the operation;

distributing, based on the multi-core neural network based code, the operation kernels and the plurality of input data to respective cores for execution; and

executing the operation through the cores.

2. The compiling method of claim 1 , wherein generating the plurality of edges further comprises assigning an additional weight to the similarity based on locations at which respective input data received by the operation kernels are stored.

3. The compiling method of claim 2 , wherein the assigning comprises assigning the additional weight to the similarity based on a distance between locations at which respective input data received by a first operation kernel and a second operation kernel of the plurality of operation kernels are stored.

4. The compiling method of claim 1 , wherein the distributing comprises:

storing, based on the multi-core-based code, respective input data corresponding to the one operation kernel.

5. The compiling method of claim 1 , wherein the graph further indicates a number of times that a first operation kernel and a second operation kernel of the plurality of operation kernels receive a same input data.

6. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the steps of:

receiving, a single-core neural network based code of an operation and a plurality of input data, each respective input data comprises an operand for the operation included in a neural network and a parameter of the neural network;

generating a graph of the single-core neural network based code, wherein the generating of the graph comprises:

generating a plurality of nodes, wherein each respective node corresponds to a respective operation kernel of a plurality of operation kernels of the single-core neural network based code; and

generating a plurality of edges connecting the nodes based on a similarity between respective input data received by operation kernels, wherein the similarity is based on the amount of similar input data shared between respective operation kernels and a distance between storage locations at which the respective input data used by the operation kernels are stored;

generating kernel clusters, based on the generated graph, wherein the kernel clusters comprise operation kernels using similar input data;

generating, based on the kernel clusters, a multi-core neural network based code for the operation;

distributing, based on the multi-core neural network based code, the operation kernels and the plurality of input data to respective cores for execution; and

executing the operation through the cores.

7. An apparatus, comprising:

a processor configured to:

receive, a single-core neural network based code of an operation and a plurality of input data, each respective input data comprises an operand for the operation included in a neural network and a parameter of the neural network;

generate, a graph of single-core neural network based code, wherein the generating of the graph comprises:

generating a plurality of nodes, wherein each respective node corresponds to a respective operation kernel of a plurality of operation kernels of the single-core neural network based code; and

generating a plurality of edges connecting the nodes based on a similarity between respective input data received by operation kernels, wherein the similarity is based on the amount of similar input data shared between respective operation kernels and a distance between storage locations at which the respective input data used by the operation kernels are stored;

generate kernel clusters, based on the generated graph, wherein the kernel clusters comprise operation kernels using similar input data;

generate, based on the kernel clusters, a multi-core neural network based code for the operation;

distribute, based on the multi-core neural network based code, the operation kernels and the plurality of input data to respective cores for execution; and

execute the operation though the cores.

8. The apparatus of claim 7 , wherein generating the plurality of edges further comprise, assign an additional weight to the similarity based on locations at which respective input data received by the operation kernels are stored.

9. The apparatus of claim 8 , wherein the processor is further configured to assign the additional weight to similarity based on a distance between locations at which respective input data received by a first operation kernel and a second operation kernel of the plurality of operation kernels are stored.

10. The apparatus of claim 7 , wherein the processor is further configured to:

store, based on the multi-core-based code, respective input data corresponding to the one operation kernel.

11. The apparatus of claim 7 , wherein the processor is further configured to:

store the respective input data corresponding to the one operation kernel in a scratchpad memory connected to the one core.

12. The apparatus of claim 7 , wherein the graph further indicates a number of times that a first operation kernel and a second operation kernel of the plurality of operation kernels receive a same input data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2021
From: PARK, KEUNMO
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 055754/0509 →
Priority Claims (1)
KR 10-2020-0115602 · Sep 9, 2020 · national
Continuity (1)
Related Publication 20220075606A1 · Mar 10, 2022