IP Library › Granted Patent US 12,614,076
Granted Patent B2
US 12,614,076 · App. 18/786,758 · Granted Apr 28, 2026

Neural network optimization device for edge device meeting on-demand instruction and method using the same

Inventor: Haeryong Jeon (Seongnam-si, KR)
Assignee: Mobilint, Inc.
G06N3/082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,076
App. No.
18/786,758
Granted
Apr 28, 2026
Kind
B2
Abstract

Disclosed is a neural network optimizing method performed by a device including performing a computation on an input value with a basic resource of a first neural network based on a customer-requested instruction for the basic resource used for computation execution of the first neural network, checking computational performance including a computational processing speed, a power consumption amount, and a chip area according to the computation execution of the first neural network, adjusting at least one of the basic resource based on the checked computational performance, and re-performing the computation on the input value based on the customer-requested instruction with a resource of the second neural network after changing to an environment of a second neural network by adjusting the at least one of the basic resource.

Claims (45)

1 . A neural network optimizing method performed by a device, the method comprising:

performing, based on a customer-requested instruction for basic resources used for computation execution of a first neural network, a first computation on an input value with the basic resources, wherein the performing of the first computation comprises inputting the customer-requested instruction into a particular resource having a greatest size, among the basic resources;

identifying a bottleneck portion caused in a process of data exchange between data blocks of the first neural network, at an intermediate branch point between a data line and a buffer;

checking a first computational performance of the first computation, wherein the first computational performance includes a computational processing speed and a power consumption amount of the bottleneck portion;

adjusting, based on the checked computational performance and a customer-requested performance according to the customer-requested instruction, at least one resource of the basic resources, adjusting a data path of the bottleneck portion, and configuring a second neural network with the adjusted at least one resource and the adjusted data path;

performing a second computation on the input value with the adjusted at least one resource and the adjusted data path of the second neural network;

checking a second computational performance of the second computation; and

when the second computational performance of the second computation minimizes consumption of the adjusted at least one resource and latency of the adjusted data path, outputting the second neural network as a final neural network,

wherein the customer-requested instruction is an instruction for generating the final neural network that meets a customer's request,

wherein the adjusting of the at least one resource of the basic resources comprises:

adjusting a first resource among the basic resources, wherein the first resource is configured to adjust internal memory capacity and a number of channels;

determining whether the customer-requested performance is satisfied with the adjusted, first resource;

when the customer-requested performance is satisfied with the adjusted first resource, further adjusting a second resource among the basic resources, wherein the second resource is configured to adjust a pipeline and a data bandwidth;

determining whether the customer-requested performance is satisfied with the adjusted second resource; and

when the customer-requested performance is satisfied with the adjusted second resource, further adjusting a third resource among the basic resources, wherein the third resource is configured to perform an additional function of the customer's request, and

wherein the method performs all of the performing of the first computation, the identifying of the bottleneck portion, the checking of the first computational performance, the adjusting of the at least one resource, the adjusting of the data path, the configuring of the second neural network, the performing of the second computation, and the outputting of the second neural network, again, for a different customer-requested instruction.

2 . A neural network optimization device comprising:

a computation execution block configured to perform, based on a customer-requested instruction for basic resources used for computation execution of a first neural network, a first computation on an input value with the basic resources, wherein the first computation is performed by inputting the customer-requested instruction into a particular resource having a greatest size, among the basic resources;

a computation performance check block configured to identify a bottleneck portion caused in a process of data exchange between data blocks of the first neural network, at an intermediate branch point between a data line and a buffer, and check a first computational performance of the first computation, wherein the first computational performance includes a computational processing speed and a power consumption amount of the bottleneck portion;

a resource adjustment block configured to adjust, based on the checked computational performance and a customer-requested performance according to the customer-requested instruction, at least one resource of the basic resources, adjust a data path of the bottleneck portion, and configure a second neural network with the adjusted at least one resource and the adjusted data path; and

an optimization execution block configured to perform a second computation on the input value with the adjusted at least one resource and the adjusted data path of the second neural network, check a second computational performance of the second computation, and output the second neural network as a final neural network when the second computational performance of the second computation minimizes consumption of the adjusted at least one resource and latency of the adjusted data path,

wherein the customer-requested instruction is an instruction for generating the final neural network that meets a customer's request,

wherein the resource adjustment block is configured to:

adjust a first resource among the basic resources, wherein the first resource is configured to adjust internal memory capacity and a number of channels,

determine whether the customer-requested performance is satisfied with the adjusted first resource;

when the customer-requested performance is satisfied with the adjusted first resource, further adjust a second resource among the basic resources, wherein the second resource is configured to adjust a pipeline and a data bandwidth;

determine whether the customer-requested performance is satisfied with the adjusted second resource; and

when the customer-requested performance is satisfied with the adjusted second resource, further adjust a third resource among the basic resources, wherein the third resource is configured to perform an additional function of the customer's request, and

wherein the neural network optimization device performs all of performing of the first computation, the identifying of the bottleneck portion, the checking of the first computational performance, the adjusting of the at least one resource, the adjusting of the data path, the configuring of the second neural network, the performing of the second computation, and the outputting of the second neural network, again, for a different customer-requested instruction.

3 . A non-transitory computer-readable recording medium that stores a computer program which, when executed by one or more processors, performs following operations for performing a neural network optimizing method, the operations including:

performing, based on a customer-requested instruction for basic resources used for computation execution of a first neural network, a first computation on an input value with the basic resources, wherein the performing of the first computation comprises inputting the customer-requested instruction into a particular resource having a greatest size, among the basic resources;

identifying a bottleneck portion caused in a process of data exchange between data blocks of the first neural network, at an intermediate branch point between a data line and a buffer;

checking a first computational performance of the first computation, wherein the first computational performance includes a computational processing speed and a power consumption amount of the bottleneck portion;

adjusting, based on the checked computational performance and a customer-requested performance according to the customer-requested instruction, at least one resource of the basic resources, adjusting a data path of the bottleneck portion, and configuring a second neural network with the adjusted at least one resource and the adjusted data path;

performing a second computation on the input value with the adjusted at least one resource and the adjusted data path of the second neural network;

checking a second computational performance of the second computation; and

when the second computational performance of the second computation minimizes consumption of the adjusted at least one resource and latency of the adjusted data path, outputting the second neural network as a final neural network,

wherein the customer-requested instruction is an instruction for generating the final neural network that meets a customer's request,

wherein the adjusting of the at least one resource of the basic resources comprises:

adjusting a first resource among the basic resources, wherein the first resource is configured to adjust internal memory capacity and a number of channels;

determining whether the customer-requested performance is satisfied with the adjusted first resource;

when the customer-requested performance is satisfied with the adjusted first resource, further adjusting a second resource among the basic resources, wherein the second resource is configured to adjust a pipeline and a data bandwidth;

determining whether the customer-requested performance is satisfied with the adjusted second resource; and

when the customer-requested performance is satisfied with the adjusted second resource, further adjusting a third resource among the basic resources, wherein the third resource is configured to perform an additional function of the customer's request, and

wherein the method performs all of the performing of the first computation, the identifying of the bottleneck portion, the checking of the first computational performance, the adjusting of the at least one resource, the adjusting of the data path, the configuring of the second neural network, the performing of the second computation, and the outputting of the second neural network, again, for a different customer-requested instruction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2024
From: JEON, HAERYONG
To: MOBILINT, INC.
Reel/Frame 068108/0828 →
Priority Claims (1)
KR 10-2022-0172462 · Dec 12, 2022 · national
Continuity (2)
Continuation PCTKR2022021348 · Dec 27, 2022
Related Publication 20240386276A1 · Nov 21, 2024
References Cited (23)
US 10685295B1 · Ross · 2020 [cited by examiner]
US 11443162B2 · Yang et al. · 2022 [cited by applicant]
US 20110320751A1 · Wang · 2011 [cited by examiner]
US 20120226809A1 · Yang · 2012 [cited by examiner]
US 20160357610A1 · Bartfai-Walcott · 2016 [cited by examiner]
US 20180285726A1 · Baum et al. · 2018 [cited by applicant]
US 20200210836A1 · Kim et al. · 2020 [cited by applicant]
US 20210056378A1 · Yang et al. · 2021 [cited by applicant]
US 20210089611A1 · Jiao · 2021 [cited by examiner]
US 20210224275A1 · Maheshwari et al. · 2021 [cited by applicant]
US 20220147801A1 · Davidi et al. · 2022 [cited by applicant]
US 20220188569A1 · Ananthanarayanan et al. · 2022 [cited by applicant]
US 20220414425A1 · Yang et al. · 2022 [cited by applicant]
KR 1020190136431A · 2019 [cited by applicant]
KR 1020200084099A · 2020 [cited by applicant]
KR 102247629B1 · 2021 [cited by applicant]
KR 1020210125911A · 2021 [cited by applicant]
KR 1020220032861A · 2022 [cited by applicant]
KR 1020220047850A · 2022 [cited by applicant]
KR 1020220061835A · 2022 [cited by applicant]
International Search Report issued in PCT/KR2022/021348; mailed Aug. 23, 2023. [cited by applicant]
“Notice of Submission of Opinions” Office Action issued in KR 10-2022-0172462; mailed by the Korean Intellectual Property Office on Sep. 21, 2024. [cited by applicant]
“Written Decision on Registration” Office Action issued in KR 10-2022-0172462; mailed by the Korean Intellectual Property Office on Mar. 7, 2025. [cited by applicant]