IP Library › Granted Patent US 12,725,004
Granted Patent B2
US 12,725,004 · App. 17/211,606 · Granted Sep 1, 2026

Neural architecture search based optimized DNN model generation for execution of tasks in electronic device

Inventors: Mayukh Das (Bengaluru, IN); Venkappa Mala (Bengaluru, IN); Brijraj Singh (Bengaluru, IN); Pradeep Nelahonne Shivamurthappa (Bengaluru, IN); Sharan Kumar Allur (Bengaluru, IN)
Assignee: Samsung Electronics Co., Ltd.
G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,004
App. No.
17/211,606
Granted
Sep 1, 2026
Kind
B2
Abstract

Embodiments herein provide a NAS method of generating an optimized DNN model for executing a task in an electronic device. The method includes identifying the task to be executed in the electronic device. The method includes estimating a performance parameter to be achieved while executing the task. The method includes determining hardware parameters of the electronic device required to execute the task based on the performance parameter and the task, and determining optimal neural blocks from a plurality of neural blocks based on the performance parameter and the hardware parameter of the electronic device. The method includes generating the optimized DNN model for executing the task based on the optimal neural blocks, and executing the task using the optimized DNN model.

Claims (96)

1 . A method of operating an electronic device, the method comprising:

identifying, by the electronic device, at least one task to be executed in the electronic device;

estimating, by the electronic device, at least one performance parameter to be achieved while executing the at least one task, wherein the at least one performance parameter is at least one of a frame rate, a resolution, and a bit rate;

determining, by the electronic device, at least one hardware parameter of the electronic device used to execute the at least one task based on the at least one performance parameter and the at least one task, wherein the at least one hardware parameter is at least one of a processor speed, a number of cores in a processor, a data transmission speed, a storage capacity of a memory, and a write/read speed at the memory;

performing, by the electronic device, a Neural Architecture Search (NAS) of a plurality of neural blocks from a Deep Neural Network (DNN) model, based on the at least one performance parameter, the at least one hardware parameter of the electronic device, and a search space including all possible choices of the plurality of neural blocks;

determining, by the electronic device, a quality of each neural block in the plurality of neural blocks based on (i) a probability distribution in executing the at least one task and (ii) a two-step truncation operation that includes truncation based on information value, and truncation based on confidence bounds;

selecting, by the electronic device, at least one optimal neural block from the plurality of neural blocks based on a result of the NAS and the quality of each neural block;

generating, by the electronic device, an optimized DNN model within the electronic device for executing the at least one task based on the at least one optimal neural block; and

executing, by the electronic device, the at least one task using the optimized DNN model,

wherein inputs to the truncation based on information value include neural choices, and a past history of usage of the neural choices and wherein inputs to the truncation based on confidence bounds include neural choices and a policy distribution over the neural choices.

2 . The method as claimed in claim 1 , wherein estimating, by the electronic device, the at least one performance parameter to be achieved while executing the at least one task comprises:

obtaining, by the electronic device, execution data for different types of DNN architectural elements from different types of hardware configuration of a plurality of electronic devices;

training, by the electronic device, a hybrid ensemble meta-model based on the execution data; and

estimating, by the electronic device, the at least one performance parameter to be achieved while executing the at least one task based on the hybrid ensemble meta-model.

3 . The method as claimed in claim 1 , wherein selecting, by the electronic device, the at least one optimal neural block from the plurality of neural blocks based on the result of the NAS comprises:

representing, by the electronic device, an intermediate DNN model using the plurality of neural blocks;

providing, by the electronic device, data inputs to the intermediate DNN model,

wherein the quality of each neural block in the plurality of neural blocks is determined based on a probability distribution in executing the at least one task using the data inputs, the at least one performance parameter and the at least one hardware parameter;

generating, by the electronic device, a standard DNN model using the at least one optimal neural block; and

optimizing, by the electronic device, the standard DNN model by modifying unsupported operations used for the execution of the at least one task with supported operations to generate the optimized DNN model.

4 . The method as claimed in claim 3 , wherein representing, by the electronic device, the intermediate DNN model using the plurality of neural blocks, comprises:

maintaining, by the electronic device, a truncated parameterized distribution over all of the plurality of neural blocks at each layer that manifests a measure of a relative value of every neural block among the plurality of neural blocks subject to the at least one hardware parameter and the at least one task;

electing, by the electronic device, useful neural elements based on the two-step truncation operation; and

representing, by the electronic device, the intermediate DNN model using the selected useful neural elements.

5 . The method as claimed in claim 3 , wherein determining, by the electronic device, the quality of each neural block in the plurality of neural blocks based on the probability distribution in executing the at least one task using the data inputs, the at least one performance parameter and the at least one hardware parameter, comprises:

encoding, by the electronic device, a layer depth and features of neural blocks;

creating, by the electronic device, an action space comprising a set of neural block choices for every learnable block;

determining, by the electronic device, based on the two-step truncation operation, a usefulness of the set of neural block choices;

adding, by the electronic device, an abstract layer with choices, from the two-step truncation operation, of the set of neural block choices with the at least one hardware parameter and the at least one task;

finding, by the electronic device, an expected latency for the set of neural block choices using a latency predictor metamodel; and

finding, by the electronic device, an expected accuracy after adding the set of neural block choices by sampling paths in the abstract layer.

6 . The method as claimed in claim 3 , wherein selecting, by the electronic device, the at least one optimal neural block from the plurality of neural blocks based on the quality of each neural block, comprises:

instantiating, by the electronic device, the intermediate DNN model;

extracting, by the electronic device, constant values for the at least one task and the at least one hardware parameter based on the intermediate DNN model; and

selecting, by the electronic device, the at least one optimal neural block from the plurality of neural blocks based on the quality of each neural block.

7 . The method as claimed in claim 3 , wherein optimizing, by the electronic device, the standard DNN model by modifying the unsupported operations used for the execution of the task with the supported operations to generate the optimized DNN model, comprises:

searching, by the electronic device, for standard operations at a knowledgebase to replace the unsupported operations, and

performing, by the electronic device, at least one of:

replacing the unsupported operations with the standard operations, and retraining at least one neural block of the plurality of neural blocks with the standard operations, when the standard operations are available; or

optimizing the unsupported operations using universal approximator Pade' Approximation Units (PAUs) for the task execution, when the standard operations are unavailable.

8 . An electronic device comprising:

a memory;

a processor; and

a Neural Architecture Search (NAS) controller, operably coupled to the memory and the processor,

wherein the processor is configured to:

identify at least one task to be executed in the electronic device;

estimate at least one performance parameter to be achieved while executing the at least one task, wherein the at least one performance parameter is at least one of a frame rate, a resolution, and a bit rate;

determine at least one hardware parameter of the electronic device used to execute the at least one task based on the at least one performance parameter and the at least one task, wherein the at least one hardware parameter is at least one of a processor speed, a number of cores in the processor, a data transmission speed, a storage capacity of the memory, and a write/read speed at the memory;

perform a Neural Architecture Search (NAS) of a plurality of neural blocks from a Deep Neural Network (DNN) model, based on the at least one performance parameter, the at least one hardware parameter of the electronic device, and a search space including all possible choices of the plurality of neural blocks;

determine a quality of each neural block in the plurality of neural blocks based on (i) a probability distribution in executing the at least one task and (ii) a two-step truncation operation that includes truncation based on information value, and truncation based on confidence bounds;

select at least one optimal neural block from the plurality of neural blocks based on a result of the NAS and the quality of each neural block;

generate an optimized DNN model within the electronic device for executing the at least one task based on the at least one optimal neural block; and

execute the at least one task using the optimized DNN model,

wherein inputs to the truncation based on information value include neural choices, and a past history of usage of the neural choices and wherein inputs to the truncation based on confidence bounds include neural choices and a policy distribution over the neural choices.

9 . The electronic device as claimed in claim 8 , wherein to estimate the at least one performance parameter to be achieved while executing the at least one task, the processor is configured to:

obtain execution data for different types of DNN architectural elements from different types of hardware configuration of a plurality of electronic devices;

train a hybrid ensemble meta-model based on the execution data; and

estimate the at least one performance parameter to be achieved while executing the at least one task based on the hybrid ensemble meta-model.

10 . The electronic device as claimed in claim 8 , wherein to select the at least one optimal neural block from the plurality of neural blocks based on the result of the NAS, the processor is configured to:

represent an intermediate DNN model using the plurality of neural blocks;

provide data inputs to the intermediate DNN model,

wherein the quality of each neural block in the plurality of neural blocks is determined based on a probability distribution in executing the at least one task using the data inputs, the at least one performance parameter and the at least one hardware parameter;

generate a standard DNN model using the at least one optimal neural block; and

optimize the standard DNN model by modifying unsupported operations used for the execution of the at least one task with supported operations to generate the optimized DNN model.

11 . The electronic device as claimed in claim 10 , wherein to represent the intermediate DNN model using the plurality of neural blocks, the processor is configured to:

maintain a truncated parameterized distribution over all of the plurality of neural blocks at each layer that manifests a measure of a relative value of every neural block among the plurality of neural blocks subject to the at least one hardware parameter and the at least one task;

select useful neural elements based on the two-step truncation operation; and

represent the intermediate DNN model using the selected useful neural elements.

12 . The electronic device as claimed in claim 10 , wherein to determine the quality of each neural block in the plurality of neural blocks based on the probability distribution in executing the at least one task using the data inputs, the at least one performance parameter and the at least one hardware parameter, the processor is configured to:

encode a layer depth and features of neural blocks;

create an action space comprising a set of neural block choices for every learnable block;

determine, based on the two-step truncation operation, a usefulness of the set of neural block choices;

add an abstract layer with choices, from the two-step truncation operation, of the set of neural block choices with the at least one hardware parameter and the at least one task;

find an expected latency for the set of neural block choices using a latency predictor metamodel; and

find an expected accuracy after adding the set of neural block choices by sampling paths in the abstract layer.

13 . The electronic device as claimed in claim 10 , wherein to select the at least one optimal neural block from the plurality of neural blocks based on the quality of each neural block, the processor is configured to:

instantiate the intermediate DNN model;

extract constant values for the at least one task and the at least one hardware parameter based on the intermediate DNN model; and

select the at least one optimal neural block from the plurality of neural blocks based on the quality of each neural block.

14 . The electronic device as claimed in claim 10 , wherein to optimize the standard DNN model by modifying the unsupported operations used for the execution of the task with the supported operations to generate the optimized DNN model, the processor is configured to:

search for standard operations at a knowledgebase to replace the unsupported operations, and

perform at least one of:

replacing the unsupported operations with the standard operations, and retraining at least one neural block of the plurality of neural blocks with the standard operations, when the standard operations are available; or

optimizing the unsupported operations, using universal approximator Pade' Approximation Units (PAUs), for the task execution, when the standard operations are unavailable.

15 . An intelligent deployment method for neural networks in a multi-device environment, comprising:

identifying, by an electronic device, a task to be executed in the electronic device;

estimating, by the electronic device, a performance threshold at a time of execution of the identified task;

identifying, by the electronic device, an operation capability of the electronic device, wherein the operation capability is at least one of a processor speed, a number of cores in a processor, a data transmission speed, a storage capacity of a memory, and a write/read speed at the memory;

performing, by the electronic device, a Neural Architecture Search (NAS) of a plurality of neural blocks from a Deep Neural Network (DNN) model, based on the operation capability of the electronic device, and a search space including all possible choices of the plurality of neural blocks;

determining, by the electronic device, a quality of each neural block in the plurality of neural blocks based on (i) a probability distribution in executing the identified task and (ii) a two-step truncation operation that includes truncation based on information value, and truncation based on confidence bounds; and

configuring, by the electronic device, a pre-trained Artificial Intelligence (AI) model to select one or more neural blocks from the plurality of neural blocks based on a result of the NAS and the quality of each neural block to optimize a performance of the identified task in the electronic device,

wherein inputs to the truncation based on information value include neural choices, and a past history of usage of the neural choices and wherein inputs to the truncation based on confidence bounds include neural choices and a policy distribution over the neural choices.

16 . The method as claimed in claim 15 , wherein the quality of each neural block is determined using a probability distribution in the task execution.

17 . The method as claimed in claim 15 , wherein a standard Deep Neural Network (DNN) model is generated using the one or more neural blocks.

18 . The method as claimed in claim 15 , wherein the performance threshold comprises an accuracy threshold, a quality threshold of image, a latency threshold, a memory consumption threshold, a power consumption threshold, and a bandwidth threshold.

19 . The method as claimed in claim 15 , wherein the operation capability of the electronic device comprises, a screen refresh rate, a sampling rate, a camera resolution, a pixel density of a screen, a frame rate, a screen resolution, single/multiple display, an audio format support, and a video format support.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2021
From: DAS, MAYUKH; MALA, VENKAPPA; SINGH, BRIJRAJ; NELAHONNE SHIVAMURTHAPPA, PRADEEP; ALLUR, SHARAN KUMAR
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 055707/0287 →
Priority Claims (2)
IN 202041019468 · May 7, 2020 · national
IN 202041019468 · Dec 15, 2020 · national
Continuity (1)
Related Publication 20210350203A1 · Nov 11, 2021
References Cited (47)
US 5787408A · Deangelis · 1998 [cited by applicant]
US 8626698B1 · Nikolaev et al. · 2014 [cited by applicant]
US 9177550B2 · Yu · 2015 [cited by applicant]
US 10496927B2 · Achin · 2019 [cited by examiner]
US 11468275B1 · Blechschmidt · 2022 [cited by examiner]
US 12353971B1 · Perumalla · 2025 [cited by examiner]
US 20140257803A1 · Yu et al. · 2014 [cited by applicant]
US 20150019214A1 · Wang · 2015 [cited by examiner]
US 20160132787A1 · Drevo · 2016 [cited by examiner]
US 20160358070A1 · Brothers et al. · 2016 [cited by applicant]
US 20180032867A1 · Son · 2018 [cited by examiner]
US 20180165597A1 · Jordan et al. · 2018 [cited by applicant]
US 20180189638A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180307987A1 · Bleiweiss et al. · 2018 [cited by applicant]
US 20190147337A1 · Yang · 2019 [cited by examiner]
US 20190188537A1 · Dutta et al. · 2019 [cited by applicant]
US 20190251440A1 · Kasiviswanathan et al. · 2019 [cited by applicant]
US 20190258964A1 · Dube et al. · 2019 [cited by applicant]
US 20190354837A1 · Zhou et al. · 2019 [cited by applicant]
US 20190362222A1 · Chen · 2019 [cited by applicant]
US 20200005135A1 · Che · 2020 [cited by applicant]
US 20200184318A1 · Minezawa · 2020 [cited by examiner]
US 20210089285A1 · Du et al. · 2021 [cited by applicant]
US 20210092035A1 · Williams · 2021 [cited by examiner]
US 20210097383A1 · Kaur · 2021 [cited by examiner]
US 20210109725A1 · Du et al. · 2021 [cited by applicant]
US 20210109726A1 · Du et al. · 2021 [cited by applicant]
US 20210109727A1 · Du et al. · 2021 [cited by applicant]
US 20210109728A1 · Du et al. · 2021 [cited by applicant]
US 20210109729A1 · Du et al. · 2021 [cited by applicant]
US 20210274251A1 · Wang · 2021 [cited by examiner]
US 20210312276A1 · Rawat · 2021 [cited by examiner]
US 20210319272A1 · Gaidon · 2021 [cited by examiner]
US 20220347583A1 · He · 2022 [cited by examiner]
US 20240273336A1 · Tan · 2024 [cited by examiner]
CN 110580527A · 2019 [cited by applicant]
CN 111159489A · 2020 [cited by applicant]
JP 2019159693A · 2019 [cited by applicant]
Molina et.al., “Padé Activation Units: End-to-End Learning of Flexible Activation Functions in Deep Networks”, Feb. 4, 2020, Published as a conference paper at ICLR 2020 (Year: 2020). [cited by examiner]
Examination report dated Dec. 20, 2021, in connection with Indian Application No. 202041019468, 8 pages. [cited by applicant]
Van Stein et al., “Automatic Configuration of Deep Neural Networks with Parallel Global Optimization”, Oct. 10, 2018, 8 pages. [cited by applicant]
Saxena et al., “Convolutional Neural Fabrics”, Computer Vision and Pattern Recognition, Jun. 8, 2016, 9 pages. [cited by applicant]
Molina et al., “Pade Activation Units: End-to-End Learning of Flexible Activation Functions in Deep Networks”, Feb. 4, 2020, 17 pages. [cited by applicant]
Cai et al., “ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware”, Feb. 23, 2019, 13 pages. [cited by applicant]
International Search Report dated Jun. 17, 2021 in connection with International Patent Application No. PCT/KR2021/002400, 4 pages. [cited by applicant]
Written Opinion of the International Searching Authority dated Jun. 17, 2021 in connection with International Patent Application No. PCT/KR2021/002400, 4 pages. [cited by applicant]
Hearing Notice issued Jan. 16, 2025, in connection with Indian Patent Application No. 202041019468, 4 pages. [cited by applicant]