Neural architecture search based optimized DNN model generation for execution of tasks in electronic device
Embodiments herein provide a NAS method of generating an optimized DNN model for executing a task in an electronic device. The method includes identifying the task to be executed in the electronic device. The method includes estimating a performance parameter to be achieved while executing the task. The method includes determining hardware parameters of the electronic device required to execute the task based on the performance parameter and the task, and determining optimal neural blocks from a plurality of neural blocks based on the performance parameter and the hardware parameter of the electronic device. The method includes generating the optimized DNN model for executing the task based on the optimal neural blocks, and executing the task using the optimized DNN model.
1 . A method of operating an electronic device, the method comprising:
identifying, by the electronic device, at least one task to be executed in the electronic device;
estimating, by the electronic device, at least one performance parameter to be achieved while executing the at least one task, wherein the at least one performance parameter is at least one of a frame rate, a resolution, and a bit rate;
determining, by the electronic device, at least one hardware parameter of the electronic device used to execute the at least one task based on the at least one performance parameter and the at least one task, wherein the at least one hardware parameter is at least one of a processor speed, a number of cores in a processor, a data transmission speed, a storage capacity of a memory, and a write/read speed at the memory;
performing, by the electronic device, a Neural Architecture Search (NAS) of a plurality of neural blocks from a Deep Neural Network (DNN) model, based on the at least one performance parameter, the at least one hardware parameter of the electronic device, and a search space including all possible choices of the plurality of neural blocks;
determining, by the electronic device, a quality of each neural block in the plurality of neural blocks based on (i) a probability distribution in executing the at least one task and (ii) a two-step truncation operation that includes truncation based on information value, and truncation based on confidence bounds;
selecting, by the electronic device, at least one optimal neural block from the plurality of neural blocks based on a result of the NAS and the quality of each neural block;
generating, by the electronic device, an optimized DNN model within the electronic device for executing the at least one task based on the at least one optimal neural block; and
executing, by the electronic device, the at least one task using the optimized DNN model,
wherein inputs to the truncation based on information value include neural choices, and a past history of usage of the neural choices and wherein inputs to the truncation based on confidence bounds include neural choices and a policy distribution over the neural choices.
2 . The method as claimed in claim 1 , wherein estimating, by the electronic device, the at least one performance parameter to be achieved while executing the at least one task comprises:
obtaining, by the electronic device, execution data for different types of DNN architectural elements from different types of hardware configuration of a plurality of electronic devices;
training, by the electronic device, a hybrid ensemble meta-model based on the execution data; and
estimating, by the electronic device, the at least one performance parameter to be achieved while executing the at least one task based on the hybrid ensemble meta-model.
3 . The method as claimed in claim 1 , wherein selecting, by the electronic device, the at least one optimal neural block from the plurality of neural blocks based on the result of the NAS comprises:
representing, by the electronic device, an intermediate DNN model using the plurality of neural blocks;
providing, by the electronic device, data inputs to the intermediate DNN model,
wherein the quality of each neural block in the plurality of neural blocks is determined based on a probability distribution in executing the at least one task using the data inputs, the at least one performance parameter and the at least one hardware parameter;
generating, by the electronic device, a standard DNN model using the at least one optimal neural block; and
optimizing, by the electronic device, the standard DNN model by modifying unsupported operations used for the execution of the at least one task with supported operations to generate the optimized DNN model.
4 . The method as claimed in claim 3 , wherein representing, by the electronic device, the intermediate DNN model using the plurality of neural blocks, comprises:
maintaining, by the electronic device, a truncated parameterized distribution over all of the plurality of neural blocks at each layer that manifests a measure of a relative value of every neural block among the plurality of neural blocks subject to the at least one hardware parameter and the at least one task;
electing, by the electronic device, useful neural elements based on the two-step truncation operation; and
representing, by the electronic device, the intermediate DNN model using the selected useful neural elements.
5 . The method as claimed in claim 3 , wherein determining, by the electronic device, the quality of each neural block in the plurality of neural blocks based on the probability distribution in executing the at least one task using the data inputs, the at least one performance parameter and the at least one hardware parameter, comprises:
encoding, by the electronic device, a layer depth and features of neural blocks;
creating, by the electronic device, an action space comprising a set of neural block choices for every learnable block;
determining, by the electronic device, based on the two-step truncation operation, a usefulness of the set of neural block choices;
adding, by the electronic device, an abstract layer with choices, from the two-step truncation operation, of the set of neural block choices with the at least one hardware parameter and the at least one task;
finding, by the electronic device, an expected latency for the set of neural block choices using a latency predictor metamodel; and
finding, by the electronic device, an expected accuracy after adding the set of neural block choices by sampling paths in the abstract layer.
6 . The method as claimed in claim 3 , wherein selecting, by the electronic device, the at least one optimal neural block from the plurality of neural blocks based on the quality of each neural block, comprises:
instantiating, by the electronic device, the intermediate DNN model;
extracting, by the electronic device, constant values for the at least one task and the at least one hardware parameter based on the intermediate DNN model; and
selecting, by the electronic device, the at least one optimal neural block from the plurality of neural blocks based on the quality of each neural block.
7 . The method as claimed in claim 3 , wherein optimizing, by the electronic device, the standard DNN model by modifying the unsupported operations used for the execution of the task with the supported operations to generate the optimized DNN model, comprises:
searching, by the electronic device, for standard operations at a knowledgebase to replace the unsupported operations, and
performing, by the electronic device, at least one of:
replacing the unsupported operations with the standard operations, and retraining at least one neural block of the plurality of neural blocks with the standard operations, when the standard operations are available; or
optimizing the unsupported operations using universal approximator Pade' Approximation Units (PAUs) for the task execution, when the standard operations are unavailable.
8 . An electronic device comprising:
a memory;
a processor; and
a Neural Architecture Search (NAS) controller, operably coupled to the memory and the processor,
wherein the processor is configured to:
identify at least one task to be executed in the electronic device;
estimate at least one performance parameter to be achieved while executing the at least one task, wherein the at least one performance parameter is at least one of a frame rate, a resolution, and a bit rate;
determine at least one hardware parameter of the electronic device used to execute the at least one task based on the at least one performance parameter and the at least one task, wherein the at least one hardware parameter is at least one of a processor speed, a number of cores in the processor, a data transmission speed, a storage capacity of the memory, and a write/read speed at the memory;
perform a Neural Architecture Search (NAS) of a plurality of neural blocks from a Deep Neural Network (DNN) model, based on the at least one performance parameter, the at least one hardware parameter of the electronic device, and a search space including all possible choices of the plurality of neural blocks;
determine a quality of each neural block in the plurality of neural blocks based on (i) a probability distribution in executing the at least one task and (ii) a two-step truncation operation that includes truncation based on information value, and truncation based on confidence bounds;
select at least one optimal neural block from the plurality of neural blocks based on a result of the NAS and the quality of each neural block;
generate an optimized DNN model within the electronic device for executing the at least one task based on the at least one optimal neural block; and
execute the at least one task using the optimized DNN model,
wherein inputs to the truncation based on information value include neural choices, and a past history of usage of the neural choices and wherein inputs to the truncation based on confidence bounds include neural choices and a policy distribution over the neural choices.
9 . The electronic device as claimed in claim 8 , wherein to estimate the at least one performance parameter to be achieved while executing the at least one task, the processor is configured to:
obtain execution data for different types of DNN architectural elements from different types of hardware configuration of a plurality of electronic devices;
train a hybrid ensemble meta-model based on the execution data; and
estimate the at least one performance parameter to be achieved while executing the at least one task based on the hybrid ensemble meta-model.
10 . The electronic device as claimed in claim 8 , wherein to select the at least one optimal neural block from the plurality of neural blocks based on the result of the NAS, the processor is configured to:
represent an intermediate DNN model using the plurality of neural blocks;
provide data inputs to the intermediate DNN model,
wherein the quality of each neural block in the plurality of neural blocks is determined based on a probability distribution in executing the at least one task using the data inputs, the at least one performance parameter and the at least one hardware parameter;
generate a standard DNN model using the at least one optimal neural block; and
optimize the standard DNN model by modifying unsupported operations used for the execution of the at least one task with supported operations to generate the optimized DNN model.
11 . The electronic device as claimed in claim 10 , wherein to represent the intermediate DNN model using the plurality of neural blocks, the processor is configured to:
maintain a truncated parameterized distribution over all of the plurality of neural blocks at each layer that manifests a measure of a relative value of every neural block among the plurality of neural blocks subject to the at least one hardware parameter and the at least one task;
select useful neural elements based on the two-step truncation operation; and
represent the intermediate DNN model using the selected useful neural elements.
12 . The electronic device as claimed in claim 10 , wherein to determine the quality of each neural block in the plurality of neural blocks based on the probability distribution in executing the at least one task using the data inputs, the at least one performance parameter and the at least one hardware parameter, the processor is configured to:
encode a layer depth and features of neural blocks;
create an action space comprising a set of neural block choices for every learnable block;
determine, based on the two-step truncation operation, a usefulness of the set of neural block choices;
add an abstract layer with choices, from the two-step truncation operation, of the set of neural block choices with the at least one hardware parameter and the at least one task;
find an expected latency for the set of neural block choices using a latency predictor metamodel; and
find an expected accuracy after adding the set of neural block choices by sampling paths in the abstract layer.
13 . The electronic device as claimed in claim 10 , wherein to select the at least one optimal neural block from the plurality of neural blocks based on the quality of each neural block, the processor is configured to:
instantiate the intermediate DNN model;
extract constant values for the at least one task and the at least one hardware parameter based on the intermediate DNN model; and
select the at least one optimal neural block from the plurality of neural blocks based on the quality of each neural block.
14 . The electronic device as claimed in claim 10 , wherein to optimize the standard DNN model by modifying the unsupported operations used for the execution of the task with the supported operations to generate the optimized DNN model, the processor is configured to:
search for standard operations at a knowledgebase to replace the unsupported operations, and
perform at least one of:
replacing the unsupported operations with the standard operations, and retraining at least one neural block of the plurality of neural blocks with the standard operations, when the standard operations are available; or
optimizing the unsupported operations, using universal approximator Pade' Approximation Units (PAUs), for the task execution, when the standard operations are unavailable.
15 . An intelligent deployment method for neural networks in a multi-device environment, comprising:
identifying, by an electronic device, a task to be executed in the electronic device;
estimating, by the electronic device, a performance threshold at a time of execution of the identified task;
identifying, by the electronic device, an operation capability of the electronic device, wherein the operation capability is at least one of a processor speed, a number of cores in a processor, a data transmission speed, a storage capacity of a memory, and a write/read speed at the memory;
performing, by the electronic device, a Neural Architecture Search (NAS) of a plurality of neural blocks from a Deep Neural Network (DNN) model, based on the operation capability of the electronic device, and a search space including all possible choices of the plurality of neural blocks;
determining, by the electronic device, a quality of each neural block in the plurality of neural blocks based on (i) a probability distribution in executing the identified task and (ii) a two-step truncation operation that includes truncation based on information value, and truncation based on confidence bounds; and
configuring, by the electronic device, a pre-trained Artificial Intelligence (AI) model to select one or more neural blocks from the plurality of neural blocks based on a result of the NAS and the quality of each neural block to optimize a performance of the identified task in the electronic device,
wherein inputs to the truncation based on information value include neural choices, and a past history of usage of the neural choices and wherein inputs to the truncation based on confidence bounds include neural choices and a policy distribution over the neural choices.
16 . The method as claimed in claim 15 , wherein the quality of each neural block is determined using a probability distribution in the task execution.
17 . The method as claimed in claim 15 , wherein a standard Deep Neural Network (DNN) model is generated using the one or more neural blocks.
18 . The method as claimed in claim 15 , wherein the performance threshold comprises an accuracy threshold, a quality threshold of image, a latency threshold, a memory consumption threshold, a power consumption threshold, and a bandwidth threshold.
19 . The method as claimed in claim 15 , wherein the operation capability of the electronic device comprises, a screen refresh rate, a sampling rate, a camera resolution, a pixel density of a screen, a frame rate, a screen resolution, single/multiple display, an audio format support, and a video format support.