IP Library › Granted Patent US 12,463,875
Granted Patent B2
US 12,463,875 · App. 17/554,964 · Granted Nov 4, 2025

Apparatus, articles of manufacture, and methods to partition neural networks for execution at distributed edge nodes

Inventors: Karthik Kumar (Chandler, AZ); Francesc Guim Bernat (Barcelona, ES)
Assignee: Intel Corporation
H04L41/5054G06N3/04H04L41/5019H04L67/34H04L43/08H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,463,875
App. No.
17/554,964
Granted
Nov 4, 2025
Kind
B2
Abstract

Methods, apparatus, systems, and articles of manufacture are disclosed to partition neural network models for executing at distributed Edge nodes. An example apparatus includes processor circuitry to perform at least one of first, second, or third operations to instantiate power consumption estimation circuitry to estimate a computation energy consumption for executing the neural network model on a first edge node, network bandwidth determination circuitry to determine a first transmission time for sending an intermediate result from the first edge node to a second or third edge node, power consumption estimation circuitry to estimate a transmission energy consumption for sending the intermediate result to the second or the third edge node, and neural network partitioning circuitry to partition the neural network model into a first portion to be executed at the first edge node and a second portion to be executed at the second or third edge node.

Claims (85)

1 . An apparatus to partition a neural network model comprising:

interface circuitry to communicate with an edge device; and

processor circuitry including one or more of:

at least one of a central processing unit, a graphic processing unit, or a digital signal processor, the at least one of the central processing unit, the graphic processing unit, or the digital signal processor having control circuitry to control data movement within the processor circuitry, arithmetic and logic circuitry to perform one or more first operations corresponding to instructions, and one or more registers to store a first result of the one or more first operations, the instructions in the apparatus;

a Field Programmable Gate Array (FPGA), the FPGA including first logic gate circuitry, a plurality of configurable interconnections, and storage circuitry, the first logic gate circuitry and interconnections to perform one or more second operations, the storage circuitry to store a second result of the one or more second operations; or

Application Specific Integrate Circuitry (ASIC) including second logic gate circuitry to perform one or more third operations;

the processor circuitry to perform at least one of the first operations, the second operations, or the third operations to instantiate:

power consumption estimation circuitry to estimate a computation energy consumption for executing the neural network model on a first edge node, the neural network model corresponding to a platform service with a service level agreement timeframe;

network bandwidth determination circuitry to determine a transmission time for sending an intermediate result from the first edge node to a second edge node or a third edge node;

the power consumption estimation circuitry to estimate a transmission energy consumption for sending the intermediate result of the neural network model to the second edge node or the third edge node; and

neural network partitioning circuitry to partition the neural network model into a first portion to be executed at the first edge node and a second portion to be executed at the second edge node or the third edge node based on at least one of the service level agreement timeframe for the platform service, the computation energy consumption, the transmission energy consumption, or the transmission time.

2 . The apparatus of claim 1 , wherein the processor circuitry is to perform at least one of the first operations, the second operations, or the third operations to instantiate the network bandwidth determination circuitry to determine the transmission time based on available network bandwidth and a payload size, the payload size including the intermediate result.

3 . The apparatus of claim 1 , wherein the transmission time is a first transmission time, and the processor circuitry is to perform at least one of the first operations, the second operations, or the third operations to instantiate the network bandwidth determination circuitry to:

determine a second transmission time for sending a final result from the second edge node or the third edge node to the first edge node; and

calculate a total transmission time based on a sum of the first transmission time and the second transmission time, the partitioning of the neural network model into the first portion and the second portion based on the total transmission time satisfying the service level agreement timeframe.

4 . The apparatus of claim 1 , wherein the processor circuitry is to perform at least one of the first operations, the second operations, or the third operations to instantiate the network bandwidth determination circuitry to:

receive temperature data from a temperature sensor;

receive wind data from a wind sensor; and

receive humidity data from a humidity sensor, the estimation of at least one of the computation energy consumption or the transmission energy consumption based on at least one of the temperature data, the wind data, or the humidity data.

5 . The apparatus of claim 1 , wherein the processor circuitry is to perform at least one of the first operations, the second operations, or the third operations to instantiate the network bandwidth determination circuitry to:

estimate at least the computation energy consumption or the transmission energy consumption based on at least one of an artificial intelligence algorithm, a set of instructions, or a lookup table; and

calculate a total power consumption based on a sum of the computation energy consumption and the transmission energy consumption, the partitioning of the neural network model into the first portion and the second portion based on the total power consumption satisfying a power consumption threshold.

6 . The apparatus of claim 1 , wherein the processor circuitry is to:

assign a first identifier to input data that the neural network model is to process;

assign a second identifier to the intermediate result, the second identifier to identify (i) the neural network model that is to be executed at the second edge node or the third edge node (ii) an initial layer of the second portion that the second edge node or the third edge node is to execute;

send the intermediate result to the second edge node or the third edge node with the second identifier; and

determine a final result based on the intermediate result and the second identifier, the second identifier to identify the second portion of the neural network model that the second edge node or the third edge node is to execute.

7 . At least one non-transitory computer readable medium comprising instructions that, when executed, cause processor circuitry to at least:

infer a computation energy consumption for executing a neural network model on a first edge node, the neural network model corresponding to a platform service with a service level agreement timeframe;

compute a transmission time for transmitting an intermediate output from the first edge node to a second edge node or a third edge node;

infer a transmission energy consumption for transmitting the intermediate output of the neural network model to the second edge node or the third edge node; and

segment the neural network model into a first portion to be executed at the first edge node and a second portion to be executed at the second edge node or the third edge node based on at least one of the service level agreement timeframe for the platform service, the computation energy consumption, the transmission energy consumption, or the transmission time.

8 . The at least one non-transitory computer readable medium of claim 7 , wherein the instructions, when executed, cause the processor circuitry to compute the transmission time based on available network bandwidth and a payload size, the payload size including the intermediate output.

9 . The at least one non-transitory computer readable medium of claim 7 , wherein the transmission time is a first transmission time, and the instructions, when executed, cause the processor circuitry to:

compute a second transmission time for sending a final result from the second edge node or the third edge node to the first edge node; and

determine a total transmission time based on a sum of the first transmission time and the second transmission time, the segmenting of the neural network model into the first portion and the second portion based on the total transmission time satisfying the service level agreement timeframe.

10 . The at least one non-transitory computer readable medium of claim 7 , wherein the instructions, when executed, cause the processor circuitry to

receive temperature data from a temperature sensor;

receive wind data from a wind sensor; and

receive humidity data from a humidity sensor, the inference of at least one of the computation energy consumption or the transmission energy consumption based on at least one of the temperature data, the wind data, or the humidity data.

11 . The at least one non-transitory computer readable medium of claim 7 , wherein the instructions, when executed, cause the processor circuitry to:

infer at least the computation energy consumption or the transmission energy consumption based on at least one of an artificial intelligence algorithm, a set of instructions, or a lookup table; and

compute a total power consumption based on a sum of the computation energy consumption and the transmission energy consumption, the segmentation of the neural network model into the first portion and the second portion based on the total power consumption satisfying a power consumption threshold.

12 . The at least one non-transitory computer readable medium of claim 7 , wherein the instructions, when executed, cause the processor circuitry to:

assign a first identifier to input data that the neural network model is to process;

assign a second identifier to the intermediate output, the second identifier to identify (i) the neural network model that is to be executed at the second edge node or the third edge node (ii) an initial layer of the second portion that the second edge node or the third edge node is to execute;

send the intermediate output to the second edge node or the third edge node with the second identifier; and

determine a final result based on the intermediate output and the second identifier, the second identifier to identify the second portion of the neural network model that the second edge node or the third edge node is to execute.

13 . An apparatus to partition a neural network model, the apparatus comprising:

means for determining a computation energy consumption for executing the neural network model on a first edge node, the neural network model corresponding to a platform service with a service level agreement timeframe, including means for determining a transmission energy consumption for sending an intermediate computation of the neural network model to a second edge node or a third edge node;

means for calculating a transmission time for sending the intermediate computation from the first edge node to the second edge node or the third edge node; and

means for dividing the neural network model into a first portion to be executed at the first edge node and a second portion to be executed at the second edge node or the third edge node based on at least one of the service level agreement timeframe for the platform service, the computation energy consumption, the transmission energy consumption, and the transmission time.

14 . The apparatus of claim 13 , wherein the means for calculating is to calculate the transmission time based on available network bandwidth and a payload size, the payload size including the intermediate computation.

15 . The apparatus of claim 13 , wherein the transmission time is a first transmission time, and the means for calculating is to:

calculate a second transmission time for sending a final result from the second edge node or the third edge node to the first edge node; and

compute a total transmission time based on a sum of the first transmission time and the second transmission time the dividing of the neural network model into the first portion and the second portion based on the total transmission time satisfying the service level agreement timeframe.

16 . The apparatus of claim 13 , wherein the means for determining is to determine at least one of the computation energy consumption or the transmission energy consumption based on received ambient data including at least one of temperature data, wind data, or humidity data.

17 . The apparatus of claim 13 , wherein the means for determining is to:

determine at least the computation energy consumption or the transmission energy consumption based on at least one of an artificial intelligence algorithm, a set of instructions, or a lookup table; and

determine a total power consumption based on a sum of the computation energy consumption and the transmission energy consumption, the dividing of the neural network model into the first portion and the second portion based on the total power consumption satisfying a power consumption threshold.

18 . The apparatus of claim 13 , further including:

means for assigning a first identifier to input data that the neural network model is to process, the means for assigning is to assign a second identifier to the intermediate computation, the second identifier to identify (i) the neural network model that is to be executed at the second edge node or the third edge node (ii) an initial layer of the second portion that the second edge node or the third edge node is to execute;

means for sending the intermediate computation to the second edge node or the third edge node with the second identifier; and

means for determining a final computation based on the intermediate computation and the second identifier, the second identifier to identify the second portion of the neural network model that the second edge node or the third edge node is to execute.

19 . A method comprising:

estimating, by executing an instruction with processor circuitry, a computation energy consumption for executing a neural network model on a first edge node, the neural network model corresponding to a platform service with a service level agreement timeframe;

determining, by executing an instruction with the processor circuitry, a transmission time for sending an intermediate result from the first edge node to a second edge node or a third edge node;

estimating by executing an instruction with the processor circuitry, a transmission energy consumption for sending the intermediate result of the neural network model to the second edge node or the third edge node; and

partitioning, by executing an instruction with the processor circuitry, the neural network model into a first portion to be executed at the first edge node and a second portion to be executed at the second edge node or the third edge node based on at least one of the service level agreement timeframe for the platform service, the computation energy consumption, the transmission energy consumption, or the transmission time.

20 . The method of claim 19 , wherein determining the transmission time is based on available network bandwidth and a payload size, the payload size including the intermediate result.

21 . The method of claim 19 , wherein the transmission time is a first transmission time, and further including:

determining a second transmission time for sending a final result from the second edge node or the third edge node to the first edge node; and

calculating a total transmission time based on a sum of the first transmission time and the second transmission time, the partitioning of the neural network model into the first portion and the second portion based on the total transmission time satisfying the service level agreement timeframe.

22 . The method of claim 19 , wherein estimating the computation energy consumption includes:

receiving temperature data from a temperature sensor;

receiving wind data from a wind sensor; and

receiving humidity data from a humidity sensor, the estimation of at least one of the computation energy consumption or the transmission energy consumption based on at least one of the temperature data, the wind data, or the humidity data.

23 . The method of claim 19 , further including:

estimating at least the computation energy consumption or the transmission energy consumption based on at least one of an artificial intelligence algorithm, a set of instructions, or a lookup table; and

calculating a total power consumption based on a sum of the computation energy consumption and the transmission energy consumption, the partitioning of the neural network model into the first portion and the second portion based on the total power consumption satisfying a power consumption threshold.

24 . The method of claim 19 , further including:

assigning a first identifier to input data that the neural network model is to process;

assigning a second identifier to the intermediate result, the second identifier to identify (i) the neural network model that is to be executed at the second edge node or the third edge node (ii) an initial layer of the second portion that the second edge node or the third edge node is to execute;

sending the intermediate result to the second edge node or the third edge node with the second identifier; and

determining a final result based on the intermediate result and the second identifier, the second identifier to identify the second portion of the neural network model that the second edge node or the third edge node is to execute.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2022
From: KUMAR, KARTHIK; GUIM BERNAT, FRANCESC
To: INTEL CORPORATION
Reel/Frame 059254/0007 →
Continuity (1)
Related Publication 20220109742A1 · Apr 7, 2022
References Cited (19)
US 9479449B2 · Breternitz · 2016 [cited by examiner]
US 20200021502A1 · Bernat · 2020 [cited by examiner]
US 20210004265A1 · Guim Bernat et al. · 2021 [cited by applicant]
US 20210360082A1 · Pinel et al. · 2021 [cited by applicant]
European Patent Office, “Extended European Search Report,” issued in connection with European Patent Application No. 22202393.9, dated Apr. 12, 2023, 11 pages. [cited by applicant]
Wang et al., “Design and implementation of an analytical framework for interference aware job scheduling on Apache Spark platform,” Cluster Computing, Springer Science+Business Media, LLC, published online Dec. 23, 2017… [cited by applicant]
Witt et al., “Predictive Performance Modeling for Distributed Computing using Black-Box Monitoring and Machine Learning,” arXiv:1805.11877v1 [cs.DC], May 30, 2018, 19 pages. [cited by applicant]
Rivas et al., “Large-Scale Video Analytics through Object-Level Consolidation,” arXiv, Nov. 30, 2021, 17 pages. [cited by applicant]
Bouwmans et al., “Scene Background Initizialization: a Taxonomy,” Accepted Manuscript, Pattern Recognition Letters, Dec. 29, 2016, 12 pages. [cited by applicant]
Cai et al., “Cascade R-CNN: Delving into High Quality Object Detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 9 pages. [cited by applicant]
Canel et al., “Scaling Video Analytics on Constrained Edge Nodes,” arXiv, May 24, 2019, 13 pages. [cited by applicant]
Eggert et al., “A Closer Look: Small Object Detection in Faster R-CNN,” in Proceedings of the IEEE International Conference on Multimedia and Expo (ICME) 2017, Jul. 10, 2017, 6 pages. [cited by applicant]
Everingham et al., “The Pascal Visual Object Classes (VOC) Challenge,” International Journal of Computer Vision, vol. 88, Sep. 9, 2009, 36 pages. [cited by applicant]
Hsieh et al., “Focus: Querying Large Video Datasets with Low Latency and Low Cost,” in Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation, Oct. 8, 2018, 19 pages. [cited by applicant]
Kang et al., “NoScope: Optimizing Neural Network Queries over Video at Scale,” arXiv, Aug. 8, 2017, 12 pages. [cited by applicant]
Lin et al., “Feature Pyramid Networks for Object Detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, 9 pages. [cited by applicant]
Oh et al., “A Large-scale Benchmark Dataset for Event Recognition in Surveillance Video,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2011, 8 pages. [cited by applicant]
Tensorflow, “TensorFlow 2 Detection Model Zoo,” retrieved from https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/tf2_ detection_zoo.md on Jun. 3, 2025, 3 pages. [cited by applicant]
Zeng et al., “Combining Background Subtraction Algorithms with Convolutional Neural Network,” arXiv, Jul. 9, 2018, 7 pages. [cited by applicant]