IP Library Granted Patent US 12,619,892
Granted Patent B2
US 12,619,892 · App. 18/060,122 · Granted May 5, 2026

System and method for managing inference model performance through proactive communication system analysis

Inventors: Ofir Ezrielev (Beer Sheva, IL); Jehuda Shemer (Kfar Saba, IL); Tomer Kushnir (Omer, IL)
Assignee: Dell Products L.P.
G06N5/043H04L41/145
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,892
App. No.
18/060,122
Granted
May 5, 2026
Kind
B2
Abstract

Methods and systems for managing execution of inference models hosted by data processing systems are disclosed. To manage execution of inference models hosted by data processing systems, a system may include an inference model manager and any number of data processing systems. The inference model manager may communication system data for the communication system linking the data processing systems. The inference model manager may use the communication system data to determine whether the communication system meets inference generation requirements of the downstream consumer. If the communication system does not meet inference generation requirements of the downstream consumer, the inference model manager may obtain an inference generation plan to return to compliance with the inference generation requirements of the downstream consumer.

Claims (81)

1 . A method of managing execution of an inference model hosted by data processing systems, the method comprising:

obtaining an execution plan for inference generation according to inference generation requirements of a downstream consumer, the execution plan designating a time interval to each of the data processing systems for transmitting operational capability data;

obtaining communication system information for a communication system connecting the data processing systems;

making a determination regarding whether the communication system information meets the inference generation requirements;

in an instance of the determination, in which the communication system information does not meet the inference generation requirements of the downstream consumer:

obtaining an inference generation path for the inference model based on the inference generation requirements of the downstream consumer and the communication system information; and

modifying a deployment of the inference model to the data processing systems based on the inference generation path.

2 . The method of claim 1 , further comprising:

prior to obtaining the communication system information:

obtaining the inference model;

obtaining characteristics of the inference model and characteristics of the data processing systems; and

obtaining portions of the inference model based on the characteristics of the data processing systems and the characteristics of the inference model; and

distributing the portions of the inference model to the data processing systems based on the execution plan,

wherein the execution plan is obtained based on the portions of the inference model, the characteristics of the data processing systems, and the inference generation requirements of the downstream consumer.

3 . The method of claim 2 , wherein the communication system information comprises:

a quantity of available communication system bandwidth between each data processing system of the data processing systems; and

a reliability of transmission between each data processing system of the data processing systems.

4 . The method of claim 3 , wherein the reliability of transmission is based on:

historical data indicating a likelihood of successful transmission of data between each data processing system of the data processing systems; or

a distance between each data processing system of the data processing systems.

5 . The method of claim 4 , wherein the inference generation requirements of the downstream consumer are based on:

an inference generation speed threshold, and

an inference generation reliability threshold.

6 . The method of claim 5 , wherein the inference generation speed threshold indicates a minimum quantity of communication bandwidth between each data processing system of the data processing systems to meet the inference generation requirements of the downstream consumer.

7 . The method of claim 6 , wherein the inference generation reliability threshold indicates a minimum likelihood of successful transmission of data between each data processing system of the data processing systems to meet the inference generation requirements of the downstream consumer.

8 . The method of claim 7 , wherein the inference generation path comprises:

a listing of instances of each of the portions of the inference model usable to generate an inference model result in compliance with the inference generation requirements of the downstream consumer; and

an ordering of the listing of the instances.

9 . The method of claim 8 , wherein modifying the deployment of the inference model comprises:

generating an updated execution plan based on the inference generation path; and

distributing the updated execution plan to the data processing systems to implement the updated execution plan.

10 . The method of claim 9 , wherein the communication system comprises:

multiple point-to-point wireless connections between the data processing systems, each point-to-point wireless connection of the multiple point-to-point wireless connections having distinct characteristics.

11 . A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations for managing execution of an inference model hosted by data processing systems, the operations comprising:

obtaining an execution plan for inference generation according to inference generation requirements of a downstream consumer, the execution plan designating a time interval to each of the data processing systems for transmitting operational capability data;

obtaining communication system information for a communication system connecting the data processing systems;

making a determination regarding whether the communication system information meets the inference generation requirements;

in an instance of the determination, in which the communication system information does not meet the inference generation requirements of the downstream consumer:

obtaining an inference generation path for the inference model based on the inference generation requirements of the downstream consumer and the communication system information; and

modifying a deployment of the inference model to the data processing systems based on the inference generation path.

12 . The non-transitory machine-readable medium of claim 11 , the operations further comprising:

prior to obtaining the communication system information:

obtaining the inference model;

obtaining characteristics of the inference model and characteristics of the data processing systems;

obtaining portions of the inference model based on the characteristics of the data processing systems and the characteristics of the inference model; and

distributing the portions of the inference model to the data processing systems based on the execution plan,

wherein the execution plan is obtained based on the portions of the inference model, the characteristics of the data processing systems, and the inference generation requirements of the downstream consumer.

13 . The non-transitory machine-readable medium of claim 12 , wherein the communication system information comprises:

a quantity of available communication system bandwidth between each data processing system of the data processing systems; and

a reliability of transmission between each data processing system of the data processing systems.

14 . The non-transitory machine-readable medium of claim 13 , wherein the reliability of transmission is based on:

historical data indicating a likelihood of successful transmission of data between each data processing system of the data processing systems; or

a distance between each data processing system of the data processing systems.

15 . The non-transitory machine-readable medium of claim 14 , wherein the inference generation requirements of the downstream consumer are based on:

an inference generation speed threshold, and

an inference generation reliability threshold.

16 . A data processing system, comprising:

a processor; and

a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations for managing execution of an inference model hosted by data processing systems, the operations comprising:

obtaining an execution plan for inference generation according to inference generation requirements of a downstream consumer, the execution plan designating a time interval to each of the data processing systems for transmitting operational capability data;

obtaining communication system information for a communication system connecting the data processing systems;

making a determination regarding whether the communication system information meets the inference generation requirements;

in an instance of the determination, in which the communication system information does not meet the inference generation requirements;

obtaining an inference generation path for the inference model based on the inference generation requirements of the downstream consumer and the communication system information; and

modifying a deployment of the inference model to the data processing systems based on the inference generation path.

17 . The data processing system of claim 16 , the operations further comprising:

prior to obtaining the communication system information:

obtaining the inference model;

obtaining characteristics of the inference model and characteristics of the data processing systems;

obtaining portions of the inference model based on the characteristics of the data processing systems and the characteristics of the inference model; and

distributing the portions of the inference model to the data processing systems based on the execution plan,

wherein the execution plan is obtained based on the portions of the inference model, the characteristics of the data processing systems, and the inference generation requirements of the downstream consumer.

18 . The data processing system of claim 17 , wherein the communication system information comprises:

a quantity of available communication system bandwidth between each data processing system of the data processing systems; and

a reliability of transmission between each data processing system of the data processing systems.

19 . The data processing system of claim 18 , wherein the reliability of transmission is based on:

historical data indicating a likelihood of successful transmission of data between each data processing system of the data processing systems; or

a distance between each data processing system of the data processing systems.

20 . The data processing system of claim 19 , wherein the inference generation requirements of the downstream consumer are based on:

an inference generation speed threshold, and

an inference generation reliability threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2022
From: EZRIELEV, OFIR; SHEMER, JEHUDA; KUSHNIR, TOMER
To: DELL PRODUCTS L.P.
Reel/Frame 061971/0997 →
Continuity (1)
Related Publication 20240177027A1 · May 30, 2024
References Cited (51)
US 7593729B2 · Barak · 2009 [cited by examiner]
US 10142174B2 · Yang et al. · 2018 [cited by applicant]
US 10341354B2 · Murugesan et al. · 2019 [cited by applicant]
US 10853129B1 · Sengupta · 2020 [cited by applicant]
US 11216482B2 · Khillar et al. · 2022 [cited by applicant]
US 11886960B2 · Kaufmann · 2024 [cited by examiner]
US 20080256167A1 · Branson · 2008 [cited by applicant]
US 20160098297A1 · Yuyitung · 2016 [cited by applicant]
US 20180365576A1 · Guttmann · 2018 [cited by applicant]
US 20190050505A1 · Nendorf · 2019 [cited by applicant]
US 20190102700A1 · Babu · 2019 [cited by applicant]
US 20190266504A1 · Apparao · 2019 [cited by applicant]
US 20200027157A1 · Xu · 2020 [cited by applicant]
US 20200104749A1 · Sartorello · 2020 [cited by applicant]
US 20200145448A1 · Vu · 2020 [cited by applicant]
US 20200193313A1 · Ghanta · 2020 [cited by applicant]
US 20200219007A1 · Byers · 2020 [cited by applicant]
US 20210255886A1 · Von Niederhausern · 2021 [cited by applicant]
US 20210319098A1 · Pogorelik · 2021 [cited by applicant]
US 20220036160A1 · Sasagawa · 2022 [cited by applicant]
US 20220156368A1 · Spyridopoulos · 2022 [cited by applicant]
US 20220269835A1 · Yang · 2022 [cited by applicant]
US 20220414503A1 · Park · 2022 [cited by applicant]
US 20230010769A1 · Umezawa · 2023 [cited by applicant]
US 20230012487A1 · Makaya · 2023 [cited by applicant]
US 20230168932A1 · Rafferty · 2023 [cited by applicant]
US 20230168950A1 · Lee · 2023 [cited by applicant]
US 20230252328A1 · Swami · 2023 [cited by applicant]
US 20230267344A1 · Chang · 2023 [cited by applicant]
US 20230273813A1 · Fong · 2023 [cited by applicant]
US 20230342203A1 · Yang · 2023 [cited by applicant]
US 20240020296A1 · Ezrielev · 2024 [cited by applicant]
US 20240020555A1 · Ezrielev · 2024 [cited by applicant]
US 20250030766A1 · Dubey · 2025 [cited by applicant]
US 20250265383A1 · Nendorf · 2025 [cited by applicant]
Liang et al., Model-driven Cluster Resource Management for AI Workloads in Edge Clouds; arXiv:2201.07312v1 [cs.DC] Jan. 18, 2022; Total pp. 23 (Year: 2022). [cited by examiner]
Filho et al., A Systematic Literature Review on Distributed Machine Learning in Edge Computing; Sensors 2022, 22, 2665. https://doi.org/10.3390/s22072665; Published Mar. 23, 2022; Total pp. 36 (Year: 2022). [cited by examiner]
Shi et al., Communication-Efficient Edge AI: Algorithms and Systems; arXiv:2002.09668v1 [cs.IT] Feb. 22, 2020; Total pp. 24 (Year: 2020). [cited by examiner]
Mao et al., A Survey on Mobile Edge Computing: The Communication Perspective; arXiv:1701.01090v4 [cs.IT] Jun. 13, 2017; Total pp. 37 (Year: 2017). [cited by examiner]
“Software performance testing”, Wikipedia, Wikimedia Foundation, https://en.wikipedia.org/wiki/Software_performance_testing (8 Pages). [cited by applicant]
“The Ultimate Guide to Performance Testing and Software Testing: Testing Types, Performance Testing Steps, Best Practices, and More”, Stackify, Apr. 16, 2021, https://stackify.com/ultimate-guide-performance-testing-and-… [cited by applicant]
“High availability software”, Wikipedia, Wikimedia Foundation, https://en.wikipedia.org/wiki/High_availability_software (5 Pages). [cited by applicant]
“Load balancing (computing)”, Wikipedia, Wikimedia Foundation, https://en.wikipedia.org/wiki/Load_balancing_(computing) (17 Pages). [cited by applicant]
Thompson, Julie D. et al., “Multiple Sequence Alignment Using ClustalW and ClustalX”, Current Protocols in Bioinformatics (2003) 2.3.1-2.3.22 (22 Pages). [cited by applicant]
Zhao, Shixiong et al., “HAMS: High Availability for Distributed Machine Learning Service Graphs”, 2020 50th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). IEEE, 2020 (13 Pages). [cited by applicant]
Hu, Chenghao et al., “Distributed Inference with Deep Learning Models across Heterogeneous Edge Devices.” IEEE Infocom 2022—IEEE Conference on Computer Communications. IEEE, 2022 (8 Pages). [cited by applicant]
Eck, B., Fusco, F., Gormally, R., Purcell, M., & Tirupathi, S. (Mar. 24, 2020). “Scalable deployment of AI time-series models for IoT.” arXiv preprint arXiv:2003.12141. (7 pages). [cited by applicant]
Chen, B., Eck, B., Fusco, F., Gormally, R., Purcell, M., Sinn, M, & Tirupathi, S. (Nov. 2018). “Castor: Contextual IoT time series data and model management at scale.” In 2018 IEEE International Conference on Data Minin… [cited by applicant]
Saeed et al., “Model Adaptation and Personalization for Physiological Stress Detection.” Oct. 2018. (Year: 2018) Retrieved from <https://tstojan.github.io/pub/DSAA2018_ModelAdaptandPersonPhysiologicaStressDetection.pdf>… [cited by applicant]
Zhang et al., “Bandwidth-Efficient Multi-Task AI Inference with Dynamic Task Importance for the Internet of Things in 4 Edge Computing.” Aug. 2022. (Year: 2022) (13 pages). [cited by applicant]
Abadi, Martin et al. (Mar. 16, 2016). “TensorFlow: Large-Scale machine learning on heterogeneous distributed systems.” arXiv.org. <https://arxiv.org/abs/1603.04467> (19 pages). [cited by applicant]