IP Library › Granted Patent US 12,273,278
Granted Patent B2
US 12,273,278 · App. 18/483,849 · Granted Apr 8, 2025

Orchestration of workloads involving an AI model

Inventors: Aladin Djuhera (Dachau, DE); Alecio Pedro Delazari Binotto (Munich, DE); Fernando Luiz Koch (Palm Beach Gardens, FL); Rob High (Round Rock, TX)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
H04L47/822H04L41/16H04L47/803
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,273,278
App. No.
18/483,849
Granted
Apr 8, 2025
Kind
B2
Abstract

The present disclosure relates to a method comprising receiving a request to execute a workload using an artificial intelligence model. A current resource utilization status in the distributed system may be determined. The current resource utilization status may be used to define a deployment configuration of the artificial intelligence model, wherein the deployment configuration is defined by: a number and structure of input blocks, a number and structure of output blocks and the intermediate block of the artificial intelligence model, a second computer system to execute the intermediate block, and one or more first computer systems to execute the input and output blocks. The artificial intelligence model may be deployed in accordance with the defined deployment configuration and the workload may be executed.

Claims (71)

1. A method for executing workloads in a distributed system using an artificial intelligence model, comprising:

the distributed system comprising a set of first computer systems configured to connect to at least one second computer system of the distributed system,

the artificial intelligence model configured to receive a specific input, process the specific input and provide a specific output, the artificial intelligence model further configured to be split into a set of one or more input blocks, an intermediate block and a set of one or more output blocks, such that the set of one or more input blocks receive the specific input and provides an intermediate output, the intermediate block receives as input the intermediate output and provides a second intermediate output, and the set of one or more output blocks receive as input the second intermediate output and provides said specific output;

the method further comprising an orchestration method comprising:

receiving a request to execute a workload using the artificial intelligence model, the workload comprising receiving the specific input;

determining a current resource utilization status in the distributed system;

using the current resource utilization status to define a deployment configuration of the artificial intelligence model, wherein the deployment configuration is defined by:

the artificial intelligence model comprising a number and structure of input blocks, a number and structure of output blocks and the intermediate block of,

a second computer system to execute the intermediate block, and

one or more first computer systems to execute the number and structure of input blocks and the number and structure of output blocks;

wherein defining the deployment configuration further comprises identifying whether an existing deployment configuration is present based on a previous deployment of the artificial intelligence model;

based on the presence of the existing deployment configuration, determining whether the existing deployment configuration is valid;

based on determining that the existing deployment configuration is valid, using the existing deployment configuration as the defined deployment configuration;

based on determining that the existing deployment configuration is not valid, changing the existing deployment configuration and using the changed existing deployment configuration as the defined deployment configuration, wherein the changing further comprises changing at least one of the one or more first computer systems, the second computer system, the number and structure of the input blocks, the number and structure of the output blocks, and the intermediate block associated with the existing deployment configuration; and

deploying the artificial intelligence model in accordance with the defined deployment configuration and executing the workload.

2. The method of claim 1 , wherein before executing the orchestration method, the method comprises deploying the artificial intelligence model in accordance with an initial deployment configuration, wherein defining the deployment configuration comprises updating the initial deployment configuration, the updating comprising at least one of:

re-splitting the artificial intelligence model;

adding the one or more first computer systems to execute the number and structure of input blocks and the number and structure of output blocks;

removing the one or more of the first computer systems of the initial deployment configuration to execute the number and structure of input blocks and the number and structure of output blocks; or

selecting the at least one second computer system to execute the intermediate block.

3. The method of claim 1 , wherein using the current resource utilization status to define the deployment configuration comprises:

performing a capacity profiling of the one or more first computer systems and the second computer system to determine whether each of the one or more first computer systems and the at least one second computer system can execute one or more blocks of the artificial intelligence model;

and defining the deployment configuration based on the capacity profiling.

4. The method of claim 1 , wherein the current resource utilization status is defined by at least one of:

utilization level of network resources of the distributed system;

utilization level of resources of the one or more first computer systems; or

utilization level of resources of the at least one second computer systems.

5. The method of claim 1 , wherein the deployment configuration is defined further using a rule, the rule requiring at least one of:

a maximum number of the one or more first computer systems to be used for deploying the number and structure of input blocks and the number and structure of output blocks;

a maximum number of blocks of the artificial intelligence model;

a fulfilment of a specific Quality of Service (QOS) requirement; or

a desired energy-efficiency (EE).

6. The method of claim 1 , wherein the distributed system includes a wireless communication system, wherein the one or more first computer systems are multi-access edge computing (MEC) nodes, and wherein the at least one second computer system is a cloud system.

7. The method of claim 1 , further comprising:

executing the deployment of the artificial intelligence model by a broadcaster of the distributed system;

sending by the broadcaster an information to an orchestrator of the distributed system, such that the orchestrator triggers the execution of the workload upon receiving the information.

8. The method of claim 1 , wherein the artificial intelligence model is a foundation model.

9. The method of claim 1 , wherein the artificial intelligence model is a deep neural network or a transformer, wherein the number and structure of input blocks represents first network layers, the intermediate block represents middle network layers; and the number and structure of output blocks represents last network layers.

10. The method of claim 1 , wherein the one or more first computer systems has an amount of processing resources which is smaller than other processing resources of the at least one second computer system.

11. The method of claim 1 , wherein the workload involves at least one of: data analytics,

sensor measurement fusion from different sources, image analysis, or processing data streams destined to a cloud.

12. The method of claim 1 , wherein the executing of the artificial intelligence model further comprises:

for each two consecutive blocks of the artificial intelligence model which are deployed on different computer systems, encoding using an encoding protocol the output of a first block of the two blocks, and sending the encoded output to a first computer system or to the at least one second computer system in order to be used as input for a second block of the two blocks.

13. The method of claim 12 , wherein the encoding protocol comprises at least one of compression or encryption.

14. A computer program product comprising a computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code configured to execute a method comprising:

receiving a request to execute a workload using an artificial intelligence model, the workload comprising receiving a specific input;

determining a current resource utilization status in a distributed system;

using the current resource utilization status to define a deployment configuration of the artificial intelligence model, wherein the deployment configuration is defined by:

a number and structure of input blocks, a number and structure of output blocks, and an intermediate block of the artificial intelligence model,

a second computer system to execute the intermediate block, and

one or more first computer systems to execute the number and structure of input blocks and the number and structure of output blocks;

wherein defining the deployment configuration further comprises identifying whether an existing deployment configuration is present based on a previous deployment of the artificial intelligence model;

based on the presence of the existing deployment configuration, determining whether the existing deployment configuration is valid;

based on determining that the existing deployment configuration is valid, using the existing deployment configuration as the defined deployment configuration;

based on determining that the existing deployment configuration is not valid, changing the existing deployment configuration and using the changed existing deployment configuration as the defined deployment configuration, wherein the changing further comprises changing at least one of the one or more first computer systems, the second computer system, the number and structure of the input blocks, the number and structure of the output blocks, and the intermediate block associated with the existing deployment configuration; and

deploying the artificial intelligence model in accordance with the defined deployment configuration and executing the workload.

15. A computer system for executing workloads in a distributed system using an artificial intelligence model, comprising:

the distributed system comprising a set of first computer systems which are configured to connect to at least one second computer system of the distributed system,

the artificial intelligence model configured to receive a specific input, process the specific input and provide a specific output, the artificial intelligence model being configured to be split into a set of one or more input blocks, an intermediate block and a set of one or more output blocks, such that the set of one or more input blocks receive the specific input and provides an intermediate output, the intermediate block receives as input the intermediate output and provides a second intermediate output, and the set of one or more output blocks receive as input the second intermediate output and provides said specific output;

the computer system comprising computer program instructions for:

receiving a request to execute a workload using the artificial intelligence model, the workload comprising receiving the specific input;

determining a current resource utilization status in the distributed system;

using the current resource utilization status to define a deployment configuration of the artificial intelligence model, wherein the deployment configuration is defined by:

a number and structure of input blocks, a number and structure of output blocks, and the intermediate block of the artificial intelligence model,

a second computer system to execute the intermediate block, and

one or more first computer systems to execute the number and structure of input blocks and the number and structure of output blocks;

wherein defining the deployment configuration further comprises identifying whether an existing deployment configuration is present based on a previous deployment of the artificial intelligence model;

based on the presence of the existing deployment configuration, determining whether the existing deployment configuration is valid;

based on determining that the existing deployment configuration is valid, using the existing deployment configuration as the defined deployment configuration;

based on determining that the existing deployment configuration is not valid, changing the existing deployment configuration and using the changed existing deployment configuration as the defined deployment configuration, wherein the changing further comprises changing at least one of the one or more first computer systems, the second computer system, the number and structure of the input blocks, the number and structure of the output blocks, and the intermediate block associated with the existing deployment configuration; and

deploying the artificial intelligence model in accordance with the defined deployment configuration and executing the workload.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 10, 2023
From: DJUHERA, ALADIN; BINOTTO, ALECIO PEDRO DELAZARI; KOCH, FERNANDO LUIZ; HIGH, ROB
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 065170/0316 →
Continuity (1)
Related Publication 20250071069A1 · Feb 27, 2025
References Cited (25)
US 10856360B1 · Sun · 2020 [cited by applicant]
US 11423254B2 · Prakash · 2022 [cited by examiner]
US 20190318240A1 · Kulkarni · 2019 [cited by examiner]
US 20200311559A1 · Chattopadhyay · 2020 [cited by applicant]
US 20210295158A1 · Jakubiuk · 2021 [cited by applicant]
US 20210382754A1 · Nimmagadda · 2021 [cited by examiner]
US 20220012607A1 · Liu · 2022 [cited by examiner]
US 20220108177A1 · Samek · 2022 [cited by applicant]
US 20220121455A1 · Hoban · 2022 [cited by examiner]
US 20220124009A1 · Metsch · 2022 [cited by examiner]
US 20230177345A1 · Vilcu · 2023 [cited by examiner]
WO 2021004478A1 · 2021 [cited by applicant]
WO 2022211981W · 2022 [cited by applicant]
UK Search Report, Application No. GB231278 I .4, Mailing Date: Feb. 21, 2024, 3 pages. [cited by applicant]
Djuhera et al., “Orchestration of Workloads Involving an AI Model”, GB application No. GB2312781.4, Filing Date: Aug. 22, 2023, 54 pages. [cited by applicant]
Disclosed Anonymously, “Methods and System for Secure and Optimal DL Split Model Deployment at Heterogeneous Edge Environment”, An IP.com Prior Art Database Technical Disclosure, IP.com No. IPCOM000267162D, IP.com Elect… [cited by applicant]
Hou et al., “DistrEdge: Speeding up Convolutional Neural Network Inference on Distributed Edge Devices”, arXiv:2202.01699v2, Feb. 8, 2022, 12 pages. [cited by applicant]
Karjee et al., “Split computing: DNN inference partition with load balancing in IoT-edge platform for beyond 5G”, ScienceDirect, Measurement: Sensors, vol. 23, Oct. 2022, 65 pages. [cited by applicant]
Kim et al., “A Bargaining Game for Personalized, Energy Efficient Split Learning over Wireless Networks”, arXiv:2212.06107v1, Dec. 12, 2022, 6 pages. [cited by applicant]
Mohammed et al., “Distributed Inference Acceleration with Adaptive DNN Partitioning and Offloading”, IEEE Infocom 2020—IEEE Conference on Computer Communications, Published Date: Jul. 6, 2020, 10 pages. [cited by applicant]
Prasad et al., “A Joint Model Provisioning and Request Dispatch Solution for Low-Latency Inference Services on Edge”, MDPI, Sensors 2021, vol. 21, pp. 1-17. [cited by applicant]
Tuli et al., “SplitPlace: AI Augmented Splitting and Placement of Large-Scale Neural Networks in Mobile Edge Environments”, IEEE Transactions on Mobile Computing, A preliminary version of this work was presented at the … [cited by applicant]
Zeng et al., “CoEdge: Cooperative DNN Inference with Adaptive Workload Partitioning over Heterogeneous Edge Devices”, arXiv:2012.03257v1, Dec. 6, 2020, 14 pages. [cited by applicant]
International Searching Authority, “Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, or Declaration,” Patent Cooperation Treaty, Nov. 28, 2… [cited by applicant]
Li et al. “Partitioning Multi-Layer Edge Network for Neural Network Collaborative Computing”, EURASIP Journal on Wireless Communications and Networking, Aug. 19, 2023, pp. 1-51, URL:https://jwcn-eurasipjournals.springer… [cited by applicant]
Cited By (1)
US 12,580,624