IP Library › Granted Patent US 12,748,628
Granted Patent B2
US 12,748,628 · App. 18/449,811 · Granted Sep 29, 2026

Distributed execution of an artificial intelligence model

Inventors: Aladin Djuhera (Dachau, DE); Alecio Pedro Delazari Binotto (Munich, DE); Fernando Luiz Koch (Greenwich, CT); Nathalie Baracaldo Angel (San Jose, CA)
Assignee: International Business Machines Corporation
G06F9/5033
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,628
App. No.
18/449,811
Granted
Sep 29, 2026
Kind
B2
Abstract

The present disclosure relates to a method for executing an artificial intelligence model, including receiving an input for execution of the AI model. An input block of the AI model can be executed by a first computer system using the received input, producing a first output. The first output can be encoded. The encoded first output can be sent to a second computer system. The second computer system can decode the encoded first output. The second computer system can execute an intermediate block of the AI model using the first output, producing a second output. The second output can be encoded. The encoded second output can be sent to the first computer system. The first computer system can decode the encoded second output. The first computer system can execute an output block of the ML model using as input the second output, producing a result output.

Claims (44)

1 . A computer-implemented method for executing an artificial intelligence model, comprising:

receiving an input for execution of the artificial intelligence model, wherein the artificial intelligence model is split into an input block, an intermediate block, and an output block, such that the input block receives specific input and provides an intermediate output, the intermediate block receives as input the intermediate output and provides another intermediate output, and the output block receives as input the another intermediate output and provides specific output;

executing the input block by a first computer system using the input, producing a first intermediate output;

encoding by the first computer system the first intermediate output using a first encoding protocol to produce an encoded first intermediate output;

sending the encoded first intermediate output to a second computer system to allow the second computer system to:

decode the encoded first intermediate output using the first encoding protocol, and execute the intermediate block using as input the first intermediate output to produce a second intermediate output,

encode the second intermediate output using a second encoding protocol to produce an encoded second intermediate output, and send the encoded second intermediate output to the first computer system; and

in response to receiving the encoded second intermediate output, decoding by the first computer system the encoded second intermediate output using the second encoding protocol, and executing the output block at the first computer system using as input the second intermediate output, producing a result output of a task performed by the artificial intelligence model as a whole by way of the input block, intermediate block, and output block.

2 . The computer-implemented method of claim 1 , further comprising: before executing the output block, deleting the input block from the first computer system.

3 . The computer-implemented method of claim 1 , further comprising: after executing the input block, deploying the output block into the first computer system.

4 . The computer-implemented method of claim 1 , wherein the splitting of the artificial intelligence model is performed based on available resources in the first computer system.

5 . The computer-implemented method of claim 4 , wherein the splitting of the artificial intelligence model is dynamically performed or performed using one of predefined splitting options which are associated with a respective amount of resources.

6 . The computer-implemented method of claim 1 , wherein the execution of the artificial intelligence model comprises execution of a succession of processing steps, wherein splitting the artificial intelligence model is performed such that the input block is configured to perform a first number of successive processing steps, the intermediate block is configured to perform a second number of successive processing steps that follow the first number of successive processing steps of the input block, and the output block is configured to perform a third number of last successive processing steps, wherein a sum of the first number of successive processing steps, the second number of successive processing steps, and the third number of last successive processing steps is a total number of processing steps in the artificial intelligence model.

7 . The computer-implemented method of claim 6 , wherein the first number of successive processing steps is smaller than the second number of successive processing steps by a first delta value, wherein the third number of last successive processing steps is smaller than the second number of successive processing steps by a second delta value.

8 . The computer-implemented method of claim 7 , further comprising determining the first and second delta values based on available resources in the first computer system.

9 . The computer-implemented method of claim 1 , wherein the first encoding protocol is the second encoding protocol.

10 . The computer-implemented method of claim 1 , wherein the first encoding protocol is different from the second encoding protocol.

11 . The computer-implemented method of claim 1 , wherein the first encoding protocol is selected from a group consisting of compression and encryption, and wherein the second encoding protocol is selected from a group consisting of compression and encryption.

12 . The computer-implemented method of claim 1 , wherein the execution of the artificial intelligence model is an inference of the artificial intelligence model which is already trained.

13 . The computer-implemented method of claim 1 , further comprising, for each further received input, training the artificial intelligence model, wherein the first computer system is further configured to compute in each iteration a loss function and to send a result to the second computer system, wherein the result is used by the first and second computer systems to update learnable parameters of the artificial intelligence model, wherein an iteration is performed until the loss function fulfils a convergence criterion.

14 . The computer-implemented method of claim 1 , wherein the first computer system has an amount of processing resources which is smaller than the processing resources of the second computer system.

15 . The computer-implemented method of claim 1 , wherein the first computer system is selected from a group consisting of an edge device and an internet of things (IoT) device.

16 . The computer-implemented method of claim 1 , wherein the second computer system is provided as a service in a cloud environment.

17 . The computer-implemented method of claim 1 , wherein the artificial intelligence model is a foundation model, and wherein the artificial intelligence model is a deep neural network where the input block represents first network layers, the intermediate block represents middle network layers, and the output block represents last network layers.

18 . The computer-implemented method of claim 1 , further comprising splitting the artificial intelligence model by a management server, and deploying by the management server the input block, the output block, and the intermediate block in the first and second computer systems.

19 . A computer program product comprising:

a non-transitory computer-readable storage media having computer-readable program code embodied therewith, the computer-readable program code configured to cause one or more processors to:

receive an input for execution of an artificial intelligence model, wherein the artificial intelligence model is split into an input block, an intermediate block, and an output block, such that the input block receives specific input and provides an intermediate output, the intermediate block receives as input the intermediate output and provides another intermediate output, and the output block receives as input the another intermediate output and provides specific output;

execute the input block by a first computer system using the input, producing a first intermediate output;

encode by the first computer system the first intermediate output using a first encoding protocol to produce an encoded first intermediate output;

send the encoded first intermediate output to a second computer system to allow the second computer system to:

decode the encoded first intermediate output using the first encoding protocol, and execute the intermediate block using as input the first intermediate output to produce a second intermediate output,

encode the second intermediate output using a second encoding protocol to produce an encoded second intermediate output, and send the encoded second intermediate output to the first computer system; and

in response to receiving the encoded second intermediate output, decode by the first computer system the encoded second intermediate output using the second encoding protocol, and execute the output block at the first computer system using as input the second intermediate output, producing a result output of a task performed by the artificial intelligence model as a whole by way of the input block, intermediate block, and output block.

20 . A system for executing an artificial intelligence model, comprising: one or more computer readable storage media storing program instructions and one or more processors which, in response to executing the program instructions, are configured to:

receive an input for execution of the artificial intelligence model, wherein the artificial intelligence model is split into an input block, an intermediate block, and an output block, such that the input block receives a specific input and provides an intermediate output, the intermediate block receives as input the intermediate output and provides another intermediate output, and the output block receives as input the another intermediate output and provides a specific output;

execute the input block using the input received for execution of the artificial intelligence model, producing a first intermediate output;

encode the first intermediate output using a first encoding protocol to produce an encoded first intermediate output;

send the encoded first intermediate output to a second computer system to allow the second computer system to perform at least:

decoding the encoded first intermediate output using the first encoding protocol;

executing the intermediate block as input the first intermediate output, resulting in a second intermediate output;

encoding the second intermediate output using a second encoding protocol to produce an encoded second intermediate output;

sending the encoded second intermediate output to the computer system; and

in response to receiving the encoded second intermediate output, decode the encoded second intermediate output using the second encoding protocol, and execute the output block using as input the second intermediate output, producing a result output of a task performed by the artificial intelligence model as a whole by way of the input block, intermediate block, and output block.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2023
From: DJUHERA, ALADIN; BINOTTO, ALECIO PEDRO DELAZARI; KOCH, FERNANDO LUIZ; BARACALDO ANGEL, NATHALIE
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 064589/0396 →
Continuity (2)
Provisional Application 63472826 · Jun 14, 2023
Related Publication 20240419498A1 · Dec 19, 2024
References Cited (20)
US 11551121B2 · Horesh · 2023 [cited by applicant]
US 20190258937A1 · Alemi · 2019 [cited by applicant]
US 20200036512A1 · Vaikuntanathan · 2020 [cited by examiner]
US 20200382273A1 · Vald · 2020 [cited by examiner]
US 20220383188A1 · Ananthanarayanan · 2022 [cited by examiner]
US 20240054404A1 · Umezawa · 2024 [cited by examiner]
US 20240154942A1 · Gharibi · 2024 [cited by examiner]
US 20250156687A1 · Liu · 2025 [cited by examiner]
CN 111600707A · 2020 [cited by applicant]
U.S. Appl. No. 63/305,845, filed Feb. 2, 2022, Liu; Miaomiao. [cited by examiner]
Baccour et al., “RL-DistPrivacy: Privacy-Aware Distributed Deep Inference for low latency IoT systems,” arXiv:2208.13032v1 [cs.LG] Aug. 27, 2022, 20 pages. [cited by applicant]
Chamikara et al., “Privacy Preserving Distributed Machine Learning with Federated Learning,” arXiv:2004.12108v2 [cs.DB] Feb. 26, 2021, 41 pages. [cited by applicant]
Chaopeng et al., “A privacy protection approach in edge-computing based on maximized dnn partition strategy with energy saving,” Journal of Cloud Computing: Advances, Systems and Applications, 2023, 16 pages. [cited by applicant]
Chinchali, et al., “Neural Networks Meet Physical Networks: Distributed Inference Between Edge Devices and the Cloud,” HotNets '18: proceedings of the 17th ACM Workshop on Hot Topics in Networks, Nov. 2018, pp. 50-56. [cited by applicant]
Gunarathne et al., “Poster : Distributing Deep Learning Inference on Edge Devices,” Association for Computing Machinery, 2020, pp. 556-557. [cited by applicant]
Kailkhura et al., “Distributed Inference in the Presence of Eavesdroppers: A Survey,” arXiv: 1502.05448v1 [cs.CR], Feb. 19, 2015, 7 pages. [cited by applicant]
Khan et al. “Split Ways: Privacy-Preserving Training of Encrypted Data Using Split Learning,” arXiv:2301.08778v1 [cs.CR] Jan. 20, 2023, 9 pages. [cited by applicant]
Phan et al., “Differential Privacy Preservation for Deep Auto-Encoders: An Application of Human Behavior Prediction,” Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, 2016, pp. 1309-1316. [cited by applicant]
Tuor et al., “Understanding information leakage of distributed inference with deep neural networks: overview of information theoretic approach and initial results,” Proc. SPIE 10635, Ground/Air Multisensor Interoperabil… [cited by applicant]
Yan et al., “Privacy-Preserving Compressive Model for Enhanced Deep-Learning-based Service Provision System in Edge Computing” IEEE Access, Jul. 8, 2019, pp. 92921-92937, vol. 7. [cited by applicant]