IP Library › Granted Patent US 12,626,163
Granted Patent B2
US 12,626,163 · App. 17/217,203 · Granted May 12, 2026

Model parameter sharing between inference application instances in processing unit of information processing system

Inventors: Jinpeng Liu (Shanghai, CN); Danqing Sha (Shanghai, CN); Zhen Jia (Shanghai, CN); Christopher S. MacLellan (Uxbridge, MA)
Assignee: EMC IP Holding Company LLC
G06N5/043G06F9/5016G06F9/544G06N20/00G06F2209/543
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,163
App. No.
17/217,203
Granted
May 12, 2026
Kind
B2
Abstract

Techniques for model parameter sharing between inference model instances are disclosed. For example, a method performed by a first process obtains a representation of an inference model for which multiple instances of the inference model are to be executed on at least one processing unit. The method determines, from the representation of the inference model, one or more model parameters that are a pre-trained type of model parameter. The method allocates a shared memory for storing the one or more model parameters that are the pre-trained type of model parameter. The method stores the one or more model parameters that are the pre-trained type of model parameter in the shared memory for access by the multiple instances of the inference model to be executed on the at least one processing unit.

Claims (73)

1 . An apparatus comprising:

at least one memory storing program code; and

at least one processing platform comprising at least one processor coupled to the at least one memory, the at least one processing platform, when executing the program code, is configured to:

obtain, via a first process, an inference model for which multiple instances of the inference model are to be executed on at least one processing unit;

differentiate, using a host manager, from the inference model, via the first process, one or more model parameters that are a pre-trained type of model parameter, the differentiating comprising:

parsing the inference model;

generating a computation table for each computation of a plurality of computations defined in the inference model, the computation table comprising a plurality of computation nodes, the computation nodes comprising computation node numbers;

generating a parameter table for each parameter used by the plurality of computations, the parameter table comprising a plurality of parameter nodes, the parameter nodes comprising parameter node numbers, wherein each computation of the plurality of computations is indicated by a computation node number of the computation table and one or more parameter node numbers of the parameter table;

generating a memory model and associating the computation table and the parameter table using the generated memory model to determine one or more parameter node numbers of the parameter table associated with each computation node number of the plurality of computations, and wherein one or more parameter nodes of the plurality of parameter nodes are associated with two or more computation nodes of the plurality of computation nodes;

inferring a shape of each parameter node of each computation; and

identifying the one or more model parameters that are the pre-trained type of model parameter;

determine, using the host manager, based on the inferred shape of the parameter nodes that are identified as the one or more model parameters that are the pre-trained type of model parameter, an amount of memory to store the one or more model parameters that are the pre-trained type of model parameter in a shared memory storage space;

allocate, via the first process, a shared memory of the shared memory storage space corresponding to the amount of memory for storing the one or more model parameters that are the pre-trained type of model parameter;

extract, using the host manager, the one or more model parameters that are the pre-trained type of model parameter from the inference model;

store, via the first process, the one or more model parameters that are the pre-trained type of model parameter in the shared memory for access by the multiple instances of the inference model to be executed on the at least one processing unit; and

permit access to the shared memory of the shared memory storage space created by the first process to allow the multiple instances of the inference model to utilize the one or more model parameters that are the pre-trained type of model parameter in the shared memory of the shared memory storage space while executing the inference model.

2 . The apparatus of claim 1 , wherein the at least one processing platform, when executing the program code, is further configured to:

obtain, via a second process associated with a given one of the multiple instances of the inference model, the inference model;

determine from the inference model, via the second process, one or more model parameters that are not the pre-trained type of model parameter;

allocate, via the second process, a local memory for storing the one or more model parameters that are not the pre-trained type of model parameter; and

store, via the second process, the one or more model parameters that are not the pre-trained type of model parameter in the local memory for the given one of the multiple instances of the inference model.

3 . The apparatus of claim 2 , wherein the at least one processing platform, when executing the program code, is further configured to adjust one or more pointers to point to the local memory.

4 . The apparatus of claim 2 , wherein the at least one processing platform, when executing the program code, is further configured to adjust pointers to point to the shared memory of the shared memory storage space.

5 . The apparatus of claim 2 , wherein the one or more model parameters that are the pre-trained type of model parameter comprise one or more immutable model parameters, and the one or more model parameters that are not the pre-trained type of model parameter comprise one or more mutable model parameters.

6 . The apparatus of claim 2 , wherein the first process comprises a host process and the second process comprises a guest process.

7 . The apparatus of claim 1 , wherein the at least one processing unit comprises at least one graphic processing unit.

8 . The apparatus of claim 7 , wherein the at least one graphic processing unit is part of an edge computing network.

9 . The apparatus of claim 1 , wherein each of the multiple instances of the inference model are configured to receive and process data sets received from multiple users.

10 . The apparatus of claim 1 , wherein the inference model comprises a deep learning model.

11 . A method, comprising:

obtaining, via a first process, an inference model for which multiple instances of the inference model are to be executed on at least one processing unit;

differentiating, using a host manager, from the inference model, via the first process, one or more model parameters that are a pre-trained type of model parameter, the differentiating comprising:

parsing the inference model;

generating a computation table for each computation of a plurality of computations defined in the inference model, the computation table comprising a plurality of computation nodes, the computation nodes comprising computation node numbers;

generating a parameter table for each parameter used by the plurality of computations, the parameter table comprising a plurality of parameter nodes, the parameter nodes comprising parameter node numbers, wherein each computation of the plurality of computations is indicated by a computation node number of the computation table and one or more parameter node numbers of the parameter table;

generating a memory model and associating the computation table and the parameter table using the generated memory model to determine one or more parameter node numbers of the parameter table associated with each computation node number of the plurality of computations, and wherein one or more parameter nodes of the plurality of parameter nodes are associated with two or more computation nodes of the plurality of computation nodes;

inferring a shape of each parameter node of each computation; and

identifying the one or more model parameters that are the pre-trained type of model parameter;

determining, using the host manager, based on the inferred shape of the parameter nodes that are identified as the one or more model parameters that are the pre-trained type of model parameter, an amount of memory to store the one or more model parameters that are the pre-trained type of model parameter in a shared memory storage space;

allocating, via the first process, a shared memory of the shared memory storage space corresponding to the amount of memory for storing the one or more model parameters that are the pre-trained type of model parameter;

extracting, using the host manager, the one or more model parameters that are the pre-trained type of model parameter from the inference model;

storing, via the first process, the one or more model parameters that are the pre-trained type of model parameter in the shared memory for access by the multiple instances of the inference model to be executed on the at least one processing unit; and

permitting access to the shared memory of the shared memory storage space created by the first process to allow the multiple instances of the inference model to utilize the one or more model parameters that are the pre-trained type of model parameter in the shared memory of the shared memory storage space while executing the inference model.

12 . The method of claim 11 , further comprising:

obtaining, via a second process associated with a given one of the multiple instances of the inference model, the inference model;

determining from the inference model, via the second process, one or more model parameters that are not the pre-trained type of model parameter;

allocating, via the second process, a local memory for storing the one or more model parameters that are not the pre-trained type of model parameter; and

storing, via the second process, the one or more model parameters that are not the pre-trained type of model parameter in the local memory for the given one of the multiple instances of the inference model.

13 . The method of claim 12 , further comprising adjusting one or more pointers to point to the local memory.

14 . The method of claim 12 , further comprising adjusting pointers to point to the shared memory of the shared memory storage space.

15 . A computer program product stored on a non-transitory computer-readable medium and comprising machine executable instructions, the machine executable instructions, when executed, causing a processing platform to:

obtain, via a first process, an inference model for which multiple instances of the inference model are to be executed on at least one processing unit;

differentiate, using a host manager, from the inference model, via the first process, one or more model parameters that are a pre-trained type of model parameter, the differentiating comprising:

parsing the inference model;

generating a computation table for each computation of a plurality of computations defined in the inference model, the computation table comprising a plurality of computation nodes, the computation nodes comprising computation node numbers;

generating a parameter table for each parameter used by the plurality of computations, the parameter table comprising a plurality of parameter nodes, the parameter nodes comprising parameter node numbers, wherein each computation of the plurality of computations is indicated by a computation node number of the computation table and one or more parameter node numbers of the parameter table;

generating a memory model and associating the computation table and the parameter table using the generated memory model to determine one or more parameter node numbers of the parameter table associated with each computation node number of the plurality of computations, and wherein one or more parameter nodes of the plurality of parameter nodes are associated with two or more computation nodes of the plurality of computation nodes;

inferring a shape of each parameter node of each computation; and

identifying the one or more model parameters that are the pre-trained type of model parameter;

determine, using the host manager, based on the inferred shape of the parameter nodes that are identified as the one or more model parameters that are the pre-trained type of model parameter, an amount of memory to store the one or more model parameters that are the pre-trained type of model parameter in a shared memory storage space;

allocate, via the first process, a shared memory of the shared memory storage space corresponding to the amount of memory for storing the one or more model parameters that are the pre-trained type of model parameter;

extract, using the host manager, the one or more model parameters that are the pre-trained type of model parameter from the inference model;

store, via the first process, the one or more model parameters that are the pre-trained type of model parameter in the shared memory for access by the multiple instances of the inference model to be executed on the at least one processing unit; and

permit access to the shared memory of the shared memory storage space created by the first process to allow the multiple instances of the inference model to utilize the one or more model parameters that are the pre-trained type of model parameter in the shared memory of the shared memory storage space while executing the inference model.

16 . The computer program product of claim 15 , wherein the processing platform is further caused to:

obtain, via a second process associated with a given one of the multiple instances of the inference model, the inference model;

determine from the inference model, via the second process, one or more model parameters that are not the pre-trained type of model parameter;

allocate, via the second process, a local memory for storing the one or more model parameters that are not the pre-trained type of model parameter; and

store, via the second process, the one or more model parameters that are not the pre-trained type of model parameter in the local memory for the given one of the multiple instances of the inference model.

17 . The computer program product of claim 16 , wherein the one or more model parameters that are the pre-trained type of model parameter comprise one or more immutable model parameters, and the one or more model parameters that are not the pre-trained type of model parameter comprise one or more mutable model parameters.

18 . The method of claim 12 , wherein the one or more model parameters that are the pre-trained type of model parameter comprise one or more immutable model parameters, and the one or more model parameters that are not the pre-trained type of model parameter comprise one or more mutable model parameters.

19 . The computer program product of claim 16 , wherein the processing platform is further caused to adjust one or more pointers to point to the local memory.

20 . The computer program product of claim 16 , wherein the processing platform is further caused to adjust pointers to point to the shared memory of the shared memory storage space.

Assignments (10)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (056295/0124) Recorded Jun 10, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
Reel/Frame 062022/0012 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (056295/0280) Recorded Jun 10, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
Reel/Frame 062022/0255 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (056295/0001) Recorded Jun 10, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
Reel/Frame 062021/0844 →
RELEASE OF SECURITY INTEREST Recorded Nov 2, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
Reel/Frame 058297/0332 →
SECURITY INTEREST Recorded May 19, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 056295/0124 →
SECURITY INTEREST Recorded May 19, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 056295/0280 →
SECURITY INTEREST Recorded May 19, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 056295/0001 →
CORRECTIVE ASSIGNMENT TO CORRECT THE MISSING PATENTS THAT WERE ON THE ORIGINAL SCHEDULED SUBMITTED BUT NOT ENTERED PREVIOUSLY RECORDED AT REEL: 056250 FRAME: 0541. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 17, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 056311/0781 →
SECURITY AGREEMENT Recorded May 14, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 056250/0541 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2021
From: LIU, JINPENG; SHA, DANQING; JIA, ZHEN; MACLELLAN, CHRISTOPHER S.
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 055769/0105 →
Continuity (1)
Related Publication 20220318656A1 · Oct 6, 2022
References Cited (48)
US 11175844B1 · Ranjan · 2021 [cited by examiner]
US 11373119B1 · Doshi · 2022 [cited by examiner]
US 11868872B1 · Minkin · 2024 [cited by examiner]
US 12093806B1 · Zejda · 2024 [cited by examiner]
US 20180203673A1 · Ravishankar · 2018 [cited by examiner]
US 20180349313A1 · Ahn · 2018 [cited by examiner]
US 20190324810A1 · Zhao et al. · 2019 [cited by applicant]
US 20190325302A1 · Savic · 2019 [cited by examiner]
US 20200012500A1 · Kern · 2020 [cited by examiner]
US 20200042856A1 · Datta · 2020 [cited by examiner]
US 20200249998A1 · Che · 2020 [cited by examiner]
US 20200278888A1 · Connor · 2020 [cited by examiner]
US 20200334083A1 · Liu et al. · 2020 [cited by applicant]
US 20200334544A1 · Liu et al. · 2020 [cited by applicant]
US 20200380374A1 · Foret · 2020 [cited by examiner]
US 20210019652A1 · Gadelrab · 2021 [cited by examiner]
US 20210034582A1 · Liu et al. · 2021 [cited by applicant]
US 20210158131A1 · Jain · 2021 [cited by examiner]
US 20210209450A1 · Cassidy · 2021 [cited by examiner]
US 20210224684A1 · Sarkar · 2021 [cited by examiner]
US 20210350205A1 · Yan · 2021 [cited by examiner]
US 20220092439A1 · Liu · 2022 [cited by examiner]
US 20220108209A1 · Pudipeddi · 2022 [cited by examiner]
US 20220300826A1 · Chauhan · 2022 [cited by examiner]
Exxact Blog, “Discover the Difference Between Deep Learning Training and Inference”, Published Aug. 10, 2017, Exxact Corporation (Year: 2017). [cited by examiner]
Dakkak et al., “TrIMS: Transparent and Isolated Model Sharing for Low Latency Deep Learning Inference in Function-as-a-Service”, Published 2019, IEEE, 2019 IEEE 12th International Conference on Cloud Computing, pp. 372-… [cited by examiner]
Wikipedia, “Intermediate Representation,” https://en.wikipedia.org/w/index.php?title=Intermediate_representation&direction=next&oldid=905361000, Jan. 24, 2020, 4 pages. [cited by applicant]
Jia et al., “Beyond Data and Model Parallelism for Deep Neural Networks,” Proceedings of the 2nd SysML Conference, Palo Alto, CA, Jul. 2018, 13 pages. [cited by applicant]
Wikipedia, “Deep Learning,” https://en.wikipedia.org/wiki/Deep_learning, Feb. 6, 2020, 33 pages. [cited by applicant]
Wikipedia, “Everything as a Service,” https://simple.wikipedia.org/wiki/Everything_as_a_service, Aug. 23, 2019, 2 pages. [cited by applicant]
L. Song et al., “HyPar: Towards Hybrid Parallelism for Deep Learning Accelerator Array,” arXiv:1901.02067v1, Jan. 7, 2019, 13 pages. [cited by applicant]
Github, “OpenNESS Architecture and Solution Overview,” https://github.com/open-ness/specs/blob/master/doc/architecture.md, accessed Jul. 7, 2020, 15 pages. [cited by applicant]
Amazon Web Services, “Machine Learning Inference with AWS IoT Greengrass Solution Accelerator” https://aws.amazon.com/iot/solutions/mli-accelerator/, Oct. 2019, 5 pages. [cited by applicant]
ETSI, “MEC in 5G Networks,” White Paper No. 28, ISBN No. 979-10-92620-22-1, Jun. 2018, 28 pages. [cited by applicant]
ETSI, “Multi-access Edge Computing (MEC); Phase 2: Use Cases and Requirements,” Group Specification MEC 002 V2.1.1, Oct. 2018, 66 pages. [cited by applicant]
ETSI, “Multi-access Edge Computing (MEC); Framework and Reference Architecture,” Group Specification MEC 003 V2.1.1, Jan. 2019, 21 pages. [cited by applicant]
ETSI, “Multi-access Edge Computing (MEC); Application Mobility Service API,” Group Specification MEC 021 V2.1.1, Jan. 2020, 47 pages. [cited by applicant]
3GPP “3rd Generation Partnership Project; Technical Specification Group Core Network and Terminals; Interface between the Control Plane and the User Plane Nodes; Stage 3,” Technical Specification 29.244 V16.2.0, Dec. 20… [cited by applicant]
3GPP “3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; System architecture for the 5G System (5GS); Stage 2,” Technical Specification 23.501 V16.3.0, Dec. 2019, 417 pages. [cited by applicant]
3GPP “3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Procedures for the 5G System (5GS); Stage 2,” Technical Specification 23.502 V16.3.0, Dec. 2019, 558 pages. [cited by applicant]
3GPP “3rd Generation Partnership Project; Technical Specification Group Core Network and Terminals; 5G System; Session Management Policy Control Service; Stage 3,” Technical Specification 29.512 V16.3.0, Dec. 2019, 178 … [cited by applicant]
3GPP “3rd Generation Partnership Project; Technical Specification Group Core Network and Terminals; 5G System; Network Exposure Function Northbound APIs; Stage 3,” Technical Specification 29.522 V16.2.0, Dec. 2019, 106 … [cited by applicant]
U.S. Appl. No. 16/789,006 filed in the name of Jin Li et al. on Feb. 12, 2020, and entitled “Scheduling Artificial Intelligence Model Partitions Based on Reversed Computation Graph.” [cited by applicant]
U.S. Appl. No. 16/823,445 filed in the name of Jinpeng Liu et al. on Mar. 19, 2020, and entitled “Task Scheduling Method, Electronic Device, and Computer Storage Medium.” [cited by applicant]
U.S. Appl. No. 16/845,682 filed in the name of Jin Li et al. on Apr. 10, 2020, and entitled “Task Processing Method, Electronic Device, and Computer Program Product.” [cited by applicant]
U.S. Appl. No. 16/925,864 filed in the name of Jinpeng Liu et al. on Jul. 10, 2020, and entitled “Managing Artificial Intelligence Model Partitions for Edge Computing Environment.” [cited by applicant]
U.S. Appl. No. 17/105,030 filed in the name of Jinpeng Liu et al. on Nov. 25, 2020, and entitled “Method, System, and Computer Program Product for Deploying Application.”. [cited by applicant]
U.S. Appl. No. 17/132,344 filed in the name of Jinpeng Liu et al. on Dec. 23, 2020, and entitled “User Context Migration Based on Computation Graph in Deep Learning Application on Edge.” [cited by applicant]