IP Library › Granted Patent US 12,481,923
Granted Patent B2
US 12,481,923 · App. 17/934,221 · Granted Nov 25, 2025

Graphics processing unit training job allocation

Inventors: Lin Dong (Beijing, CN); Jun Feng Liu (Ontario, CA)
Assignee: International Business Machines Corporation
G06N20/00G06F9/5027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,481,923
App. No.
17/934,221
Granted
Nov 25, 2025
Kind
B2
Abstract

A computer-implemented method for training a machine learning model includes receiving a first training job at a processing device of a computer system having a plurality of graphics processing unit (GPU) resources, the first training job being part of a set of training jobs, and determining an amount of available memory in each GPU resource of the plurality of GPU resources. The method also includes loading the training job into one or more GPU resources with at least one second training job. The loading includes determining a cost model indicating an efficiency cost of each of a plurality of packing patterns, and packing the first training job and the second training job into the one or more GPU resources according to a packing pattern associated with a lowest efficiency cost.

Claims (34)

1 . A computer-implemented method for training a machine learning model, the method comprising:

receiving a first training job at a processing device of a computer system having a plurality of graphics processing unit (GPU) resources, the first training job being part of a set of training jobs;

determining an amount of available memory in each GPU resource of the plurality of GPU resources; and

loading the training job into one or more GPU resources with at least one second training job, wherein the loading includes:

determining a cost model indicating an efficiency cost of each of a plurality of packing patterns, and packing the first training job and the second training job into the one or more GPU resources according to a packing pattern associated with a lowest efficiency cost.

2 . The method of claim 1 , wherein determining the cost model includes performing a test run for the first training job, the test run including a selected number of iterations, and collecting runtime metrics for the first training job.

3 . The method of claim 2 , wherein determining the cost model includes generating a packing pair for each iteration, each packing pair including an iteration for the first training job and an iteration for the second training job.

4 . The method of claim 3 , wherein determining the cost model includes calculating the efficiency cost for each packing pair, and selecting the packing pattern associated with a packing pair having the lowest efficiency cost.

5 . The method of claim 2 , further comprising scaling in an existing job to reduce a number of GPU resources executing the existing job, to provide sufficient memory to perform the test run.

6 . The method of claim 1 , wherein determining the cost model includes splitting a training model associated with the first training job.

7 . The method of claim 1 , wherein the loading includes adaptively packing the first training job into the one or more GPU resources based on the packing pattern associated with the lowest efficiency cost, wherein adaptively packing includes scaling up the first training job based on there being insufficient memory to hold the first training job and the second training job in a single GPU resource.

8 . The method of claim 1 , wherein the first training job is configured for elastic distributed training.

9 . A system for training a machine learning model, the system comprising:

a processor in communication with a computer system having a plurality of graphics processing unit (GPU) resources, the processor configured to perform:

receiving a first training job, the first training job being part of a set of training jobs;

determining an amount of available memory in each GPU resource of the plurality of GPU resources; and

loading the training job into one or more GPU resources with at least one second training job, wherein the loading includes:

determining a cost model indicating an efficiency cost of each of a plurality of packing patterns, and packing the first training job and the second training job into the one or more GPU resources according to a packing pattern associated with a lowest efficiency cost.

10 . The system of claim 9 , wherein determining the cost model includes performing a test run for the first training job, the test run including a selected number of iterations, and collecting runtime metrics for the first training job.

11 . The system of claim 10 , wherein determining the cost model includes generating a packing pair for each iteration, each packing pair including an iteration for the first training job and an iteration for the second training job.

12 . The system of claim 11 , wherein determining the cost model includes calculating the efficiency cost for each packing pair, and selecting the packing pattern associated with a packing pair having the lowest efficiency cost.

13 . The system of claim 10 , wherein the processor is further configured to perform scaling in an existing job to reduce a number of GPU resources executing the existing job, to provide sufficient memory to perform the test run.

14 . The system of claim 9 , wherein the loading includes adaptively packing the first training job into the one or more GPU resources based on the packing pattern associated with the lowest efficiency cost, wherein adaptively packing includes scaling up the first training job based on there being insufficient memory to hold the first training job and the second training job in a single GPU resource.

15 . A computer program product for training a machine learning model, the computer program product comprising:

a non-transitory storage medium readable by a processing circuit and storing instructions for execution by the processing circuit for performing a method comprising:

receiving a first training job at a processing device of a computer system having a plurality of graphics processing unit (GPU) resources, the first training job being part of a set of training jobs;

determining an amount of available memory in each GPU resource of the plurality of GPU resources; and

loading the training job into one or more GPU resources with at least one second training job, wherein the loading includes:

determining a cost model indicating an efficiency cost of each of a plurality of packing patterns, and packing the first training job and the second training job into the one or more GPU resources according to a packing pattern associated with a lowest efficiency cost.

16 . The computer program product of claim 15 , wherein determining the cost model includes performing a test run for the first training job, the test run including a selected number of iterations, and collecting runtime metrics for the first training job.

17 . The computer program product of claim 16 , wherein determining the cost model includes generating a packing pair for each iteration, each packing pair including an iteration for the first training job and an iteration for the second training job.

18 . The computer program product of claim 17 , wherein determining the cost model includes calculating the efficiency cost for each packing pair, and selecting the packing pattern associated with a packing pair having the lowest efficiency cost.

19 . The computer program product of claim 16 , wherein the method further includes scaling in an existing job to reduce a number of GPU resources executing the existing job, to provide sufficient memory to perform the test run.

20 . The computer program product of claim 15 , wherein the loading includes adaptively packing the first training job into the one or more GPU resources based on the packing pattern associated with the lowest efficiency cost, wherein adaptively packing includes scaling up the first training job based on there being insufficient memory to hold the first training job and the second training job in a single GPU resource.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2022
From: DONG, LIN; LIU, JUN FENG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061176/0646 →
Continuity (1)
Related Publication 20240104418A1 · Mar 28, 2024
References Cited (17)
US 10884795B2 · Liu et al. · 2021 [cited by applicant]
US 20030158906A1 · Hayes · 2003 [cited by examiner]
US 20200051317A1 · Muthler · 2020 [cited by examiner]
US 20210124998A1 · Jaenisch · 2021 [cited by examiner]
CN 106575246A · 2017 [cited by examiner]
CN 106663224A · 2017 [cited by examiner]
CN 110383296A · 2019 [cited by examiner]
CN 113470179A · 2021 [cited by examiner]
CN 113935886A · 2022 [cited by examiner]
CN 114078076A · 2022 [cited by examiner]
CN 117632447A · 2024 [cited by examiner]
Juncheng Gu et al.; “Tiresias: A GPU Cluster Manager for Distributed Deep Learning”, Proceedings of the 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI '19), Feb. 26-28, 2019, 17 pages. [cited by applicant]
Lui et al.; “Elastic Distributed Training in Watson Machine Learning Accelerator”, IBM, Mar. 16, 2020, 8 pages. [cited by applicant]
Lukas Biewald; “Monitor and Improve GPU Usage for Training Deep Learning Models”, Towards data science, Mar. 27, 2019, 10 pages. [cited by applicant]
Rui Pan, “Salus: Fine-Grained GPU Sharing Primitives for Deep Learning Applications”, rui pan's blog,2020, 1 page. [cited by applicant]
Wencong Xiao, et al.; “Gandiva: Introspective Cluster Scheduling for Deep Learning”, Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI '18), Oct. 8-10, 2018, 17 pages. [cited by applicant]
Wencong Xiao, et al; “AntMan: Dynamic Scaling on GPU Clusters for Deep Learning”, Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation, Nov. 4-6, 2020, 17 pages. [cited by applicant]