IP Library › Granted Patent US 12,737,189
Granted Patent B2
US 12,737,189 · App. 18/765,440 · Granted Sep 15, 2026

Job allocations to fractions of parallel processing units (PPUs)

Inventors: Diman Zad Tootaghaj (Milpitas, CA); Yunming Xiao (Ypsilanti, MI); Aditya Dhakal (Santa Clarita, CA); Lianjie Cao (Milpitas, CA); Puneet Sharma (Palo Alto, CA)
Assignee: Hewlett Packard Enterprise Development LP
G06F9/3885G06F9/4875G06T1/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,189
App. No.
18/765,440
Granted
Sep 15, 2026
Kind
B2
Abstract

In some examples, a controller receives a request to schedule a first job in a system including a plurality of physical parallel processing units (PPUs), where a physical PPU of the plurality of physical PPUs includes multiple PPU fractions. The controller allocates the first job to a first collection of PPU fractions of the multiple PPU fractions based on an operational cost reduction objective to reduce a cost associated with a usage of the plurality of physical PPUs. The controller triggers processing of the first job according to the allocation of the first job to the first collection of PPU fractions, where data isolation is provided between the first job allocated to the first collection of PPU fractions and a second job allocated to a second collection of PPU fractions of the multiple PPU fractions.

Claims (39)

1 . A non-transitory machine-readable storage medium comprising instructions, which upon execution, cause a controller to:

receive a request to schedule a first job in a system comprising a plurality of physical parallel processing units (PPUs), wherein a physical PPU of the plurality of physical PPUs comprises multiple PPU fractions, and wherein a first PPU fraction of the multiple PPU fractions comprises a first PPU compute resource and a first PPU memory resource that is separate and isolated from a second PPU compute resource and a second PPU memory resource of a second PPU fraction of the multiple PPU fractions;

allocate the first job to a first collection of PPU fractions of the multiple PPU fractions based on an operational cost reduction objective to reduce a cost associated with a usage of the plurality of physical PPUs; and

trigger processing of the first job according to the allocation of the first job to the first collection of PPU fractions, wherein data isolation is provided between the first job allocated to the first collection of PPU fractions and a second job allocated to a second collection of PPU fractions of the multiple PPU fractions.

2 . The non-transitory machine-readable storage medium of claim 1 , wherein a compute capacity of the first PPU fraction of the multiple PPU fractions is different from a compute capacity of the second PPU fraction of the multiple PPU fractions.

3 . The non-transitory machine-readable storage medium of claim 1 , wherein a memory capacity of the first PPU fraction of the multiple PPU fractions is different from a memory capacity of the second PPU fraction of the multiple PPU fractions.

4 . The non-transitory machine-readable storage medium of claim 1 , wherein the data isolation is based on isolation of PPU compute resources and PPU memory resources between the first collection of PPU fractions and the second collection of PPU fractions.

5 . The non-transitory machine-readable storage medium of claim 1 , wherein the controller is accessible by a plurality of tenants to use the plurality of physical PPUs, wherein the first job is requested by a first tenant, and the second job is requested by a second tenant different from the first tenant, and wherein tenant isolation is provided by allocating the first job to the first collection of PPU fractions of the physical PPU, and allocating the second job to the second collection of PPU fractions of the physical PPU.

6 . The non-transitory machine-readable storage medium of claim 5 , wherein the allocating of the first job to the first collection of PPU fractions is further based on a tenant isolation constraint to provide tenant isolation wherein a single tenant of the plurality of tenants is to use a PPU fraction of the physical PPU at a time.

7 . The non-transitory machine-readable storage medium of claim 6 , wherein the tenant isolation constraint comprises a tenant-job variable to indicate whether a respective PPU fraction of a physical PPU of the plurality of physical PPUs has been allocated to a respective tenant of the plurality of tenants.

8 . The non-transitory machine-readable storage medium of claim 7 , wherein the tenant-job variable is based on variables indicating whether corresponding jobs of the respective tenant have been allocated to the respective PPU fraction.

9 . The non-transitory machine-readable storage medium of claim 8 , wherein the tenant-job variable is based on a sum of the variables indicating whether corresponding jobs of the respective tenant have been allocated to the respective PPU fraction.

10 . The non-transitory machine-readable storage medium of claim 9 , wherein the tenant-job variable is set to a specified value if any job of the respective tenant is assigned to the respective PPU fraction.

11 . The non-transitory machine-readable storage medium of claim 1 , wherein the allocating of the first job to the first collection of PPU fractions is further based on a migration cost reduction objective to reduce a cost associated with migrating jobs between physical PPUs.

12 . The non-transitory machine-readable storage medium of claim 1 , wherein the allocating of the first job to the first collection of PPU fractions is further based on a constraint to ensure that cumulative resources allocated to one or more jobs within a given PPU fraction does not exceed a total capacity of the given PPU fraction.

13 . The non-transitory machine-readable storage medium of claim 1 , wherein the allocating of the first job to the first collection of PPU fractions is further based on a constraint to ensure cumulative resources allocated to one or more jobs across the multiple PPU fractions of the physical PPU does not exceed a total capacity of the physical PPU.

14 . The non-transitory machine-readable storage medium of claim 1 , wherein the controller to execute the instructions is part of an adapter that is separate from a central processing unit (CPU) of a computing system including the plurality of PPUs.

15 . The non-transitory machine-readable storage medium of claim 14 , wherein the adapter is to:

transfer first job data of the first job using a direct memory access (DMA) transfer from the adapter to a memory of the first collection of PPU fractions, and

transfer second job data of the second job using a DMA transfer from the adapter to a memory of the second collection of PPU fractions.

16 . The non-transitory machine-readable storage medium of claim 15 , wherein the adapter is to receive the first job data and the second job data from clients in remote DMA (RDMA) transfers over a network.

17 . An adapter for a system comprising a plurality of physical processing units (PPUs), the adapter comprising:

a network interface to communicate over a network; and

an adapter controller to:

receive, over the network, a request from a first tenant to schedule a first job in the system, wherein a physical PPU of the plurality of physical PPUs comprises multiple PPU fractions, and wherein a first PPU fraction of the multiple PPU fractions comprises a first PPU compute resource and a first PPU memory resource that is separate and isolated from a second PPU compute resource and a second PPU memory resource of a second PPU fraction of the multiple PPU fractions;

allocate the first job to a first collection of PPU fractions of the multiple PPU fractions based on:

an operational cost reduction objective to reduce a cost associated with a usage of the plurality of physical PPUs, and

a tenant isolation constraint to provide tenant isolation wherein a single tenant of a plurality of tenants including the first tenant is to use a PPU fraction of the physical PPU at a time; and

trigger processing of the first job according to the allocation of the first job to the first collection of PPU fractions, wherein data isolation is provided between the first job allocated to the first collection of PPU fractions and a second job of a second tenant allocated to a second collection of PPU fractions of the multiple PPU fractions.

18 . The adapter of claim 17 , wherein the adapter controller is to:

allocate multiple jobs of the first tenant to a common PPU fraction.

19 . A method comprising:

receiving, by a job scheduler executed on a controller, the method, a request from a first tenant to schedule a first job in a system including a plurality of physical processing units (PPUs), wherein a physical PPU of the plurality of physical PPUs comprises multiple PPU fractions, and wherein a first PPU fraction of the multiple PPU fractions comprises a first PPU compute resource and a first PPU memory resource that is separate and isolated from a second PPU compute resource and a second PPU memory resource of a second PPU fraction of the multiple PPU fractions;

allocating, by the job scheduler, the first job to a first collection of PPU fractions of the multiple PPU fractions based on:

an operational cost reduction objective to reduce a cost associated with a usage of the plurality of physical PPUs,

a migration cost reduction objective to reduce a cost associated with migrating jobs between physical PPUs, and

a tenant isolation constraint to provide tenant isolation wherein a single tenant of a plurality of tenants including the first tenant is to use a PPU fraction of a physical PPU at a time; and

processing the first job according to the allocation of the first job to the first collection of PPU fractions, wherein data isolation is provided between the first job allocated to the first collection of PPU fractions and a second job of a second tenant allocated to a second collection of PPU fractions of the multiple PPU fractions.

20 . The method of claim 19 , wherein the plurality of physical PPUs comprise a plurality of graphics processing units (GPUs), and wherein a physical GPU of the plurality of GPUs comprises a first GPU compute resource and a first GPU memory resource of a first GPU fraction that is separate and isolated from a second GPU compute resource and a second GPU memory resource of a second GPU fraction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2024
From: ZAD TOOTAGHAJ, DIMAN; XIAO, YUNMING; DHAKAL, ADITYA; CAO, LIANJIE; SHARMA, PUNEET
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 067923/0303 →
Continuity (1)
Related Publication 20260010375A1 · Jan 8, 2026
References Cited (109)
US 7620530B2 · Bordes · 2009 [cited by examiner]
US 9128739B1 · Juels et al. · 2015 [cited by applicant]
US 10313261B1 · Walton, III · 2019 [cited by applicant]
US 10325343B1 · Zhao et al. · 2019 [cited by applicant]
US 10430991B2 · Ma et al. · 2019 [cited by applicant]
US 11055809B2 · Shah et al. · 2021 [cited by applicant]
US 11449963B1 · Beeler et al. · 2022 [cited by applicant]
US 11494228B2 · Suh · 2022 [cited by examiner]
US 11651470B2 · Tootaghaj et al. · 2023 [cited by applicant]
US 11915041B1 · Tanach · 2024 [cited by examiner]
US 20040194055A1 · Galloway et al. · 2004 [cited by applicant]
US 20100287560A1 · Neft · 2010 [cited by applicant]
US 20110041131A1 · Srivatsa et al. · 2011 [cited by applicant]
US 20120017213A1 · Hunt et al. · 2012 [cited by applicant]
US 20120130554A1 · Jain et al. · 2012 [cited by applicant]
US 20120254439A1 · Yamasaki et al. · 2012 [cited by applicant]
US 20130238785A1 · Hawk et al. · 2013 [cited by applicant]
US 20130275382A1 · Chi et al. · 2013 [cited by applicant]
US 20140298338A1 · Doi · 2014 [cited by applicant]
US 20150039793A1 · Rossetti · 2015 [cited by applicant]
US 20150052531A1 · Helak et al. · 2015 [cited by applicant]
US 20160019937A1 · Arora et al. · 2016 [cited by applicant]
US 20160077871A1 · Kaplan et al. · 2016 [cited by applicant]
US 20160246625A1 · Katsura et al. · 2016 [cited by applicant]
US 20170255496A1 · Deng et al. · 2017 [cited by applicant]
US 20180060996A1 · Tunuguntla et al. · 2018 [cited by applicant]
US 20180089881A1 · Johnson · 2018 [cited by applicant]
US 20180373570A1 · Xu et al. · 2018 [cited by applicant]
US 20190050163A1 · Dewey et al. · 2019 [cited by applicant]
US 20190324822A1 · Gottin et al. · 2019 [cited by applicant]
US 20200019311A1 · Zolotow et al. · 2020 [cited by applicant]
US 20200090001A1 · Zargahi et al. · 2020 [cited by applicant]
US 20200219464A1 · Callway et al. · 2020 [cited by applicant]
US 20200409733A1 · Sankaran et al. · 2020 [cited by applicant]
US 20210004273A1 · You et al. · 2021 [cited by applicant]
US 20210026672A1 · Kurkure et al. · 2021 [cited by applicant]
US 20210026696A1 · Chen et al. · 2021 [cited by applicant]
US 20210110506A1 · Prakash et al. · 2021 [cited by applicant]
US 20210216365A1 · Zhao et al. · 2021 [cited by applicant]
US 20210216375A1 · Sivaraman et al. · 2021 [cited by applicant]
US 20210256418A1 · Creedon et al. · 2021 [cited by applicant]
US 20210373972A1 · Kurkure et al. · 2021 [cited by applicant]
US 20220091894A1 · Xia et al. · 2022 [cited by applicant]
US 20220188965A1 · Li et al. · 2022 [cited by applicant]
US 20220229701A1 · Zhang et al. · 2022 [cited by applicant]
US 20220237014A1 · Kurkure et al. · 2022 [cited by applicant]
US 20220414817A1 · Zad Tootaghaj et al. · 2022 [cited by applicant]
US 20230089925A1 · Cho et al. · 2023 [cited by applicant]
US 20230099950A1 · Porter · 2023 [cited by applicant]
US 20230102063A1 · Wong et al. · 2023 [cited by applicant]
US 20230155958A1 · An et al. · 2023 [cited by applicant]
US 20230195972A1 · Desai et al. · 2023 [cited by applicant]
US 20230297406A1 · Rogers et al. · 2023 [cited by applicant]
US 20230297421A1 · Cowperthwaite et al. · 2023 [cited by applicant]
US 20230418467A1 · Ezrielev et al. · 2023 [cited by applicant]
US 20230418826A1 · Tanigawa et al. · 2023 [cited by applicant]
US 20240056491A1 · Kamaraju et al. · 2024 [cited by applicant]
US 20250355716A1 · Duluk, Jr. · 2025 [cited by examiner]
Nvidia, “The Run:ai Scheduler: concepts and principles”, available online at <http://web.archive.org/web/20231209221226/https://docs.run.ai/latest/Researcher/scheduling/the-runai-scheduler/>, Dec. 9, 2023 23 pages. [cited by applicant]
Cho et al., SLA-Driven ML Inference Framework for Clouds with Heterogeneous Accelerators, Proceedings of Machine Learning and Systems (MLSys), 2022, 13 pages. [cited by applicant]
Crankshaw et al., “Clipper: A low-latency online prediction serving system.” NDSI, 2017, 17 pages. [cited by applicant]
Dakkak et al., “Trims: Transparent and isolated model sharing for low latency deep learning inference in function-as-a-service.” CLOUD, 2018, 13 pages. [cited by applicant]
DEEPOMATIC, Fork of the NVIDIA device plugin for Kubernetes with support for shared GPUs by declaring GPUs multiple times downloaded Feb. 6, 2023 (6 pages). [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, 2019, 16 pages. [cited by applicant]
DPDK Project, About DPDK downloaded Jun. 22, 2024 (5 pages). [cited by applicant]
Github, alibaba/clusterdata, clusterdata/cluster-trace-gpu-v2020 downloaded Jun. 22, 2024 (20 pages). [cited by applicant]
Github, AliyunContainerService / gpushare-device-plugin, GPU Sharing Device Plugin for Kubernetes Cluster downloaded Jun. 22, 2024 (2 pages). [cited by applicant]
Github, AliyunContainerService / gpushare-device-plugin, GPU Sharing Device Plugin in Kuberntes downloaded Feb. 6, 2023 (2 pages). [cited by applicant]
Github, Deepomatic / shared-gpu-nvidia-k8s-device-plugin downloaded Jun. 22, 2024 (5 pages). [cited by applicant]
Github, intel/linux-intel-its downloaded Jun. 22, 2024 (5 pages). [cited by applicant]
Google for Developers, About OR-Tools, Jan. 2023 (2 pages). [cited by applicant]
Hayashi et al., “ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit”, 2020, 5 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, In Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, 9 pages. [cited by applicant]
Hsu et al., Simultaneous and Heterogenous Multithreading, MICRO '23, Oct. 28-Nov. 1, 2023 (16 pages). [cited by applicant]
ImageNet, About ImageNet downloaded Jun. 22, 2024 (2 pages). [cited by applicant]
Intel, CPU vs. GPU: Making the Most of Both downloaded Feb. 6, 2023 (4 pages). [cited by applicant]
Junguk, et al., “SLA-Driven ML Inference Framework For Clouds With Heterogeneous Accelerators.” Proceedings of Machine Learning and Systems 4 (2022), 13 pages. [cited by applicant]
Kubernetes, Manage clusters with different types of GPUs downloaded Jun. 22, 2024 (3 pages). [cited by applicant]
Marvell, Marvell LiquidIO III, Sep. 2020 (3 pages). [cited by applicant]
Murray, et al., “tf. data: a machine learning data processing framework”, Proceedings of the VLDB Endowment, 2021, 16 pages. [cited by applicant]
NVIDIA BLUEFIELD-3 DPU Programmable Data Center Infrastructure On-A-Chip, Dec. 2021 (2 pages). [cited by applicant]
NVIDIA BLUEFIELD-3 Networking Platform Datasheet, Nov. 2023 (2 pages). [cited by applicant]
NVIDIA Converted Accelerators downloaded Jun. 22, 2024 (8 pages). [cited by applicant]
NVIDIA Corporation, “NVIDIA TensorRT”, available online at <https://web.archive.org/web/20230921054928/https://developer.nvidia.com/tensorrt>, Sep. 21, 2023, 5 pages. [cited by applicant]
NVIDIA DOCA Comm Channel, Programming Guide, May 2023 (31 pages). [cited by applicant]
NVIDIA DOCA DMA docs downloaded Jun. 22, 2024 (16 pages). [cited by applicant]
NVIDIA DOCA RDMA downloaded Jun. 22, 2024 (59 pages). [cited by applicant]
NVIDIA GPUDirect, Enhancing Data Movementand Access for GPUs downloaded Jun. 22, 2024 (5 pages). [cited by applicant]
NVIDIA Mellanox Innova-2 Flex Open Programmable SmartNIC downloaded Jun. 22, 2024 (6 pages). [cited by applicant]
NVIDIA Multi-Instance GPU User Guide, Mar. 2024 (58 pages). [cited by applicant]
NVIDIA, “NVIDIA Triton Inference Server”, available online at <https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/index.html>, Jan. 18, 2024, 2 pages. [cited by applicant]
NVIDIA, Multi-Process Service, Feb. 2024 (38 pages). [cited by applicant]
NVIDIA, Multi-Process Service, Oct. 2022 (37 pages). [cited by applicant]
NVIDIA, NVIDIA BLUEFIELD-2 DPU Datasheet, Data Center Infrastructure on a Chip, Nov. 2023 (2 pages). [cited by applicant]
Paras Jain et al., “Dynamic Space-Time Scheduling for GPU Inference.” arXiv preprint, 2018, 9 Pages. [cited by applicant]
Paras Jain et al., “The OoO VLIW JIT Compiler for GPU Inference.” arXiv preprint, 2019, 7 pages. [cited by applicant]
Romero et al., “INFaaS: A Model-less and Managed Inference Serving System”, Dec. 15, 2020, 16 pages. [cited by applicant]
Sapio et al., “Scaling Distributed Machine Learning with In-Network Aggregation”, Proceedings of the 18th USENIX Symposium on Networked Systems Design and Implementation, 2021, 25 pages. [cited by applicant]
Schedule GPUs_Kubernetes last modified Oct. 18, 2022 (3 pages). [cited by applicant]
Sengupta et al., “Multi-tenancy on GPGPU-based servers.” 7th international workshop on Virtual-ization technologies in distributed computing, 2013, 8 Pages. [cited by applicant]
Skolnick et al., “AlphaFold 2: Why It Works and Its Implications for Understanding the Relationships of Protein Sequence, Structure, and Function”, 2021, 8 pages. [cited by applicant]
Tootaghaj et al., U.S. Appl. No. 18/299,855 entitled Job Allocations to Graphics Processing Units With Tenant Isolation filed Apr. 13, 2023 (38 pages). [cited by applicant]
Tootaghaj et al., U.S. Appl. No. 18/765,440 entitled Job Allocations to Fractions of Parallel Processing Units (PPUs) filed Jul. 8, 2024 (63 pages). [cited by applicant]
Tootaghaj et al., U.S. Appl. No. 18/765,445 entitled DMA Transfers of Job Data From an Adapter to Parallel Processing Unit (PPU) Fractions filed Jul. 8, 2024 (64 pages). [cited by applicant]
Weng et al., “MLaaS in the Wild:Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters”, Apr. 4-6, 2022, 17 pages. [cited by applicant]
Wikipedia, CUDA, Stable release: 12.1.0 / Mar. 1, 2023 (23 pages). [cited by applicant]
Xiao et al., “Conspirator: SmartNIC-Aided Control Framework for ML Workloads Orchestration”, 2023, 6 pages. [cited by applicant]
Yeh et. al., “Pagoda: Fine-grained gpu resource virtualization for narrow tasks.” CM SIGPLAN Notices, 2017, 13 pages. [cited by applicant]
Zhou, et al., “Deep interest network for click-through rate prediction”, In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery data mining, 2018, 9 pages. [cited by applicant]