IP Library › Granted Patent US 12,737,230
Granted Patent B2
US 12,737,230 · App. 18/765,445 · Granted Sep 15, 2026

DMA transfers of job data from an adapter to parallel processing unit (PPU) fractions

Inventors: Diman Zad Tootaghaj (Milpitas, CA); Yunming Xiao (Ypsilanti, MI); Aditya Dhakal (Santa Clarita, CA); Puneet Sharma (Palo Alto, CA); Lianjie Cao (Milpitas, CA)
Assignee: Hewlett Packard Enterprise Development LP
G06F9/5038G06F9/5016G06F15/17331
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,230
App. No.
18/765,445
Granted
Sep 15, 2026
Kind
B2
Abstract

In some examples, an adapter is for use with a system including a plurality of physical processing units (PPUs). The adapter includes a network interface to communicate over a network an adapter controller to receive, over the network, job data for multiple jobs to be executed on PPU fractions of one or more physical PPUs, and determine that first job data of a first job of the multiple jobs is to be provided to a first PPU fraction of the PPU fractions, and that second job data of a second job of the multiple jobs is to be provided to a second PPU fraction of the PPU fractions. The adapter controller initiates a direct memory access (DMA) transfer of the first job data to a first PPU memory buffer of the first PPU fraction, and a DMA transfer of the second job data to a second PPU memory buffer of the second PPU fraction.

Claims (50)

1 . An adapter for a system comprising a plurality of physical parallel processing units (PPUs), the adapter comprising:

a network interface to communicate over a network; and

an adapter controller to:

receive, over the network, job data for multiple jobs to be executed on PPU fractions of one or more physical PPUs,

determine that first job data of a first job of the multiple jobs is to be provided to a first PPU fraction of the PPU fractions, and that second job data of a second job of the multiple jobs is to be provided to a second PPU fraction of the PPU fractions, and

initiate a direct memory access (DMA) transfer of the first job data to a first PPU memory buffer of the first PPU fraction, and a DMA transfer of the second job data to a second PPU memory buffer of the second PPU fraction.

2 . The adapter of claim 1 , comprising a job scheduler executable by the adapter controller to allocate the multiple jobs to the PPU fractions, wherein data isolation is provided between the first job allocated to the first PPU fraction and the second job allocated to the second PPU fraction.

3 . The adapter of claim 2 , wherein the job scheduler is executable by the adapter controller to allocate the multiple jobs to the PPU fractions based on one or more objectives, the one or more objectives selected from among an operational cost reduction objective to reduce a cost associated with a usage of the one or more physical PPUs, or a migration cost reduction objective to reduce a cost associated with migrating jobs between physical PPUs.

4 . The adapter of claim 1 , further comprising:

an adapter memory,

wherein the adapter controller is to provide adapter memory buffers in the adapter memory for respective clients that submitted the multiple jobs, and

wherein a first adapter memory buffer of the adapter memory buffers is to receive the first job data of the first job from a first client, and a second adapter memory buffer of the adapter memory buffers is to receive the second job data of the second job from a second client.

5 . The adapter of claim 4 , wherein the first job data is received in a Remote Direct Memory Access (RDMA) transfer from the first client to the first adapter memory buffer, and the second job data is received in a RDMA transfer from the second client to the second adapter memory buffer.

6 . The adapter of claim 1 , wherein the adapter is separate from a central processing unit (CPU) of the system, and the adapter controller is to:

responsive to a completion of the DMA transfer of the first job data of the first job to the first PPU memory buffer of the first PPU fraction, notify the CPU of the completion to cause invocation of machine-readable instructions by the CPU at the one or more physical PPUs to process the first job data.

7 . The adapter of claim 6 , wherein the adapter controller is to:

retrieve, at the adapter from the CPU, a result of the processing of the first job data, the result retrieved using a DMA transfer from the first PPU memory buffer of the first PPU fraction.

8 . The adapter of claim 7 , wherein the adapter controller is to:

receive an indication from a physical PPU of the one or more physical PPUs or from the CPU that the result is available at the first PPU memory buffer; and

initiate the DMA transfer from the first PPU memory buffer in response to the indication.

9 . The adapter of claim 1 , wherein the adapter controller is to:

receive a first memory address of the first PPU memory buffer of the first PPU fraction reserved by a central processing unit (CPU) of the system, and

receive a second memory address of the second PPU memory buffer of the second PPU fraction reserved by the CPU.

10 . The adapter of claim 9 , wherein the adapter controller is to:

store the first memory address and the second memory address in respective entries of mapping information that contain buffer identifiers of respective PPU memory buffers and memory addresses of the respective PPU memory buffers, the mapping information tracking allocations of PPU memory buffers in the one or more physical PPUs, and

responsive to receiving the first job data of the first job, perform a lookup of the mapping information to obtain the first memory address for accessing the first PPU memory buffer, wherein the DMA transfer of the first job data to the first PPU memory buffer uses the first memory address obtained from the mapping information.

11 . The adapter of claim 10 , wherein each respective entry of the entries of the mapping information further comprises a mapping value derived by applying a function on a memory address and a buffer identifier of a PPU memory buffer, wherein the first job data is associated with metadata comprising a first mapping value, and wherein the lookup of the mapping information uses the first mapping value to retrieve an entry of the mapping information.

12 . The adapter of claim 1 , comprising a smart network interface controller (NIC) including the network interface and the adapter controller.

13 . The adapter of claim 1 , wherein the adapter controller is to:

receive, over the network, the first job data of the first job according to a first format, and

translate the first job data according to the first format to converted job data according to a second format different from the first format,

wherein the DMA transfer of the first job data to the first PPU memory buffer comprises a DMA transfer of the converted job data to the first PPU memory buffer.

14 . The adapter of claim 13 , wherein the first format is a serial format, and the first job data comprises a serial stream of data, and

wherein the translating comprises deserializing the first job data into the converted job data according to the second format.

15 . The adapter of claim 1 , wherein the adapter is separate from a central processing unit (CPU) of the system, and the CPU is not in a data path of the DMA transfer of the first job data to the first PPU memory buffer of the first PPU fraction, and the DMA transfer of the second job data to the second PPU memory buffer of the second PPU fraction.

16 . The adapter of claim 1 , comprising one or more of logic to compress and decompress data associated with jobs executed by the one or more physical PPUs, or logic to encrypt or decrypt data associated with the jobs.

17 . A non-transitory machine-readable storage medium storing instructions that upon execution cause a host central processing unit (CPU) of a system to:

allocate parallel processing unit (PPU) memory buffers in respective PPU fractions of one or more physical PPUs;

send, to an adapter, references to the PPU memory buffers for association in mapping information to jobs from clients, wherein the adapter is separate from the host CPU;

receive, from the adapter, an indication of a direct memory access (DMA) transfer of job data of a job from the adapter to a first PPU memory buffer of a first PPU fraction of the PPU fractions; and

based on the indication, invoke machine-readable instructions in the first PPU fraction to process the job data in the first PPU memory buffer.

18 . The non-transitory machine-readable storage medium of claim 17 , wherein the mapping information is populated with the references to the PPU memory buffers and identifiers indicating the clients based on job scheduling of jobs to the PPU fractions by a job scheduler executed by the adapter.

19 . A method comprising:

receiving, by an adapter over a network:

first job data of a first job transferred from a first client in a first remote direct memory access (RDMA) transfer to a first adapter memory buffer in an adapter memory of the adapter, and

second job data of a second job transferred from a second client in a second RDMA transfer to a second adapter memory buffer in the adapter memory;

determining, by the adapter, that the first job data is to be provided to a first parallel processing unit (PPU) fraction of one or more physical PPUs, and that the second job data is to be provided to a second PPU fraction of the one or more physical PPUs;

performing a direct memory access (DMA) transfer of the first job data from the adapter to a first PPU memory buffer of the first PPU fraction, and a DMA transfer of the second job data from the adapter to a second PPU memory buffer of the second PPU fraction; and

receiving, by the adapter, a first result of processing of the first job data by a first compute resource in the first PPU fraction, and a second result of processing of the second job data by a second compute resource in the second PPU fraction.

20 . The method of claim 19 , wherein the determining that the first job data is to be provided to the first PPU fraction and that the second job data is to be provided to the second PPU fraction is based on looking up a hash map at the adapter, the hash map comprising a plurality of entries that track PPU memory buffers to jobs from clients.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2024
From: ZAD TOOTAGHAJ, DIMAN; XIAO, YUNMING; DHAKAL, ADITYA; SHARMA, PUNEET; CAO, LIANJIE
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 067923/0405 →
Continuity (1)
Related Publication 20260010406A1 · Jan 8, 2026
References Cited (110)
US 7533384B2 · Chan · 2009 [cited by examiner]
US 7620530B2 · Bordes · 2009 [cited by examiner]
US 9128739B1 · Juels et al. · 2015 [cited by applicant]
US 10313261B1 · Walton, III · 2019 [cited by applicant]
US 10325343B1 · Zhao et al. · 2019 [cited by applicant]
US 10430991B2 · Ma et al. · 2019 [cited by applicant]
US 11055809B2 · Shah et al. · 2021 [cited by applicant]
US 11449963B1 · Beeler et al. · 2022 [cited by applicant]
US 11494228B2 · Suh · 2022 [cited by examiner]
US 11651470B2 · Zad Tootaghaj et al. · 2023 [cited by applicant]
US 11915041B1 · Tanach · 2024 [cited by examiner]
US 20040194055A1 · Galloway et al. · 2004 [cited by applicant]
US 20100287560A1 · Neft · 2010 [cited by applicant]
US 20110041131A1 · Srivatsa et al. · 2011 [cited by applicant]
US 20120017213A1 · Hunt et al. · 2012 [cited by applicant]
US 20120130554A1 · Jain et al. · 2012 [cited by applicant]
US 20120254439A1 · Yamasaki et al. · 2012 [cited by applicant]
US 20130238785A1 · Hawk et al. · 2013 [cited by applicant]
US 20130275382A1 · Chi et al. · 2013 [cited by applicant]
US 20140298338A1 · Doi · 2014 [cited by applicant]
US 20150039793A1 · Rossetti · 2015 [cited by applicant]
US 20150052531A1 · Helak et al. · 2015 [cited by applicant]
US 20160019937A1 · Arora et al. · 2016 [cited by applicant]
US 20160077871A1 · Kaplan et al. · 2016 [cited by applicant]
US 20160246625A1 · Katsura et al. · 2016 [cited by applicant]
US 20170255496A1 · Deng et al. · 2017 [cited by applicant]
US 20180060996A1 · Tunuguntla et al. · 2018 [cited by applicant]
US 20180089881A1 · Johnson · 2018 [cited by applicant]
US 20180373570A1 · Xu et al. · 2018 [cited by applicant]
US 20190050163A1 · Dewey et al. · 2019 [cited by applicant]
US 20190324822A1 · Gottin et al. · 2019 [cited by applicant]
US 20200019311A1 · Zolotow et al. · 2020 [cited by applicant]
US 20200090001A1 · Zargahi et al. · 2020 [cited by applicant]
US 20200219464A1 · Callway et al. · 2020 [cited by applicant]
US 20200409733A1 · Sankaran et al. · 2020 [cited by applicant]
US 20210004273A1 · You et al. · 2021 [cited by applicant]
US 20210026672A1 · Kurkure et al. · 2021 [cited by applicant]
US 20210026696A1 · Chen et al. · 2021 [cited by applicant]
US 20210110506A1 · Prakash et al. · 2021 [cited by applicant]
US 20210216365A1 · Zhao et al. · 2021 [cited by applicant]
US 20210216375A1 · Sivaraman et al. · 2021 [cited by applicant]
US 20210256418A1 · Creedon et al. · 2021 [cited by applicant]
US 20210373972A1 · Kurkure et al. · 2021 [cited by applicant]
US 20220091894A1 · Xia et al. · 2022 [cited by applicant]
US 20220188965A1 · Li et al. · 2022 [cited by applicant]
US 20220229701A1 · Zhang et al. · 2022 [cited by applicant]
US 20220237014A1 · Kurkure et al. · 2022 [cited by applicant]
US 20220414817A1 · Zad Tootaghaj et al. · 2022 [cited by applicant]
US 20230089925A1 · Cho et al. · 2023 [cited by applicant]
US 20230099950A1 · Porter · 2023 [cited by applicant]
US 20230102063A1 · Wong et al. · 2023 [cited by applicant]
US 20230155958A1 · An et al. · 2023 [cited by applicant]
US 20230195972A1 · Desai et al. · 2023 [cited by applicant]
US 20230297406A1 · Rogers et al. · 2023 [cited by applicant]
US 20230297421A1 · Cowperthwaite et al. · 2023 [cited by applicant]
US 20230418467A1 · Ezrielev et al. · 2023 [cited by applicant]
US 20230418826A1 · Tanigawa et al. · 2023 [cited by applicant]
US 20240056491A1 · Kamaraju et al. · 2024 [cited by applicant]
US 20250355716A1 · Duluk, Jr. · 2025 [cited by examiner]
Cho et al., SLA-Driven ML Inference Framework for Clouds with Heterogeneous Accelerators, Proceedings of Machine Learning and Systems (MLSys), 2022, 13 pages. [cited by applicant]
Crankshaw et. al., “Clipper: A low-latency online prediction serving system.” NDSI, 2017, 17 pages. [cited by applicant]
Dakkak et. al., “Trims: Transparent and isolated model sharing for low latency deep learning inference in function-as-a-service.” CLOUD, 2018, 13 pages. [cited by applicant]
Deepomatic, Fork of the NVIDIA device plugin for Kubernetes with support for shared GPUs by declaring GPUs multiple times downloaded Feb. 6, 2023 (6 pages). [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, 2019, 16 pages. [cited by applicant]
DPDK Project, About DPDK downloaded Jun. 22, 2024 (5 pages). [cited by applicant]
Github, alibaba/clusterdata, clusterdata/cluster-trace-gpu-v2020 downloaded Jun. 22, 2024 (20 pages). [cited by applicant]
Github, AliyunContainerService / gpushare-device-plugin, GPU Sharing Device Plugin for Kubernetes Cluster downloaded Jun. 22, 2024 (2 pages). [cited by applicant]
Github, AliyunContainerService / gpushare-device-plugin, GPU Sharing Device Plugin in Kuberntes downloaded Feb. 6, 2023 (2 pages). [cited by applicant]
Github, Deepomatic / shared-gpu-nvidia-k8s-device-plugin downloaded Jun. 22, 2024 (5 pages). [cited by applicant]
Github, intel/linux-intel-its downloaded Jun. 22, 2024 (5 pages). [cited by applicant]
Google for Developers, About OR-Tools, Jan. 2023 (2 pages). [cited by applicant]
Hayashi et al., “ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit”, 2020, 5 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, In Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, 9 pages. [cited by applicant]
Hsu et al., Simultaneous and Heterogenous Multithreading, MICRO '23, Oct. 28-Nov. 1, 2023 (16 pages). [cited by applicant]
ImageNet, About ImageNet downloaded Jun. 22, 2024 (2 pages). [cited by applicant]
Intel, CPU vs. GPU: Making the Most of Both downloaded Feb. 6, 2023 (4 pages). [cited by applicant]
Junguk, et al., “SLA-Driven ML Inference Framework For Clouds With Heterogeneous Accelerators.” Proceedings of Machine Learning and Systems 4 (2022), 13 pages. [cited by applicant]
Kubernetes, Manage clusters with different types of GPUs downloaded Jun. 22, 2024 (3 pages). [cited by applicant]
Marvell, Marvell LiquidIO III, Sep. 2020 (3 pages). [cited by applicant]
Murray, et al., “tf. data: a machine learning data processing framework”, Proceedings of the VLDB Endowment, 2021, 16 pages. [cited by applicant]
NVIDIA Bluefield-3 DPU Programmable Data Center Infrastructure On-A-Chip, Dec. 2021 (2 pages). [cited by applicant]
NVIDIA Bluefield-3 Networking Platform Datasheet, Nov. 2023 (2 pages). [cited by applicant]
NVIDIA Converted Accelerators downloaded Jun. 22, 2024 (8 pages). [cited by applicant]
NVIDIA Corporation, “NVIDIA TensorRT”, available online at <https://web.archive.org/web/20230921054928/https://developer.nvidia.com/tensorrt>, Sep. 21, 2023, 5 pages. [cited by applicant]
NVIDIA DOCA Comm Channel, Programming Guide, May 2023 (31 pages). [cited by applicant]
NVIDIA DOCA DMA docs downloaded Jun. 22, 2024 (16 pages). [cited by applicant]
NVIDIA DOCA RDMA downloaded Jun. 22, 2024 (59 pages). [cited by applicant]
NVIDIA GPUDirect, Enhancing Data Movementand Access for GPUs downloaded Jun. 22, 2024 (5 pages). [cited by applicant]
NVIDIA Mellanox Innova-2 Flex Open Programmable SmartNIC downloaded Jun. 22, 2024 (6 pages). [cited by applicant]
NVIDIA Multi-Instance GPU User Guide, Mar. 2024 (58 pages). [cited by applicant]
NVIDIA, “NVIDIA Triton Inference Server”, available online at <https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/index.html>, Jan. 18, 2024, 2 pages. [cited by applicant]
NVIDIA, Multi-Process Service, Feb. 2024 (38 pages). [cited by applicant]
NVIDIA, Multi-Process Service, Oct. 2022 (37 pages). [cited by applicant]
NVIDIA, NVIDIA Bluefield-2 DPU Datasheet, Data Center Infrastructure on a Chip, Nov. 2023 (2 pages). [cited by applicant]
Paras Jain et. al., “Dynamic Space-Time Scheduling for GPU Inference.” arXiv preprint, 2018, 9 Pages. [cited by applicant]
Paras Jain et. al., “The OoO VLIW JIT Compiler for GPU Inference.” arXiv preprint, 2019, 7 pages. [cited by applicant]
Romero et al., “INFaaS: A Model-less and Managed Inference Serving System”, Dec. 15, 2020, 16 pages. [cited by applicant]
Sapio et al., “Scaling Distributed Machine Learning with In-Network Aggregation”, Proceedings of the 18th USENIX Symposium on Networked Systems Design and Implementation, 2021, 25 pages. [cited by applicant]
Schedule GPUs_Kubernetes last modified Oct. 18, 2022 (3 pages). [cited by applicant]
Sengupta et. al., “Multi-tenancy on GPGPU-based servers.” 7th international workshop on Virtual-ization technologies in distributed computing, 2013, 8 Pages. [cited by applicant]
Skolnick et al., “AlphaFold 2: Why It Works and Its Implications for Understanding the Relationships of Protein Sequence, Structure, and Function”, 2021, 8 pages. [cited by applicant]
Tootaghaj et al., U.S. Appl. No. 18/299,855 entitled Job Allocations to Graphics Processing Units With Tenant Isolation filed Apr. 13, 2023 (38 pages). [cited by applicant]
Tootaghaj et al., U.S. Appl. No. 18/765,440 entitled Job Allocations to Fractions of Parallel Processing Units (PPUs) filed Jul. 8, 2024 (63 pages). [cited by applicant]
Tootaghaj et al., U.S. Appl. No. 18/765,445 entitled Dma Transfers of Job Data From an Adapter to Parallel Processing Unit (PPU) Fractions filed Jul. 8, 2024 (64 pages). [cited by applicant]
Weng et al., “MLaaS in the Wild:Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters”, Apr. 4-6, 2022, 17 pages. [cited by applicant]
Wikipedia, CUDA, Stable release: 12.1.0 / Mar. 1, 2023 (23 pages). [cited by applicant]
Xiao et al., “Conspirator: SmartNIC-Aided Control Framework for ML Workloads Orchestration”, 2023, 6 pages. [cited by applicant]
Yeh et. al., “Pagoda: Fine-grained gpu resource virtualization for narrow tasks.” CM SIGPLAN Notices, 2017, 13 pages. [cited by applicant]
Zhou, et al., “Deep interest network for click-through rate prediction”, In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery data mining, 2018, 9 pages. [cited by applicant]
NVIDIA, “The Run:ai Scheduler: concepts and principles”, available online at <http://web.archive.org/web/20231209221226/https://docs.run.ai/latest/Researcher/scheduling/the-runai-scheduler/>, Dec. 9, 2023 23 pages. [cited by applicant]