IP Library › Granted Patent US 12,277,080
Granted Patent B2
US 12,277,080 · App. 18/460,043 · Granted Apr 15, 2025

Smart network interface card control plane for distributed machine learning workloads

Inventors: Yunming Xiao (Evanston, IL); Diman Zad Tootaghaj (San Jose, CA); Aditya Dhakal (Santa Clarita, CA); Puneet Sharma (Palo Alto, CA)
Assignee: Hewlett Packard Enterprise Development LP
G06F13/36G06F2213/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,080
App. No.
18/460,043
Filed
Sep 1, 2023
Granted
Apr 15, 2025
Kind
B2
Examiner
MAMO, ELIAS
Art Unit
2184
USPC
710/105
Abstract

In certain embodiments, a method includes receiving, at an interface of a Smart network interface card (SmartNIC) of a computing device, via a network, a network data unit; processing, by a data allocator of a SmartNIC subsystem of the SmartNIC, the network data unit to make a determination that data included in the network data unit is intended for processing by an accelerator of the computing device, wherein the accelerator is configured to execute a machine learning algorithm; storing, by the data allocator and based on the determination, the data in a local buffer of the SmartNIC subsystem; identifying, by the data allocator, a memory resource associated with the accelerator; and transferring the data from the local buffer to the memory resource.

Claims (47)

1. A computing device, comprising:

one or more processors;

a Smart network interface card (SmartNIC) comprising a SmartNIC subsystem comprising a data allocator and a local buffer;

an accelerator; and

one or more non-transitory computer readable media storing instruction which, when executed by the SmartNIC subsystem, cause the SmartNIC subsystem to:

receive, at an interface of the SmartNIC, via a network, a network data unit;

process, by the data allocator, the network data unit to make a determination that data included in the network data unit is intended for processing by the accelerator, wherein the accelerator is configured to execute a machine learning algorithm;

store, by the data allocator and based on the determination, the data in the local buffer;

identify, by the data allocator, a memory resource associated with the accelerator;

transfer the data from the local buffer to the memory resource; and

the instructions further cause the accelerator to execute the machine learning algorithm using the data.

2. The computing device of claim 1 , wherein the instructions further cause the SmartNIC to:

obtain, before receiving the network data unit, accelerator configuration information from the one or more processors.

3. The computing device of claim 2 , wherein the instructions further cause the SmartNIC to be configured to execute a bin packing algorithm.

4. The computing device of claim 2 , wherein the instructions further cause the SmartNIC to be configured to request that the one or more processors execute a bin packing algorithm.

5. The computing device of claim 2 , wherein the instructions further cause the SmartNIC to:

obtain, before receiving the network data unit, accelerator service rate information corresponding to the accelerator.

6. The computing device of claim 2 , wherein the accelerator configuration information includes accelerator virtualization information comprising identification of a plurality of accelerator slices of the accelerator.

7. The computing device of claim 6 , wherein the memory resource corresponds to an accelerator slice of the plurality of accelerator slices,

alternative heuristic algorithm.

8. The computing device of claim 2 , wherein the instructions further cause the SmartNIC to be configured to execute an alternative heuristic algorithm.

9. A method, comprising:

receiving, at an interface of a Smart network interface card (SmartNIC) of a computing device, via a network, a network data unit;

processing, by a data allocator of a SmartNIC subsystem of the SmartNIC, the network data unit to make a determination that data included in the network data unit is intended for processing by an accelerator of the computing device, wherein the accelerator is configured to execute a machine learning algorithm;

storing, by the data allocator and based on the determination, the data in a local buffer of the SmartNIC subsystem;

identifying, by the data allocator, a memory resource associated with the accelerator; and

transferring the data from the local buffer to the memory resource.

10. The method of claim 9 , further comprising

obtaining, before receiving the network data unit, accelerator configuration information from one or more processors of the computing device.

11. The method of claim 10 , further comprising configuring the SmartNIC to execute a bin packing algorithm.

12. The method of claim 10 , further comprising configuring the SmartNIC to request that the one or more processors execute a bin packing algorithm.

13. The method of claim 10 , further comprising:

obtaining, by the SmartNIC and before receiving the network data unit, accelerator service rate information corresponding to the accelerator.

14. The method of claim 10 , wherein the accelerator configuration information includes accelerator virtualization information comprising identification of a plurality of accelerator slices of the accelerator.

15. The method of claim 14 , wherein the memory resource corresponds to an accelerator slice of the plurality of accelerator slices.

16. A non-transitory computer-readable medium storing programming for execution by one or more processing units of a Smart network interface card (SmartNIC), the programming comprising instructions to:

receive, at an interface of the SmartNIC of a computing device, via a network, a network data unit;

process, by a data allocator of a SmartNIC subsystem of the SmartNIC, the network data unit to make a determination that data included in the network data unit is intended for processing by an accelerator of the computing device, wherein the accelerator is configured to execute a machine learning algorithm;

store, by the data allocator and based on the determination, the data in a local buffer of the SmartNIC subsystem;

identify, by the data allocator, a memory resource associated with the accelerator; and

transfer the data from the local buffer to the memory resource.

17. The non-transitory computer-readable medium of claim 16 , wherein the programming comprises further instructions to:

obtain, before receiving the network data unit, accelerator configuration information from one or more processors of the computing device.

18. The non-transitory computer-readable medium of claim 17 , wherein the programming comprises further instructions to configure the SmartNIC to execute a bin packing algorithm.

19. The non-transitory computer-readable medium of claim 17 , wherein the programming comprises further instructions to:

obtain, by the SmartNIC and before receiving the network data unit, accelerator service rate information corresponding to the accelerator.

20. The non-transitory computer-readable medium of claim 17 , wherein the accelerator configuration information includes accelerator virtualization information comprising identification of a plurality of accelerator slices of the accelerator, and the memory resource corresponds to an accelerator slice of the plurality of accelerator slices.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 1, 2023
From: XIAO, YUNMING; ZAD TOOTAGHAJ, DIMAN; DHAKAL, ADITYA; SHARMA, PUNEET
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 064776/0671 →
Continuity (1)
Related Publication 20250077456A1 · Mar 6, 2025
References Cited (66)
US 9680774B2 · Pirko · 2017 [cited by examiner]
US 10430991B2 · Ma et al. · 2019 [cited by applicant]
US 11055809B2 · Shah et al. · 2021 [cited by applicant]
US 11651470B2 · Zad Tootaghaj · 2023 [cited by examiner]
US 11948050B2 · Creedon · 2024 [cited by examiner]
US 12190405B2 · Rimmer · 2025 [cited by examiner]
US 20180373570A1 · Xu et al. · 2018 [cited by applicant]
US 20190324822A1 · Gottin et al. · 2019 [cited by applicant]
US 20210110506A1 · Prakash et al. · 2021 [cited by applicant]
US 20210216365A1 · Zhao · 2021 [cited by examiner]
US 20210256418A1 · Creedon et al. · 2021 [cited by applicant]
US 20210373972A1 · Kurkure et al. · 2021 [cited by applicant]
US 20220188965A1 · Li et al. · 2022 [cited by applicant]
US 20220237014A1 · Kurkure et al. · 2022 [cited by applicant]
US 20220414817A1 · Zad Tootaghaj et al. · 2022 [cited by applicant]
US 20230089925A1 · Cho et al. · 2023 [cited by applicant]
DPDK Project, About DPDK downloaded Jun. 22, 2024 (5 pages). [cited by applicant]
Github, alibaba/clusterdata, clusterdata/cluster-trace-gpu-v2020 downloaded Jun. 22, 2024 (20 pages). [cited by applicant]
Github, AliyunContainerService / gpushare-device-plugin, GPU Sharing Device Plugin for Kubernetes Cluster downloaded Jun. 22, 2024 (2 pages). [cited by applicant]
Github, Deepomatic / shared-gpu-nvidia-k8s-device-plugin downloaded Jun. 22, 2024 (5 pages). [cited by applicant]
Github, intel/linux-intel-its downloaded Jun. 22, 2024 (5 pages). [cited by applicant]
Google for Developers, About OR-Tools, Jan. 2023 (2 pages). [cited by applicant]
Hsu et al., Simultaneous and Heterogenous Multithreading, MICRO '23, Oct. 28-Nov. 1, 2023 (16 pages). [cited by applicant]
Imagenet, About ImageNet downloaded Jun. 22, 2024 (2 pages). [cited by applicant]
Kubernetes, Manage clusters with different types of GPUs downloaded Jun. 22, 2024 (3 pages). [cited by applicant]
Marvell, Marvell LiquidIO III, Sep. 2020 (3 pages). [cited by applicant]
NVIDIA Bluefield-3 DPU Programmable Data Center Infrastructure On-a-Chip, Dec. 2021 (2 pages). [cited by applicant]
NVIDIA Bluefield-3 Networking Platform Datasheet, Nov. 2023 (2 pages). [cited by applicant]
NVIDIA Converted Accelerators downloaded Jun. 22, 2024 (8 pages). [cited by applicant]
NVIDIA DOCA Comm Channel, Programming Guide, May 2023 (31 pages). [cited by applicant]
NVIDIA DOCA DMA docs downloaded Jun. 22, 2024 (16 pages). [cited by applicant]
NVIDIA DOCA RDMA downloaded Jun. 22, 2024 (59 pages). [cited by applicant]
NVIDIA GPUDirect, Enhancing Data Movementand Access for GPUs downloaded Jun. 22, 2024 (5 pages). [cited by applicant]
NVIDIA Mellanox Innova-2 Flex Open Programmable SmartNIC downloaded Jun. 22, 2024 (6 pages). [cited by applicant]
NVIDIA Multi-Instance GPU User Guide, Mar. 2024 (58 pages). [cited by applicant]
NVIDIA, Multi-Process Service, Feb. 2024 (38 pages). [cited by applicant]
NVIDIA, NVIDIA Bluefield-2 DPU Datasheet, Data Center Infrastructure on a Chip, Nov. 2023 (2 pages). [cited by applicant]
Tootaghaj et al., U.S. Appl. No. 18/299,855 entitled Job Allocations to Graphics Processing Units With Tenant Isolation filed Apr. 13, 2023 (38 pages). [cited by applicant]
Tootaghaj et al., U.S. Appl. No. 18/765,440 entitled Job Allocations to Fractions of Parallel Processing Units (PPUs) filed Jul. 8, 2024 (63 pages). [cited by applicant]
Tootaghaj et al., U.S. Appl. No. 18/765,445 entitled DMA Transfers of Job Data From an Adapter to Parallel Processing Unit (PPU) Fractions filed Jul. 8, 2024 (64 pages). [cited by applicant]
Cho et al., SLA-Driven ML Inference Framework for Clouds with Heterogeneous Accelerators, Proceedings of Machine Learning and Systems (MLSys), 2022, 13 pages. [cited by applicant]
Crankshaw et al., “Clipper: A low-latency online prediction serving system.” NDSI, 2017, 17 pages. [cited by applicant]
Dakkak et al., “Trims: Transparent and isolated model sharing for low latency deep learning inference in function-as-a-service.” CLOUD, 2018, 13 pages. [cited by applicant]
Deepomatic, Fork of the NVIDIA device plugin for Kubernetes with support for shared GPUs by declaring GPUs multiple times downloaded Feb. 6, 2023 (6 pages). [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, 2019, 16 pages. [cited by applicant]
Github, AliyunContainerService / gpushare-device-plugin, GPU Sharing Device Plugin in Kuberntes downloaded Feb. 6, 2023 (2 pages). [cited by applicant]
Hayashi et al., “ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit”, 2020, 5 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, In Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, 9 pages. [cited by applicant]
Intel, CPU vs. GPU: Making the Most of Both downloaded Feb. 6, 2023 (4 pages). [cited by applicant]
Junguk, et al., “SLA-Driven ML Inference Framework For Clouds With Heterogeneous Accelerators.” Proceedings of Machine Learning and Systems 4 (2022), 13 pages. [cited by applicant]
Murray, et al., “tf. data: a machine learning data processing framework”, Proceedings of the VLDB Endowment, 2021, 16 pages. [cited by applicant]
NVIDIA Corporation, “NVIDIA TensorRT”, available online at <https://web.archive.org/web/20230921054928/https://developer.nvidia.com/tensorrt>, Sep. 21, 2023, 5 pages. [cited by applicant]
NVIDIA, “NVIDIA Triton Inference Server”, available online at <https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/index.html>, Jan. 18, 2024, 2 pages. [cited by applicant]
NVIDIA, Multi-Process Service, Oct. 2022 (37 pages). [cited by applicant]
Paras Jain et al., “Dynamic Space-Time Scheduling for GPU Inference.” arXiv preprint, 2018, 9 Pages. [cited by applicant]
Paras Jain et al., “The OoO VLIW JIT Compiler for GPU Inference.” arXiv preprint, 2019, 7 pages. [cited by applicant]
Romero et al., “INFaaS: a Model-less and Managed Inference Serving System”, Dec. 15, 2020, 16 pages. [cited by applicant]
Sapio et al., “Scaling Distributed Machine Learning with In-Network Aggregation”, Proceedings of the 18th USENIX Symposium on Networked Systems Design and Implementation, 2021, 25 pages. [cited by applicant]
Schedule GPUs_Kubernetes last modified Oct. 18, 2022 (3 pages). [cited by applicant]
Sengupta et al., “Multi-tenancy on GPGPU-based servers.” 7th international workshop on Virtual-ization technologies in distributed computing, 2013, 8 Pages. [cited by applicant]
Skolnick et al., “AlphaFold 2: Why It Works and Its Implications for Understanding the Relationships of Protein Sequence, Structure, and Function”, 2021, 8 pages. [cited by applicant]
Weng et al., “MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters”, Apr. 4-6, 2022, 17 pages. [cited by applicant]
Wikipedia, CUDA, Stable release: 12.1.0 / Mar. 1, 2023 (23 pages). [cited by applicant]
Xiao et al., “Conspirator: SmartNIC-Aided Control Framework for ML Workloads Orchestration”, 2023, 6 pages. [cited by applicant]
Yeh et al., “Pagoda: Fine-grained gpu resource virtualization for narrow tasks.” CM SIGPLAN Notices, 2017, 13 pages. [cited by applicant]
Zhou, et al., “Deep interest network for click-through rate prediction”, In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery data mining, 2018, 9 pages. [cited by applicant]