IP Library Granted Patent US 12,335,142
Granted Patent B2
US 12,335,142 · App. 18/417,570 · Granted Jun 17, 2025

Network interface for data transport in heterogeneous computing environments

Inventors: Pratik M. Marolia (Hillsboro, OR); Rajesh M. Sankaran (Portland, OR); Ashok Raj (Portland, OR); Nrupal Jani (Hillsboro, OR); Parthasarathy Sarangam (Portland, OR); Robert O. Sharp (Austin, TX)
Assignee: Intel Corporation
H04L45/742G06F12/1081G06F13/28H04L45/60H04L49/9068
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,335,142
App. No.
18/417,570
Granted
Jun 17, 2025
Kind
B2
Abstract

A network interface controller can be programmed to direct write received data to a memory buffer via either a host-to-device fabric or an accelerator fabric. For packets received that are to be written to a memory buffer associated with an accelerator device, the network interface controller can determine an address translation of a destination memory address of the received packet and determine whether to use a secondary head. If a translated address is available and a secondary head is to be used, a direct memory access (DMA) engine is used to copy a portion of the received packet via the accelerator fabric to a destination memory buffer associated with the address translation. Accordingly, copying a portion of the received packet through the host-to-device fabric and to a destination memory can be avoided and utilization of the host-to-device fabric can be reduced for accelerator bound traffic.

Claims (157)

1. Network interface controller circuitry configurable for use in a host node and in association with a device driver, the host node comprising at least one graphics processing unit (GPU)-accessible memory, at least one host memory, and at least one host fabric, the host node to be communicatively coupled via at least one multi-switch fabric to a remote system, the remote system comprising at least one other GPU-accessible memory, at least one other host memory, and at least one other host fabric, the network interface controller circuitry comprising:

network interface circuitry for use in Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) packet data communication with the remote system via the at least one multi-switch fabric, the ROCE packet data communication to indicate at least one RDMA write to the host node from the remote system and/or at least one RDMA read from the host node to the remote system, the ROCE packet data communication to be initiated in response, at least in part, to at least one host application request; and

programmable circuitry to perform operations comprising:

in event that the ROCE packet data communication indicates the at least one RDMA write, directly writing, via the at least one host fabric, received packet data to the at least one GPU-accessible memory;

in event that the ROCE packet data communication indicates the at least one RDMA read, directly reading, via the at least one host fabric, other data from the at least one GPU-accessible memory that is to be provided to the remote system via the ROCE packet data communication; and

encryption, decryption, and compression-related host central processing unit (CPU) offload operations;

wherein:

the writing and the reading are to be performed in a manner that bypasses both (1) host CPU and/or host operating system (OS) in the writing and the reading, and (2) copying of the received packet data and the other data to the at least one host memory of the host node;

the writing and/or the reading are configurable to comprise use of direct data placement (DDP);

the writing and/or the reading are configurable to comprise use of address translation;

the address translation is to be implemented, at least in part, using the device driver;

portions of the received packet data and/or the other data are to be routed to their destinations via respective fabric-associated routings;

the respective fabric-associated routings are configurable to be mutually different from each other, at least in part; and

the at least one multi-switch fabric is to communicatively couple multiple switches associated with the host node and the remote system.

2. The network interface controller circuitry of claim 1 , wherein:

prior to being received by the programmable circuitry, the received packet data is to be directly read from the at least one other GPU-accessible memory via the at least one other host fabric in a manner that bypasses both (1) remote system CPU and/or remote system OS in the remote system, and (2) copying of the received packet data to the at least one other host memory.

3. The network interface controller circuitry of claim 2 , wherein:

the at least one host fabric comprises Peripheral Component Interconnect Express (PCIe) interconnect; and

the network interface controller circuitry is to be comprised in a circuit board that is to be communicatively coupled to the PCIe interconnect.

4. The network interface controller circuitry of claim 3 , wherein:

the received packet data and/or the other data are for use in association with artificial intelligence and/or machine learning.

5. The network interface controller circuitry of claim 4 , wherein:

the host node and the remote system each comprise multiple respective graphics processing units;

the at least one GPU-accessible memory is accessible by the multiple respective graphics processing units of the host node; and

the at least one other GPU-accessible memory is accessible by the multiple respective graphics processing units of the remote system.

6. The network interface controller circuitry of claim 1 , wherein:

the programmable circuitry is also for use in association with memory isolation.

7. The network interface controller circuitry of claim 1 , wherein:

at least one application specific integrated circuit (ASIC) comprises the programmable circuitry; and

the at least one host fabric comprises at least one accelerator fabric.

8. A method to be implemented using network interface controller circuitry, the network interface controller circuitry to be configured for use in a host node and in association with a device driver, the host node comprising at least one graphics processing unit (GPU)-accessible memory, at least one host memory, and at least one host fabric, the host node to be communicatively coupled via at least one multi-switch fabric to a remote system, the remote system comprising at least one other GPU-accessible memory, at least one other host memory, and at least one other host fabric, the network interface controller circuitry comprising network interface circuitry and programmable circuitry, the method comprising:

using the network interface circuitry in Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) packet data communication with the remote system via the at least one multi-switch fabric, the ROCE packet data communication to indicate at least one RDMA write to the host node from the remote system and/or at least one RDMA read from the host node to the remote system, the ROCE packet data communication to be initiated in response, at least in part, to at least one host application request; and

using the programmable circuitry to perform operations comprising:

in event that the ROCE packet data communication indicates the at least one RDMA write, directly writing, via the at least one host fabric, received packet data to the at least one GPU-accessible memory;

in event that the ROCE packet data communication indicates the at least one RDMA read, directly reading, via the at least one host fabric, other data from the at least one GPU-accessible memory that is to be provided to the remote system via the ROCE packet data communication; and

encryption, decryption, and compression-related host central processing unit (CPU) offload operations;

wherein:

the writing and the reading are to be performed in a manner that bypasses both (1) host CPU and/or host operating system (OS) in the writing and the reading, and (2) copying of the received packet data and the other data to the at least one host memory of the host node;

the writing and/or the reading are configurable to comprise use of direct data placement (DDP);

the writing and/or the reading are configurable to comprise use of address translation;

the address translation is to be implemented, at least in part, using the device driver;

portions of the received packet data and/or the other data are to be routed to their destinations via respective fabric-associated routings;

the respective fabric-associated routings are configurable to be mutually different from each other, at least in part; and

the at least one multi-switch fabric is to communicatively couple multiple switches associated with the host node and the remote system.

9. The method of claim 8 , wherein:

prior to being received by the programmable circuitry, the received packet data is to be directly read from the at least one other GPU-accessible memory via the at least one other host fabric in a manner that bypasses both (1) remote system CPU and/or remote system OS in the remote system, and (2) copying of the received packet data to the at least one other host memory.

10. The method of claim 9 , wherein:

the at least one host fabric comprises Peripheral Component Interconnect Express (PCIe) interconnect; and

the network interface controller circuitry is to be comprised in a circuit board that is to be communicatively coupled to the PCIe interconnect.

11. The method of claim 10 , wherein:

the received packet data and/or the other data are for use in association with artificial intelligence and/or machine learning.

12. The method of claim 11 , wherein:

the host node and the remote system each comprise multiple respective graphics processing units;

the at least one GPU-accessible memory is accessible by the multiple respective graphics processing units of the host node; and

the at least one other GPU-accessible memory is accessible by the multiple respective graphics processing units of the remote system.

13. The method of claim 8 , wherein:

the programmable circuitry is also for use in association with memory isolation.

14. The method of claim 8 , wherein:

at least one application specific integrated circuit (ASIC) comprises the programmable circuitry; and

the at least one host fabric comprises at least one accelerator fabric.

15. At least one non-transitory machine-readable storage medium storing instructions to be executed by at least one machine associated with network interface controller circuitry, the network interface controller circuitry to be configured for use in a host node and in association with a device driver, the host node comprising at least one graphics processing unit (GPU)-accessible memory, at least one host memory, and at least one host fabric, the host node to be communicatively coupled via at least one multi-switch fabric to a remote system, the remote system comprising at least one other GPU-accessible memory, at least one other host memory, and at least one other host fabric, the network interface controller circuitry comprising network interface circuitry and programmable circuitry, the instructions, when executed by the at least one machine, resulting in performance of operations comprising:

using the network interface circuitry in Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) packet data communication with the remote system via the at least one multi-switch fabric, the ROCE packet data communication to indicate at least one RDMA write to the host node from the remote system and/or at least one RDMA read from the host node to the remote system, the ROCE packet data communication to be initiated in response, at least in part, to at least one host application request; and

using the programmable circuitry to perform a set of operations comprising:

in event that the ROCE packet data communication indicates the at least one RDMA write, directly writing, via the at least one host fabric, received packet data to the at least one GPU-accessible memory;

in event that the ROCE packet data communication indicates the at least one RDMA read, directly reading, via the at least one host fabric, other data from the at least one GPU-accessible memory that is to be provided to the remote system via the ROCE packet data communication; and

encryption, decryption, and compression-related host central processing unit (CPU) offload operations;

wherein:

the writing and the reading are to be performed in a manner that bypasses both (1) host CPU and/or host operating system (OS) in the writing and the reading, and (2) copying of the received packet data and the other data to the at least one host memory of the host node;

the writing and/or the reading are configurable to comprise use of direct data placement (DDP);

the writing and/or the reading are configurable to comprise use of address translation;

the address translation is to be implemented, at least in part, using the device driver;

portions of the received packet data and/or the other data are to be routed to their destinations via respective fabric-associated routings;

the respective fabric-associated routings are configurable to be mutually different from each other, at least in part; and

the at least one multi-switch fabric is to communicatively couple multiple switches associated with the host node and the remote system.

16. The at least one non-transitory machine-readable storage medium of claim 15 , wherein:

prior to being received by the programmable circuitry, the received packet data is to be directly read from the at least one other GPU-accessible memory via the at least one other host fabric in a manner that bypasses both (1) remote system CPU and/or remote system OS in the remote system, and (2) copying of the received packet data to the at least one other host memory.

17. The at least one non-transitory machine-readable storage medium of claim 16 , wherein:

the at least one host fabric comprises Peripheral Component Interconnect Express (PCIe) interconnect; and

the network interface controller circuitry is to be comprised in a circuit board that is to be communicatively coupled to the PCIe interconnect.

18. The at least one non-transitory machine-readable storage medium of claim 17 , wherein:

the received packet data and/or the other data are for use in association with artificial intelligence and/or machine learning.

19. The at least one non-transitory machine-readable storage medium of claim 18 , wherein:

the host node and the remote system each comprise multiple respective graphics processing units;

the at least one GPU-accessible memory is accessible by the multiple respective graphics processing units of the host node; and

the at least one other GPU-accessible memory is accessible by the multiple respective graphics processing units of the remote system.

20. The at least one non-transitory machine-readable storage medium of claim 15 , wherein:

the programmable circuitry is also for use in association with memory isolation.

21. The at least one non-transitory machine-readable storage medium of claim 15 , wherein:

at least one application specific integrated circuit (ASIC) comprises the programmable circuitry; and

the at least one host fabric comprises at least one accelerator fabric.

22. A host system to be communicatively coupled via at least one multi-switch fabric to a remote system, the host system comprising:

at least one graphics processing unit (GPU)-accessible memory;

at least one host memory;

at least one host fabric; and

network interface controller circuitry comprising:

network interface circuitry for use in Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) packet data communication with the remote system via the at least one multi-switch fabric, the ROCE packet data communication to indicate at least one RDMA write to the host system from the remote system and/or at least one RDMA read from the host system to the remote system, the RoCE packet data communication to be initiated in response, at least in part, to at least one host application request; and

programmable circuitry to perform operations comprising:

in event that the ROCE packet data communication indicates the at least one RDMA write, directly writing, via the at least one host fabric, received packet data to the at least one GPU-accessible memory;

in event that the ROCE packet data communication indicates the at least one RDMA read, directly reading, via the at least one host fabric, other data from the at least one GPU-accessible memory that is to be provided to the remote system via the ROCE packet data communication; and

encryption, decryption, and compression-related host central processing unit (CPU) offload operations;

wherein:

the writing and the reading are to be performed in a manner that bypasses both (1) host CPU and/or host operating system (OS) in the writing and the reading, and (2) copying of the received packet data and the other data to the at least one host memory of the host system;

the writing and/or the reading are configurable to comprise use of direct data placement (DDP);

the writing and/or the reading are configurable to comprise use of address translation;

the address translation is to be implemented, at least in part, using the device driver;

portions of the received packet data and/or the other data are to be routed to their destinations via respective fabric-associated routings;

the respective fabric-associated routings are configurable to be mutually different from each other, at least in part; and

the at least one multi-switch fabric is to communicatively couple multiple switches associated with the host system and the remote system.

23. The host system of claim 22 , wherein:

the at least one host fabric comprises Peripheral Component Interconnect Express (PCIe) interconnect; and

the network interface controller circuitry is to be comprised in a circuit board that is to be communicatively coupled to the PCIe interconnect.

24. The host system of claim 23 , wherein:

the received packet data and/or the other data are for use in association with artificial intelligence and/or machine learning.

25. The host system of claim 24 , wherein:

the host system and the remote system each comprise multiple respective graphics processing units;

the at least one GPU-accessible memory is accessible by the multiple respective graphics processing units of the host system; and

at least one other GPU-accessible memory of the remote system is accessible by the multiple respective graphics processing units of the remote system.

26. The host system of claim 22 , wherein:

the programmable circuitry is also for use in association with memory isolation.

27. The host system of claim 22 , wherein:

at least one application specific integrated circuit (ASIC) comprises the programmable circuitry; and

the at least one host fabric comprises at least one accelerator fabric.

28. A data center system comprising:

at least one multi-switch fabric;

a remote system; and

a host system to be communicatively coupled via the at least one multi-switch fabric to the remote system, the host system comprising:

at least one graphics processing unit (GPU)-accessible memory;

at least one host memory;

at lea one host fabric; and

network interface controller circuitry comprising:

network interface circuitry for use in Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) packet data communication with the remote system via the at least one multi-switch fabric, the ROCE packet data communication to indicate at least one RDMA write to the host system from the remote system and/or at least one RDMA read from the host system to the remote system, the ROCE packet data communication to be initiated in response, at least in part, to at least one host application request; and

programmable circuitry to perform operations comprising:

in event that the ROCE packet data communication indicates the at least one RDMA write, directly writing, via the at least one host fabric, received packet data to the at least one GPU-accessible memory;

in event that the ROCE packet data communication indicates the at least one RDMA read, directly reading, via the at least one host fabric, other data from the at least one GPU-accessible memory that is to be provided to the remote system via the ROCE packet data communication; and

encryption, decryption, and compression-related host central processing unit (CPU) offload operations;

wherein:

the writing and the reading are to be performed in a manner that bypasses both (1) host CPU and/or host operating system (OS) in the writing and the reading, and (2) copying of the received packet data and the other data to the at least one host memory of the host system;

the writing and/or the reading are configurable to comprise use of direct data placement (DDP);

the writing and/or the reading are configurable to comprise use of address translation;

the address translation is to be implemented, at least in part, using the device driver;

portions of the received packet data and/or the other data are to be routed to their destinations via respective fabric-associated routings;

the respective fabric-associated routings are configurable to be mutually different from each other, at least in part; and

the at least one multi-switch fabric is to communicatively couple multiple switches associated with the host system and the remote system.

29. The data center system of claim 28 , wherein:

the at least one host fabric comprises Peripheral Component Interconnect Express (PCIe) interconnect; and

the network interface controller circuitry is to be comprised in a circuit board that is to be communicatively coupled to the PCIe interconnect.

30. The data center system of claim 29 , wherein:

the received packet data and/or the other data are for use in association with artificial intelligence and/or machine learning.

31. The data center system of claim 30 , wherein:

the host system and the remote system each comprise multiple respective graphics processing units;

the at least one GPU-accessible memory is accessible by the multiple respective graphics processing units of the host system; and

at least one other GPU-accessible memory of the remote system is accessible by the multiple respective graphics processing units of the remote system.

32. The data center system of claim 28 , wherein:

the programmable circuitry is also for use in association with memory isolation.

33. The data center system of claim 28 , wherein:

at least one application specific integrated circuit (ASIC) comprises the programmable circuitry; and

the at least one host fabric comprises at least one accelerator fabric.

Continuity (3)
Continuation 17129756 · Dec 21, 2020
Division 16435328 · Jun 7, 2019
Related Publication 20240314072A1 · Sep 19, 2024
References Cited (84)
US 7945721B1 · Johnsen et al. · 2011 [cited by applicant]
US 8914556B2 · Magro et al. · 2014 [cited by applicant]
US 10459847B1 · Shah et al. · 2019 [cited by applicant]
US 10509764B1 · Izenberg · 2019 [cited by examiner]
US 10523675B2 · Bhabbur · 2019 [cited by examiner]
US 10579591B1 · Diamant et al. · 2020 [cited by applicant]
US 10880204B1 · Shalev · 2020 [cited by examiner]
US 11444886B1 · Stawitzky · 2022 [cited by examiner]
US 20020032876A1 · Okagaki et al. · 2002 [cited by applicant]
US 20020144001A1 · Collins et al. · 2002 [cited by applicant]
US 20060034310A1 · Connor · 2006 [cited by applicant]
US 20060149919A1 · Arizpe et al. · 2006 [cited by applicant]
US 20070143395A1 · Uehara et al. · 2007 [cited by applicant]
US 20070162641A1 · Oztaskin et al. · 2007 [cited by applicant]
US 20090138645A1 · Chun et al. · 2009 [cited by applicant]
US 20110179199A1 · Chen et al. · 2011 [cited by applicant]
US 20120287944A1 · Pandit · 2012 [cited by examiner]
US 20130138758A1 · Cohen et al. · 2013 [cited by applicant]
US 20150163014A1 · Birrittella · 2015 [cited by applicant]
US 20150180782A1 · Rimmer et al. · 2015 [cited by applicant]
US 20150189047A1 · Naaman · 2015 [cited by examiner]
US 20150324306A1 · Chudgar et al. · 2015 [cited by applicant]
US 20160077976A1 · Raikin et al. · 2016 [cited by applicant]
US 20160154756A1 · Dodson et al. · 2016 [cited by applicant]
US 20160188527A1 · Cherian · 2016 [cited by examiner]
US 20160342547A1 · Liss et al. · 2016 [cited by applicant]
US 20170075857A1 · Izenberg et al. · 2017 [cited by applicant]
US 20170249079A1 · Mutha et al. · 2017 [cited by applicant]
US 20170371789A1 · Blaner et al. · 2017 [cited by applicant]
US 20180285288A1 · Bernat et al. · 2018 [cited by applicant]
US 20180329650A1 · Guim Bernat · 2018 [cited by examiner]
US 20190043064A1 · Chin · 2019 [cited by examiner]
US 20190044738A1 · Liu · 2019 [cited by examiner]
US 20190079895A1 · Kim · 2019 [cited by examiner]
US 20190087218A1 · Loftus et al. · 2019 [cited by applicant]
US 20190098432A1 · Carlson · 2019 [cited by examiner]
US 20190102311A1 · Gupta et al. · 2019 [cited by applicant]
US 20190171612A1 · Shahar et al. · 2019 [cited by applicant]
US 20190236022A1 · Gopal et al. · 2019 [cited by applicant]
US 20190297015A1 · Marolia et al. · 2019 [cited by applicant]
US 20190095343A1 · Gopal · 2019 [cited by applicant]
US 20190332879A1 · Rosenkrantz et al. · 2019 [cited by applicant]
US 20200233717A1 · Smith et al. · 2020 [cited by applicant]
US 20200257545A1 · Ogawa et al. · 2020 [cited by applicant]
US 20200293499A1 · Kohli et al. · 2020 [cited by applicant]
US 20200387611A1 · Yao · 2020 [cited by examiner]
US 20210078571A1 · Zhu · 2021 [cited by examiner]
US 20210389741A1 · Iwami · 2021 [cited by applicant]
EP 2722767A1 · 2014 [cited by applicant]
JP 2014509106A · 2014 [cited by applicant]
KR 20180134745A · 2018 [cited by applicant]
WO 2018102416A1 · 2018 [cited by applicant]
Indian Examination Report for Indian Patent Application No. 202147041932, Mailed Apr. 4, 2024, 8 pages. [cited by applicant]
Alam, Sadaf, Bianco Mauro, “Use Cases and Status of GPUDirect: A High Productivity & Performance Feature for GPU Supercomputing”, Swiss National Supercomputing Centre (CSCS), 2012, 25 pages. [cited by applicant]
European First Office Action, (EP Exam Report Article 94(3) EPC), for Patent Application No. 20165324.3, Mailed Jun. 21, 2022, 5 pages. [cited by applicant]
European Second Office Action, (EP Exam Report Article 94(3) EPC), for Patent Application No. 20165324.3, Mailed Jan. 4, 2024, 4 pages. [cited by applicant]
Extended European Search Report for Patent Application No. 20165324.3, Mailed Sep. 18, 2020, 10 pages. [cited by applicant]
Final Office Action for U.S. Appl. No. 17/129,756, Mailed Mar. 17, 2023, 24 pages. [cited by applicant]
First Office Action for U.S. Appl. No. 16/435,328, Mailed Jul. 23, 2020, 29 pages. [cited by applicant]
First Office Action for U.S. Appl. No. 17/129,756, Mailed Oct. 12, 2022, 20 pages. [cited by applicant]
Humair's Blogs, “Separate Control Plane and Data Plane”, 6WIND-From Data Plane Acceleration to Virtual Appliances, Humair Ahmed.com, retrieved from: http://humairahmed.com/blog/?tag=separate-control-plane-and-data-plane… [cited by applicant]
International Search Report and Written Opinion for PCT Patent Application No. PCT/US20/23934, Mailed Jul. 9, 2020, 9 pages. [cited by applicant]
Microsoft Docs, “Scale up a Service Fabric cluster primary node type”, retrieved from https://https://docs.microsoft.com/en-us/azure/service-fabric/service-fabric-scale-up-node-type, Feb. 12, 2019, 4 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/435,328, Mailed Feb. 1, 2021, 20 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/129,756, Mailed Oct. 17, 2023, 15 pages. [cited by applicant]
Recio, R. et al., “A Remote Direct Memory Access Protocol Specification”, Network Working Group, RFC 5040, Standards Track, Oct. 2007, 66 pages. [cited by applicant]
Recio, Renato, “A Tutorial of the RDMA Model”, retrieved from https://www.hpcwire.com/2006/09/15/a_tutorial_of_the_rdma_model-1/, Sep. 15, 2006, 20 pages. [cited by applicant]
Russell, Robert D., “Introduction to RDMA Programming”, InterOperability Laboratory and Computer Science Department, University of New Hampshire, New Hampshire, USA, Iol, 2012, 76 pages. [cited by applicant]
SDN Tutorials, “Difference Between Control Plane and Data Plane”, Software Defined Networking for Beginners, http://sdntutorials.com/difference-between-control-plane-and-data-plane/, Copyright 2019, 6 pages. [cited by applicant]
Techopedia, “Definition-What does Control Plane mean?”, https://www.techopedia.com/definition/32317/control-plane, 2 pages. [cited by applicant]
Wiki-Github, “RDMA (Remote Direst Memory Access) Tutorial”, retrieved from: https://github.com/jcxue/RDMA- Tutorial/wiki, Mar. 2017, 11 pages. [cited by applicant]
Wikipedia, “Control Plane”, retrieved from: https://en.wikipedia.org/w/index.php?title=Control_plane&oldid=885661998, last edited on Mar. 1, 2019, 3 pages. [cited by applicant]
Notice of Allowance from European Patent Application No. 20165324.3 notified May 27, 2024, 8 pgs. [cited by applicant]
“CUDA Toolkit Archive”, <https://developer.nvidia.com/cuda-toolkit-archive> NVIDIA Developer, 3 pgs. [cited by applicant]
“Does the NVIDIA RDMA GPUDirect always operate only physical addresses (in physical address space of the CPU) ?”, <https://stackoverflow.com/questions/19841815/does-the-nvidia-rdma-gpudirect-always-operate-only-physical… [cited by applicant]
“Mellanox Introduces ConnectX-3 10/40GbE Adapters with Multiple Physical Functions for VMware ESXi 5.0”, Aug. 29, 2011 Press Release, 4 pgs. [cited by applicant]
“Mellanox OFED GPUDirect RDMA”, 2018 Mellanox Technologies, Software Product Brief, 2 pgs. [cited by applicant]
“NVIDIA Mellanox Quantum Hdr 200G Infiniband Switch Silicon”, Mellanox Technologies, Product Brief, Sep. 2020, 3 pgs. [cited by applicant]
“QM8700 Mellanox Quantum HDR Edge Switch”, 2018 Mellanox Technologies, Switch system Product Brief, 2 pgs. [cited by applicant]
“RDMA Aware Networks Programming User Manual”, Rev. 1.7, 2015 Mellanox Technologies, 216 pgs. [cited by applicant]
“Running ROCE over L2 Network Enabled with PFC”, Mellanox Technologies, Rev. 2.0, May 2016, 3 pgs. [cited by applicant]
Daoud, Feras, et al., “Asynchronous Peer-to-Peer Device Communication”, OpenFabrics Alliance Workshop, Mar. 28, 2017, 28 pgs. [cited by applicant]
Rossetti, Davide , “Benchmarking GPUDirect RDMA on Modern Server Platforms”, NVIDIA Technical Blog, Oct. 7, 2014, 12 pgs. [cited by applicant]
Extended European Search Report from European Patent Application No. 24191588.3 notified Nov. 14, 2024, 14 pgs. [cited by applicant]