IP Library Granted Patent US 12,242,413
Granted Patent B2
US 12,242,413 · App. 17/460,146 · Granted Mar 4, 2025

Methods, systems and computer readable media for improving remote direct memory access performance

Inventors: Nitesh Kumar Singh (Mukundapur, IN); Rakhahari Bhunia (Kolkata, IN); Abhijit Singha (Kolkata, IN); Sujoy Nandy (Howrah, IN)
Assignee: KEYSIGHT TECHNOLOGIES, INC.
G06F15/17331G06F13/1621G06F13/1642G06F15/167
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,413
App. No.
17/460,146
Granted
Mar 4, 2025
Kind
B2
Abstract

The subject matter described herein includes methods, systems, and computer readable media for improving remote direct memory access (RDMA) performance. A method for improving RDMA performance occurs at an RDMA node utilizing a user space and a kernel space for executing software. The method includes posting, by an application executing in the user space, an RDMA work request including a data element indicating a plurality of RDMA requests associated with the RDMA work request to be generated by software executing in the kernel space; and generating and sending, by the software executing in the kernel space, the plurality of RDMA requests to or via a system under test (SUT).

Claims (25)

1. A method for improving remote direct memory access (RDMA) performance, the method comprising:

at a remote direct memory access (RDMA) client utilizing a user space and a kernel space for executing software:

posting, by an application executing in the user space, a single RDMA work request including a data element indicating a plurality of RDMA requests associated with the single RDMA work request to be generated by software executing in the kernel space, wherein the data element includes a work queue entry loop count parameter comprising a positive integer indicating the number of RDMA requests in the plurality of RDMA requests;

in response to the single RDMA work request and, using the work queue entry loop count parameter of the single RDMA work request, generating and sending, by the software executing in the kernel space, the plurality of RDMA requests to an RDMA server via a system under test (SUT), wherein the plurality of RDMA requests are identical, the SUT comprises at least one network switch located between the RDMA client and the RDMA server, and sending the RDMA requests to the RDMA server via the SUT includes stress testing the at least one network switch by flooding the at least one network switch with data packets carrying the RDMA requests; and

after determining, by the software executing in kernel space, that operations requested by the plurality of RDMA requests have been completed, posting, by the software executing in kernel space, a notification indicating that processing of the single RDMA work request is completed.

2. The method of claim 1 wherein determining that operations requested by the plurality of RDMA requests have been completed includes receiving a corresponding response for each of the plurality of RDMA requests.

3. The method of claim 1 wherein each of the plurality of RDMA requests is generated and sent without multiple context-switches between the kernel space and the user space.

4. The method of claim 1 wherein the RDMA work request is posted using an InfiniBand (IB) verbs related application programming interface (API).

5. The method of claim 1 wherein the plurality of RDMA requests includes a read operation request or a write operation request.

6. A system for improving remote direct memory access (RDMA) performance, the system comprising:

at least one processor; and

a remote direct memory access (RDMA) client implemented using the at least one processor, the RDMA node utilizing a user space and a kernel space for executing software, wherein the RDMA node is configured for:

posting, by an application executing in the user space, a single RDMA work request including a data element indicating a plurality of RDMA requests associated with the single RDMA work request to be generated by software executing in the kernel space, wherein the data element includes a work queue entry loop count parameter comprising a positive integer indicating the number of RDMA requests in the plurality of RDMA requests;

in response to the single RDMA work request and, using the work queue entry loop count parameter of the single RDMA work request, generating and sending, by the software executing in the kernel space, the plurality of RDMA requests to an RDMA server via a system under test (SUT), wherein the plurality of RDMA requests are identical, the SUT comprises at least one network switch located between the RDMA client and the RDMA server, and sending the RDMA requests to the RDMA server via the SUT includes stress testing the at least one network switch by flooding the at least one network switch with data packets carrying the RDMA requests; and

after determining, by the software executing in kernel space, that operations requested by the plurality of RDMA requests have been completed, posting, by the software executing in kernel space, a notification indicating that processing of the single RDMA work request is completed.

7. The system of claim 6 wherein determining that operations requested by the plurality of RDMA requests have been completed includes receiving a corresponding response for each of the plurality of RDMA requests.

8. The system of claim 6 wherein each of the plurality of RDMA requests is generated and sent without multiple context-switches between the kernel space and the user space.

9. The system of claim 6 wherein the RDMA work request is posted using an InfiniBand (IB) verbs related application programming interface (API).

10. The system of claim 6 wherein the plurality of RDMA requests includes a read operation request or a write operation request.

11. A non-transitory computer readable medium having stored thereon executable instructions that when executed by at least one processor of at least one computer cause the at least one computer to perform steps comprising:

at a remote direct memory access (RDMA) client utilizing a user space and a kernel space for executing software:

posting, by an application executing in the user space, a single RDMA work request including a data element indicating a plurality of RDMA requests associated with the single RDMA work request to be generated by software executing in the kernel space, wherein the data element includes a work queue entry loop count parameter comprising a positive integer indicating the number of RDMA requests in the plurality of RDMA requests;

in response to the single RDMA work request and, using the work queue entry loop count parameter of the single RDMA work request, generating and sending, by the software executing in the kernel space, the plurality of RDMA requests to an RDMA server via a system under test (SUT), wherein the plurality of RDMA requests are identical, the SUT comprises at least one network switch located between the RDMA client and the RDMA server, and sending the RDMA requests to the RDMA server via the SUT includes stress testing the at least one network switch by flooding the at least one network switch with data packets carrying the RDMA requests; and

after determining, by the software executing in kernel space, that operations requested by the plurality of RDMA requests have been completed, posting, by the software executing in kernel space, a notification indicating that processing of the single RDMA work request is completed.

12. The non-transitory computer readable medium of claim 11 wherein each of the plurality of RDMA requests is generated and sent without multiple context-switches between the kernel space and the user space.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: SINGH, NITESH KUMAR; BHUNIA, RAKHAHARI; SINGHA, ABHIJIT; NANDY, SUJOY
To: KEYSIGHT TECHNOLOGIES, INC.
Reel/Frame 059101/0225 →
Continuity (1)
Related Publication 20230066835A1 · Mar 2, 2023
References Cited (98)
US 7448049B1 · Xing · 2008 [cited by examiner]
US 8526938B1 · Gardner · 2013 [cited by examiner]
US 9037640B2 · Boddukuri et al. · 2015 [cited by applicant]
US 9842083B2 · Tsirkin · 2017 [cited by applicant]
US 10534719B2 · Beard · 2020 [cited by examiner]
US 10983704B1 · Van Gaasbeck · 2021 [cited by examiner]
US 11258719B1 · Sommers · 2022 [cited by examiner]
US 20020141424A1 · Gasbarro · 2002 [cited by examiner]
US 20050223274A1 · Bernick · 2005 [cited by examiner]
US 20050246578A1 · Bruckert · 2005 [cited by examiner]
US 20050246587A1 · Bernick · 2005 [cited by examiner]
US 20060020852A1 · Bernick · 2006 [cited by examiner]
US 20060031600A1 · Ellis · 2006 [cited by examiner]
US 20060047867A1 · Craddock · 2006 [cited by examiner]
US 20060075067A1 · Blackmore · 2006 [cited by examiner]
US 20070083680A1 · King · 2007 [cited by examiner]
US 20070282967A1 · Fineberg · 2007 [cited by examiner]
US 20080005385A1 · Lubbers · 2008 [cited by examiner]
US 20080013448A1 · Horie · 2008 [cited by examiner]
US 20080126509A1 · Subramanian · 2008 [cited by examiner]
US 20080147822A1 · Benhase · 2008 [cited by examiner]
US 20090106771A1 · Benner · 2009 [cited by examiner]
US 20100281201A1 · O'Brien · 2010 [cited by examiner]
US 20110106905A1 · Frey · 2011 [cited by examiner]
US 20110145211A1 · Gerber · 2011 [cited by examiner]
US 20110242993A1 · Gotou · 2011 [cited by examiner]
US 20120066179A1 · Saika · 2012 [cited by examiner]
US 20120331243A1 · Aho · 2012 [cited by examiner]
US 20130066923A1 · Sivasubramanian · 2013 [cited by examiner]
US 20140181232A1 · Manula · 2014 [cited by examiner]
US 20150280972A1 · Sivan · 2015 [cited by examiner]
US 20160026605A1 · Pandit · 2016 [cited by examiner]
US 20160119422A1 · Aslam · 2016 [cited by examiner]
US 20160248628A1 · Pandit · 2016 [cited by examiner]
US 20170103039A1 · Shamis · 2017 [cited by examiner]
US 20170149890A1 · Shamis · 2017 [cited by examiner]
US 20170168986A1 · Sajeepa · 2017 [cited by examiner]
US 20180102978A1 · Shen · 2018 [cited by examiner]
US 20180314657A1 · Chen · 2018 [cited by examiner]
US 20190121889A1 · Gold · 2019 [cited by examiner]
US 20190219999A1 · Lam · 2019 [cited by examiner]
US 20190354406A1 · Ganguli et al. · 2019 [cited by applicant]
US 20190386924A1 · Srinivasan · 2019 [cited by examiner]
US 20200067792A1 · Aktas · 2020 [cited by examiner]
US 20200097323A1 · Nider · 2020 [cited by examiner]
US 20200120029A1 · Sankaran et al. · 2020 [cited by applicant]
US 20200145480A1 · Sohail · 2020 [cited by examiner]
US 20200241927A1 · Yang · 2020 [cited by examiner]
US 20200296303A1 · Yamashita · 2020 [cited by examiner]
US 20200319812A1 · He · 2020 [cited by examiner]
US 20200326971A1 · Yang · 2020 [cited by applicant]
US 20200358701A1 · Su · 2020 [cited by examiner]
US 20200366608A1 · Pan et al. · 2020 [cited by applicant]
US 20200396288A1 · Zhang · 2020 [cited by examiner]
US 20210004165A1 · Benisty · 2021 [cited by examiner]
US 20210112002A1 · Pan et al. · 2021 [cited by applicant]
US 20220060422A1 · Sommers · 2022 [cited by examiner]
US 20220103484A1 · Penaranda Cebrian · 2022 [cited by examiner]
US 20220103536A1 · Kida · 2022 [cited by examiner]
US 20220164510A1 · Reid · 2022 [cited by examiner]
US 20220165423A1 · Luber · 2022 [cited by examiner]
US 20220179809A1 · Venkataramani · 2022 [cited by examiner]
US 20220197552A1 · Corbetta · 2022 [cited by examiner]
US 20220229768A1 · Martin · 2022 [cited by examiner]
US 20230061873A1 · Margolin · 2023 [cited by examiner]
US 20230066835A1 · Singh · 2023 [cited by examiner]
US 20230305747A1 · Subramanian · 2023 [cited by examiner]
US 20230401005A1 · Muthiah · 2023 [cited by examiner]
Notice of Allowance and Fee(s) Due for U.S. Appl. No. 17/001,614 (Sep. 29, 2021). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/001,614 (Apr. 23, 2021). [cited by applicant]
Commonly-assigned, co-pending U.S. Appl. No. 17/001,614 for “Methods, Systems and Computer Readable Media for Network Congestion Control Tuning,” (Unpublished, filed Aug. 24, 2020). [cited by applicant]
Beltman et al., “Collecting telemetry data using P4 and RDMA,” University of Amsterdam, pp. 1-12 (2020). [cited by applicant]
Liu et al., “HPCC++: Enhanced High Precision Congestion Control,” Network Working Group, pp. 1-15 (Jun. 17, 2020). [cited by applicant]
“Traffic Management User Guide (QFX Series and EX4600 Switches),” Juniper Networks, pp. 1-1121 (Mar. 18, 2020). [cited by applicant]
“H3C S6850 Series Data Center Switches,” New H3C Technologies Co., Limited, pp. 1-13 (Mar. 2020). [cited by applicant]
Even et al., “Data Center Fast Congestion Management,” pp. 1-15 (Oct. 23, 2019). [cited by applicant]
Li et al., “HPCC: High Precision Congestion Control,” SIGCOMM '19, pp. 1-15 (Aug. 19-23, 2019). [cited by applicant]
“RoCE Congestion Control Interoperability Perception vs. Reality,” Broadcom White Paper, pp. 1-8 (Jul. 23, 2019). [cited by applicant]
“What is RDMA?,” Mellanox, pp. 1-3 (Apr. 7, 2019). [cited by applicant]
Mandal, “In-band Network Telemetry—More Insight into the Network,” Ixia, https://www.ixiacom.com/company/blog/band-network-telemetry-more-insight-network, pp. 1-9 (Mar. 1, 2019). [cited by applicant]
Geng et al., “P4QCN: Congestion Control Using P4-Capable Device in Data Center Networks,” Electronics, vol. 8, No. 280, pp. 1-17 (Mar. 2, 2019). [cited by applicant]
“Understanding DC-QCN Algorithm for RoCE Congestion Control,” Mellanox, pp. 1-4 (Dec. 5, 2018). [cited by applicant]
“Data Center Quantized Congestion Notification (DCQCN),” Juniper Networks, pp. 1-7 (Oct. 4, 2018). [cited by applicant]
“Understanding RoCEv2 Congestion Management,” Mellanox, https://community.mellanox.com/s/article/understanding-rocev2-congestion-management, pp. 1-6 (Dec. 3, 2018). [cited by applicant]
Mittal et al., “Revisiting Network Support for RDMA,” SIGCOMM '18, pp. 1-14 (Aug. 20-25, 2018). [cited by applicant]
Varadhan et al., “Validating ROCEV2 in the Cloud Datacenter,” OpenFabrics Alliance, 13th Annual Workshop 2017, pp. 1-17 (Mar. 31, 2017). [cited by applicant]
Zhu et al., “ECN or Delay: Lessons Learnt from Analysis of DCQCN and TIMELY,” CoNEXT '16, pp. 1-15 (Dec. 12-15, 2016). [cited by applicant]
Kim et al., “In-band Network Telemetry (INT),” pp. 1-28 (Jun. 2016). [cited by applicant]
Zhu et al., “Congestion Control for Large-Scale RDMA Deployments,” SIGCOMM '15, pp. 1-14 (Aug. 17-21, 2015). [cited by applicant]
Zhu et al., “Packet-Level Telemetry in Large Datacenter Networks,” SIGCOMM '15, pp. 1-13 (Aug. 17-21, 2015). [cited by applicant]
Mittal et al., “TIMELY: RTT-based Congestion Control for the Datacenter,” SIGCOMM '15, pp. 1-14 (Aug. 17-21, 2015). [cited by applicant]
“RoCE in the Data Center,” Mellanox Technologies, White Paper, pp. 1-3 (Oct. 2014). [cited by applicant]
Barak, “Introduction to Remote Direct Memory Access (RDMA),” http://www.rdmamojo.com/2014/03/31/remote-direct-memory-access-rdma/, pp. 1-14 (Mar. 31, 2014). [cited by applicant]
“Quick Concepts Part 1—Introduction to RDMA,” ZCopy, Education and Sample Code for RDMA Programming, pp. 1-5 (Oct. 8, 2010). [cited by applicant]
Alizadeh et al., “Data Center TCP (DCTCP),” SIGCOMM '10, pp. 1-12 (Aug. 30-Sep. 3, 2010). [cited by applicant]
Grochla, “Simulation comparison of active queue management algorithms in TCP/IP networks,” Telecommunication Systems, pp. 1-9 (Oct. 2008). [cited by applicant]
Chen et al., “Data Center Congestion Management requirements,” https://tools.ietf.org/id/draft-yueven-tsvwg-dccm-requirements-01.html, pp. 1-7 (Jul. 2019). [cited by applicant]
Kalia et al., “Using RDMA Efficiently for Key-Value Services,” SIGCOMM'14, pp. 1-12 (Aug. 17-22, 2014). [cited by applicant]