IP Library Granted Patent US 10,298,494
Granted Patent B2
US 10,298,494 · App. 14/936,712 · Granted May 21, 2019

Reducing short-packet overhead in computer clusters

Inventors: Liaz Kamper (Ra'anana, IL); Vadim Suraev (Ariel, IL)
Assignee: Strato Scale Ltd.
H04L45/74H04L12/4633H04L47/50H04L47/70H04L69/166H04L67/2861
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,298,494
App. No.
14/936,712
Granted
May 21, 2019
Kind
B2
Abstract

A method includes, in a computing system that includes multiple compute nodes that run workloads and are connected by a network, establishing a dedicated Transport Control Protocol (TCP) connection over the network between a first compute node and a second compute node. Packets, which originate from one or more source workloads on the first compute node and are destined to one or more destination workloads on the second compute node, are identified and queued in the first compute node. The queued packets are aggregated in the first compute node into one or more TCP segments, and the TCP segments are sent over the dedicated TCP connection to the second compute node. In the second compute node, the TCP segments are received over the dedicated TCP connection, the packets are extracted from the received TCP segments, and the extracted packets are forwarded to the destination workloads.

Claims (28)

1. A method, comprising:

in a computing system that includes multiple compute nodes, which run workloads and are connected by a network, establishing a dedicated Transport Control Protocol (TCP) connection over the network between a first compute node and a second compute node;

in the first compute node, identifying and queuing packets that originate from one or more source workloads on the first compute node and are destined to one or more destination workloads on the second compute node;

aggregating the queued packets in the first compute node into one or more TCP segments, and sending the TCP segments over the dedicated TCP connection to the second compute node; and

in the second compute node, receiving the TCP segments over the dedicated TCP connection, extracting the packets from the received TCP segments, and forwarding the extracted packets to the destination workloads,

wherein aggregating the queued packets in the first compute node comprises, in response to identifying among the packets at least a predefined number of duplicate TCP acknowledgements belonging to a same TCP connection, terminating aggregation of the packets into a current TCP segment.

2. The method according to claim 1 , wherein aggregating the queued packets comprises generating a TCP segment that jointly encapsulates at least two packets that are destined to different destination workloads on the second compute node.

3. The method according to claim 1 , wherein aggregation of the queued packets into the TCP segments is performed by a Central Processing Unit (CPU) of the first compute node.

4. The method according to claim 1 , wherein aggregation of the queued packets into the TCP segments is performed by a Network Interface Controller (NIC) in the first compute node, while offloading a Central Processing Unit (CPU) of the first compute node.

5. The method according to claim 1 , wherein aggregating the queued packets comprises deciding to terminate aggregation of the packets into a current TCP segment, by evaluating a predefined criterion.

6. The method according to claim 1 , wherein identifying the packets destined to the second compute node comprises querying a mapping that maps Medium Access Control (MAC) addresses to compute-node identifiers.

7. The method according to claim 1 , wherein queuing the packets comprises identifying an Address Resolution Protocol (ARP) packet that duplicates another ARP packet that is already queued, and discarding the identified ARP packet.

8. The method according to claim 1 , wherein queuing the packets comprises identifying a TCP packet that is a retransmission of another TCP packet that is already queued, and discarding the identified TCP packet.

9. The method according to claim 1 , wherein aggregating the queued packets comprises, in response to identifying a TCP packet that is a retransmission of another TCP packet that is already queued, terminating aggregation of the packets into a current TCP segment and sending the current TCP segment to the second compute node.

10. The method according to claim 5 , wherein deciding to terminate the aggregation comprises constraining a maximum delay incurred by queuing the packets.

11. The method according to claim 5 , wherein deciding to terminate the aggregation comprises constraining a maximum total data volume of the queued packets.

12. The method according to claim 6 , wherein querying the mapping comprises extracting destination MAC addresses from the packets generated in the first compute node, and selecting, using the mapping, the packets whose destination MAC addresses are mapped to a compute-node identifier of the second compute node.

13. A method, comprising:

in a computing system that includes multiple compute nodes, which run workloads and are connected by a network, establishing a dedicated Transport Control Protocol (TCP) connection over the network between a first compute node and a second compute node;

in the first compute node, identifying and queuing packets that originate from one or more source workloads on the first compute node and are destined to one or more destination workloads on the second compute node;

aggregating the queued packets in the first compute node into one or more TCP segments, and sending the TCP segments over the dedicated TCP connection to the second compute node; and

in the second compute node, receiving the TCP segments over the dedicated TCP connection, extracting the packets from the received TCP segments, and forwarding the extracted packets to the destination workloads,

wherein extracting the packets in the second compute node comprises, in response to identifying among the packets at least a predefined number of duplicate TCP acknowledgements belonging to a same TCP connection, delivering the duplicate TCP acknowledgements immediately to a respective destination workload.

14. The method according to claim 13 , wherein aggregating the queued packets in the first compute node comprises, in response to identifying among the packets at least a predefined number of duplicate TCP acknowledgements belonging to a same TCP connection, terminating aggregation of the packets into a current TCP segment.

15. A compute node, comprising:

a Network Interface Controller (NIC) for communicating over a network; and

a Central Processing Unit (CPU), which is configured to run one or more source workloads, to establish a dedicated Transport Control Protocol (TCP) connection over the network with a peer compute node, to identify and queue packets that originate from the source workloads and are destined to one or more destination workloads on the peer compute node, and, in conjunction with the NIC, to aggregate the queued packets into one or more TCP segments and send the TCP segments over the dedicated TCP connection to the peer compute node,

wherein the CPU is configured, in response to identifying among the packets at least a predefined number of duplicate TCP acknowledgements belonging to a same TCP connection, to terminate aggregation of the packets into a current TCP segment.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2020
From: STRATO SCALE LTD.
To: MELLANOX TECHNOLOGIES, LTD.
Reel/Frame 053184/0620 →
SECURITY INTEREST Recorded Jan 24, 2019
From: STRATO SCALE LTD.
To: KREOS CAPITAL VI (EXPERT FUND) L.P.
Reel/Frame 048115/0134 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2015
From: KAMPER, LIAZ; SURAEV, VADIM
To: STRATO SCALE LTD.
Reel/Frame 036997/0736 →
Continuity (3)
Continuation PCTIB2015058524 · Nov 4, 2015
Provisional Application 62081593 · Nov 19, 2014
Related Publication 20160142307A1 · May 19, 2016
Cited By (2)
US 12,206,578 US 12,425,365