Reliably forwarding endpoint processing unit computations through a network fabric
Some embodiments provide a method of executing a distributed application with multiple endpoint processing units (EPUs) that perform computations for the distributed application. The EPUs are connected through a network having multiple network elements. The method iteratively provides instructions to the EPUs to perform computations associated with the distributed application. The method stores a result of each EPU computation at the EPU until the result has been forwarded to a destination in the network and a confirmation has been received that the result has successfully been received at its destination in the network.
1 . A method of executing a distributed application with a plurality of graphics processing units (GPUs) that perform computations for the distributed application, the GPUs connected through a network comprising a plurality of network elements, the method comprising:
configuring each GPU to store a result of each GPU computation in a memory of the GPU until the result has been forwarded to a destination in the network and a confirmation has been received that the result has successfully been received at its destination in the network, wherein configuring each GPU comprises configuring a process operating on the GPU (i) to notify a network interface of the GPU each time that the GPU completes a computation and stores the result of the computation in the GPU memory, for the network interface to retrieve the result from the GPU memory and to forward the result through the network to the result's destination, and (ii) to discard each result from the GPU memory after receiving confirmation that the result has successfully been received at its destination in the network; and
iteratively providing instructions to the GPUs to perform computations associated with the distributed application.
2 . The method of claim 1 further comprising configuring each GPU's network interface with forwarding records that specify the network interface's forwarding of the GPU's results through the network.
3 . The method of claim 1 further comprising configuring each GPU's network interface to notify the GPU's process that the result of a GPU's computation has successfully been received at its destination in the network.
4 . The method of claim 1 , wherein for each result, the confirmation is received as an acknowledgment from the destination that the result has been completely received at the destination.
5 . The method of claim 4 , wherein for each result, the acknowledgment is sent from a network interface that connects the destination to the network.
6 . A method of executing a distributed application with a plurality of endpoint processing units (EPUs) that perform computations for the distributed application, the EPUs connected through a network comprising a plurality of network elements, the method comprising:
iteratively providing instructions to the EPUs to perform computations associated with the distributed application; and
storing a result of each EPU computation at the EPU until the result has been forwarded to a destination in the network and an acknowledgment has been received from the destination of the result in the network that the result has been completely received at the destination, wherein for each result, the acknowledgment is sent from a network interface that connects the destination to the network,
wherein an EPU interface of each EPU is configured to send each result of each computation of the EPU as a plurality of segments in payloads of data messages in a data message flow, each segment sent with a segment identifier,
wherein the network interface of each destination is configured to use the segment identifiers of each data message flow for each result to determine when the destination network interface has received all the segments of the result and to send the acknowledgment for the result after determining that all segments of the result have been received.
7 . The method of claim 6 , wherein the network interface of each destination is further configured to identify any segment that has not been received for the network interface of the EPU that was a source of the result to retransmit the identified segment, said retransmission making forwarding of the EPU computation results through the network reliable as the retransmission ensures that the forwarded computation results are fully received at their destinations before being discarded.
8 . The method of claim 1 , wherein each EPU is configured to discard each stored result after the confirmation has been received for the result.
9 . The method of claim 8 , wherein by discarding the result only after the confirmation is received for the result, the result does not get lost or does not have to be maintained at one or more intermediate nodes in the network.
10 . A method of executing a distributed application with a plurality of endpoint processing units (EPUs) that perform computations for the distributed application, the EPUs connected through a network comprising a plurality of network elements, the method comprising:
iteratively providing instructions to the EPUs to perform computations associated with the distributed application; and
storing a result of each EPU computation at the EPU until the result has been forwarded to a destination in the network and a confirmation has been received that the result has successfully been received at its destination in the network,
wherein for a first result computed by a first EPU, a network interface of the first EPU that connects the first EPU with the network is configured to forward the first result to a first destination in the network after receiving a first notification that the first result has been stored in a memory of the first EPU,
wherein for a second result computed by the first EPU, the first EPU's network interface is configured to forward the second result to a second destination in the network only after (i) receiving a second notification that the second result has been stored in first EPU's memory and (ii) after receiving the second notification, requesting and then receiving scheduling parameters for governing at least one of timing or rate of the forwarding of the second result through the network to the second destination.
11 . A non-transitory machine readable medium storing a program that when executed by a processor reliably forwards results of a particular graphics processing unit (GPU) through a network that connects a plurality of GPUs, the program comprising sets of instructions for:
detecting that the particular GPU has performed a computation that has produced a result stored in a memory of the particular GPU;
communicating with a network interface of the particular GPU to direct the network interface to forward the result to a destination in the network and to receive confirmation from the network interface that the result has been successfully received at the destination; and
maintaining the result in the particular GPU's memory until the result has been forwarded to the destination in the network and a confirmation has been received that the result has successfully been received at its destination in the network,
wherein the network interface is configured (i) to send the result as a plurality of segments in payloads of data messages in a data message flow, each segment sent with a segment identifier, and (ii) to provide the confirmation after receiving a confirmation from the destination that each segment has been received at the destination.
12 . The non-transitory machine readable medium of claim 11 , wherein the program further comprises a set of instructions for discarding the stored result from the particular GPU's memory after receiving confirmation that the result has successfully been received at its destination in the network.
13 . The non-transitory machine readable medium of claim 11 , wherein the program is a driver or kernel process executed by the particular GPU or a control unit processor of the particular GPU.
14 . The method of claim 6 , wherein the EPUs are graphics processing units (GPUs).
15 . The method of claim 10 , wherein the EPUs are graphics processing units (GPUs).