IP Library › Granted Patent US 12,739,192
Granted Patent B2
US 12,739,192 · App. 19/227,462 · Granted Sep 15, 2026

Reliably forwarding endpoint processing unit computations through a network fabric

Inventors: Daniel P. Daly (Santa Barbara, CA); Edward V. E. Doe (Palo Alto, CA); Alain J. E. Gravel (Thousand Oaks, CA)
Assignee: DELOS DATA INC.
H04L45/24H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,739,192
App. No.
19/227,462
Granted
Sep 15, 2026
Kind
B2
Abstract

Some embodiments provide a method of executing a distributed application with multiple endpoint processing units (EPUs) that perform computations for the distributed application. The EPUs are connected through a network having multiple network elements. The method iteratively provides instructions to the EPUs to perform computations associated with the distributed application. The method stores a result of each EPU computation at the EPU until the result has been forwarded to a destination in the network and a confirmation has been received that the result has successfully been received at its destination in the network.

Claims (29)

1 . A method of executing a distributed application with a plurality of graphics processing units (GPUs) that perform computations for the distributed application, the GPUs connected through a network comprising a plurality of network elements, the method comprising:

configuring each GPU to store a result of each GPU computation in a memory of the GPU until the result has been forwarded to a destination in the network and a confirmation has been received that the result has successfully been received at its destination in the network, wherein configuring each GPU comprises configuring a process operating on the GPU (i) to notify a network interface of the GPU each time that the GPU completes a computation and stores the result of the computation in the GPU memory, for the network interface to retrieve the result from the GPU memory and to forward the result through the network to the result's destination, and (ii) to discard each result from the GPU memory after receiving confirmation that the result has successfully been received at its destination in the network; and

iteratively providing instructions to the GPUs to perform computations associated with the distributed application.

2 . The method of claim 1 further comprising configuring each GPU's network interface with forwarding records that specify the network interface's forwarding of the GPU's results through the network.

3 . The method of claim 1 further comprising configuring each GPU's network interface to notify the GPU's process that the result of a GPU's computation has successfully been received at its destination in the network.

4 . The method of claim 1 , wherein for each result, the confirmation is received as an acknowledgment from the destination that the result has been completely received at the destination.

5 . The method of claim 4 , wherein for each result, the acknowledgment is sent from a network interface that connects the destination to the network.

6 . A method of executing a distributed application with a plurality of endpoint processing units (EPUs) that perform computations for the distributed application, the EPUs connected through a network comprising a plurality of network elements, the method comprising:

iteratively providing instructions to the EPUs to perform computations associated with the distributed application; and

storing a result of each EPU computation at the EPU until the result has been forwarded to a destination in the network and an acknowledgment has been received from the destination of the result in the network that the result has been completely received at the destination, wherein for each result, the acknowledgment is sent from a network interface that connects the destination to the network,

wherein an EPU interface of each EPU is configured to send each result of each computation of the EPU as a plurality of segments in payloads of data messages in a data message flow, each segment sent with a segment identifier,

wherein the network interface of each destination is configured to use the segment identifiers of each data message flow for each result to determine when the destination network interface has received all the segments of the result and to send the acknowledgment for the result after determining that all segments of the result have been received.

7 . The method of claim 6 , wherein the network interface of each destination is further configured to identify any segment that has not been received for the network interface of the EPU that was a source of the result to retransmit the identified segment, said retransmission making forwarding of the EPU computation results through the network reliable as the retransmission ensures that the forwarded computation results are fully received at their destinations before being discarded.

8 . The method of claim 1 , wherein each EPU is configured to discard each stored result after the confirmation has been received for the result.

9 . The method of claim 8 , wherein by discarding the result only after the confirmation is received for the result, the result does not get lost or does not have to be maintained at one or more intermediate nodes in the network.

10 . A method of executing a distributed application with a plurality of endpoint processing units (EPUs) that perform computations for the distributed application, the EPUs connected through a network comprising a plurality of network elements, the method comprising:

iteratively providing instructions to the EPUs to perform computations associated with the distributed application; and

storing a result of each EPU computation at the EPU until the result has been forwarded to a destination in the network and a confirmation has been received that the result has successfully been received at its destination in the network,

wherein for a first result computed by a first EPU, a network interface of the first EPU that connects the first EPU with the network is configured to forward the first result to a first destination in the network after receiving a first notification that the first result has been stored in a memory of the first EPU,

wherein for a second result computed by the first EPU, the first EPU's network interface is configured to forward the second result to a second destination in the network only after (i) receiving a second notification that the second result has been stored in first EPU's memory and (ii) after receiving the second notification, requesting and then receiving scheduling parameters for governing at least one of timing or rate of the forwarding of the second result through the network to the second destination.

11 . A non-transitory machine readable medium storing a program that when executed by a processor reliably forwards results of a particular graphics processing unit (GPU) through a network that connects a plurality of GPUs, the program comprising sets of instructions for:

detecting that the particular GPU has performed a computation that has produced a result stored in a memory of the particular GPU;

communicating with a network interface of the particular GPU to direct the network interface to forward the result to a destination in the network and to receive confirmation from the network interface that the result has been successfully received at the destination; and

maintaining the result in the particular GPU's memory until the result has been forwarded to the destination in the network and a confirmation has been received that the result has successfully been received at its destination in the network,

wherein the network interface is configured (i) to send the result as a plurality of segments in payloads of data messages in a data message flow, each segment sent with a segment identifier, and (ii) to provide the confirmation after receiving a confirmation from the destination that each segment has been received at the destination.

12 . The non-transitory machine readable medium of claim 11 , wherein the program further comprises a set of instructions for discarding the stored result from the particular GPU's memory after receiving confirmation that the result has successfully been received at its destination in the network.

13 . The non-transitory machine readable medium of claim 11 , wherein the program is a driver or kernel process executed by the particular GPU or a control unit processor of the particular GPU.

14 . The method of claim 6 , wherein the EPUs are graphics processing units (GPUs).

15 . The method of claim 10 , wherein the EPUs are graphics processing units (GPUs).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2025
From: DALY, DANIEL P.; DOE, EDWARD V. E.; GRAVEL, ALAIN J. E.
To: DELOS DATA INC.
Reel/Frame 072010/0943 →
Continuity (6)
Provisional Application 63777883 · Mar 26, 2025
Provisional Application 63714118 · Oct 30, 2024
Provisional Application 63707702 · Oct 15, 2024
Provisional Application 63698036 · Sep 23, 2024
Provisional Application 63697485 · Sep 21, 2024
Related Publication 20260089202A1 · Mar 26, 2026
References Cited (94)
US 6101549A · Baugher et al. · 2000 [cited by applicant]
US 7627744B2 · Maher et al. · 2009 [cited by applicant]
US 8738684B1 · Ellis · 2014 [cited by applicant]
US 8743888B2 · Casado et al. · 2014 [cited by applicant]
US 8743889B2 · Koponen et al. · 2014 [cited by applicant]
US 8817620B2 · Koponen et al. · 2014 [cited by applicant]
US 8966035B2 · Casado et al. · 2015 [cited by applicant]
US 9083609B2 · Casado et al. · 2015 [cited by applicant]
US 10200235B2 · Pfaff et al. · 2019 [cited by applicant]
US 10275851B1 · Zhao et al. · 2019 [cited by applicant]
US 10382401B1 · Lee et al. · 2019 [cited by applicant]
US 10728091B2 · Zhao et al. · 2020 [cited by applicant]
US 11080225B2 · Borikar et al. · 2021 [cited by applicant]
US 11477114B2 · Hu · 2022 [cited by applicant]
US 12603829B2 · Clark · 2026 [cited by applicant]
US 20010034771A1 · Hutsch et al. · 2001 [cited by applicant]
US 20020110119A1 · Fredette et al. · 2002 [cited by applicant]
US 20040105440A1 · Strachan et al. · 2004 [cited by applicant]
US 20080059602A1 · Matsuda et al. · 2008 [cited by applicant]
US 20120131252A1 · Rau · 2012 [cited by applicant]
US 20130329549A1 · Kusama et al. · 2013 [cited by applicant]
US 20140269379A1 · Holbrook et al. · 2014 [cited by applicant]
US 20150319009A1 · Zhao · 2015 [cited by applicant]
US 20160124852A1 · Shachar et al. · 2016 [cited by applicant]
US 20170351555A1 · Coffin · 2017 [cited by applicant]
US 20180322387A1 · Sridharan · 2018 [cited by examiner]
US 20190109789A1 · Friedman et al. · 2019 [cited by applicant]
US 20190123894A1 · Yuan · 2019 [cited by applicant]
US 20190132150A1 · Ramachandran et al. · 2019 [cited by applicant]
US 20190158371A1 · Dillon et al. · 2019 [cited by applicant]
US 20190182143A1 · Kim et al. · 2019 [cited by applicant]
US 20190182149A1 · Kim et al. · 2019 [cited by applicant]
US 20200177629A1 · Hooda et al. · 2020 [cited by applicant]
US 20200302568A1 · Li et al. · 2020 [cited by applicant]
US 20210191774A1 · Kanteti et al. · 2021 [cited by applicant]
US 20210219175A1 · Xu et al. · 2021 [cited by applicant]
US 20210318878A1 · Zhao et al. · 2021 [cited by applicant]
US 20210328902A1 · Huselton et al. · 2021 [cited by applicant]
US 20210334234A1 · Yudanov · 2021 [cited by applicant]
US 20220085916A1 · Debbage · 2022 [cited by examiner]
US 20220086096A1 · Sugiyama et al. · 2022 [cited by applicant]
US 20220188688A1 · Launay et al. · 2022 [cited by applicant]
US 20220247696A1 · He · 2022 [cited by examiner]
US 20220351326A1 · Rimmer · 2022 [cited by examiner]
US 20220361262A1 · Liu et al. · 2022 [cited by applicant]
US 20230034757A1 · Wu et al. · 2023 [cited by applicant]
US 20230064808A1 · Thubert et al. · 2023 [cited by applicant]
US 20230214345A1 · Taylor · 2023 [cited by applicant]
US 20230297292A1 · Potyraj et al. · 2023 [cited by applicant]
US 20230325265A1 · Balle et al. · 2023 [cited by applicant]
US 20230327988A1 · Rennie et al. · 2023 [cited by applicant]
US 20240061796A1 · Dastidar et al. · 2024 [cited by applicant]
US 20240069978A1 · Singh et al. · 2024 [cited by applicant]
US 20240086258A1 · Lal · 2024 [cited by examiner]
US 20240098033A1 · Hong et al. · 2024 [cited by applicant]
US 20240129234A1 · Farrokhbakht et al. · 2024 [cited by applicant]
US 20240152409A1 · Brar et al. · 2024 [cited by applicant]
US 20240220336A1 · Punniyamurthy et al. · 2024 [cited by applicant]
US 20240256475A1 · Chen et al. · 2024 [cited by applicant]
US 20240406093A1 · Chuang et al. · 2024 [cited by applicant]
US 20250126071A1 · Brar et al. · 2025 [cited by applicant]
US 20250274380A1 · Barkai et al. · 2025 [cited by applicant]
US 20250291760A1 · Ensey et al. · 2025 [cited by applicant]
US 20250292355A1 · Heinecke et al. · 2025 [cited by applicant]
US 20250321795A1 · Narayanaswamy et al. · 2025 [cited by applicant]
US 20260046317A1 · Crabtree et al. · 2026 [cited by applicant]
US 20260052100A1 · Holl et al. · 2026 [cited by applicant]
US 20260058900A1 · Kommula et al. · 2026 [cited by applicant]
US 20260089203A1 · Daly · 2026 [cited by examiner]
CN 117157953A · 2023 [cited by applicant]
CN 118503194A · 2024 [cited by applicant]
HU E033041T2 · 2017 [cited by applicant]
Author Unknown, “Source Routing,” Wikipedia, Sep. 5, 2024, 3 pages, Wikipedia.com. [cited by applicant]
Bosshart, Pat, et al., “P4: Programming Protocol-Independent Packet Processors,” ACM SIGCOMM Computer Communication Review, vol. 44, No. 3, Jul. 2014, pp. 88-95. [cited by applicant]
Cai, Zheng, et al., “The Preliminary Design and Implementation of the Maestro Network Control Platform,” Oct. 1, 2008, 17 pages, NSF. [cited by applicant]
Casado, Martin, et al. “Ethane: Taking Control of the Enterprise,” SIGCOMM'07, Aug. 27-31, 2007, 12 pages, ACM, Kyoto, Japan. [cited by applicant]
Casado, Martin, et al., “Rethinking Packet Forwarding Hardware,” Seventh ACM SIGCOMM' HotNets Workshop, Nov. 2008, 6 pages, ACM. [cited by applicant]
Casado, Martin, et al., “SANE: A Protection Architecture for Enterprise Networks,” Proceedings of the 15th USENIX Security Symposium, Jul. 31-Aug. 4, 2006, 15 pages, USENIX, Vancouver, Canada. [cited by applicant]
Casado, Martin, et al., “Scaling Out: Network Virtualization Revisited,” Month Unknown 2010, 8 pages. [cited by applicant]
Casado, Martin, et al., “Virtualizing the Network Forwarding Plane,” Dec. 2010, 6 pages. [cited by applicant]
Das, Saurav, et al., “Unifying Packet and Circuit Switched Networks with OpenFlow,” Dec. 7, 2009, 10 pages, available at https://yuba.stanford.edu/~nickm/papers/UArch_Cam_Rdy.pdf. [cited by applicant]
Filsfils, Clarence, et al., “SRv6”, Dec. 5, 2017, Cisco, available at https://www.segment-routing.net/tutorials/2017-12-05-srv6-introduction. [cited by applicant]
Greenberg, Albert, et al., “VL2: A Scalable and Flexible Data Center Network,” SIGCOMM '09, Aug. 17-21, 2009, 12 pages, ACM, Barcelona, Spain. [cited by applicant]
Gude, Natasha, et al., “NOX: Towards an Operating System for Networks,” ACM SIGCOMM Computer Communication Review, Jul. 2008, 6 pages, Vo. 38, No. 3, ACM. [cited by applicant]
Guo, Chanxiong, et al., “BCube: A High Performance, Server-centric Network Architecture for Modular Data Centers,” SIGCOMM'09, Aug. 17-21, 2009, 12 pages, ACM, Barcelona, Spain. [cited by applicant]
Kim, Changhoon, et al., “Floodless in Seattle: A Scalable Ethernet Architecture for Large Enterprises,” SIGCOMM'08, Aug. 17-22, 2008, 12 pages, ACM, Seattle, Washington, USA. [cited by applicant]
Koponen, Teemu, et al., “Onix: A Distributed Control Platform for Large-scale Production Networks,” In Proc. OSDI, Oct. 2010, 14 pages. [cited by applicant]
PCT International Search Report and Written Opinion of Commonly Owned International Patent Application PCT/US2025/047269, mailing date Nov. 25, 2025, 17 pages, International Searching Authority (US). [cited by applicant]
Sherwood, Rob, et al., “FlowVisor: A Network Virtualization Layer,” Oct. 14, 2009, 15 pages, Openflow—TR-2009-1. [cited by applicant]
Tavakoli, Arsalan, et al., “Applying NOX to the Datacenter,” Proc. HotNets, Month Unknown 2009, 6 pages. [cited by applicant]
Yang, L., et al., “Forwarding and Control Element Separation (ForCES) Framework,” Apr. 2004, 41 pages, The Internet Society. [cited by applicant]
Yu, Minlan, et al., “Scalable Flow-Based Networking with DIFANE,” In Proc. SIGCOMM, Aug. 2010, 16 pages. [cited by applicant]
Asterfuison, “AOC, DAC, ACC, Aec Modules: The most Complete Overview”, May 20, 2024, Asterfusion Data Technologies, Co., Ltd., Shenzhen, China, 17 pages. [cited by applicant]
Lebeane, Michael, et al., “GPU Triggered Networking for Intra-Kernel Communications,” The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC '17), Nov. 12-17, 2017, Denver, Co… [cited by applicant]