IP Library › Granted Patent US 12,405,843
Granted Patent B2
US 12,405,843 · App. 17/134,324 · Granted Sep 2, 2025

Infrastructure processing unit

Inventors: Kshitij A. Doshi (Tempe, AZ); Johan Van De Groenendaal (Portland, OR); Edmund Chen (Sunnyvale, CA); Ravi Sahita (Portland, OR); Andrew J. Herdrich (Hillsboro, OR); Debra Bernstein (Sudbury, MA); Christine E. Severns-Williams (Deephaven, MN); Uri V. Cummings (Los Gatos, CA); Utkarsh Y. Kakaiya (Folsom, CA)
Assignee: Intel Corporation
G06F9/522G06F9/5072G06F9/547H04L43/08H04L67/1014H04L67/148
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,405,843
App. No.
17/134,324
Filed
Dec 26, 2020
Granted
Sep 2, 2025
Kind
B2
Art Unit
2195
USPC
718/106
Abstract

Examples described herein relate to an Infrastructure Processing Unit (IPU) that comprises: interface circuitry to provide a communicative coupling with a platform; network interface circuitry to provide a communicative coupling with a network medium; and circuitry to expose infrastructure services to be accessed by microservices for function composition and to selectively provide a barrier to halt operation of at least one microservice based on event data from a composite node that performs the at least one microservice.

Claims (36)

1. A system comprising:

a network interface device that comprises:

a direct memory access (DMA) circuitry;

interface circuitry to provide a communicative coupling with a first host platform;

network interface circuitry to provide a communicative coupling with a network medium to transmit and receive Ethernet packets;

circuitry to:

provide access to distributed nodes to microservices executed by the first host platform, wherein the distributed nodes comprises circuitry of a second host platform and a second network interface device and wherein the second host platform and the second network interface devices are communicatively coupled to the network interface device via the interface circuitry or the network interface circuitry; and

selectively provide, from the network interface device, a barrier to halt operation of at least one microservice of the microservices based on event data from the distributed nodes and based on the event data from the distributed nodes, release the barrier to permit operation of the at least one microservice of the microservices, wherein the release the barrier comprises permit devices of the distributed nodes to execute awaiting microservices.

2. The system of claim 1 , wherein the circuitry is to selectively provide a barrier to halt operation of at least one microservice based on the event data from the distributed nodes to determine a malfunction that could negatively impact performance of one or more currently executed microservices or one or more microservices to be executed.

3. The system of claim 1 , wherein the event data comprises one or more physical events at a device in the distributed nodes.

4. The system of claim 1 , wherein the event data comprises one or more of: voltage droop at a device in the distributed nodes, too high a memory error correction rate at a device in the distributed nodes, excess temperature of a device in the distributed nodes, or intrusion detection at a device in the distributed nodes.

5. The system of claim 1 , wherein the event data comprises one or more performance extrema of an operation of the distributed nodes.

6. The system of claim 1 , wherein the event data comprises one or more of: heartbeat signal timeout from an XPU of the distributed nodes, timeouts reported from hardware that is part of the distributed nodes, or excess memory errors from the distributed nodes.

7. The system of claim 1 , wherein the barrier to halt operation of at least one microservice comprises an acknowledgement that devices in the distributed nodes have entered active standby mode.

8. The system of claim 1 , wherein during application of the barrier, the circuitry is to place devices of the distributed nodes into active-standby mode to execute commands from the network interface device.

9. The system of claim 1 , wherein in response to the barrier, the circuitry is to cause migration of microservices and data from the distributed nodes to another distributed nodes.

10. The system of claim 9 , comprising:

a managed node comprising one or more distributed nodes, wherein the distributed nodes comprises computing resources that are logically coupled, wherein the computing resources include memory devices, data storage devices, accelerator devices, or general purpose processors, wherein resources from the one or more distributed nodes are collectively utilized in execution of a workload and wherein the workload execution is implemented by a group of microservices.

11. A method comprising:

at a network interface device comprising a direct memory access (DMA) circuitry, an interface to a first platform, and a network interface to transmit and receive Ethernet packets:

providing access to distributed nodes to microservices executed by the first platform, wherein the distributed nodes comprises circuitry of a second host platform and a second network interface device and wherein the second host platform and the second network interface device are communicatively coupled to the network interface device via the interface or the network interface;

selectively providing, by the network interface device, a barrier to halt operation of at least one microservice of the microservices based on event data from the distributed nodes; and

based on the event data from the distributed nodes, releasing the barrier and permitting devices of the distributed nodes to execute operation of the at least one microservice of the microservices.

12. The method of claim 11 , wherein the selectively providing a barrier to halt operation of at least one microservice based on the event data from the distributed nodes to determine a malfunction that could negatively impact performance of one or more currently executed microservices or one or more microservices to be executed.

13. The method of claim 11 , wherein the event data comprises one or more physical events at a device in the distributed nodes.

14. The method of claim 11 , wherein the event data comprises one or more of: voltage droop at a device in the distributed nodes, too high a memory error correction rate at a device in the distributed nodes, excess temperature of a device in the distributed nodes, or intrusion detection at a device in the distributed nodes.

15. The method of claim 11 , wherein the event data comprises one or more performance extrema of an operation of the distributed nodes.

16. The method of claim 11 , wherein the event data comprises one or more of: heartbeat signal timeout from an XPU of the distributed nodes, timeouts reported from hardware that is part of the distributed nodes, or excess memory errors from the distributed nodes.

17. The method of claim 11 , wherein during application of the barrier, the network interface device placing devices of the distributed nodes into active-standby mode to execute commands from the network interface device.

18. The method of claim 11 , comprising: the network interface device releasing the barrier and permitting devices of the distributed nodes to execute awaiting microservices.

19. The method of claim 11 , wherein in response to the barrier, the network interface device causing migration of microservices and data from the distributed nodes to another distributed nodes.

20. An apparatus comprising:

a network interface device, comprising a direct memory access (DMA) circuitry, host interface, network interface, and circuitry, wherein the circuitry is to:

manage, at the network interface device, forward progress of microservices executed by distributed nodes by halting forward progress of microservices executed by the distributed nodes and based on event data from the distributed nodes, release a barrier and permit resuming progress of the microservices executed by the distributed nodes based on event data from the distributed nodes.

21. The apparatus of claim 20 , wherein to manage forward progress of microservices executed by distributed nodes, the circuitry is to issue a barrier to halt progress of a first process.

22. The apparatus of claim 20 , wherein the manage forward progress of microservices executed by distributed nodes is based on telemetry data comprising one or more physical events in the distributed nodes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2021
From: DOSHI, KSHITIJ A.; VAN DE GROENENDAAL, JOHAN; CHEN, EDMUND; SAHITA, RAVI; HERDRICH, ANDREW J.; BERNSTEIN, DEBRA; SEVERNS-WILLIAMS, CHRISTINE E.; CUMMINGS, URI V.; KAKAIYA, UTKARSH Y.
To: INTEL CORPORATION
Reel/Frame 055853/0539 →
Continuity (2)
Provisional Application 63087218 · Oct 3, 2020
Related Publication 20210117249A1 · Apr 22, 2021
References Cited (30)
US 4608688A · Hansen · 1986 [cited by examiner]
US 6181677B1 · Valli · 2001 [cited by examiner]
US 10754666B1 · Sorani et al. · 2020 [cited by applicant]
US 20160124742A1 · Rangasamy et al. · 2016 [cited by applicant]
US 20180150343A1 · Bernat et al. · 2018 [cited by applicant]
US 20180152392A1 · Reed et al. · 2018 [cited by applicant]
US 20180234459A1 · Kung et al. · 2018 [cited by applicant]
US 20180331905A1 · Toledo · 2018 [cited by examiner]
US 20180343208A1 · Narkier et al. · 2018 [cited by applicant]
US 20190004871A1 · Sukhomlinov · 2019 [cited by examiner]
US 20190220601A1 · Sood et al. · 2019 [cited by applicant]
US 20190253518A1 · Nachimuthu et al. · 2019 [cited by applicant]
US 20190297150A1 · Takahashi et al. · 2019 [cited by applicant]
US 20190347125A1 · Sankaran et al. · 2019 [cited by applicant]
US 20190370263A1 · Nucci · 2019 [cited by examiner]
US 20190394081A1 · Tahhan et al. · 2019 [cited by applicant]
US 20200401457A1 · Singhal et al. · 2020 [cited by applicant]
US 20210117242A1 · Groenendaal et al. · 2021 [cited by applicant]
US 20210117249A1 · Doshi et al. · 2021 [cited by applicant]
US 20210342188A1 · Novakovic et al. · 2021 [cited by applicant]
US 20220050897A1 · Gaddam · 2022 [cited by examiner]
International Search Report and Written Opinion for PCT Patent Application No. PCT/US21/48092, Mailed Dec. 22, 2021, 12 pages. [cited by applicant]
French and English Translation of First Office Action for Patent Application No. 2110069, Mailed May 20, 2022, 16 pages. [cited by applicant]
Netherlands Search Report for Patent Application No. 2029116, Mailed Apr. 4, 2022, 9 pages. [cited by applicant]
“NVIDIA® Mellanox@ BlueField® SmartNIC for Ethernet High Performance Ethernet Network Adapter Cards”, Mellanox Technologies, May 17, 2020, 3 pages. [cited by applicant]
Kennedy, Patrick, “Mellanox Bluefield-2 IPU SmartNIC Bringing AWS-like Features to VMware”, STH, https://www.servethehome.com/mellanox-bluefield-2-ipu-smartnic-bringing-aws-like-features-to-vmware/, Sep. 2, 2019, 3 page… [cited by applicant]
Kit, Ariel, “Programming the Entire Data Center Infrastructure with the NVIDIA DOCA SDK”, https://developer.nvidia.com/blog/programming-the-entire-data-center-infrastructure-with-the-nvidia-doca-sdk/, Oct. 5, 2020, 4 pa… [cited by applicant]
Yuan, Yifan, et al., “HALO: Accelerating Flow Classification for Scalable Packet Processing in NFV”, 2019 Association for Computing Machinery, https://dl.acm.org/doi/10.1145/3307650.3322272, ISCA '19, Jun. 22-26, 2019, … [cited by applicant]
First Office Action for U.S. Appl. No. 17/134,321, Mailed Dec. 5, 2023, 17 pages. [cited by applicant]
Final Office Action for U.S. Appl. No. 17/134,321, Mailed Jul. 30, 2024, 23 pages. [cited by applicant]