IP Library › Granted Patent US 12,468,578
Granted Patent B2
US 12,468,578 · App. 17/559,833 · Granted Nov 11, 2025

Infrastructure managed workload distribution

Inventors: Francesc Guim Bernat (Barcelona, ES); Karthik Kumar (Chandler, AZ); Alexander Bachmutsky (Sunnyvale, CA); Marcos E. Carranza (Portland, OR); Rita H. Wouhaybi (Portland, OR)
Assignee: Intel Corporation
G06F9/5083G06F9/4881
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,578
App. No.
17/559,833
Granted
Nov 11, 2025
Kind
B2
Abstract

System and techniques for infrastructure managed workload distribution are described herein. An infrastructure processing unit (IPU) receives a workload that includes a workload definition. The workload definition includes stages of the workload and a performance expectation. The IPU provides the workload, for execution, to a processing unit of a compute node to which the IPU belongs. The IPU monitors execution of the workload to determine that a stage of the workload is performing outside of the performance expectation from the workload definition. In response, the IPU modifies the execution of the workload.

Claims (51)

1 . An infrastructure processing unit (IPU) comprising:

a network connector;

a Compute Express Link (CXL) connector; and

processing circuitry configured to:

receive a workload via the network connector or the CXL connector, the workload including a workload definition including a plurality of pipelined stages of the workload and a respective performance expectation for each of the plurality of pipelined stages;

provide, via the CXL connector, the workload to a processor of a compute node for execution;

access execution stacks from pooled memory of the compute node, wherein the execution stacks indicate that a stage of the workload is performing below its respective performance expectation based on the workload execution; and

modify the execution of the stage of the workload in response to the indication that the stage of the workload is performing below its respective performance expectation, wherein the modifying results in performance of the workload meeting a predetermined service level agreement by intercepting, during the execution and via the CXL connector, a resource request of the stage of the workload to the compute node and redirecting the resource request to additional hardware utilized by the stage of the workload during the execution.

2 . The IPU of claim 1 , wherein, to modify the execution of the stage of the workload, the processing circuitry is further configured to:

stop the stage of the workload on the compute node; and

transfer the stage of the workload to a second compute node via the network connector or the CXL connector.

3 . The IPU of claim 2 , wherein the second compute node is one of multiple compute nodes accessible via the network connector or via the CXL connector.

4 . The IPU of claim 3 , wherein the multiple compute nodes each have a corresponding resource availability.

5 . The IPU of claim 4 , wherein the processing circuitry is further configured to:

receive, at predefined intervals, resource messages from each of the multiple compute nodes, the resource messages indicating an available resource at a respective compute node; and

store records for each of the multiple compute nodes of the available resource from the resource messages.

6 . The IPU of claim 5 , wherein the records are stored according to a predetermined order.

7 . The IPU of claim 1 , wherein the IPU further comprises:

tracking circuitry;

migration circuitry;

pipeline circuitry; and

telemetry circuitry.

8 . The IPU of claim 1 , wherein the processing circuitry is further configured to receive, from the compute node, interrupts or instruction fetches associated with the execution of the workload indicating that the stage of the workload is performing below its respective performance expectation.

9 . The IPU of claim 1 , wherein the processing circuitry is further configured to store the workload definition in a memory of the IPU.

10 . The IPU of claim 9 , wherein the memory is a register, dynamic random access memory (DRAM), synchronous DRAM (SDRAM), or static RAM (SRAM).

11 . The IPU of claim 1 , wherein the respective performance expectation is a respective expected time to complete the execution of the respective stage, and wherein performing below the respective performance expectation includes taking longer to complete the respective stage than the respective expected time to complete included in the workload definition.

12 . The IPU of claim 1 , wherein the workload definition includes a workload identification (ID) field, a stage ID field, an expected performance field, and a resource ID field.

13 . At least one non-transitory machine readable medium including instructions that when executed by processing circuitry, cause the processing circuitry to execute operations comprising:

receiving, at an infrastructure processing unit (IPU) and via a network connector or a CXL connector of a compute node, a workload including a workload definition including a plurality of pipelined stages of the workload and a respective performance expectation for each of the plurality of pipelined stages;

providing, via the CXL connector, the workload to a processor of the compute node for execution;

accessing execution stacks from pooled memory of the compute node, wherein the execution stacks indicate that a stage of the workload is performing below its respective performance expectation based on the workload execution; and

modifying the execution of the stage of the workload in response to the indication that the stage of the workload is performing below its respective performance expectation, wherein the modifying results in performance of the workload meeting a predetermined service level agreement by intercepting, during the execution and via the CXL connector, a resource request of the stage of the workload to the compute node and redirecting the resource request to additional hardware utilized by the stage of the workload during the execution.

14 . The at least one non-transitory machine readable medium of claim 13 , wherein modifying the execution of the stage of the workload further includes:

stopping the stage of the workload on the compute node; and

transferring the stage of the workload to a second compute node via the network connector or the CXL connector.

15 . The at least one machine readable medium of claim 14 , wherein the second compute node is one of multiple compute nodes accessible via the network connector or via the CXL connector.

16 . The at least one non-transitory machine readable medium of claim 15 , wherein the multiple compute nodes each have a corresponding resource availability.

17 . The at least one non-transitory machine readable medium of claim 16 , wherein the operations further comprise:

receiving, periodically, resource messages from each of the multiple compute nodes, the resource messages indicating an available resource at a respective compute node; and

storing records for each of the multiple compute nodes of the available resource from the resource messages.

18 . The at least one non-transitory machine readable medium of claim 17 , wherein the records are stored according to a predetermined order.

19 . The at least one non-transitory machine readable medium of claim 13 , wherein the IPU further comprises:

tracking circuitry;

migration circuitry;

pipeline circuitry; and

telemetry circuitry.

20 . The at least one non-transitory machine readable medium of claim 13 , wherein the operations further comprise receiving, from the compute node, interrupts or instruction fetches associated with the execution of the workload indicating that the stage of the workload is performing below its respective performance expectation.

21 . The at least one non-transitory machine readable medium of claim 13 , wherein the operations further comprise storing the workload definition in a memory of the IPU.

22 . The at least one non-transitory machine readable medium of claim 21 , wherein the memory is a register, dynamic random access memory (DRAM), synchronous DRAM (SDRAM), or static RAM (SRAM).

23 . The at least one non-transitory machine readable medium of claim 13 , wherein the respective performance expectation is a respective expected time to complete the execution of the respective stage, and wherein performing below the respective performance expectation includes taking longer to complete the respective stage than the respective expected time to complete included in the workload definition.

24 . The at least one non-transitory machine readable medium of claim 13 , wherein the workload definition includes a workload identification (ID) field, a stage ID field, an expected performance field, and a resource ID field.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2022
From: GUIM BERNAT, FRANCESC; KUMAR, KARTHIK; BACHMUTSKY, ALEXANDER; CARRANZA, MARCOS E.; WOUHAYBI, RITA H.
To: INTEL CORPORATION
Reel/Frame 059199/0463 →
Continuity (1)
Related Publication 20220114032A1 · Apr 14, 2022
References Cited (66)
US 5212782A · Asato · 1993 [cited by examiner]
US 7669029B1 · Mishra · 2010 [cited by examiner]
US 8978034B1 · Goodson · 2015 [cited by examiner]
US 9495222B1 · Jackson · 2016 [cited by examiner]
US 9547484B1 · Frazier · 2017 [cited by examiner]
US 10057122B1 · Andersen · 2018 [cited by examiner]
US 10254970B1 · Martin · 2019 [cited by examiner]
US 10291488B1 · Srinivasan · 2019 [cited by examiner]
US 10382380B1 · Suzani · 2019 [cited by examiner]
US 11057318B1 · Matthews · 2021 [cited by examiner]
US 11134013B1 · Allen · 2021 [cited by examiner]
US 11216314B2 · Harwood · 2022 [cited by examiner]
US 11372689B1 · Allen · 2022 [cited by examiner]
US 11886926B1 · Gadalin · 2024 [cited by examiner]
US 20030204706A1 · Kim · 2003 [cited by examiner]
US 20050066148A1 · Luick · 2005 [cited by examiner]
US 20070124684A1 · Riel · 2007 [cited by examiner]
US 20070198679A1 · Duyanovich · 2007 [cited by examiner]
US 20090037554A1 · Herington · 2009 [cited by examiner]
US 20090252047A1 · Coffey · 2009 [cited by examiner]
US 20100223619A1 · Jaquet · 2010 [cited by examiner]
US 20110016214A1 · Jackson · 2011 [cited by examiner]
US 20110173470A1 · Tran · 2011 [cited by examiner]
US 20110239220A1 · Gibson et al. · 2011 [cited by applicant]
US 20120179824A1 · Jackson · 2012 [cited by examiner]
US 20130055262A1 · Lubsey · 2013 [cited by examiner]
US 20130073724A1 · Parashar · 2013 [cited by examiner]
US 20130086404A1 · Sankar · 2013 [cited by examiner]
US 20130198386A1 · Srikanth · 2013 [cited by examiner]
US 20130290976A1 · Cherkasova · 2013 [cited by examiner]
US 20140359353A1 · Chen · 2014 [cited by examiner]
US 20150106522A1 · Ryan · 2015 [cited by examiner]
US 20150199141A1 · Faulkner · 2015 [cited by examiner]
US 20150244595A1 · Oberlin · 2015 [cited by examiner]
US 20160050294A1 · Kruse · 2016 [cited by examiner]
US 20160259665A1 · Gaurav · 2016 [cited by examiner]
US 20160381128A1 · Pai · 2016 [cited by examiner]
US 20170295200A1 · Mirza · 2017 [cited by examiner]
US 20180129503A1 · Narayan · 2018 [cited by examiner]
US 20180189101A1 · Xu · 2018 [cited by examiner]
US 20180219899A1 · Joy · 2018 [cited by examiner]
US 20190158537A1 · Miriyala · 2019 [cited by examiner]
US 20200026579A1 · Bahramshahry · 2020 [cited by examiner]
US 20200026631A1 · Reed · 2020 [cited by examiner]
US 20200303060A1 · Haemel · 2020 [cited by examiner]
US 20210117307A1 · MacNamara · 2021 [cited by examiner]
US 20210135685A1 · Kumar · 2021 [cited by examiner]
US 20210136122A1 · Crabtree · 2021 [cited by examiner]
US 20210240354A1 · Shiraki · 2021 [cited by examiner]
US 20210373951A1 · Malladi · 2021 [cited by examiner]
US 20210374056A1 · Malladi · 2021 [cited by examiner]
US 20210405913A1 · Stonelake · 2021 [cited by examiner]
US 20210406075A1 · Illikkal · 2021 [cited by examiner]
US 20220066813A1 · Taher · 2022 [cited by examiner]
US 20220100573A1 · Allen · 2022 [cited by examiner]
US 20220141099A1 · Prasanna Kumar · 2022 [cited by examiner]
US 20220179695A1 · Dawkins · 2022 [cited by examiner]
US 20220179697A1 · Kumar · 2022 [cited by examiner]
US 20220276914A1 · Kundu · 2022 [cited by examiner]
US 20220377612A1 · Radunovic · 2022 [cited by examiner]
US 20230056042A1 · Vichare · 2023 [cited by examiner]
US 20230195452A1 · Bregman · 2023 [cited by examiner]
EP 3929746 · 2021 [cited by applicant]
“European Application Serial No. 22204294.7, Extended European Search Report mailed May 4, 2023”, 9 pgs. [cited by applicant]
Kennedy, Patrick, “Intel IPU is an Exotic Answer to the Industry DPU—ServeTheHome”, XP093041586, [Online]. Retrieved from the Internet: URL: https: www.servethehome.com intel-ipu-exotic-answer-to-industry-dpu , (Jun. 14… [cited by applicant]
Kuriata, Andrzej, “Predictable Performance for QoS-Sensitive, Scalable, Multi-tenant Function-as-a-Service Deployments”, Abstract : XP 2020 Workshops, Copenhagen, Denmark, Jun. 8and#8211;12, 2020, Revised Selected Paper… [cited by applicant]