IP Library › Granted Patent US 12,645,634
Granted Patent B2
US 12,645,634 · App. 17/549,727 · Granted Jun 2, 2026

Technologies for hardware microservices accelerated in XPU

Inventors: Susanne M. Balle (Hudson, NH); Duane E. Galbi (Wayland, MA); Andrzej Kuriata (Gdansk, PL); Sundar Nadathur (Cupertino, CA); Nagabhushan Chitlur (Portland, OR); Francesc Guim Bernat (Barcelona, ES); Alexander Bachmutsky (Sunnyvale, CA)
Assignee: Intel Corporation
G06F15/7889G06F9/5044G06F2209/5019G06F2209/509
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,634
App. No.
17/549,727
Granted
Jun 2, 2026
Kind
B2
Abstract

Methods, apparatus, and software and for hardware microservices accelerated in other processing units (XPUs). The apparatus may be a platform including a System on Chip (SOC) and an XPU, such as a Field Programmable Gate Array (FPGA). The FPGA is configured to implement one or more Hardware (HW) accelerator functions associated with HW microservices. Execution of microservices is split between a software front-end that executes on the SOC and a hardware backend comprising the HW accelerator functions. The software front-end offloads a portion of a microservice and/or associated workload to the HW microservice backend implemented by the accelerator functions. An XPU or FPGA proxy is used to provide the microservice front-ends with shared access to HW accelerator functions, and schedules/multiplexes access to the HW accelerator functions using, e.g., telemetry data generated by the microservice front-ends and/or the HW accelerator functions. The platform may be an infrastructure processing unit (IPU) configured to accelerate infrastructure operations.

Claims (46)

1 . A platform comprising:

a circuit board to which multiple components are operatively coupled, including,

a System On Chip (SOC), having a plurality of cores;

an Other Processing Unit (XPU), communicatively coupled to the SOC via interconnect circuitry in the circuit board;

an input/output (I/O) interface coupled to at least one of the SOC and XPU; and

memory, in which host software to be executed on one or more of the plurality of cores in the SOC is stored, including software code to implement one or more hardware (HW) microservice front-ends,

wherein the XPU is configured to implement one or more accelerator functions used to accelerate HW microservice backend operations that are offloaded from the one or more HW microservice front-ends, and

wherein the platform is configured to be coupled, via the I/O interface, to a compute host including a central processing unit (CPU) used to host software-based workloads.

2 . The platform of claim 1 , wherein the host software includes an XPU proxy that is configured to enable HW microservices front-ends to share the one or more accelerator functions.

3 . The platform of claim 2 , wherein the XPU proxy is configured to:

receive telemetry data generated from at least one of the HW microservice front-ends and the one or more accelerator functions; and

schedule access to the accelerator functions by the HW microservice front-ends based, at least in part, on the telemetry data.

4 . The platform of claim 2 , wherein the host software includes software code for implementing a software microservice, and wherein the XPU proxy is configured to determine that a microservice should be implemented as a software microservice or a HW microservice.

5 . The platform of claim 1 , wherein the XPU comprises a Field Programmable Gate Array (FPGA), and the one or more accelerated functions comprise FPGA kernels.

6 . The platform of claim 5 , wherein the host software is configured to:

provision accelerator functions in the FPGA by sending FPGA kernel bitstreams to the FPGA; and

reprovision FPGA bits for at least one accelerator function by sending a new FPGA kernel bitstream to the FPGA.

7 . The platform of claim 1 , wherein the platform comprises an infrastructure processing unit (IPU).

8 . The platform of claim 7 , wherein the XPU comprises a Field Programmable Gate Array (FPGA), wherein the IPU comprises a Peripheral Component Interconnect Express (PCIe) card, and the FPGA is configured to implement one or more PCIe interfaces, one of which is the I/O interface.

9 . The platform of claim 1 , wherein the platform is configured to be installed in a server comprising the compute host.

10 . A method implemented on a platform including a System on Chip (SOC) operatively coupled to a circuit board and communicatively coupled to an other processing unit (XPU) operatively coupled to the circuit board, comprising:

configuring the XPU to implement one or more accelerator functions;

splitting execution of a microservice between a hardware front-end and a hardware backend, wherein the hardware front-end is executed via execution of software on the SOC, and the hardware backend is used to execute an offloaded portion of the microservice on an accelerator function in the XPU,

wherein the platform is coupled to a compute host having a central processing unit (CPU) hosting software-based workloads.

11 . The method of claim 10 , wherein the XPU comprises a Field Programmable Gate Array (FPGA), further comprising programming an accelerator kernel in the FPGA to implement an accelerator function.

12 . The method of claim 10 , further comprising:

executing a plurality of hardware front-ends on the SOC; and

sharing the one or more accelerator functions among the plurality of hardware front-ends.

13 . The method of claim 12 , further comprising:

obtaining telemetry data generated from at least one of the hardware front-ends and the one or more accelerator functions; and

scheduling access to the accelerator functions by the hardware front-ends based, at least in part, on the telemetry data.

14 . The method of claim 10 , wherein the platform comprises an infrastructure processing unit (IPU) and the XPU comprises a Field Programmable Gate Array (FPGA).

15 . The method of claim 10 , wherein the method is implemented on a platform in a data center, and wherein the hardware front-end is run in a server worker pod in the data center.

16 . A non-transitory machine-readable medium having software instructions stored thereon configured to be executed on one or more cores on a System on Chip (SOC) in a platform including the SOC and an other processing unit (XPU) communicatively coupled to the SOC, the XPU configured to implement one or more accelerator functions, wherein execution of the instructions on the SOC enables the platform to:

implement a hardware (HW) microservice front-end that performs a first portion of a microservice;

determine a second portion of the microservice to offload to a first accelerator function; and

offload the second portion of the microservice to the XPU to execute the first accelerator function.

17 . The non-transitory machine-readable medium of claim 16 , wherein execution of the software instructions enables the platform to implement a plurality of HW microservice front-ends, wherein the software instructions include a software instruction for implementing an XPU proxy service, wherein the XPU proxy service facilitates sharing of the one or more accelerator functions among the plurality of HW microservice front-ends.

18 . The non-transitory machine-readable medium of claim 17 , wherein execution of the software instructions further enabled the platform to:

receive telemetry data generated from at least one of the HW microservice front-ends and the one or more accelerator functions; and

schedule access to an accelerator function by a HW microservice front-end based, at least in part, on the telemetry data.

19 . The non-transitory machine-readable medium of claim 17 , wherein execution of the software instructions enables the XPU proxy service to:

schedule access to the one or more accelerator functions based on pre-existing characteristics of the one or more accelerator functions and feedback from one or more of the plurality of HW microservice front-ends.

20 . The non-transitory machine-readable medium of claim 16 , wherein the XPU comprises a Field Programmable Gate Array (FPGA), and wherein execution of the software instructions further enables the platform to:

provision accelerator functions in the FPGA by sending FPGA kernel bitstreams to the FPGA; and

reprovision FPGA bits for at least one accelerator function by sending a new FPGA kernel bitstream to the FPGA.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2021
From: BALLE, SUSANNE M.; GALBI, DUANE E.; KURIATA, ANDRZEJ; NADATHUR, SUNDAR; CHITLUR, NAGABHUSHAN; GUIM BERNAT, FRANCESC; BACHMUTSKY, ALEXANDER
To: INTEL CORPORATION
Reel/Frame 058493/0377 →
Continuity (1)
Related Publication 20230185760A1 · Jun 15, 2023
References Cited (25)
US 10691597B1 · Akkary · 2020 [cited by examiner]
US 10754666B1 · Sorani · 2020 [cited by examiner]
US 11194707B2 · Stalzer · 2021 [cited by applicant]
US 20190004871A1 · Sukhomlinov · 2019 [cited by examiner]
US 20190050522A1 · Alvarez et al. · 2019 [cited by applicant]
US 20190068693A1 · Bernat · 2019 [cited by examiner]
US 20190155239A1 · Salhuana · 2019 [cited by examiner]
US 20200142735A1 · Maciocco · 2020 [cited by examiner]
US 20200371828A1 · Chiou · 2020 [cited by examiner]
US 20210042254A1 · Marolia et al. · 2021 [cited by applicant]
US 20210117249A1 · Doshi et al. · 2021 [cited by applicant]
US 20210152659A1 · Cai · 2021 [cited by examiner]
US 20210209035A1 · Galbi et al. · 2021 [cited by applicant]
US 20210319307A1 · Dhruvanarayan et al. · 2021 [cited by applicant]
David Ojika [cited by examiner]
Julien Lallet, Andrea Enrici and Anfel Saffar, “FPGA-based System for the Acceleration of Cloud Microservices”, Aug. 16, IEEE, pp. 1-5 (Year: 2018). [cited by examiner]
A. M. Caulfield et al., “A Cloud-Scale Acceleration Architecture,” in 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), Oct. 2016, 13 pages. [cited by applicant]
Eric Chung, et al., “Serving DNNs in Real Time at Datacenter Scale with Project Brainwave,” IEEE Micro, vol. 38, Mar. 2018, 13 pages. [cited by applicant]
Fungible Storage Cluster, downloaded from website https://www.fungible.com/, Dec. 21, 2021, 6 pages. [cited by applicant]
J. Lallet et al., “FPGA based system for the acceleration of Cloud Microservices,” in IEEE International Symposium on Broadband Multimedia Systems and Broadcasting (BMSB), Jun. 2018, 2 pages. [cited by applicant]
M. Branscombe, “FPGAs and the New Era of Cloud-based 'Hardware Microservices',” https://thenewstack.io/developers-fpgas-cloud/, Jun. 8, 2017, 9 pages. [cited by applicant]
Mazen Ezzeddine et al., “RESTful Hardware Microservices Using Reconfigurable Networked Accelerators in Cloud and Edge Datacenters”, 2018 IEEE 7th International Conference on Cloud Networking (CloudNet), Oct. 2018, 4 pag… [cited by applicant]
Susanne M.Balle, et al., “Inter-Kernal Links for Direct Inter-FPGA Communication,” White Paper, Distributed Direct Inter-FPGA Communication Framework, Multi-FPGA Dep Learning Inference Application and Model Parallelism,… [cited by applicant]
International Search Report and Written Opinion for PCT Patent Application No. PCT/US22/48101, Mailed Feb. 21, 2023, 10 pages. [cited by applicant]
Extended European Search Report for Patent Application No. 22908182.3, Mailed Dec. 17, 2025, 9 pages. [cited by applicant]