IP Library › Granted Patent US 12,724,642
Granted Patent B2
US 12,724,642 · App. 17/683,524 · Granted Sep 1, 2026

Distributing workloads to hardware accelerators during transient workload spikes

Inventors: Diman Zad Tootaghaj (Milpitas, CA); Anu Mercian (Santa Clara, CA); Puneet Sharma (Palo Alto, CA)
Assignee: Hewlett Packard Enterprise Development LP
G06F9/505H04L47/805H04L47/83G06F2209/5019G06F2209/508
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,724,642
App. No.
17/683,524
Granted
Sep 1, 2026
Kind
B2
Abstract

Systems and methods are provided for strategically harvesting untapped compute capacity of hardware accelerators to manage transient workload spikes at computing systems, are provided. Examples provide a low-cost and scalable computing system which orchestrates seamless offloading of workloads to hardware accelerators during transient workload spikes. By utilizing hardware accelerators as short-term emergency buffers, examples improve upon existing approaches which deploy more expensive, and often significantly under-utilized servers for these emergency purposes. Accordingly, examples may reduce the occurrence of SLA violations while minimizing capital expenditure in computing power.

Claims (37)

1 . A method comprising:

predicting a transient workload spike based on monitored historical data regarding past workloads received by a computing system, wherein the computing system includes a server and a hardware accelerator, and wherein, prior to the transient workload spike, a service time of the server in responding to incoming workloads is less than a service time of the hardware accelerator in responding to the incoming workloads, wherein a window-based prediction model is used to predict the transient workload spike and a window size of the window-based prediction model is increased when workload variation in the monitored historical data is less than a first percentage and is decreased when the workload variation is more than a second percentage, wherein the first percentage is less than the second percentage;

monitoring values of the service time of the server in servicing the incoming workloads distributed to the server;

predicting that a value of the service time of the server will exceed a threshold value for the service time at a time prior to or during the predicted transient workload spike;

responsive to predicting that the value of the service time will exceed the threshold value for the service time, determining that the service time of the server, in responding to the incoming workloads distributed to the server, will exceed the service time of the hardware accelerator during the predicted transient workload spike; and

responsive to the determination, offloading at least one incoming workloads of the incoming workload distributed to the server to the hardware accelerator and executing the at least one incoming workload on the hardware accelerator.

2 . The method of claim 1 , wherein the hardware accelerator comprises a System on a Chip (SOC) based Smart Network Interface Card (SmartNIC).

3 . The method of claim 1 , wherein the monitored historical data is specific to an application and distributing the at least one incoming workload is for the application.

4 . The method of claim 3 , wherein the threshold value for the service time corresponds to a specification in a service level agreement for the application.

5 . The method of claim 1 , wherein:

the past workloads received by the computing system are past serverless queries received by the computing system;

the at least one incoming workload is at least one incoming serverless query; and

the at least one incoming serverless query is executed within a workload container at the hardware accelerator.

6 . The method of claim 5 , wherein the workload container at the hardware accelerator is started before the predicted transient workload spike.

7 . The method of claim 6 , wherein the starting the workload container comprising starting a containerized runtime environment on the hardware accelerator prior to arrival of the predicted transient spike.

8 . The method of claim 1 , wherein the transient workload spike is predicted based on a support vector regression (SVR) prediction model.

9 . A computing system comprising:

a plurality of processing resources associated with the computing system; and

a non-transitory computer-readable medium, coupled to the plurality of processing resources, having stored therein instructions that when executed by the processing resources cause the computing system to:

predict a transient workload spike based on monitored historical data regarding past workloads received by the computing system, wherein the computing system includes a server and a hardware accelerator, and wherein, prior to the transient workload spike, a service time of the server in responding to incoming workloads is less than a service time of the hardware accelerator in responding to the incoming workloads, wherein a window-based prediction model is used to predict the transient workload spike and a window size of the window-based prediction model is increased when workload variation in the monitored historical data is less than a first percentage and is decreased when the workload variation is more than a second percentage, wherein the first percentage is less than the second percentage;

predict that a value of a service time of the computing system will exceed a threshold value of the service time at a time prior to or during the predicted transient workload spike unless at least one incoming workload is distributed to the hardware accelerator;

start a workload container at the hardware accelerator prior to the predicted transient workload spike; and

based on responsive to predicting that the value of the service time will exceed the threshold value for the service time, determining that the service time of the server will exceed the service time of the hardware accelerator during the predicted transient workload spike; and

responsive to the determination, offloading at least one incoming workload distributed to the server to a workload container at the hardware accelerator and executing the at least one incoming workload by the workload container at the hardware accelerator.

10 . The computing system of claim 9 , wherein the threshold value for the service time corresponds to a specification in a service level agreement.

11 . The computing system of claim 9 , wherein the hardware accelerator comprises a network accelerator.

12 . The computing system of claim 11 , wherein the network accelerator comprises a System on a Chip (SOC) based Smart Network Interface Card (SmartNIC).

13 . The computing system of claim 9 , wherein the transient workload spike is predicted based on a support vector regression (SVR) prediction model.

14 . The computing system of claim 9 , wherein the monitored historical data is specific to an application and distributing the at least one incoming workload is for the application.

15 . A non-transitory computer-readable medium storing instructions, which when executed by a plurality of processing resources of an edge-computing system, cause the edge-computing system to:

receive a query from an Application Programming Interface (API) gateway of the edge-computing system, wherein the edge-computing system includes an edge server and a hardware accelerator;

predict a transient workload spike using a window-based prediction model, a window size of the window-based prediction model is increased when workload variation in monitored historical data is less than a first percentage and is decreased when the workload variation is more than a second percentage, wherein the first percentage is less than the second percentage;

based on a determination that the transient workload spike exceeds a threshold value of a service time of the edge server, determine a distribution of queries over a time horizon which includes the predicted transient workload spike, and determine that, during the predicted transient workload spike, the service time of the edge server in responding to the distribution of queries will exceed a service time of the hardware accelerator in responding to the distribution of queries, wherein the service time of the edge server is less than the service time of the hardware accelerator prior to the transient workload spike; and

based on the determined distribution of queries and on the determination that the service time of the edge server will exceed the service time of the hardware accelerator during the predicted transient workload spike, offload the query from the edge server to a workload container at the hardware accelerator and execute the query by the workload container at the hardware accelerator.

16 . The non-transitory computer-readable medium of claim 15 , wherein the hardware accelerator comprises a System on a Chip (SOC) based Smart Network Interface Card (SmartNIC).

17 . The non-transitory computer-readable medium of claim 15 , wherein the workload container at the hardware accelerator is started before the predicted transient workload spike.

18 . The non-transitory computer-readable medium of claim 17 , wherein the monitored historical data is specific to an application.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2022
From: ZAD TOOTAGHAJ, DIMAN; MERCIAN, ANU; SHARMA, PUNEET
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 059131/0710 →
Continuity (1)
Related Publication 20230281052A1 · Sep 7, 2023
References Cited (67)
US 8930948B2 · Shanmuganathan · 2015 [cited by examiner]
US 9680774B2 · Pirko · 2017 [cited by applicant]
US 9692642B2 · Pirko · 2017 [cited by applicant]
US 10484301B1 · Shukla · 2019 [cited by examiner]
US 11303534B2 · Tootaghaj · 2022 [cited by examiner]
US 11436054B1 · Zad Tootaghaj · 2022 [cited by examiner]
US 11513828B1 · Zelenov · 2022 [cited by examiner]
US 11526376B2 · Zhang · 2022 [cited by examiner]
US 12136040B2 · Benmakrelouf · 2024 [cited by examiner]
US 20050050187A1 · Freimuth · 2005 [cited by examiner]
US 20050177549A1 · Hornick · 2005 [cited by examiner]
US 20110179057A1 · Wojcik · 2011 [cited by examiner]
US 20150253837A1 · Sukonik et al. · 2015 [cited by applicant]
US 20160128059A1 · Hsu · 2016 [cited by examiner]
US 20170083382A1 · LeBeane · 2017 [cited by examiner]
US 20170090992A1 · Bivens · 2017 [cited by examiner]
US 20180027062A1 · Bernat · 2018 [cited by examiner]
US 20180307535A1 · Suzuki · 2018 [cited by examiner]
US 20190188043A1 · Jalai · 2019 [cited by examiner]
US 20190222518A1 · Bernat et al. · 2019 [cited by applicant]
US 20200026576A1 · Kaplan · 2020 [cited by applicant]
US 20200097810A1 · Hetherington · 2020 [cited by examiner]
US 20200310886A1 · Rajamani · 2020 [cited by examiner]
US 20210042661A1 · St-Onge · 2021 [cited by examiner]
US 20210064432A1 · Benmakrelouf · 2021 [cited by examiner]
US 20210144517A1 · Guim Bernat · 2021 [cited by examiner]
US 20210184942A1 · Tootaghaj · 2021 [cited by examiner]
US 20220321434A1 · Kuriata · 2022 [cited by examiner]
US 20220360645A1 · Badic · 2022 [cited by examiner]
US 20220377612A1 · Radunovic · 2022 [cited by examiner]
US 20220405127A1 · Huo · 2022 [cited by examiner]
US 20220413895A1 · Nizar · 2022 [cited by examiner]
US 20230195485A1 · Nagasundaram · 2023 [cited by examiner]
US 20230274160A1 · Mujumdar · 2023 [cited by examiner]
Islam et al.; “Empirical prediction models for adaptive resource provisioning in the cloud”; Future Generation Computer Systems 28 (2012) 155-162; 2011 Elsevier B.V. All rights reserved; doi:10.1016/j.future.2011.05.027… [cited by examiner]
Alelaiwi et al.; “An efficient method of computation offloading in an edge cloud platform”; Journal of Parallel and Distributed Computing 127 (2019) 58-64; 2019 Elsevier Inc.; https://doi.org/10.1016/j.jpdc.2019.01.003 … [cited by examiner]
Miao et al.; “Intelligent task prediction and computation offloading based on mobile-edge cloud computing”; Future Generation Computer Systems 102 (2020) 925-931; 2019 Elsevier Inc.; https://doi.org/10.1016/j.future.201… [cited by examiner]
Baig et al.; “Adaptive sliding windows for improved estimation of data center resource utilization”; 2019 published by Elsevier B.V.; https://doi.org/10.1016/j.future.2019.10.026; (Baig_2019.pdf) (Year: 2019). [cited by examiner]
Abdullah et al., “Burst-Aware Predictive Autoscaling for Container”, IEEE Transactions on Services Computing, May 2020, 14 pages. [cited by applicant]
Choi et al., “Lambda -NIC: Interactive Serverless Compute on Programmable SmartNICs”, Sep. 26, 2019, pp. 1-15. [cited by applicant]
D. Coulter et al., “Performance Tuning Network Adapters”, Microsoft Docs, Dec. 23, 2019, 13 pages. [cited by applicant]
Firestone et al., “Azure Accelerated Networking: SmartNICs in the Public Cloud”, 15th USENIX Symposium on Networked Systems Design and Implementation, Apr. 9-11, 2018, 15 pages. [cited by applicant]
Kim et al., “HyperLoop: Group-Based NIC-Offloading to Accelerate Replicated Transactions in Multi-Tenant Storage Systems”, ACM, 2018, pp. 297-312. [cited by applicant]
Krishnamurthy et al., “E3: Energy-Efficient Microservices on SmartNIC-Accelerated Servers”, Proceedings of the 2019 USENIX Annual Technical Conference, Jul. 10-12, 2019, 17 pages. [cited by applicant]
Liu et al., “iPipe: A Framework for Building Distributed Applications on Multicore SoC SmartNICs”, SIGCOMM'19, 2019, pp. 1-18. [cited by applicant]
Liu et al., “Offloading Distributed Applications onto SmartNICs using iPipe”, ACM, Aug. 19-23, 2019, 16 pages. [cited by applicant]
Phothilimthana et al., “Floem: A Programming System for NIC-Accelerated Network Applications”, Usenix, Oct. 8-10, 2018, 18 pages. [cited by applicant]
Wikipedia, “PID controller”, available online at <https://en.wikipedia.org/wiki/PID_controller#Overview_of_tuning_methods>, Mar. 25, 2021, 26 pages. [cited by applicant]
Ziegler et al., “Optimum Settings for Automatic Controllers,” Transactions of the A.S.M.E., Nov. 1942, pp. 759-768. [cited by applicant]
Agilio CX SmartNICs', available online at <http://web.archive.org/web/20170914210510/https://netronome.com/products/agilio-cx/>, Sep. 14, 2017, 10 pages. [cited by applicant]
“AMD Pensando™ Infrastructure Accelerators”, available online at <http://web.archive.org/web/20240406203031/https://www.amd.com/en/accelerators/pensando>, Apr. 6, 2024, 12 pages. [cited by applicant]
“SPIFFE: Secure Production Identity Framework for Everyone”, available online at <http://web.archive.org/web/20200301121129/https://spiffe.io/>, Mar. 1, 2020, 5 pages. [cited by applicant]
Akkus et al., “SAND: Towards High-Performance Serverless Computing”, 2018 USENIX Annual Technical Conference (USENIX ATC '18). Jul. 11-13, 2018, 14 pages. [cited by applicant]
Alrowaily et al., “Secure Edge Computing in IoT Systems: Review and Case Studies”, In 2018 IEEE/ACM Symposium on Edge Computing (SEC). IEEE, 2018, 5 pages. [cited by applicant]
Broadcom, “Ethernet Network Adapters”, available online at <https://www.broadcom.com/products/ethernet-connectivity/network-adapters>, 2005, 4 pages. [cited by applicant]
Chien et al., “Moore's Law: The First Ending and a New Beginning”, IEEE, 2013, 6 Pages. [cited by applicant]
Github, “Azure/AzurePublicDataset”, available online at <http://web.archive.org/web/20180604111946/https://github.com/Azure/AzurePublicDataset/>, Jun. 4, 2018, 2 pages. [cited by applicant]
Hendrickson et al., “Serverless Computation with OpenLambda”, 2016, 7 pages. [cited by applicant]
Liu et al., “E3: energy-efficient microservices on SmartNIC-accelerated servers”, 2019 USENIX Annual Technical Conference. Jul. 10-12, 2019, 17 pages. [cited by applicant]
Liu et al., “iPipe: A FrameworkforBuilding Distributed Applications onMulticoreSoCSmartNICs”, In Proceedings of the ACM Special Interest Group on Data Communication (SIGCOMM), 2019, 18 pages. [cited by applicant]
Lloyd et al., “Serverless Computing: An Investigation of Factors Influencing Microservice Performance”, 2018, 11 pages. [cited by applicant]
Nvidia, “Nvidia Mellanox DPU & NPU”, available online at <http:/web.archive.org/web/20201002054710/https://www.nvidia.com/en-us/networking/products/data-processing-unit/>, Oct. 2, 2020, 3pages. [cited by applicant]
Qiu et al., “Clara: Performance Clarity for SmartNIC Offloading”, ACM, 2020, 7 Pages. [cited by applicant]
Shahrad et al. “Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider”, In USENIX Annual Technical Conference (USENIX ATC), 2020, 15 pages. [cited by applicant]
Valuates, “Edge Computing Market—Global Opportunity Analysis”, available online at <https://reports.valuates.com/market-reports/360I-Auto-3U70/edge-computing-market/>, 2018, 5 pages. [cited by applicant]
Wang et al., “Edge Computing: Applications, State-of-the-Art and Challenges ”, Advances in Networks 2019; 7(1): 8-15. [cited by applicant]
Wang et al., “Peeking Behind the Curtains of Serverless Platforms”, 2018 USENIX Annual Technical Conference (USENIX ATC '18). Jul. 11-13, 2018, 14 pages. [cited by applicant]