IP Library › Granted Patent US 12,513,207
Granted Patent B2
US 12,513,207 · App. 18/668,099 · Granted Dec 30, 2025

System and method for coordinated resource scaling in microservice-based and serverless applications

Inventors: Chitra Subramanian (Mahopac, NY); Pavithra Harsha (Pleasantville, NY); Shivaram Subramanian (Frisco, TX)
Assignee: International Business Machines Corporation
H04L67/1008H04L67/1012
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,513,207
App. No.
18/668,099
Granted
Dec 30, 2025
Kind
B2
Abstract

A computer-implemented method for trace-driven call-graph-aware proactive coordinated autoscaling of component microservices in an application includes generating performance-resource elasticity models of endpoints of the component microservices of the application. Workload levels of the endpoint of the component microservices is predicted based on user traffic observed at a front end service. A trace-level performance of the application is predicted for different microservice replica scaling based on the performance-resource elasticity models at end points, the ends points on the trace call graph and the predicted workload levels. A microservice replica scaling is recommended for each of the component microservices to meet predefined trace-level user service level objectives.

Claims (43)

1 . A computer-implemented method for trace-driven call-graph-aware proactive coordinated autoscaling of component microservices in an application, the method comprising:

generating performance-resource elasticity models of endpoints of the component microservices of the application;

predicting workload levels of the endpoint of the component microservices based on user traffic observed at a front end service;

predicting a trace-level performance of the application for different microservice replica scaling based on the performance-resource elasticity models of the component microservice endpoints, the end points on a trace call graph and the predicted workload levels; and

recommending a microservice replica scaling for each of the component microservices to meet predefined trace-level user service level objectives.

2 . The computer-implemented method of claim 1 , wherein the recommended microservice replica scaling is based on a microservice endpoint self-latency aggregation.

3 . The computer-implemented method of claim 2 , wherein the microservice endpoint self-latency is defined as an isolated latency contribution of a given endpoint of a given component microservice to an overall trace latency for each component microservice used in the application.

4 . The computer-implemented method of claim 1 , wherein the trace-level service level objectives include at least one of latency or throughput targets.

5 . The computer-implemented method of claim 1 , wherein the performance-resource elasticity models predict self-latency performance at each of the endpoints of a given component microservice as a function of a vector of workload levels to all endpoints that belong to the given microservice.

6 . The computer-implemented method of claim 1 , wherein the predicted trace-level performance is based on an aggregated self-latency of the component endpoints identified in a trace call graph.

7 . The computer-implemented method of claim 1 , further comprising using a machine learning model for generating the performance-resource elasticity models.

8 . The computer-implemented method of claim 1 , wherein the predicted workload levels are based on a predicted load, a currently observed load, or a combination thereof.

9 . The computer-implemented method of claim 1 , further comprising predicting workload levels at each endpoints of the component microservice used in the application.

10 . The computer-implemented method of claim 1 , further comprising using a machine learning model for generating the predicted workload levels.

11 . The computer-implemented method of claim 1 , further comprising leveraging mixed-integer programming scaling optimization for predicting the trace-level performance of the application for different microservice replica scaling.

12 . The computer-implemented method of claim 1 , further comprising learning a pattern of cascading calls to predict workload levels across multiple traces.

13 . A system comprising:

a processor;

a memory coupled to the processor; and

a computer readable storage embodying a computer program code, the computer program code comprising instructions for trace-driven call-graph-aware proactive coordinated autoscaling of component microservices in an application, the instructions executable by the processor and configured to:

generate performance-resource elasticity models of endpoints of the component microservices of the application;

predict workload levels of the endpoints of the component microservices based on user traffic observed at a front end service;

predict a trace-level performance of the application for different microservice replica scaling based on the performance-resource elasticity models and the predicted workload levels; and

recommend a microservice replica scaling for each of the component microservices to meet predefined trace-level user service level objectives.

14 . The system of claim 13 , wherein:

the recommended microservice replica scaling is based on a microservice endpoint self-latency aggregation; and

the microservice endpoint self-latency is defined as an isolated latency contribution of a given endpoint of a given component microservice to an overall trace latency for each component microservice used in the application.

15 . The system of claim 13 , wherein the trace-level service level objectives include at least one of latency or throughput targets.

16 . The system of claim 13 , wherein:

the performance-resource elasticity models predict self-latency performance at each of the endpoints of any given component microservice as a function of a vector of workload levels to all endpoints that belong to the given microservice; and

the predicted trace-level performance is based on an aggregated self-latency of the component microservices identified in a trace graph.

17 . The system of claim 13 , wherein the instructions are further configured to:

use a machine learning model for generating the performance-resource elasticity models; and

use a machine learning model for generating the predicted workload levels.

18 . The system of claim 13 , wherein the instructions are further configured to predict workload levels at each endpoints of the component microservice used in the application, wherein the predicted workload levels are based on a predicted load, a currently observed load, or a combination thereof.

19 . A computer program product for trace-driven call-graph-aware proactive coordinated autoscaling of component microservices in an application, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to:

generate performance-resource elasticity models of endpoints of the component microservices of the application;

predict workload levels of the endpoint of the component microservices based on user traffic observed at a front end service;

predict a trace-level performance of the application for different microservice replica scaling based on the performance-resource elasticity models and the predicted workload levels; and

recommend a microservice replica scaling for each of the component microservices to meet predefined trace-level user service level objectives.

20 . The computer program product of claim 19 , wherein:

the recommended microservice replica scaling is based on a microservice endpoint self-latency aggregation; and

the microservice endpoint self-latency is defined as an isolated latency contribution of a given endpoint of a given component microservice to an overall trace latency for each component microservice used in the application.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2024
From: SUBRAMANIAN, CHITRA; HARSHA, PAVITHRA; SUBRAMANIAN, SHIVARAM
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 067457/0639 →
Continuity (1)
Related Publication 20250358331A1 · Nov 20, 2025
References Cited (25)
US 11966788B2 · Macdonald · 2024 [cited by applicant]
US 20100299437A1 · Moore · 2010 [cited by examiner]
US 20180337980A1 · Schreter · 2018 [cited by examiner]
US 20190342379A1 · Shukla · 2019 [cited by examiner]
US 20200389517A1 · Eloy · 2020 [cited by examiner]
US 20220046083A1 · Nair · 2022 [cited by examiner]
US 20220353201A1 · Navali · 2022 [cited by examiner]
US 20230308398A1 · Vedam · 2023 [cited by examiner]
US 20250013504A1 · Misur · 2025 [cited by examiner]
CN 115858318A · 2023 [cited by applicant]
CN 115913967A · 2023 [cited by applicant]
WO 202348609A1 · 2023 [cited by applicant]
Barnhart, C. et al., “Branch-and-Price: Column Generation for Solving Huge Integer Programs”, Operations Research (1970), 34 pgs. [cited by applicant]
Irnich, S. et al., “Shortest Path Problems with Resource Constraints”, In Column generation, Springer, 2004, 31 pgs. [cited by applicant]
Wang, Z. et al., “DeepScaling: Microservices Autoscaling for Stable CPU Utilization in Large Scale Cloud Systems”, In Proceedings of the 13th Symposium on Cloud Computing, SoCC '22, p. 16-30, New York, NY, USA, 2022. As… [cited by applicant]
Zhang, Y. et al., “Analytically-Driven Resource Management for Cloud-Native Microservices”, IEEE (2024), 16 pgs. [cited by applicant]
Barnhart, C. et al., “Branch-and-Price: Column Generation for Solving Huge Integer Programs”, INFORMS (2014) 15, pgs. [cited by applicant]
Subramanian, S. et al., “Constrained Prescriptive Trees via Column Generation”, arXiv:2207.10163v1 (2022), 22 pgs. [cited by applicant]
Authors (Disclosed without attribution), “Multi-Dimensional Pod Autoscaling with Reinforcement Learning,” IPCOM000272749D, IP.com, Jul. 31, 2023, 6 pages. [cited by applicant]
Authors (Disclosed without attribution), “Setting and Aligning Service Level Objectives (SLOs) in Distributed Applications,” IPCOM000270080D, IP.com, Jun. 1, 2022, 6 pages. [cited by applicant]
Abdullah, Muhammad, et al., “Learning Predictive Autoscaling Policies for Cloud-Hosted Microservices Using Trace-Driven Modeling,” 2019 IEEE International Conference on Cloud Computing Technology and Science (CloudCom),… [cited by applicant]
Vu, Dinh-Dai, et al., “Predictive Hybrid Autoscaling for Containerized Applications,” IEEE Access 10 (2022): 109768-109778. [cited by applicant]
List of IBM Patents or Patent Applications Treated as Related (2024) 2 pgs. [cited by applicant]
International Searching Authority, “Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, or Declaration,” Patent Cooperation Treaty Sep. 18, 20… [cited by applicant]
Park et al., “A Graph Neural Network based Proactive Resource Allocation Framework for SLO-Oriented Microservices”, CoNEXT '21: Proceedings of the 17th International Conference on emerging Networking Experiments and Tec… [cited by applicant]