IP Library Granted Patent US 12,608,261
Granted Patent B2
US 12,608,261 · App. 18/127,711 · Granted Apr 21, 2026

Estimating propagation time for an injected fault

Inventors: Larisa Shwartz (Greenwich, CT); Saurabh Jha (White Plains, NY); Jesus Maria Rios Aliaga (Philadelphia, PA); Eitan Daniel Farchi (Pardes Hanna-Karkur, IL); Frank Bagehorn (Dottikon, CH); Robert Filepp (Westport, CT)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F11/0775G06F11/079
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,261
App. No.
18/127,711
Granted
Apr 21, 2026
Kind
B2
Abstract

A method, system, and computer program product for estimating propagation time for an injected fault are configured to: determine normal execution times of respective services in a call graph of an application; determine normal execution times of respective network communications between ones of the services; determine faulty execution times of respective ones of the services; and generate a propagation time for a particular type of fault injected at a particular fault injection location in the call graph based on the determined normal execution times of respective services, the determined normal execution times of respective network communications, and the determined faulty execution times of respective ones of the services.

Claims (41)

1 . A method, comprising:

determining, by a processor set, normal execution times of respective services in a call graph of an application;

determining, by the processor set, normal execution times of respective network communications between ones of the services;

determining, by the processor set, faulty execution times of respective ones of the services; and

generating, by the processor set, a propagation time for a particular type of fault injected at a particular fault injection location in the call graph based on the determined normal execution times of respective services, the determined normal execution times of respective network communications, and the determined faulty execution times of respective ones of the services,

wherein the faulty execution times of the respective ones of the services are determined by estimating a probability of a tail event representing a service level objective (SLO) violation using Extreme Value Theory (EVT).

2 . The method of claim 1 , where a fault is realized in different ways by optimization of combinatorial test design.

3 . The method of claim 1 , further comprising repeating the generating for each of plural different types of faults and each of plural different fault injection locations.

4 . The method of claim 3 , further comprising generating a list comprising multiple entries each comprising:

a respective one of the plural different types of faults;

a respective one of the plural different fault injection locations; and

the generated propagation time for a combination of the respective one of the plural different types of faults and the respective one of the plural different fault injection locations.

5 . The method of claim 4 , further comprising adjusting one or more entries of the list based on feedback.

6 . The method of claim 1 , wherein:

the application comprises hybrid cloud application that includes multiple microservices;

the services in the call graph correspond to respective ones of the multiple microservices and network communications; and

the network communications include directionality of execution between the respective ones of the multiple microservices and network.

7 . The method of claim 1 , wherein the normal execution times of the respective services are determined using tracing.

8 . The method of claim 7 , wherein the normal execution times of the respective services and network nodes comprise statistical distributions.

9 . The method of claim 1 , wherein the normal execution times of the respective network communications are determined using tracing and comprise statistical distributions.

10 . A method, comprising:

determining, by a processor set, normal execution times of respective services in a call graph of an application;

determining, by the processor set, normal execution times of respective network communications between ones of the services;

determining, by the processor set, faulty execution times of respective ones of the services; and

generating, by the processor set, a propagation time for a particular type of fault injected at a particular fault injection location in the call graph based on the determined normal execution times of respective services, the determined normal execution times of respective network communications, and the determined faulty execution times of respective ones of the services, wherein the generating the propagation time is further based on one or more parallelization numbers.

11 . A computer program product comprising one or more computer readable storage media having program instructions collectively stored on the one or more computer readable storage media, the program instructions executable to:

determine normal execution times of respective services in a call graph of an application, wherein nodes of the call graph represent microservices in the application and paths between the nodes represent direction of request flow between the microservices;

determine normal execution times of respective network communications between ones of the services;

determine faulty execution times of respective ones of the services by running fault injection simulations; and

generate a propagation time for a particular type of fault injected at a particular fault injection location in the call graph based on the determined normal execution times of respective services, the determined normal execution times of respective network communications, and the determined faulty execution times of respective ones of the services,

wherein the program instructions are further executable to:

repeat the generating for each of plural different types of faults and each of plural different fault injection locations; and

generate a list comprising multiple entries each comprising:

a respective one of the plural different types of faults;

a respective one of the plural different fault injection locations; and

the generated propagation time for a combination of the respective one of the plural different types of faults and the respective one of the plural different fault injection locations; and

wherein the program instructions are further executable to:

generate training data based on the list; and

train a machine learning model using the training data, wherein the machine learning model is configured to predict a fault type and a fault location of the application based on an input comprising log data of the application.

12 . The computer program product of claim 11 , wherein:

the normal execution times of the respective services and the normal execution times of the respective network communications are determined using tracing.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2023
From: SHWARTZ, LARISA; JHA, SAURABH; RIOS ALIAGA, JESUS MARIA; FARCHI, EITAN DANIEL; BAGEHORN, FRANK; FILEPP, ROBERT
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 063140/0726 →
Continuity (1)
Related Publication 20240330093A1 · Oct 3, 2024
References Cited (17)
US 9270557B1 · Shpilyuck · 2016 [cited by examiner]
US 11934258B1 · Kasaiezadeh Mahabadi · 2024 [cited by examiner]
US 20090089805A1 · Srinivasan · 2009 [cited by examiner]
US 20100017791A1 · Finkler · 2010 [cited by examiner]
US 20170024299A1 · Deng et al. · 2017 [cited by applicant]
US 20220188219A1 · Hicks · 2022 [cited by examiner]
US 20230040564A1 · Wang et al. · 2023 [cited by applicant]
US 20230208855A1 · Sheriff · 2023 [cited by examiner]
CN 110262972 · 2020 [cited by applicant]
Fault propagation timing (Year: 1988). [cited by examiner]
Shin et al, ault propagation timing (Year: 1988). [cited by examiner]
Lee et al., “Eadro: An End-to-End Troubleshooting Framework for Microservices on Multi-source Data”, arXiv preprint arXiv:2302.05092, Feb. 10, 2023, 13 pages. [cited by applicant]
Anonymous, “A Method to Inject Faults Into an Application”, IP.com No. IPCOM000140338D, Sep. 8, 2006, 4 pages. [cited by applicant]
Bagehorn et al., “A fault injection platform for learning AIOps models”, In 37th IEEE/ACM International Conference on Automated Software Engineering, 2022, 5 pages. [cited by applicant]
Grabarnik et al., “Management of Service Process QoS in a Service Provider—Service Supplier Environment”, In The 9th IEEE International Conference on E-Commerce Technology and The 4th IEEE International Conference on En… [cited by applicant]
Rios et al., “Localizing and Explaining Faults in Microservices Using Distributed Tracing”, 2022, 11 pages. [cited by applicant]
Garcia Solla, “What is a Call Graph? And How to Generate them Automatically”, https://www.freecodecamp.org/news/how-to-automate-call-graph-creation/, archived on Mar. 7, 2023, 24 pages. [cited by applicant]