IP Library Granted Patent US 12,554,606
Granted Patent B2
US 12,554,606 · App. 18/202,079 · Granted Feb 17, 2026

Resiliency testing of applications and compute infrastructures

Inventors: Sandeep Hans (New Delhi, IN); Mudit Verma (New Delhi, IN); Samuel Solomon Ackerman (Haifa, IL); Diptikalyan Saha (Bangalore, IN); Eitan Daniel Farchi (Pardes Hana, IL); Praveen Jayachandran (Bangalore, IN)
Assignee: International Business Machines Corporation
G06F11/263
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,554,606
App. No.
18/202,079
Filed
May 25, 2023
Granted
Feb 17, 2026
Kind
B2
Art Unit
2113
USPC
714/33
Abstract

A computer-implemented method, according to one embodiment, includes: intentionally causing faults to be injected in a compute infrastructure, and determining whether the injected faults cause application failures. Weights are also assigned to the injected faults based on severity of the respective application failures. The weighted faults are compared, and changes to the compute infrastructure are recommended based on the comparison. Moreover, the changes that are recommended are configured to prevent the application failures. Other systems, methods, and computer program products are described in additional embodiments.

Claims (46)

1 . A computer-implemented method, comprising:

intentionally causing faults to be injected in a compute infrastructure;

determining whether the injected faults cause application failures;

assigning weights to the injected faults based on severity of the respective application failures, wherein a negative or zero weight is assigned to ones of the injected faults that cause a change to compute infrastructure behavior, and do not cause an application error;

comparing the weighted faults;

using an artificial intelligence (AI) model that has been trained to generate physical and/or logical changes to the compute infrastructure based at least in part on the comparison, the injected faults, and/or performance of the compute infrastructure and application(s) in response to the faults being injected, wherein the physical and/or logical changes are configured to prevent the application failures; and

causing the physical and/or logical changes to be implemented in the compute infrastructure.

2 . The computer-implemented method of claim 1 , wherein assigning weights to the injected faults includes:

creating a heat map having entries identifying whether respective ones of the injected faults cause application errors;

determining weights for the injected faults identified in the heat map as causing application errors; and

correlating the determined weights with the respective injected faults identified in the heat map as causing the application errors.

3 . The computer-implemented method of claim 2 , wherein the weights are determined based at least in part on an amount of time the application experienced respective errors, wherein the weights increase as the amount of time the compute infrastructure experiences respective errors increases.

4 . The computer-implemented method of claim 3 , wherein the injected faults include infrastructure faults that intentionally subject the compute infrastructure to strain.

5 . The computer-implemented method of claim 2 , wherein the weights are determined based on the outcome and distribution difference of microservice internal calls.

6 . The computer-implemented method of claim 1 ,

wherein the AI model is trained to generate the physical and/or logical changes to the compute infrastructure based on the comparison, the injected faults, and the performance of the compute infrastructure and application in response to injecting the faults.

7 . The computer-implemented method of claim 1 , wherein the injected faults are selected based on faults previously injected in the compute infrastructure, wherein the injected faults relate to parameters of the compute infrastructure, the parameters being selected from the group consisting of: system errors, fault codes, test workloads, application topologies, and key performance indicators.

8 . A computer program product, comprising a computer readable storage medium having program instructions embodied therewith, the program instructions readable by a processor, executable by the processor, or readable and executable by the processor, to cause the processor to:

intentionally cause faults to be injected in a compute infrastructure;

determine whether the injected faults cause application failures;

assign weights to the injected faults based on severity of the respective application failures, wherein a negative or zero weight is assigned to ones of the injected faults that cause a change to compute infrastructure behavior, and do not cause an application error;

compare the weighted faults;

use an artificial intelligence (AI) model that has been trained to generate physical and/or logical changes to the compute infrastructure based at least in part on the comparison, the injected faults, and/or performance of the compute infrastructure and application(s) in response to the faults being injected, wherein the physical and/or logical changes are configured to prevent the application failures; and

cause the physical and/or logical changes to be implemented in the compute infrastructure.

9 . The computer program product of claim 8 , wherein assigning weights to the injected faults includes:

creating a heat map having entries identifying whether respective ones of the injected faults cause application errors;

determining weights for the injected faults identified in the heat map as causing application errors; and

correlating the determined weights with the respective injected faults identified in the heat map as causing the application errors.

10 . The computer program product of claim 9 , wherein the weights are determined based at least in part on an amount of time the application experienced respective errors.

11 . The computer program product of claim 10 , wherein the weights increase as the amount of time the compute infrastructure experiences respective errors increases.

12 . The computer program product of claim 9 , wherein the weights are determined based on the outcome and distribution difference of microservice internal calls.

13 . The computer program product of claim 8 , wherein

the AI model is trained to generate the physical and/or logical changes to the compute infrastructure based on the comparison, the injected faults, and the performance of the compute infrastructure and application in response to injecting the faults.

14 . The computer program product of claim 8 , wherein the injected faults are selected based on faults previously injected in the compute infrastructure, wherein the injected faults are selected from the group consisting of: system errors, fault codes, test workloads, application topologies, and key performance indicators.

15 . The computer program product of claim 14 , wherein the injected faults relate to parameters of the compute infrastructure, the parameters being selected from the group consisting of: system errors, fault codes, test workloads, application topologies, and key performance indicators.

16 . A system, comprising:

a processor; and

logic integrated with the processor, executable by the processor, or integrated with and executable by the processor, the logic being configured to:

intentionally cause faults to be injected in a compute infrastructure;

determine whether the injected faults cause application failures;

assign weights to the injected faults based on severity of the respective application failures, wherein a negative or zero weight is assigned to ones of the injected faults that cause a change to compute infrastructure behavior, and do not cause an application error;

compare the weighted faults;

use an artificial intelligence (AI) model that has been trained to generate physical and/or logical changes to the compute infrastructure based at least in part on the comparison, the injected faults, and/or performance of the compute infrastructure and application(s) in response to the faults being injected, wherein the physical and/or logical changes are configured to prevent the application failures; and

cause the physical and/or logical changes to be implemented in the compute infrastructure.

17 . The system of claim 16 , wherein

the AI model is trained to generate the physical and/or logical changes to the compute infrastructure based on the comparison, the injected faults, and the performance of the compute infrastructure and application in response to injecting the faults.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 31, 2023
From: HANS, SANDEEP; VERMA, MUDIT; ACKERMAN, SAMUEL SOLOMON; SAHA, DIPTIKALYAN; FARCHI, EITAN DANIEL; JAYACHANDRAN, PRAVEEN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 063809/0739 →
Continuity (1)
Related Publication 20240394162A1 · Nov 28, 2024
References Cited (17)
US 10684940B1 · Kayal et al. · 2020 [cited by applicant]
US 11755402B1 · Chavan · 2023 [cited by examiner]
US 11934258B1 · Kasaiezadeh Mahabadi · 2024 [cited by examiner]
US 20060271825A1 · Keaffaber · 2006 [cited by examiner]
US 20110239048A1 · Andrade · 2011 [cited by examiner]
US 20170242784A1 · Heorhiadi et al. · 2017 [cited by applicant]
US 20190149622A1 · Kumar et al. · 2019 [cited by applicant]
US 20220231904A1 · Di Martino · 2022 [cited by applicant]
CN 113935178A · 2022 [cited by applicant]
Hernandez-Serrato et al., “Applying Machine Learning with Chaos Engineering,” IEEE International Symposium on Software Reliability Engineering Workshops (ISSREW), 2020, pp. 151-152. [cited by applicant]
Basiri et al., “Automating chaos experiments in production,” arXiv, 2019, 10 pages, retrieved from https://arxiv.org/pdf/1905.04648.pdf. [cited by applicant]
Zhang et al., “Chaos Engineering of Ethereum Blockchain Clients,” arXiv, 2021, 14 pages, retrieved from https://arxiv.org/pdf/2111.00221.pdf. [cited by applicant]
Jernberg et al., “Getting Started with Chaos Engineering—design of an implementation framework in practice,” Empirical Software Engineering and Measurement Conference (ESEM), Oct. 2020, 10 pages. [cited by applicant]
Torkura et al., “CloudStrike: Chaos Engineering for Security and Resiliency in Cloud Infrastructure,” IEEE Access, vol. 8, Jul. 6, 2020, pp. 123044-123060. [cited by applicant]
Canonico et al., “Human-AI Partnerships for Chaos Engineering,” IEEE/ACM 42nd International Conference on Software Engineering Workshops (ICSEW), 2020, pp. 499-053. [cited by applicant]
Anonymous, “Principles of chaos engineering,” Last update Mar. 2019, 3 pages, retrieved from http://principlesofchaos.org/. [cited by applicant]
“Middleware”, 24th ACM/IFIP International Middleware Conference, Dec. 11-15, 2023, 03 pages, https://middleware-conf.github.io/2023/. [cited by applicant]