IP Library › Granted Patent US 9,483,383
Granted Patent B2
US 9,483,383 · App. 14/097,713 · Granted Nov 1, 2016

Injecting faults at select execution points of distributed applications

Inventors: Salman A. Baset (New York, NY); Cuong M. Pham (Urbana, IL); Harigovind V. Ramasamy (Ossining, NY); Manas Singh (Chappaqua, NY); Byung Chul Tak (Peekskill, NY); Chunqiang Tang (Ossining, NY); Long Wang (White Plains, NY)
Assignee: International Business Machines Corporation
G06F11/3644G06F11/2023G06F11/2028G06F11/2033G06F11/3006G06F11/3612H04L43/0805
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,483,383
App. No.
14/097,713
Granted
Nov 1, 2016
Kind
B2
Abstract

Methods, systems, and articles of manufacture for injecting faults at select execution points of distributed applications are provided herein. A method includes monitoring a run-time state of each of multiple components of a distributed application to determine one or more sequence of events that triggers a fault injection point at one of the multiple components; defining a fault injection scenario in a specification based on said monitoring, wherein said fault injection scenario comprises a description of one or more sequence of events during which an intended fault is to be injected to a target component of the multiple components at one selected event; and executing the fault injection defined in the specification to perform injection of the intended fault during run-time of the distributed application.

Claims (25)

1. A method comprising:

monitoring a run-time state of each of multiple distributed components of a distributed application that comprises a variable number of multiple states, based on collected execution traces of the distributed application, to determine one or more sequence of events that triggers a fault injection point at each of the multiple components, wherein said one or more sequence of events comprises a sequence of log events;

defining a scenario of multiple fault injections in a specification based on said monitoring, wherein said scenario comprises a description of one or more sequence of events during which each of the multiple faults is to be injected across the multiple distributed components at one selected event, and wherein said defining comprises defining the scenario based on one or more event dependencies discovered from said monitoring; and

executing the multiple fault injections defined in the specification concurrently across the multiple distributed components of the distributed application during run-time of the distributed application;

wherein said monitoring, said defining, and said executing are carried out by at least one computing device.

2. The method of claim 1 , wherein said monitoring comprises monitoring the run-time state of each of the multiple distributed components of the distributed application under multiple workloads.

3. The method of claim 1 , wherein each of the multiple faults comprises at least one of an application crash, network congestion, a lack of memory and/or disk space, a memory, disk or network corruption and a configuration error.

4. The method of claim 1 , wherein each of the multiple distributed components comprises at least one of a given function, a given process, a virtual machine, and a physical machine.

5. The method of claim 1 , comprising:

analyzing results from said multiple fault injections.

6. The method of claim 5 , wherein said analyzing comprises collecting multiple logs derived from said multiple fault injections.

7. The method of claim 6 , comprising:

displaying said multiple logs to a user.

8. The method of claim 5 , wherein said analyzing comprises generating statistics pertaining to said multiple fault injections.

9. An article of manufacture comprising a non-transitory computer readable storage medium having computer readable instructions tangibly embodied thereon which, when implemented, cause a computer to carry out a plurality of method steps comprising:

monitoring a run-time state of each of multiple distributed components of a distributed application that comprises a variable number of multiple states, based on collected execution traces of the distributed application, to determine one or more sequence of events that triggers a fault injection point at each of the multiple components, wherein said one or more sequence of events comprises a sequence of log events;

defining a scenario of multiple fault injections in a specification based on said monitoring, wherein said scenario comprises a description of one or more sequence of events during which each of the multiple faults is to be injected across the multiple distributed components at one selected event, and wherein said defining comprises defining the scenario based on one or more event dependencies discovered from said monitoring; and

executing the multiple fault injections defined in the specification concurrently across the multiple distributed components of the distributed application during run-time of the distributed application.

10. The article of manufacture of claim 9 , wherein said monitoring comprises monitoring the run-time state of each of the multiple distributed components of the distributed application under multiple workloads.

11. A system comprising:

a memory; and

at least one processor coupled to the memory and configured for:

monitoring a run-time state of each of multiple distributed components of a distributed application that comprises a variable number of multiple states, based on collected execution traces of the distributed application, to determine one or more sequence of events that triggers a fault injection point at each of the multiple components, wherein said one or more sequence of events comprises a sequence of log events;

defining a scenario of multiple fault injections in a specification based on said monitoring, wherein said scenario comprises a description of one or more sequence of events during which each of the multiple faults is to be injected across the multiple distributed components at one selected event, and wherein said defining comprises defining the scenario based on one or more event dependencies discovered from said monitoring; and

executing the multiple fault injections defined in the specification concurrently across the multiple distributed components of the distributed application during run-time of the distributed application.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2014
From: BASET, SALMAN A.; PHAM, CUONG M.; RAMASAMY, HARIGOVIND V.; SINGH, MANAS; TAK, BYUNG CHUL; TANG, CHUNQIANG; WANG, LONG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 031905/0120 →
Continuity (1)
Related Publication 20150161025A1 · Jun 11, 2015