IP Library › Granted Patent US 11,573,848
Granted Patent B2
US 11,573,848 · App. 17/094,666 · Granted Feb 7, 2023

Identification and/or prediction of failures in a microservice architecture for enabling automatically-repairing solutions

Inventors: Nicholas Linck (San Jose, CA); Sangeetha Seshadri (Plano, CA); Paul Henri Muench (San Jose, CA); Umesh Deshpande (San Jose, CA); Priyaranjan Behera (Santa Clara, CA); Wilfred Edmund Plouffe, Jr. (San Jose, CA)
Assignee: International Business Machines Corporation
G06F11/079G06F11/0772G06F11/0787G06F11/326G06K9/6262
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,573,848
App. No.
17/094,666
Filed
Nov 10, 2020
Granted
Feb 7, 2023
Kind
B2
Art Unit
2113
USPC
714/37
Abstract

A computer-implemented method according to one embodiment includes causing a failure event in each of a plurality of microservices of a system and collecting failure effect data associated with the caused failure events. A mapping is created detailing transition of the microservices between different states and the collected failure effect data is analyzed for creating the mapping. The method further includes outputting a predetermined notification in response to a determination that a first of the microservices is close to experiencing a predicted failure event, and outputting a suggested solution for repairing the system in response to a determination that the system has failed, using the mapping to identify a root cause of the system failure. Using the mapping to identify the root cause of the system failure includes identifying the microservices that caused the system failure.

Claims (42)

1. A computer-implemented method, comprising:

causing a failure event in each of a plurality of microservices of a system;

collecting failure effect data associated with the caused failure events;

creating a mapping detailing transition of the microservices between different states, wherein the collected failure effect data is analyzed for creating the mapping, wherein creating the mapping includes executing a reinforcement learning algorithm;

in response to a determination that a first of the microservices is close to experiencing a predicted failure event, outputting a predetermined notification; and

in response to a determination that the system has failed, using the mapping to identify a root cause of the system failure, and outputting a suggested solution for repairing the system, wherein using the mapping to identify the root cause of the system failure includes identifying the microservices that caused the system failure.

2. The computer-implemented method of claim 1 , wherein the failure effect data indicates the states of each of the microservices after each of the caused failure events, wherein identifying the microservices that caused the system failure includes comparing the states of each of the microservices in the failure effect data with states of the microservices after the system failure.

3. The computer-implemented method of claim 2 , wherein the state of at least one of the microservices after a caused failure event matching the state of a second of the microservices after the system failure identifies the second microservice as having caused the system failure.

4. The computer-implemented method of claim 1 , comprising: instructing a forced micro-reset of the first microservice in response to the determination that the first microservice is close to experiencing the predicted failure event.

5. The computer-implemented method of claim 1 , wherein the determination that the first microservice is close to experiencing the predicted failure event is based on the collected failure effect data and includes comparing a current state of the first microservice to potential actions of the first microservice previously determined to cause the first microservice to change to a different state.

6. The computer-implemented method of claim 1 , comprising: identifying avoidance of a failure event by one of the microservices; and logging information detailing the failure event avoidance and success rates of the failure event avoidance.

7. The computer-implemented method of claim 1 , wherein the reinforcement learning algorithm analyses the collected failure effect data to create the mapping showing how the microservices move between the states.

8. The computer-implemented method of claim 1 , wherein the suggested solution includes a summary of the identified root cause of the system failure.

9. A computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions readable and/or executable by a computer to cause the computer to:

cause, by the computer, a failure event in each of a plurality of microservices of a system;

collect, by the computer, failure effect data associated with the caused failure events;

create, by the computer, a mapping detailing transition of the microservices between different states, wherein the collected failure effect data is analyzed for creating the mapping, wherein creating the mapping includes executing a reinforcement learning algorithm;

in response to a determination that a first of the microservices is close to experiencing a predicted failure event, output, by the computer, a predetermined notification; and

in response to a determination that the system has failed, use, by the computer, the mapping to identify a root cause of the system failure, and output, by the computer, a suggested solution for repairing the system, wherein using the mapping to identify the root cause of the system failure includes identifying the microservices that caused the system failure.

10. The computer program product of claim 9 , wherein the failure effect data indicates the states of each of the microservices after each of the caused failure events, wherein identifying the microservices that caused the system failure includes comparing the states of each of the microservices in the failure effect data with states of the microservices after the system failure.

11. The computer program product of claim 10 , wherein the state of at least one of the microservices after a caused failure event matching the state of a second of the microservices after the system failure identifies the second microservice as having caused the system failure.

12. The computer program product of claim 9 , wherein the reinforcement learning algorithm analyses the collected failure effect data to create the mapping showing how the microservices move between the states, and the program instructions readable and/or executable by the computer to cause the computer to:

instruct, by the computer, a forced micro-reset of the first microservice in response to the determination that the first microservice is close to experiencing the predicted failure event.

13. The computer program product of claim 9 , wherein the determination that the first microservice is close to experiencing the predicted failure event is based on the collected failure effect data and includes comparing a current state of the first microservice to potential actions of the first microservice capable of causing the first microservice to change to a different state.

14. The computer program product of claim 9 , the program instructions readable and/or executable by the computer to cause the computer to:

identify, by the computer, avoidance of a failure event by one of the microservices; and

log, by the computer, information detailing the failure event avoidance and success rates of the failure event avoidance.

15. A system, comprising:

a processor; and

logic integrated with the processor, executable by the processor, or integrated with and executable by the processor, the logic being configured to:

cause a failure event in each of a plurality of microservices of the system;

collect failure effect data associated with the caused failure events;

create a mapping detailing transition of the microservices between different states, wherein the collected failure effect data is analyzed for creating the mapping, wherein creating the mapping includes executing a reinforcement learning algorithm;

in response to a determination that a first of the microservices is close to experiencing a predicted failure event, output a predetermined notification; and

in response to a determination that the system has failed, use the mapping to identify a root cause of the system failure, and output a suggested solution for repairing the system,

wherein using the mapping to identify the root cause of the system failure includes identifying the microservices that caused the system failure.

16. The system of claim 15 , wherein the failure effect data indicates the states of each of the microservices after each of the caused failure events, wherein identifying the microservices that caused the system failure includes comparing the states of each of the microservices in the failure effect data with states of the microservices after the system failure.

17. The system of claim 16 , wherein the state of at least one of the microservices after a caused failure event matching the state of a second of the microservices after the system failure identifies the second microservice as having caused the system failure.

18. The system of claim 15 , the logic being configured to: instruct a forced micro-reset of the first microservice in response to the determination that the first microservice is close to experiencing the predicted failure event.

19. The system of claim 15 , wherein the determination that the first microservice is close to experiencing the predicted failure event is based on the collected failure effect data and includes comparing a current state of the first microservice to potential actions of the first microservice capable of causing the first microservice to change to a different state.

20. The system of claim 15 , the logic being configured to:

identify avoidance of a failure event by one of the microservices; and log information detailing the failure event avoidance and success rates of the failure event avoidance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2020
From: LINCK, NICHOLAS; SESHADRI, SANGEETHA; MUENCH, PAUL HENRI; DESHPANDE, UMESH; BEHERA, PRIYARANJAN; PLOUFFE, WILFRED EDMUND, JR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054593/0888 →
Continuity (1)
Related Publication 20220147409A1 · May 12, 2022
Cited By (33)
US 12,197,859 US 12,198,030 US 12,204,323 US 12,292,811 US 12,299,140 US 12,321,862 US 12,346,820 US 12,361,334 US 12,361,335 US 12,367,292 US 12,443,894 US 12,450,494 US 12,505,291 US 12,505,352 US 12,517,724 US 12,524,508 US 12,587,489 US 12,587,490 US 12,592,897 US 12,596,738 US 12,596,813 US 12,602,418 US 12,602,624 US 12,608,486 US 12,614,124 US 12,621,253 US 12,650,893 US 12,675,360 US 12,681,830 US 12,694,133 US 12,694,343 US 12,725,095 US 12,748,827