IP Library › Granted Patent US 11,609,811
Granted Patent B2
US 11,609,811 · App. 17/004,566 · Granted Mar 21, 2023

Automatic root cause analysis and prediction for a large dynamic process execution system

Inventors: Sanjay Ramanujan (Sammamish, WA); Andrew Tianze Wang (Potomac, MD); Marwan Elias Jubran (Kirkland, WA); Weiping Hu (Bellevue, WA); Xiaoguang Fan (Kirkland, WA)
G06F11/079G06F9/4881G06F11/0721G06F11/0793
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,609,811
App. No.
17/004,566
Granted
Mar 21, 2023
Kind
B2
Abstract

An automated root-cause analysis (RCA) system may provide a fully automated platform that provides dependency and execution order modeling for tasks included in a capacity provisioning process, anomaly detection, ticket correlation, root-cause analysis, monitoring and feedback, and data visualization. The automated RCA system may continuously collect and store data for use in determining a root cause of a blockage on a capacity provisioning process. The blockage may be identified in a ticket generated by a cloud-computing system. The automated RCA system may receive the ticket and attempt to determine the root cause of the blockage based on root causes associated with previous tickets generated by the cloud-computing system. The automated RCA system may identify a true root cause, recommend repair items based on the true root cause, identify one or more responsible teams to drive a fix, and provide an estimated time for completion.

Claims (69)

1. A system for automatically determining a root cause of a failure associated with a capacity provisioning process in a cloud-computing system, the system comprising:

one or more processors;

memory in electronic communication with the one or more processors;

a data store, wherein the data store includes historical tickets and root-cause records associated with one or more of the historical tickets; and

instructions stored in the memory, the instructions being executable by the one or more processors to:

receive events emitted by tasks associated with the capacity provisioning process, wherein the events describe a status of the tasks;

receive health signals from one or more external systems, wherein the health signals indicate whether the one or more external systems are operational;

receive a ticket, wherein the ticket identifies the failure associated with the capacity provisioning process;

determine the root cause of the failure based on the events, the health signals, the historical tickets, and the root-cause records, wherein determining the root cause of the failure comprises pattern-matching the ticket against the historical tickets and analyzing an execution graph that indicates an execution order of tasks included in the capacity provisioning process, dependencies among the tasks, and dependencies of the tasks on cloud-computing subsystems;

associate the root cause of the failure with the ticket; and

provide the root cause of the failure to operations.

2. The system of claim 1 , wherein the instructions to determine the root cause of the failure based on the events, the health signals, the historical tickets, and the root-cause records comprise instructions executable by the one or more processors to:

identify a matching ticket among the historical tickets, wherein the matching ticket has one or more similarities to the ticket;

identify a root-cause record associated with the matching ticket, wherein the root-cause record associated with the matching ticket identifies a root cause of the matching ticket; and

determine the root cause of the failure based in part on the root-cause record associated with the matching ticket.

3. The system of claim 2 , wherein the instructions are further executable by the one or more processors to:

identify additional matching tickets among the historical tickets, wherein the additional matching tickets have one or more similarities to the ticket;

identify one or more root-cause records associated with the additional matching tickets, wherein the one or more root cause records associated with the additional matching tickets identify root causes for the additional matching tickets; and

prune the one or more root-cause records based on time and spatial correlation.

4. The system of claim 2 , wherein the instructions to identify the matching ticket among the historical tickets comprise instructions executable by the one or more processors to identify the matching ticket among the historical tickets using a machine-learning engine.

5. The system of claim 2 , wherein the instructions are further executable by the one or more processors to:

determine repairs for the failure based on the root-cause record associated with the matching ticket;

identify one or more repair teams for the repairs; and

provide the root cause of the failure and the repairs to the one or more repair teams.

6. The system of claim 5 , wherein the instructions are further executable by the one or more processors to:

prioritize the repairs versus other repairs associated with the cloud-computing system.

7. The system of claim 1 , wherein the instructions are further executable by the one or more processors to:

determine that the root-cause records do not contain the root cause of the failure;

generate the root cause of the failure based on the events or the health signals; and

store the root cause of the failure in the data store.

8. The system of claim 7 , wherein the instructions are further executable by the one or more processors to:

collect tags associated with one or more of the historical tickets, wherein a tag comprises information about a root cause of a historical ticket and was entered according to a defined schema and wherein generating the root cause of the failure is based in part on the tags.

9. The system of claim 1 , wherein the events emitted by the tasks associated with the capacity provisioning process for the cloud-computing system and the health signals from the one or more external systems are received according to a data contract.

10. The system of claim 1 , wherein the instructions are further executable by the one or more processors to:

store the events on the data store; and

store the health signals on the data store.

11. The system of claim 1 , wherein the instructions are further executable by the one or more processors to:

display the execution graph for the capacity provisioning process on a presentation system; and

display the root cause of the failure on the presentation system.

12. The system of claim 1 , wherein the instructions are further executable by the one or more processors to:

query the external systems for metadata in response to receiving the health signals; and

store the metadata on the data store.

13. A method executed by a computing device for automatically determining a root cause of a failure associated with a capacity provisioning process, wherein the capacity provisioning process is for a cloud-computing system, the method comprising:

receiving a ticket, the ticket indicating the failure;

identifying zero or more potential root causes of the failure, wherein the zero or more potential root causes are described in root-cause records associated with one or more historical tickets and execution graphs of capacity provisioning processes that indicate an execution order of tasks included in the capacity provisioning process, dependencies among the tasks, and dependencies of the tasks on cloud-computing subsystems;

assigning scores to the zero or more potential root causes of the failure;

determining whether the zero or more potential root causes include the root cause based on the scores;

determining, if the zero or more potential root causes do not include the root cause of the failure, whether the root cause of the failure can be generated without human intervention;

determining whether repairs for the failure are known based on the root-cause records associated with the one or more historical tickets;

seeking human intervention if either the root cause or the repairs are unknown;

attaching, if the root cause is known, the root cause to the ticket; and

identifying, if the repairs are known, one or more repair teams to complete the repairs.

14. The method of claim 13 further comprising generating the root cause of the failure based on events received from tasks associated with the capacity provisioning process, health signals received from external systems that support the cloud-computing system, and tags associated with the historical tickets.

15. The method of claim 13 , wherein the one or more historical tickets are included in a set of historical tickets and identifying the zero or more potential root causes for the failure comprises:

identifying the one or more historical tickets based on comparing the ticket to the set of historical tickets; and

identifying the root-cause records associated with the one or more historical tickets.

16. The method of claim 15 , wherein identifying the one or more historical tickets based on comparing the ticket to the set of historical tickets comprises identifying the one or more historical tickets based on performing a similarity analysis between the ticket and the set of historical tickets.

17. The method of claim 15 , wherein identifying the one or more historical tickets based on comparing the ticket to the set of historical tickets comprises identifying the one or more historical tickets from among the set of historical tickets using a machine-learning engine.

18. The method of claim 13 , wherein the scores assigned to the zero or more potential root causes for the failure are based on events received from tasks associated with the capacity provisioning process, health signals received from external systems that support the cloud-computing system, and tags associated with the historical tickets.

19. A non-transitory computer-readable medium comprising instructions that are executable by one or more processors to cause a computing system to:

receive events emitted by tasks associated with a capacity provisioning process for a cloud-computing system, wherein the events describe a status of the tasks;

receive health signals from one or more external systems, wherein the health signals indicate whether the one or more external systems are operational;

receive a ticket, wherein the ticket identifies a failure associated with the capacity provisioning process;

access a data store storing historical tickets of the cloud-computing system, wherein one or more of the historical tickets have associated root-cause records;

determine a root cause of the failure based on the events, the health signals, the historical tickets, and the associated root-cause records, wherein determining the root cause of the failure comprises pattern-matching the ticket against the historical tickets and analyzing an execution graph that indicates an execution order of tasks included in the capacity provisioning process, dependencies among the tasks, and dependencies of the tasks on cloud-computing subsystems;

determine necessary repairs for the failure based on the associated root-cause records; and

determine an estimated time for the necessary repairs to be complete.

20. The non-transitory computer-readable medium of claim 19 further comprising instructions that are executable by the one or more processors to cause the computing system to:

perform the necessary repairs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2020
From: RAMANUJAN, SANJAY; WANG, ANDREW TIANZE; JUBRAN, MARWAN ELIAS; HU, WEIPING; FAN, XIAOGUANG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 053616/0890 →
Continuity (1)
Related Publication 20220066852A1 · Mar 3, 2022
Cited By (4)
US 12,455,958 US 12,462,024 US 12,511,186 US 12,513,039