IP Library › Granted Patent US 11,962,456
Granted Patent B2
US 11,962,456 · App. 17/382,072 · Granted Apr 16, 2024

Automated cross-service diagnostics for large scale infrastructure cloud service providers

Inventors: Zhangwei Xu (Redmond, WA); Xiaofeng Gao (Redmond, WA); Cary L. Mitchell (Bellevue, WA); Steve J. Lunsman (Redmond, WA); Tony V. Perez (Bothell, WA); Arvind Narasimhan (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
H04L41/065H04L41/0654H04L41/0686H04L41/22H04L43/091
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,962,456
App. No.
17/382,072
Granted
Apr 16, 2024
Kind
B2
Abstract

Example aspects include techniques for employing cross-service diagnostics for cloud service providers. These techniques may include dynamically generating a workflow of one or more diagnostic modules based on relationship information between an origin service experiencing an incident and one or more related services that the origin service depends on, and executing the workflow of one or more diagnostic modules to determine a root cause of the incident, each of the one or more diagnostic modules implemented by an individual service of the one or more related services in accordance with a schema. In addition, the techniques may include determining a diagnostic action based on the root cause, and transmitting, based on the diagnostic action, an engagement notification to a responsible entity.

Claims (69)

1. A cloud computing platform comprising:

a memory storing instructions; and

at least one processor coupled with the memory and configured to execute the instructions to:

in response to detecting an incident affecting an origin service implemented by the cloud computing platform, invoke a first diagnostic module implemented by the origin service, wherein invoking the first diagnostic module generates a health status of the origin service and identifies one or more related services of the cloud computing platform upon which the origin service depends, the health status indicating whether the origin service is healthy;

based on the health status of the origin service, dynamically generate a workflow including the first diagnostic module and a plurality of diagnostic modules implemented by the one or more related services, the workflow being based on relationship information that indicates a programmatic dependency between the origin service and the one or more related services, the programmatic dependency indicating that the origin service uses the one or more related services to execute a task of the origin service;

determine a priority order for invoking the plurality of diagnostic modules based on the programmatic dependency between the origin service and the one or more related services;

invoke the plurality of diagnostic modules according to the priority order to generate output information indicating respective health statuses of the one or more related services and respective programmatic dependencies on additional services related to the one or more related services, the respective health statuses indicating whether the one or more related services are healthy;

determine a root cause of the incident based on the output information;

determine a diagnostic action based on the root cause; and

transmit, based on the diagnostic action, an engagement notification to a responsible entity.

2. The cloud computing platform of claim 1 , wherein to dynamically generate the workflow, the at least one processor is configured to:

receive at least a portion of the relationship information from the first diagnostic module, the portion of the relationship information identifying a first related service of the one or more related services that the origin service depends on; and

add a second diagnostic module to the workflow based at least in part on the portion of the relationship information, the second diagnostic module implemented by the first related service.

3. The cloud computing platform of claim 1 , wherein to invoke the plurality of diagnostic modules, the at least one processor is configured to:

send incident information and a resource identifier of a resource affected by the incident to the plurality of diagnostic modules.

4. The cloud computing platform of claim 1 , wherein to determine the diagnostic action based on the root cause, the at least one processor is configured to:

determine an accuracy value of the root cause; and

compare the accuracy value to a preconfigured threshold.

5. The cloud computing platform of claim 1 , wherein to determine the diagnostic action, the at least one processor is configured to:

identify a source service that is associated with the root cause;

identify action preference information associated with the source service; and

determine the diagnostic action based at least in part on the action preference information.

6. The cloud computing platform of claim 1 , wherein the diagnostic action includes at least one of an incident enrichment action, an incident transfer action, an incident linking action, or an incident auto-invite action.

7. The cloud computing platform of claim 1 , wherein the root cause is a first root cause, the diagnostic action is a first diagnostic action, the engagement notification is a first engagement notification, the responsible entity is a first responsible entity, and the at least one processor is configured to:

receive, via a user interface, an update to incident information used to at least one of:

dynamically generate the workflow; or

execute the plurality of diagnostic modules;

re-invoke, based on the update to the incident information, the plurality of diagnostic modules to determine a second root cause;

determine a second diagnostic action based on the second root cause; and

transmit, based on the second diagnostic action, a second engagement notification to a second responsible entity.

8. The cloud computing platform of claim 1 , wherein the root cause is a first root cause, the diagnostic action is a first diagnostic action, the engagement notification is a first engagement notification, the responsible entity is a first responsible entity, and the at least one processor is configured to:

identify one or more potential errors within initial incident information based on a user update to the initial incident information;

correct the one or more potential errors within the initial incident information to determine updated incident information;

re-invoke, based on the updated incident information, the plurality of diagnostic modules to determine a second root cause;

determine a second diagnostic action based on the second root cause; and

transmit, based on the second diagnostic action, a second engagement notification to a second responsible entity.

9. The cloud computing platform of claim 1 , wherein the at least one processor is further configured to:

present, via a graphical user interface, a visual representation of the workflow, the visual representation displaying each of the plurality of diagnostic modules with a health status.

10. A method comprising:

in response to detecting an incident affecting an origin service implemented by the cloud computing platform, invoke a first diagnostic module implemented by the origin service, wherein invoking the first diagnostic module generates a health status of the origin service and identifies one or more related services of the cloud computing platform upon which the origin service depends, the health status indicating whether the origin service is healthy;

based on the health status of the origin service, dynamically generating a workflow of the first diagnostic module and a plurality of diagnostic modules implemented by the one or more related services, the workflow being based on relationship information that indicates a programmatic dependency between the origin service and the one or more related services, the programmatic dependency indicating that the origin service uses the one or more related services to execute a task of the origin service;

determine a priority order for invoking the plurality of diagnostic modules based on the programmatic dependency between the origin service and the one or more related services;

invoking the plurality of diagnostic modules according to the priority order to generate output information indicating respective health statuses of the one or more related services and respective programmatic dependencies on additional services related to the one or more related services, the respective health statuses indicating whether the one or more related services are healthy;

determine a root cause of the incident based on the output information;

determining a diagnostic action based on the root cause; and

transmitting, based on the diagnostic action, an engagement notification to a responsible entity.

11. The method of claim 10 , wherein dynamically generating the workflow comprises:

requesting the relationship information from the one or more related services.

12. The method of claim 10 , wherein invoking the plurality of diagnostic modules, comprises:

sending incident information and a resource identifier of a resource affected by the incident to a diagnostic module of the plurality of diagnostic modules.

13. The method of claim 10 , wherein invoking the plurality of diagnostic modules comprises:

determining an accuracy value of the root cause; and

comparing the accuracy value and a predefined threshold.

14. The method of claim 10 , wherein the diagnostic action includes at least one of an incident enrichment action, an incident transfer action, an incident linking action, or an incident auto-invite action.

15. The method of claim 10 , presenting, via a graphical user interface, a visual representation of the workflow, the visual representation displaying each of the plurality of diagnostic modules with a health status.

16. A non-transitory computer-readable device having instructions thereon that, when executed by at least one computing device, causes the at least one computing device to perform operations comprising:

in response to detecting an incident affecting an origin service implemented by the cloud computing platform, invoking a first diagnostic module implemented by the origin service, wherein invoking the first diagnostic module generates a health status of the origin service and identifies one or more related services of the cloud computing platform upon which the origin service depends, the health status indicating whether the origin service is healthy;

based on the health status of the origin service, dynamically generating a workflow of the first diagnostic module and a plurality of diagnostic modules implemented by the one or more related services, the workflow being based on relationship information that indicates a programmatic dependency between the origin service and the one or more related services, the programmatic dependency indicating that the origin service uses the one or more related services to execute a task of the origin service;

determining a priority order for invoking the plurality of diagnostic modules based on the programmatic dependency between the origin service and the one or more related services;

invoking the plurality of diagnostic modules according to the priority order to generate output information indicating respective health statuses of the one or more related services and respective programmatic dependencies on additional services related to the one or more related services, the respective health statuses indicating whether the one or more related services are healthy;

determining a root cause of the incident based on the output information;

determining a diagnostic action based on the root cause; and

transmitting, based on the diagnostic action, an engagement notification to a responsible entity.

17. The non-transitory computer-readable device of claim 16 , wherein dynamically generating the workflow comprises:

requesting the relationship information from the one or more related services.

18. The non-transitory computer-readable device of claim 16 , wherein invoking the plurality of diagnostic modules comprises:

sending incident information and a resource identifier of a resource affected by the incident to a diagnostic module of the plurality of diagnostic modules.

19. The non-transitory computer-readable device of claim 16 , wherein the diagnostic action includes at least one of an incident enrichment action, an incident transfer action, an incident linking action, or an incident auto-invite action.

20. The non-transitory computer-readable device of claim 16 , further comprising presenting, via a graphical user interface, at least one of a visual representation of the workflow or a dependency graph of the origin service and the one or more related services, the visual representation displaying each of the plurality of diagnostic modules with a health status.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 21, 2021
From: XU, ZHANGWEI; GAO, XIAOFENG; MITCHELL, CARY L.; LUNSMAN, STEVE J.; PEREZ, TONY V.; NARASIMHAN, ARVIND
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 056939/0446 →
Continuity (1)
Related Publication 20230026283A1 · Jan 26, 2023
Cited By (1)
US 12,542,706