IP Library › Granted Patent US 11,582,087
Granted Patent B2
US 11,582,087 · App. 16/718,178 · Granted Feb 14, 2023

Node health prediction based on failure issues experienced prior to deployment in a cloud computing system

Inventors: Sanjay Ramanujan (Issaquah, WA); Luke Rafael Rodriguez (Seattle, WA); Muhammad Khizar Qazi (Seattle, WA); Aleksandr Mikhailovich Gershaft (Redmond, WA); Marwan Elias Jubran (Kirkland, WA); Saurabh Agarwal (Redmond, WA)
Assignee: MICROSOFT TECHNOLOGY LICENSING, LLC
H04L41/0645H04L41/064H04L41/0654H04L41/0686H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,582,087
App. No.
16/718,178
Granted
Feb 14, 2023
Kind
B2
Abstract

To improve the reliability of nodes that are utilized by a cloud computing provider, information about the entire lifecycle of nodes can be collected and used to predict when nodes are likely to experience failures based at least in part on early lifecycle errors. In one aspect, a plurality of failure issues experienced by a plurality of production nodes in a cloud computing system during a pre-production phase can be identified. A subset of the plurality of failure issues can be selected based at least in part on correlation with service outages for the plurality of production nodes during a production phase. A comparison can be performed between the subset of the plurality of failure issues and a set of failure issues experienced by a pre-production node during the pre-production phase. A risk score for the pre-production node can be calculated based at least in part on the comparison.

Claims (58)

1. A method, comprising:

identifying a plurality of failure issues experienced by a plurality of production nodes in a cloud computing system during a pre-production phase;

selecting a subset of the plurality of failure issues based at least in part on correlation with service outages for the plurality of production nodes during a production phase;

performing a comparison between the subset of the plurality of failure issues experienced by the plurality of production nodes during the pre-production phase that correlate with service outages during the production phase and a set of failure issues experienced by a pre-production node during the pre-production phase before the pre-production node enters the production phase;

calculating a risk score for the pre-production node based at least in part on the comparison; and

performing corrective action with respect to the pre-production node based at least in part on the risk score, wherein the corrective action is performed before the pre-production node enters the production phase.

2. The method of claim 1 , wherein selecting the subset comprises:

determining, for each failure issue of the plurality of failure issues, an average out of service metric for the plurality of production nodes that experienced the failure issue during the pre-production phase; and

selecting any failure issues whose average out of service metric satisfies a defined condition.

3. The method of claim 2 , wherein:

the average out of service metric for a failure issue is an average value of an out of service metric calculated for the plurality of production nodes that experienced the failure issue; and

the out of service metric calculated for a production node indicates how often the production node has been out of service since entering the production phase.

4. The method of claim 1 , wherein identifying the plurality of failure issues comprises:

obtaining a first set of test results from a first set of tests that a system integrator performs on the plurality of production nodes; and

obtaining a second set of test results from a second set of tests that a cloud computing provider performs on the plurality of production nodes.

5. The method of claim 1 , further comprising determining, for each failure issue of the plurality of failure issues, a frequency of occurrence metric that indicates how many of the plurality of production nodes experienced the failure issue during the pre-production phase.

6. The method of claim 1 , further comprising classifying the plurality of failure issues into a plurality of categories corresponding to different hardware components.

7. The method of claim 1 , wherein the corrective action comprises at least one of:

repairing the pre-production node;

replacing the pre-production node;

replacing a component within the pre-production node; or

placing the pre-production node in a state of probation.

8. The method of claim 1 , further comprising:

determining, for each failure issue of the plurality of failure issues, a mean time to repair metric; and

prioritizing repairs based at least in part on the mean time to repair metric.

9. The method of claim 1 , wherein selecting the subset comprises:

selecting failure issues that are correlated with out of service rates for the plurality of production nodes that exceed a defined minimum value.

10. The method of claim 1 , wherein the correlation with the service outages identifies failure issues that resulted in the plurality of production nodes going out of service frequently during the production phase.

11. A device, comprising:

a memory to store data and instructions; and

a processor in communication with the memory, wherein the processor is operable to:

identify a plurality of failure issues experienced by a plurality of production nodes in a cloud computing system during a pre-production phase;

select a subset of the plurality of failure issues based at least in part on correlation with service outages for the plurality of production nodes during a production phase;

perform a comparison between the subset of the plurality of failure issues experienced by the plurality of production nodes during the pre-production phase that correlate with service outages during the production phase and a set of failure issues experienced by a pre-production node during the pre-production phase before the pre-production node enters the production phase;

calculate a risk score for the pre-production node based at least in part on the comparison; and

perform corrective action with respect to the pre-production node based at least in part on the risk score, wherein the corrective action is performed before the pre-production node enters the production phase.

12. The device of claim 11 , wherein selecting the subset comprises:

determining, for each failure issue of the plurality of failure issues, an average out of service metric for the plurality of production nodes that experienced the failure issue during the pre-production phase; and

selecting any failure issues whose average out of service metric satisfies a defined condition.

13. The device of claim 12 , wherein:

the average out of service metric for a failure issue is an average value of an out of service metric calculated for the plurality of production nodes that experienced the failure issue; and

the out of service metric calculated for a production node indicates how often the production node has been out of service since entering the production phase.

14. The device of claim 11 , wherein identifying the plurality of failure issues comprises:

obtaining a first set of test results from a first set of tests that a system integrator performs on the plurality of production nodes; and

obtaining a second set of test results from a second set of tests that a cloud computing provider performs on the plurality of production nodes.

15. The device of claim 11 , further comprising determining, for each failure issue of the plurality of failure issues, a frequency of occurrence metric that indicates how many of the plurality of production nodes experienced the failure issue during the pre-production phase.

16. The device of claim 11 , further comprising classifying the plurality of failure issues into a plurality of categories corresponding to different hardware components.

17. The device of claim 11 , wherein the corrective action comprises at least one of:

repairing the pre-production node;

replacing the pre-production node;

replacing a component within the pre-production node; or

placing the pre-production node in a state of probation.

18. The device of claim 11 , further comprising:

determining, for each failure issue of the plurality of failure issues, a mean time to repair metric; and

prioritizing repairs based at least in part on the mean time to repair metric.

19. The device of claim 11 , wherein selecting the subset comprises:

selecting failure issues that are correlated with out of service rates for the plurality of production nodes that exceed a defined minimum value.

20. The device of claim 11 , wherein the correlation with the service outages identifies failure issues that resulted in the plurality of production nodes going out of service frequently during the production phase.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 24, 2019
From: RAMANUJAN, SANJAY; RODRIGUEZ, LUKE RAFAEL; QAZI, MUHAMMAD KHIZAR; GERSHAFT, ALEKSANDR MIKHAILOVICH; JUBRAN, MARWAN ELIAS; AGARWAL, SAURABH
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 051362/0965 →
Continuity (1)
Related Publication 20210184916A1 · Jun 17, 2021