IP Library Granted Patent US 9,772,920
Granted Patent B2
US 9,772,920 · App. 14/700,083 · Granted Sep 26, 2017

Dynamic service fault detection and recovery using peer services

Inventors: Sajithkumar Kizhakkiniyil (Pleasanton, CA); Anil Maipady (San Jose, CA); Krishnam Chapa (Bangalore, IN); Narender Vattikonda (San Jose, CA); Jeevan Pingali (Bangalore, IN); Rahul Kumar (Bangalore, IN)
Assignee: Apollo Education Group, Inc.
G06F11/3006G06F11/3051G06F11/3055G06F11/3433G06F2201/81G06F2201/875
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,772,920
App. No.
14/700,083
Granted
Sep 26, 2017
Kind
B2
Abstract

Techniques are described for identifying unhealthy nodes in a multi-node system. One or more parameters of each node is monitored, then compared with the values for the same parameter running on other nodes in the multi-node system. Based on the comparison, a determination is made whether a node is healthy. If the multi-node system comprises one or more nodes with differing capabilities, an adjustment is performed to account for the differing capabilities of each respective node. Further provided are methods of taking remedial action upon a determination that a node is unhealthy. A tuner is used to modify values of health parameters until the node is performing similarly to its peers.

Claims (66)

1. A method of identifying unhealthy nodes in a multi-node system, comprising:

monitoring one or more health parameters of each node of a plurality of nodes in the multi-node system;

wherein the plurality of nodes include a first node and one or more other nodes;

determining expected performance of the first node based on the one or more health parameters of the one or more other nodes;

determining actual performance of the first node based on the one or more health parameters of the first node;

performing a comparison between the expected performance of the first node and the actual performance of the first node;

based at least in part on the comparison, determining whether the first node is an unhealthy node; and

responsive to determining that the first node is an unhealthy node, performing a remedial action relative to the first node;

wherein the method is performed automatically by one or more computing devices.

2. The method of claim 1 wherein:

the one or more other nodes have different capabilities than the first node; and

the method includes, based on the different capabilities, adjusting, prior to performing the comparison, at least one of

the one or more health parameters of the first node, or

the one or more health parameters of the one or more other nodes.

3. The method of claim 1 wherein:

the step of determining whether the first node is an unhealthy node is further based on whether the one or more health parameters of the first node exceed predetermined minimum health parameters; and

the first node is determined to be unhealthy if either:

the comparison indicates that the actual performance of the first node is a threshold amount below the expected performance of the first node; or

the one or more health parameters of the first node do not exceed the predetermined minimum health parameters.

4. The method of claim 1 wherein:

the first node is executing a particular service; and

the one or more other nodes are nodes that are peer nodes that are executing the particular service.

5. The method of claim 1 wherein performing a remedial action includes:

determining which nodes of the one or more other nodes are healthy based on the one or more health parameters of the one or more other nodes, and

adjusting one or more configuration values of the first node based on one or more configuration values used by at least one healthy node of the one or more other nodes.

6. The method of claim 5 , wherein determining which nodes of the one or more other nodes are healthy includes:

determining an average value of the one or more health parameters of the one or more other nodes; and

comparing the average value of the one or more health parameters with the one or more health parameters of the first node.

7. The method of claim 5 , wherein determining which nodes of the one or more other nodes are healthy includes:

determining a median value of the one or more health parameters of the one or more other nodes; and

comparing the median value of the one or more health parameters with the one or more health parameters of the first node.

8. The method of claim 5 , wherein determining which nodes of the one or more other nodes are healthy includes:

determining a high value of the one or more health parameters of the one or more other nodes; and

comparing the high value of the one or more health parameters with the one or more health parameters of the first node.

9. One or more non-transitory computer-readable media storing one or more sequences of instructions for identifying unhealthy nodes in a multi-node system which, when executed by one or more processors, cause:

monitoring one or more health parameters of each node of a plurality of nodes in the multi-node system;

wherein the plurality of nodes include a first node and one or more other nodes;

determining expected performance of the first node based on the one or more health parameters of the one or more other nodes;

determining actual performance of the first node based on the one or more health parameters of the first node;

performing a comparison between the expected performance of the first node and the actual performance of the first node;

based at least in part on the comparison, determining whether the first node is an unhealthy node; and

responsive to determining that the first node is an unhealthy node, performing a remedial action relative to the first node.

10. The one or more non-transitory computer-readable media of claim 9 wherein:

the one or more other nodes have different capabilities than the first node; and

the one or more sequences of instructions further comprise instructions which, when executed by one or more processors, cause: based on the different capabilities, adjusting, prior to performing the comparison, at least one of the one or more health parameters of the first node, or

the one or more health parameters of the one or more other nodes.

11. The one or more non-transitory computer-readable media of claim 9 wherein:

determining whether the first node is an unhealthy node is further based on whether the one or more health parameters of the first node exceed predetermined minimum health parameters; and

the first node is determined to be unhealthy if either:

the comparison indicates that the actual performance of the first node is a threshold amount below the expected performance of the first node; or

the one or more health parameters of the first node do not exceed the predetermined minimum health parameters.

12. The one or more non-transitory computer-readable media of claim 9 wherein:

the first node is executing a particular service; and

the one or more other nodes are nodes that are peer nodes that are executing the particular service.

13. The one or more non-transitory computer-readable media of claim 9 wherein performing a remedial action includes:

determining which nodes of the one or more other nodes are healthy based on the one or more health parameters of the one or more other nodes, and

adjusting one or more configuration values of the first node based on one or more configuration values used by at least one healthy node of the one or more other nodes.

14. The one or more non-transitory computer-readable media of claim 13 , wherein determining which nodes of the one or more other nodes are healthy includes:

determining an average value of the one or more health parameters of the one or more other nodes; and

comparing the average value of the one or more health parameters with the one or more health parameters of the first node.

15. The one or more non-transitory computer-readable media of claim 13 , wherein determining which nodes of the one or more other nodes are healthy includes:

determining a median value of the one or more health parameters of the one or more other nodes; and

comparing the median value of the one or more health parameters with the one or more health parameters of the first node.

16. The one or more non-transitory computer-readable media of claim 13 , wherein determining which nodes of the one or more other nodes are healthy includes:

determining a high value of the one or more health parameters of the one or more other nodes; and

comparing the high value of the one or more health parameters with the one or more health parameters of the first node.

Assignments (5)
SECURITY INTEREST Recorded Nov 13, 2025
From: THE UNIVERSITY OF PHOENIX, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 072896/0972 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2020
From: APOLLO EDUCATION GROUP, INC.
To: THE UNIVERSITY OF PHOENIX, INC.
Reel/Frame 053308/0512 →
RELEASE OF SECURITY INTEREST Recorded Jul 15, 2019
From: EVEREST REINSURANCE COMPANY
To: APOLLO EDUCATION GROUP, INC.
Reel/Frame 049753/0187 →
SECURITY INTEREST Recorded Feb 14, 2017
From: APOLLO EDUCATION GROUP, INC.
To: EVEREST REINSURANCE COMPANY
Reel/Frame 041750/0137 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 20, 2015
From: KIZHAKKINIYIL, SAJITHKUMAR; MAIPADY, ANIL; CHAPA, KRISHNAM; VATTIKONDA, NARENDER; PINGALI, JEEVAN; KUMAR, RAHUL
To: APOLLO EDUCATION GROUP, INC.
Reel/Frame 035684/0833 →
Continuity (1)
Related Publication 20160321147A1 · Nov 3, 2016