IP Library Granted Patent US 7,849,368
Granted Patent B2
US 7,849,368 · App. 12/100,970 · Granted Dec 7, 2010

Method for monitoring server sub-system health

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,849,368
App. No.
12/100,970
Granted
Dec 7, 2010
Kind
B2
Abstract

A server self health monitor (SHM) system monitors the health of the server it resides on. The health of a server is determined by the health of all of a server's sub-systems and deployed applications. The SHM may make health check inquiries to server sub-systems periodically or based on external trigger events. The sub-systems perform self health checks on themselves and provide sub-system health information to requesting entities such as the SHM. Sub-systems self health updates may be based on internal events such as counters or changes in status or based on external entity requests. Corrective action may be performed upon sub-systems by the SHM depending on their health status or the health status of the server. Corrective action may also be performed by a sub-system upon itself.

Claims (33)

1. A method of monitoring the health of a server comprising:

registering a sub-system of the server with a server health monitor;

monitoring the sub-system by the server health monitor;

receiving a health update from the sub-system to the server health monitor wherein the health update provides information regarding the health of the sub-system;

determining health of the server by the server health monitor based at least in part upon the health update received from monitoring the sub-system;

performing a corrective action upon the sub-system, wherein the corrective action is based on the health of the sub-system; and

wherein a parameter specifies the maximum number of times the server can be restarted within a period of time specified by another parameter.

2. The method as claimed in claim 1 further comprising:

configuring the sub-system to be monitored by the server; and

transmitting information to the server health monitor indicating the sub-system is to be monitored by the server health monitor.

3. The method as claimed in claim 2 wherein said configuring the sub-system includes configuring a MBean residing in the sub-system.

4. The method as claimed in claim 1 wherein said monitoring the sub-system includes:

detecting a triggering event; and

performing a health status update on the sub-system by the sub-system, the sub-system determining a health information for the sub-system as a result of the health status update; and

processing the health information.

5. The method as claimed in claim 4 wherein the triggering event is a change in the status of the sub-system.

6. The method as claimed in claim 4 wherein the triggering event is an event occurring external to the sub-system, and the occurrence of the event is communicated to the sub-system.

7. The method as claimed in claim 4 wherein said monitoring the sub-system further comprises:

providing sub-system health information to an entity, the entity having requested the sub-system's health information from the sub-system.

8. The method as claimed in claim 2 further comprising:

performing shutdown of the sub-system, and the sub-system unregistering from being monitored by the server health monitor.

9. The method of claim 1 , wherein a node manager enables an administrator to start and kill servers remotely.

10. The method of claim 1 , wherein a health update indicates that the sub-system is at one of multiple pre-defined health levels.

11. The method of claim 10 , wherein health levels correspond to conditions, the conditions including good, failed, and between good and failed.

12. The method of claim 10 , wherein the sub-system is set to the critical level if a minimum number of transactions have timed out.

13. The method of claim 1 , wherein the health update is triggered periodically by a counter.

14. The method of claim 1 , wherein the server health monitor determines that the server has failed if any critical sub-system has failed.

15. The method of claim 1 , wherein the server health monitor determines that a sub-system has failed if the sub-system fails to respond.

16. The method of claim 1 , wherein the server health monitor can shut down a failed sub-system.

17. The method of claim 1 , wherein the server health monitor can shut down the server if a critical sub-system failed.

18. The method of claim 1 , wherein all communication between a node manager and an administrative server is encrypted.

19. The method of claim 1 , wherein the sub-system performs a health check upon itself and provides sub-system health information to requesting entities.

20. The method of claim 1 , wherein the sub-system health updates are triggered by external entity requests, internal events such as counters, or changes in status.