IP Library Granted Patent US 9,164,864
Granted Patent B1
US 9,164,864 · App. 13/338,543 · Granted Oct 20, 2015

Minimizing false negative and duplicate health monitoring alerts in a dual master shared nothing database appliance

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,164,864
App. No.
13/338,543
Granted
Oct 20, 2015
Kind
B1
Abstract

A primary master node and a standby master node monitor the health of a shared nothing database appliance to afford high availability while minimizing false negatives and duplicate alerts by executing continuously in parallel complimentary processes that determine whether the database is running, and which master node is the active database master node. The active database master node monitors the health of the components of the database appliance by polling each component to detect failures and warnings, and the other master node monitors the status of the active master node. Upon detecting a failure of the active master node, the other node takes over health monitoring. If the database is not running, the designated primary master node performs health monitoring.

Claims (55)

1. A method of monitoring the health of a database appliance comprising a database distributed on a plurality of database nodes, and having redundant master nodes including a primary master node and a standby master node, the database being active on and controlled by only one of said redundant master nodes at a time, the method comprising:

executing concurrently in parallel and independently on both said redundant master nodes a database monitoring process, said database monitoring process comprising a resolution process and a health monitor process, the resolution process resolving on which one of said redundant master nodes said database is active at said time, said one node being designated the primary master node, and confirming that said primary master node is executing said health monitor process to monitor hardware and software components of said database and report alerts, the other redundant master node being the standby master node and not issuing alerts;

resolving by executing said resolution process in parallel by said primary master node and said standby master node whether the database is running on said primary master node, including:

attempting, by the primary master node, a first login to the database on the primary master node; and

attempting concurrently, by the standby master node, a second login to the database on the primary master node;

upon said first and second logins being successful, resolving by the resolution process on the primary master node that the database is running on the primary master node and that the primary master node is executing said health monitor process of hardware and software components of the database;

monitoring by the standby master node the status of the primary master node to detect a failure of the primary master node, including:

attempting, by the standby master node, a third login to the database on the primary master node after a first predetermined period of time;

upon identifying that the third login attempt is unsuccessful, determining that the primary master node has failed based on the unsuccessful third login attempt; and

upon determining said failure of the primary master node by the standby master node:

attempting, by the standby master node, a fourth login to the database on the standby master node;

upon the fourth login attempt by the standby master node being successful, determining that the database is active on said standby master node; and

executing said health monitor process of said components of said database by the standby master node in response to determining that the fourth login attempt was successful.

2. The method of claim 1 , where attempting said first login comprises issuing a database query to the primary master node, and determining that the database is running upon receiving a response to the query.

3. The method of claim 1 further comprising waiting a second predetermined period of time following said executing said health monitor process by said primary master node, and repeating said determining and health monitor process by said primary master node.

4. The method of claim 1 , wherein said first predetermined period of time is of the order of about 1-5 minutes and said second predetermined period of time is of the order of about 1-5 seconds.

5. The method of claim 1 , further comprising:

attempting, by the standby master node, a fifth login to the database on the standby master node;

determining that said fifth login to the database on the standby master node is unsuccessful;

in response to determining that said fifth login attempt was unsuccessful, attempting, by the standby master node, to contact a health monitor process on the primary master node; and, if said contact is unsuccessful, performing said health monitor process by said standby master node.

6. The method of claim 1 , wherein said health monitor process comprises polling said hardware and software components to determine their status, and issuing alerts upon detecting a failure or a warning.

7. The method of claim 1 further comprising determining, upon said first and second logins being unsuccessful, that said database is not operating, and issuing an alert by one of said redundant master nodes.

8. Non-transitory computer readable media storing executable instructions for controlling the operation of a computer to perform health monitoring of a database appliance that includes a database distributed on a plurality of database nodes, and having redundant master nodes including a primary master node and a standby master node, the database being active on and controlled by only one of said redundant master nodes at a time, comprising instructions for:

executing concurrently in parallel and independently on both said redundant master nodes a database monitoring process, said database monitoring process comprising a resolution process and a health monitor process, the resolution process resolving on which one of said redundant master nodes said database is active at said time, said one node being designated the primary master node, and confirming that said primary master node is executing said health monitor process to monitor hardware and software components of said database and report alerts, the other redundant master node being the standby master node and not issuing alerts;

resolving by executing said resolution process in parallel by said primary master node and said standby master node whether the database is running on said primary master node, including:

attempting, by the primary master node, a first login to the database on the primary master node; and

attempting concurrently, by the standby master node, a second login to the database on the primary master node;

upon said first and second logins being successful, resolving by the resolution process on the primary master node that the database is running on the primary master node and that the primary master node is executing said health monitor process of hardware and software components of the database;

monitoring by the standby master node the status of the primary master node to detect a failure of the primary master node, including:

attempting, by the standby master node, a third login to the database on the primary master node after a first predetermined period of time;

upon identifying that the third login attempt is unsuccessful,

determining that the primary master node has failed based on the unsuccessful third login attempt; and

upon determining said failure of the primary master node by the standby master node:

attempting, by the standby master node, a fourth login to the database on the standby master node;

upon the fourth login attempt by the standby master node being successful, determining that the database is active on said standby master node; and

executing said health monitor process of said components of said database by the standby master node in response to determining that the fourth login attempt was successful.

9. Non-transitory computer readable media according to claim 8 , wherein said instructions for attempting said first login comprise instructions for issuing a database query to the primary master node, and determining that the database is running upon receiving a response to the query.

10. Non-transitory computer readable media according to claim 8 further comprising instructions for waiting a second predetermined period of time following said executing said health monitor process by said primary master node, and repeating said determining and health monitor process by said primary master node.

11. Non-transitory computer readable media according to claim 8 , further comprising instructions for:

attempting, by the standby master node, a fifth login to the database on the standby master node;

determining that said fifth login to the database on the standby master node is unsuccessful;

in response to determining that said fifth login attempt was unsuccessful, attempting, by the standby master node, to contact a health monitor process on the primary master node; and, if said contact is unsuccessful, performing said health monitor process by said standby master node.

12. Non-transitory computer readable media according to claim 8 , wherein said health monitor process comprises instructions for polling said hardware and software components to determine their status, and for issuing alerts upon detecting a failure or a warning.

13. A method of monitoring the health of a database appliance comprising a database distributed on a plurality of database nodes, and having redundant master nodes including a primary master node and a standby master node, the database being active on and controlled by only one of said redundant master nodes at a time, the method comprising:

executing concurrently in parallel and independently on both said redundant master nodes a database monitoring process, said database monitoring process comprising a resolution process and a health monitor process, the resolution process resolving on which one of said redundant master nodes said database is active at said time, said one node being designated the primary master node, and confirming that said primary master node is executing said health monitor process to monitor hardware and software components of said database and report alerts, the other redundant master node being the standby master node and not issuing alerts;

resolving by executing said resolution process in parallel and concurrently by said redundant master nodes on which one of said redundant master nodes the database is currently active;

designating said one redundant master node on which the database is currently active as the primary master node, and designating the other redundant master node as the standby master node;

executing substantially continuously by said primary master node and by said standby master node said resolution processes to confirm that the database remains active on the primary master node;

upon confirming the database remains active on the primary master node, executing by the primary master node the health monitor process to monitor the health of said database, and executing concurrently by said standby master node said health monitor process to monitor the health of said primary master node;

upon the health monitor process executing on the primary master node determining that the database has failed, issuing a first alert and executing by said standby master node said health monitor process of said database to confirm said failure;

upon determining by said primary master node said database failure, issuing a second alert; and

otherwise upon the health monitor process on the standby master node determining that either or both the primary master node or the database has failed, issuing by the standby master node a third alert.

14. The method of claim 13 , wherein said resolving comprises issuing by each of said redundant master nodes a database query to both of said redundant master nodes, and resolving the master node on which the database is active by a response to said query.

15. The method of claim 13 , wherein said standby master node confirms the failure of the database on the primary master node by attempting unsuccessfully to login to the database on the primary master node.

16. The method of claim 13 , wherein said standby master node monitors the health of said primary master node by attempting to contact the health monitor process on the primary master node, and determines that the primary master node has failed if said contact is unsuccessful.

Assignments (14)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (045455/0001) Recorded May 20, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO ASAP SOFTWARE EXPRESS, INC.); DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL USA L.P.; DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC CORPORATION (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO MAGINATICS LLC); EMC IP HOLDING COMPANY LLC (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO MOZY, INC.); SCALEIO LLC
Reel/Frame 061753/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (040136/0001) Recorded Apr 26, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO ASAP SOFTWARE EXPRESS, INC.); DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL USA L.P.; DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC CORPORATION (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO MAGINATICS LLC); EMC IP HOLDING COMPANY LLC (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO MOZY, INC.); SCALEIO LLC
Reel/Frame 061324/0001 →
RELEASE OF SECURITY INTEREST Recorded Nov 3, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: ASAP SOFTWARE EXPRESS, INC.; AVENTAIL LLC; CREDANT TECHNOLOGIES, INC.; DELL USA L.P.; DELL INTERNATIONAL, L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL SOFTWARE INC.; DELL SYSTEMS CORPORATION; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; FORCE10 NETWORKS, INC.; MAGINATICS LLC; MOZY, INC.; SCALEIO LLC; WYSE TECHNOLOGY L.L.C.
Reel/Frame 058216/0001 →
SECURITY AGREEMENT Recorded Apr 22, 2020
From: CREDANT TECHNOLOGIES INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 053546/0001 →
SECURITY AGREEMENT Recorded Mar 21, 2019
From: CREDANT TECHNOLOGIES, INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 049452/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2016
From: EMC CORPORATION
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 040203/0001 →
SECURITY AGREEMENT Recorded Sep 21, 2016
From: ASAP SOFTWARE EXPRESS, INC.; AVENTAIL LLC; CREDANT TECHNOLOGIES, INC.; DELL USA L.P.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL SOFTWARE INC.; DELL SYSTEMS CORPORATION; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; FORCE10 NETWORKS, INC.; MAGINATICS LLC; MOZY, INC.; SCALEIO LLC; SPANNING CLOUD APPS LLC; WYSE TECHNOLOGY L.L.C.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 040134/0001 →
SECURITY AGREEMENT Recorded Sep 21, 2016
From: ASAP SOFTWARE EXPRESS, INC.; AVENTAIL LLC; CREDANT TECHNOLOGIES, INC.; DELL USA L.P.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL SOFTWARE INC.; DELL SYSTEMS CORPORATION; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; FORCE10 NETWORKS, INC.; MAGINATICS LLC; MOZY, INC.; SCALEIO LLC; SPANNING CLOUD APPS LLC; WYSE TECHNOLOGY L.L.C.
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 040136/0001 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED ON REEL 032800 FRAME 0004. ASSIGNOR(S) HEREBY CONFIRMS THE CORRECT MISPELLING OF ASSIGNEE TO "EMC CORPORATION". Recorded May 2, 2014
From: PIVOTAL SOFTWARE, INC.
To: EMC CORPORATION
Reel/Frame 032815/0301 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2014
From: PIVOTAL SOFTWARE, INC.
To: EMC CORPORATIOLN
Reel/Frame 032800/0004 →
CHANGE OF NAME Recorded Apr 1, 2014
From: GOPIVOTAL, INC.
To: PIVOTAL SOFTWARE, INC.
Reel/Frame 032588/0795 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2013
From: EMC CORPORATION
To: GOPIVOTAL, INC.
Reel/Frame 030488/0384 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2012
From: NOVICK, IVAN D.; HEATH, TIMOTHY; KALA, SHARAD
To: EMC CORPORATION
Reel/Frame 027570/0839 →