IP Library Granted Patent US 9,460,183
Granted Patent B2
US 9,460,183 · App. 14/078,488 · Granted Oct 4, 2016

Split brain resistant failover in high availability clusters

Inventor: Michael W. Dalton (San Francisco, CA)
Assignee: ZETTASET, INC.
G06F17/30581G06F11/1425G06F11/187G06F11/2028G06F11/2097
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,460,183
App. No.
14/078,488
Granted
Oct 4, 2016
Kind
B2
Abstract

Method and high availability clusters that support synchronous state replication to provide for failover between nodes, and more precisely, between the master candidate machines at the corresponding nodes. There are at least two master candidates (m=2) in the high availability cluster and the election of the current master is performed by a quorum-based majority vote among quorum machines, whose number n is at least three and odd (n≧3 and n is odd). The current master is issued a current time-limited lease to be measured off by the current master's local clock. In setting the duration or period of the lease, a relative clock skew is used to bound the duration to an upper bound, thus ensuring resistance to split brain situations during failover events.

Claims (48)

1. A high availability cluster with failover capability between nodes comprising machines of said high availability cluster without split brain situations, said high availability cluster comprising:

a) a number m of master candidates identified among said machines, where said number m is at least two;

b) a number n of quorum machines among said machines, where said number n is at least three and is odd;

c) a local area network for synchronously replicating and updating states among said number m of master candidates to maintain a current state;

d) a quorum-based majority vote protocol among said quorum machines for electing a current master from among said number m of master candidates;

e) a mechanism for issuing a current time-limited lease to said current master, said current time-limited lease to be measured off by a local clock belonging to said current master;

f) a physical parameter for bounding a relative clock skew of said current time-limited lease to an upper bound, said relative clock skew estimated by utilizing the Network Time Protocol (NTP);

g) said high availability cluster providing a service to at least one network client;

wherein a failure of said current master triggers failover to a new master from among said number m of master candidates and issuance of a new time-limited lease to said new master, thereby preventing split brain situations between said master candidates.

2. The high availability cluster of claim 1 , wherein said service is a Kerberos authentication service.

3. The high availability cluster of claim 1 , wherein said master candidate machines provide a Kerberos ticket granting service.

4. The high availability cluster of claim 1 , wherein said service is a database.

5. The high availability cluster of claim 4 , wherein said database is a relational database.

6. The high availability cluster of claim 4 , wherein said database is PostgreSQL.

7. The high availability cluster of claim 1 , wherein said nodes comprise a Hadoop cluster.

8. The high availability cluster of claim 1 , wherein said master candidate machines are Hadoop NameNode machines.

9. The high availability cluster of claim 1 , wherein said nodes comprise a Hadoop Distributed File System (HDFS).

10. The high availability cluster of claim 1 , wherein said nodes belong to a distributed file system.

11. The high availability cluster of claim 1 , wherein said nodes belong to a distributed computing environment.

12. The high availability cluster of claim 1 , wherein said service is selected from the group consisting of an electronic mail service, a financial transactions service, a service involving interactions with Domain Name Servers (DNS) and a metadata service.

13. The high availability cluster of claim 1 , wherein said upper bound is chosen based on a performance objective of said high availability cluster.

14. The high availability cluster of claim 1 , wherein said service comprises a legacy application and said step (c) utilizes in said replicating and said updating of said states, a Distributed Replicated Block Device (DRBD) of Linux operating system.

15. A method of operating a high availability cluster serving at least one network client to provide for failover between nodes comprising machines of said high availability cluster without split brain situations, said method comprising:

a) identifying a number m of master candidates among said machines, where said number m is at least two;

b) identifying a number n of quorum machines among said machines, where said number n is at least three and is odd;

c) synchronously updating each of said m master candidates to maintain a current state;

d) electing a current master from said number m of master candidates through a quorum-based majority vote among said quorum machines;

e) issuing a current time-limited lease to said current master, said current timed-limited lease to be measured off by a local clock belonging to said current master, said current master running a service requested by said at least one network client while holding said current time-limited lease;

f) bounding a relative clock skew of said current time-limited lease to an upper bound, said relative clock skew estimated by utilizing the Network Time Protocol (NTP);

wherein a failure of said current master triggers failover to a new master from among said number m of master candidates and issuance of a new time-limited lease to said new master, thereby preventing split brain situations between said master candidates.

16. The method of claim 15 , wherein said number m of master candidates is two, said method further comprising:

a) identifying one machine of said two master candidates as said current master, and the other machine of said two master candidates as a backup master;

b) performing said identification prior to said failure of said current master; and

c) in response to said failure of said current master, triggering said failover to said backup master and issuing said new time-limited lease to said backup master.

17. The method of claim 15 , wherein said service is selected from the group consisting of an electronic mail service, a financial transactions service, a service involving interactions with Domain Name Servers (DNS) and a metadata service.

18. The method of claim 15 , wherein said high availability cluster services Kerberos authentication to said at least one network client.

19. The method of claim 15 , wherein said nodes belong to a distributed system.

20. The method of claim 19 , wherein said distributed system persists its state via block device writes.

21. The method of claim 20 , wherein said distributed system is a PostgreSQL database.

22. A method of operating a high availability Hadoop cluster to provide for failover between nodes comprising machines of said high availability Hadoop cluster without split brain situations, said method comprising:

a) identifying a number m of master candidates among said machines to operate as Hadoop NameNode machines, where said number m is at least two;

b) identifying a number n of quorum machines among said machines, where said number n is at least three and is odd;

c) synchronously updating each of said m master candidates to maintain a current state;

d) electing a current master from said number m of master candidates through a quorum-based majority vote among said quorum machines;

e) issuing a current time-limited lease to said current master, said current timed-limited lease to be measured off by a local clock belonging to said current master, said current master running a service requested by said at least one network client while holding said current time-limited lease;

f) bounding a relative clock skew of said current time-limited lease to an upper bound, said relative clock skew estimated by utilizing the Network Time Protocol (NTP);

wherein a failure of said current master triggers failover to a new master from among said number m of master candidates and issuance of a new time-limited lease to said new master, thereby preventing split brain situations between said master candidates.

23. The method of claim 22 , wherein said number n of quorum machines are Hadoop Datallode machines.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2025
From: ZETTASET, INC.; TPK INVESTMENTS, LLC
To: TPK INVESTMENTS, LLC; CYBER CASTLE, INC.
Reel/Frame 070572/0348 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 18, 2014
From: DALTON, MICHAEL W.
To: ZETTASET, INC.
Reel/Frame 032500/0306 →
Continuity (2)
Continuation 13317803 · Oct 28, 2011
Related Publication 20140188794A1 · Jul 3, 2014