IP Library Granted Patent US 12,265,858
Granted Patent B1
US 12,265,858 · App. 17/827,563 · Granted Apr 1, 2025

Implementing a split-brain prevention strategy when configuring automatic cluster manager failover

Inventors: Sayantan Bhattacharyya (Sydney, AU); Wendi Qiu (Norwest, AU); How Yin Tan (Sydney, AU); Amritpal Singh Bath (Alamo, CA)
Assignee: Splunk Inc.
G06F9/5072G06F9/44505G06F2209/505
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,858
App. No.
17/827,563
Granted
Apr 1, 2025
Kind
B1
Abstract

A method of dynamic cluster manager failover includes routing data traffic associated with managing a plurality of indexers in a cluster to a first cluster manager, wherein the first cluster manager is associated with an active role and is operable to manage the plurality of indexers in the cluster. The method also includes transmitting periodic heartbeat request messages from a second cluster manager of the cluster to the first cluster manager, wherein the second cluster manager is associated with a standby role. Further, the method includes detecting, at the second cluster manager, a loss of heartbeat response messages from the first cluster manager. Also, the method includes receiving information from a set of indexers regarding a status of the first cluster manager and in response to a determination that the status of the first cluster manager is offline, promoting the second cluster manager to switch over to the active role.

Claims (46)

1. A computer-implemented method of managing dynamic cluster manager failover, the method comprising:

routing data traffic associated with managing a plurality of indexers in a cluster to a first cluster manager, wherein the first cluster manager is associated with an active role and is operable to manage the plurality of indexers in the cluster in a capacity of the active role;

transmitting periodic heartbeat request messages from a second cluster manager of the cluster to the first cluster manager, wherein the second cluster manager is associated with a standby role;

detecting, at the second cluster manager, a loss of heartbeat response messages from the first cluster manager, wherein the heartbeat response messages are transmitted from the first cluster manager to the second cluster manager in response to the periodic heartbeat request message and are operable to communicate a status of the first cluster manager to the second cluster manager;

selecting, at the second cluster manager, a set of indexers from the plurality of indexers to establish communication with the first cluster manager;

receiving information regarding the first cluster manager from the set of indexers, wherein the information indicates a status of the first cluster manager; and

in response to a determination that the status of the first cluster manager is offline, promoting the second cluster manager to switch over to the active role, otherwise, maintaining the first cluster manage in the active role.

2. The method of claim 1 , wherein the promoting comprises:

re-routing the data traffic associated with managing the plurality of indexers to the second cluster manager.

3. The method of claim 1 , wherein the detecting comprises determining that the loss of the heartbeat response messages has exceeded a predetermined duration of time.

4. The method of claim 1 , wherein the detecting comprises determining that a predetermined number of heartbeat request messages did not receive corresponding heartbeat response messages from the second cluster manager.

5. The method of claim 1 , wherein the selecting the set of indexers comprises non-deterministically choosing between active indexers indicated on a status map received from the first cluster manager.

6. The method of claim 1 , wherein the plurality of indexers are dispersed over multiple sites, and wherein the set of indexers comprises one or more indexers from each of the multiple sites.

7. The method of claim 1 , wherein the determination that the status of the first cluster manager is online is based on the information indicating that at least one indexer from the set of indexers recently established communication with the first cluster manager.

8. The method of claim 1 , wherein the promoting comprises re-routing the data traffic associated with managing the plurality of indexers to the second cluster manager, and wherein the determination that the status of the first cluster manager is offline is based on the information indicating at least half the indexers from the set of indexers voted that the first cluster manager is offline.

9. The method of claim 1 , further comprising: wherein the determining further comprises:

responsive to a determination that the set of indexers has not achieved consensus regarding the status of the first cluster manager, retrying the selecting and the receiving using exponentially backed off intervals, wherein a non-deterministic selection of indexers from the set of indexers is used at each exponentially backed off interval.

10. The method of claim 1 , further comprising:

responsive to a determination that the set of indexers cannot be contacted, maintaining the first cluster manager in the active role.

11. The method of claim 1 , wherein the routing is performed using a proxy server or a load balancer.

12. A computing device, comprising:

a processor; and

a non-transitory computer-readable medium having stored thereon instructions that, when executed by the processor, cause the processor to perform operations including:

routing data traffic associated with managing a plurality of indexers in a cluster to a first cluster manager, wherein the first cluster manager is associated with an active role and is operable to manage the plurality of indexers in the cluster in a capacity of the active role;

transmitting periodic heartbeat request messages from a second cluster manager of the cluster to the first cluster manager, wherein the second cluster manager is associated with a standby role;

detecting, at the second cluster manager, a loss of heartbeat response messages from the first cluster manager, wherein the heartbeat response messages are transmitted from the first cluster manager to the second cluster manager in response to the periodic heartbeat request message and are operable to communicate a status of the first cluster manager to the second cluster manager;

selecting, at the second cluster manager, a set of indexers from the plurality of indexers to establish communication with the first cluster manager;

receiving information regarding the first cluster manager from the set of indexers, wherein the information indicates a status of the first cluster manager; and

in response to a determination that the status of the first cluster manager is offline, promoting the second cluster manager to switch over to the active role, otherwise, maintaining the first cluster manage in the active role.

13. The computing device of claim 12 , wherein the promoting comprises:

re-routing data traffic associated with the managing the plurality of indexers to the second cluster manager.

14. The computing device of claim 12 , wherein the selecting the set of indexers comprises non-deterministically choosing between active indexers indicated on a status map received from the first cluster manager.

15. The computing device of claim 12 , wherein the plurality of indexers are dispersed over multiple sites, and wherein the set of indexers comprises one or more indexers from each of the multiple sites.

16. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processor to perform operations including:

routing data traffic associated with managing a plurality of indexers in a cluster to a first cluster manager, wherein the first cluster manager is associated with an active role and is operable to manage the plurality of indexers in the cluster in a capacity of the active role;

transmitting periodic heartbeat request messages from a second cluster manager of the cluster to the first cluster manager, wherein the second cluster manager is associated with a standby role;

detecting, at the second cluster manager, a loss of heartbeat response messages from the first cluster manager, wherein the heartbeat response messages are transmitted from the first cluster manager to the second cluster manager in response to the periodic heartbeat request message and are operable to communicate a status of the first cluster manager to the second cluster manager;

selecting, at the second cluster manager, a set of indexers from the plurality of indexers to establish communication with the first cluster manager;

receiving information regarding the first cluster manager from the set of indexers, wherein the information indicates a status of the first cluster manager; and

in response to a determination that the status of the first cluster manager is offline, promoting the second cluster manager to switch over to the active role, otherwise, maintaining the first cluster manage in the active role.

17. The non-transitory computer-readable medium of claim 16 , wherein the promoting comprises:

re-routing data traffic associated with the managing the plurality of indexers to the second cluster manager.

18. The non-transitory computer-readable medium of claim 16 , wherein the promoting comprises re-routing the data traffic associated with managing the plurality of indexers to the second cluster manager, and wherein the determination that the status of the first cluster manager is offline is based on the information indicating at least half the indexers from the set of indexers voted that the first cluster manager is offline.

19. The non-transitory computer-readable medium of claim 16 , wherein the operations further comprise:

responsive to a determination that the set of indexers has not achieved consensus regarding the status of the first cluster manager, retrying the selecting and the receiving using exponentially backed off intervals, wherein a non-deterministic selection of indexers from the set of indexers is used at each exponentially backed off interval.

20. The non-transitory computer-readable medium of claim 16 , wherein the plurality of indexers are dispersed over multiple sites, and wherein the set of indexers comprises one or more indexers from each of the multiple sites.

Assignments (4)
CHANGE OF NAME Recorded Jul 22, 2025
From: SPLUNK INC.
To: SPLUNK LLC
Reel/Frame 072170/0599 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2025
From: SPLUNK LLC
To: CISCO TECHNOLOGY, INC.
Reel/Frame 072173/0058 →
CHANGE OF NAME Recorded Jan 6, 2025
From: SPLUNK INC.
To: SPLUNK LLC
Reel/Frame 069826/0065 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 27, 2022
From: BHATTACHARYYA, SAYANTAN; QIU, WENDI; TAN, HOW YIN; BATH, AMRITPAL SINGH
To: SPLUNK INC.
Reel/Frame 060044/0100 →
References Cited (6)
US 8041798B1 · Pabla · 2011 [cited by examiner]
US 20040153700A1 · Nixon · 2004 [cited by examiner]
US 20110289344A1 · Bae · 2011 [cited by examiner]
US 20230083450A1 · Kaitha · 2023 [cited by examiner]
Non Final Office Action received for U.S. Appl. No. 17/827,526 dated Mar. 25, 2024, 17 pages. [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 17/827,526 dated Jul. 17, 2024, 9 pages. [cited by applicant]
Cited By (5)
US 12,591,459 US 12,596,618 US 12,639,191 US 12,693,904 US 12,717,636