IP Library Granted Patent US 11,809,395
Granted Patent B1
US 11,809,395 · App. 17/444,173 · Granted Nov 7, 2023

Load balancing, failover, and reliable delivery of data in a data intake and query system

Inventors: Jeff Fan (Vancouver, CA); Daniel Ferstay (Vancouver, CA); Denis Vergnes (North Vancouver, CA)
Assignee: Splunk Inc.
G06F16/2228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,809,395
App. No.
17/444,173
Filed
Jul 30, 2021
Granted
Nov 7, 2023
Kind
B1
Examiner
LIN, ALLEN S
Art Unit
2153
USPC
707/741
Abstract

Systems and methods are described for balancing workloads and reliably delivering data to a plurality of indexing systems in a data intake and query system. A topic-based indexing system load balancer may receive event data from various data sources, each of which may be associated with a topic. The event data may be entirely unparsed, unparsed but divided into events, or parsed into events. The topic-based indexing system load balancer may distribute the received event data on a per-topic or per-event basis to a set of indexing systems, and may distribute topics and events based on the volume received. Unparsed data may be divided into portions, and the topic-based indexing system load balancer may ensure that portions data associated with the same topic are delivered to the same indexer so that events split between two portions may be recombined and indexed.

Claims (66)

1. A method comprising:

receiving at an intermediary system, from an intake system, a start block of unparsed data associated with a group of related blocks from a first data source of a plurality of data sources, wherein the first data source is associated with a first topic of a plurality of topics;

obtaining a partition key for the first topic;

obtaining workload information regarding workloads of individual indexing systems of a plurality of indexing systems, the workload information including associations between individual topics of the plurality of topics and the individual indexing systems;

selecting a first indexing system of the plurality of indexing systems to associate with the partition key;

transmitting, from the intermediary system, the start block of unparsed data to the first indexing system;

receiving at the intermediary system, from the intake system subsequent to transmitting the start block, an end block of unparsed data associated with the group of related blocks from the first data source, wherein the intermediary system is configured to transmit the end block to the first indexing system when the first indexing system is determined to be available;

determining that the first indexing system associated with the partition key is unavailable to receive the end block of unparsed data; and

responsive to determining that the first indexing system associated with the partition key is unavailable to receive the end block of unparsed data, transmitting from the intermediary system to a second indexing system, selected for association with the partition key, the group of related blocks, including both the end block of unparsed data that the first indexing system is determined to be unavailable to receive and the start block of unparsed data previously transmitted from the intermediary system to the first indexing system.

2. The method of claim 1 , wherein the start block contains the start of an event, and wherein the end block contains an end of the event.

3. The method of claim 1 , wherein the first data source comprises a file.

4. The method of claim 1 , wherein the start block corresponds to the start of a file, and wherein the end block corresponds to the current end of the file.

5. The method of claim 1 further comprising associating the first data source with the first topic based at least in part on a characteristic of the first data source.

6. The method of claim 1 further comprising:

determining that the end block corresponds to the end of a file;

receiving, from the intake system, one or more further blocks of unparsed data from the first data source, the one or more further blocks corresponding to data appended to the file after the end block was transmitted; and

associating the one or more further blocks of unparsed data with a second topic.

7. The method of claim 1 further comprising:

determining that the end block satisfies a criterion for being associated with the first topic;

receiving, from the intake system, one or more further blocks of unparsed data from the first data source;

determining that the one or more further blocks of unparsed data do not satisfy the criterion for being associated with the first topic; and

associating the one or more further blocks of unparsed data with a second topic.

8. The method of claim 1 further comprising obtaining historical information regarding one or more of a volume of data received from the first data source, a volume of data received that is associated with the first topic, or volumes of data received that are associated with individual topics of the plurality of topics.

9. The method of claim 1 , wherein selecting the first indexing system to associate with the partition key is based at least in part on historical information regarding volumes of data received for individual topics of the plurality of topics.

10. The method of claim 1 , wherein the determining whether the first indexing system is available to receive the end block of unparsed data comprises determining, whether the first indexing system has sufficient capacity to receive and process the end block of unparsed data.

11. The method of claim 1 further comprising identifying a block of unparsed data received from the first data source as the end block.

12. The method of claim 1 , wherein the first topic corresponds to one or more of the first data source, a file associated with the first data source, or an event associated with the first data source.

13. The method of claim 1 , wherein obtaining the partition key for the first topic comprises generating the partition key based at least in part on the first topic.

14. The method of claim 1 further comprising:

receiving at the intermediary system, from the intake system, parsed event data from a second data source of the plurality of data sources;

selecting, based at least in part on the workload information, a third indexing system of the plurality of indexing systems; and

transmitting the parsed event data from the intermediary system to the third indexing system.

15. The method of claim 1 further comprising transmitting parsed event data to individual indexing systems of the plurality of indexing systems in accordance with one of more of a round-robin distribution, random distribution, least-data-received distribution, or longest-time-idle distribution.

16. The method of claim 1 further comprising transmitting parsed event data to individual indexing systems of the plurality of indexing systems based at least in part on a set of weighting factors.

17. The method of claim 1 , wherein the start block and the end block are consecutive blocks received from the first data source.

18. The method of claim 1 , wherein transmitting at least the end block of unparsed data to the available indexing system comprises transmitting the start block, the end block, and one or more blocks that were received from the first data source in between the start block and the end block.

19. A system comprising:

a data store including computer-executable instructions; and

one or more processors configured to execute the computer-executable instructions, wherein execution of the computer-executable instructions causes the system to:

receive at the system, from an intake system, a start block of unparsed data associated with a group of related blocks from a first data source of a plurality of data sources, wherein the first data source is associated with a first topic of a plurality of topics;

obtain a partition key for the first topic;

obtain workload information regarding workloads of individual indexing systems of a plurality of indexing systems, the workload information including associations between individual topics of the plurality of topics and the individual indexing systems;

select a first indexing system of the plurality of indexing systems to associate with the partition key;

transmit the start block of unparsed data from the system to the first indexing system;

receive at the system, from the intake system subsequent to transmitting the start block, an end block of unparsed data associated with the group of related blocks from the first data source, wherein the end block is transmitted to the first indexing system when the first indexing system is determined to be available;

determine that the first indexing system associated with the partition key is unavailable to receive the end block of unparsed data; and

responsive to a determination that the first indexing system associated with the partition key is unavailable to receive the end block of unparsed data, transmit from the system to a second indexing system, selected for association with the partition key, the group of related blocks, including both the end block of unparsed data that the first indexing system is determined to be unavailable to receive and the start block of unparsed data previously transmitted to the first indexing system.

20. Non-transitory computer-readable media including computer-executable instructions that, when executed by an intermediary system, cause the computing system to:

receive at the intermediary system, from an intake system, a start block of unparsed data associated with a group of related blocks from a first data source of a plurality of data sources, wherein the first data source is associated with a first topic of a plurality of topics;

obtain a partition key for the first topic;

obtain workload information regarding workloads of individual indexing systems of a plurality of indexing systems, the workload information including associations between individual topics of the plurality of topics and the individual indexing systems;

select a first indexing system of the plurality of indexing systems to associate with the partition key;

transmit the start block of unparsed data from the intermediary system to the first indexing system;

receive at the intermediary system, from the intake system subsequent to transmitting the start block, an end block of unparsed data from the first data source, wherein the end block is transmitted to the first indexing system when the first indexing system is determined to be available;

determine that the first indexing system associated with the partition key is unavailable to receive the end block of unparsed data; and

responsive to a determination that the first indexing system associated with the partition key is unavailable to receive the end block of unparsed data, transmit from the intermediary system to a second indexing system, selected for association with the partition key, the group of related blocks, including both the end block of unparsed data that the first indexing system is determined to be unavailable to receive and the start block of unparsed data previously transmitted to the first indexing system.

21. The non-transitory computer-readable media of claim 20 further comprising computer-executable instructions to:

determine that the end block corresponds to the end of a file;

receive at the intermediary system, from the intake system, one or more further blocks of unparsed data from the first data source, the one or more further blocks corresponding to data appended to the file after the end block was transmitted; and

associate the one or more further blocks of unparsed data with a second topic.

22. The non-transitory computer-readable media of claim 20 further comprising computer-executable instructions to:

determine that the end block satisfies a criterion for being associated with the first topic;

receiving at the intermediary system, from the intake system, one or more further blocks of unparsed data from the first data source;

determine that the one or more further blocks of unparsed data do not satisfy the criterion for being associated with the first topic; and

associate the one or more further blocks of unparsed data with a second topic.

23. The non-transitory computer-readable media of claim 20 wherein the first indexing system to associate with the partition key is selected based at least in part on historical information regarding volumes of data received for individual topics of the plurality of topics.

Assignments (3)
CHANGE OF NAME Recorded Jul 22, 2025
From: SPLUNK INC.
To: SPLUNK LLC
Reel/Frame 072170/0599 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2025
From: SPLUNK LLC
To: CISCO TECHNOLOGY, INC.
Reel/Frame 072173/0058 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 24, 2021
From: FAN, JEFF; FERSTAY, DANIEL; VERGNES, DENIS
To: SPLUNK, INC.
Reel/Frame 058208/0268 →
Continuity (1)
Provisional Application 63203283 · Jul 15, 2021
Cited By (7)
US 12,299,508 US 12,321,396 US 12,373,414 US 12,613,864 US 12,639,379 US 12,670,170 US 12,711,032