IP Library › Granted Patent US 11,126,632
Granted Patent B2
US 11,126,632 · App. 16/051,203 · Granted Sep 21, 2021

Subquery generation based on search configuration data from an external data system

Inventors: Sourav Pal (Foster City, CA); Arindam Bhattacharjee (Fremont, CA)
Assignee: Splunk Inc.
G06F16/2471G06F16/211G06F16/27G06F16/951G06F40/205
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,126,632
App. No.
16/051,203
Filed
Jul 31, 2018
Granted
Sep 21, 2021
Kind
B2
Art Unit
2169
USPC
707/722
Abstract

Systems and methods are disclosed for executing a query that includes an indication to process data managed by an external data system. The system identifies the external data system that manages the data to be processed, and obtained search configuration data from the external system. The system uses the search configuration data to generate a subquery for the external data system. The system also generates instructions for one or more worker nodes to receive and process results of the subquery from the external data system.

Claims (118)

1. A computer-implemented method, comprising:

receiving, at a data intake and query system, a query identifying a set of data to be processed and a manner of processing the set of data, wherein the data intake and query system comprises at least one computing device;

determining that the set of data includes at least a subset of data associated with an external data system;

defining, by the data intake and query system, a query processing scheme for obtaining and processing the set of data, wherein the defining the query processing scheme comprises:

obtaining search configuration data from the external data system,

determining a subquery for the external data system based on the search configuration data, the subquery identifying the at least a subset of data and a manner of processing the at least a subset of data,

determining a data ingest estimate based on an analysis of the subquery, wherein the data ingest estimate includes an estimate of an amount of data to be received from the external data system as a result of the external data system executing the subquery, and

generating, based on the data ingest estimate, instructions for one or more worker nodes to receive and process results of the subquery to form processed results and to provide the processed results to the data intake and query system; and

executing the query based on the defining the query processing scheme, wherein the executing the query comprises communicating the generated instructions to the one or more worker nodes.

2. The method of claim 1 , wherein obtaining the search configuration data comprises:

assigning a worker node of the one or more worker nodes to request a version identifier from the external data system;

receiving the version identifier from the worker node; and

based on the version identifier, requesting the search configuration data from the external data system.

3. The method of claim 1 , wherein obtaining the search configuration data comprises:

assigning a worker node of the one or more worker nodes to request a version identifier from the external data system;

receiving, at the worker node, the version identifier; and

based on the version identifier, requesting the search configuration data from the external data system.

4. The method of claim 1 , wherein the at least a subset of data is a second subset of data, and the processed results are second processed results, the method further comprising:

determining that the set of data includes a first subset of data associated with the data intake and query system,

wherein defining the query processing scheme, further comprises:

generating a subquery for the data intake and query system, the subquery for the data intake and query system identifying the first subset of data and a manner of processing the first subset of data, and

generating instructions for one or more worker nodes to receive and process results of the subquery for the data intake and query system to form first processed results and to provide the first processed results to the data intake and query system.

5. The method of claim 1 , wherein the at least a subset of data is a second subset of data, and the processed results are second processed results, the method further comprising:

determining that the set of data includes a first subset of data associated with the data intake and query system,

wherein defining the query processing scheme, further comprises:

generating a subquery for the data intake and query system, the subquery for the data intake and query system identifying the first subset of data and a manner of processing the first subset of data, and

generating instructions for one or more worker nodes to receive and process results of the subquery for the data intake and query system to generate first processed results, to combine and process the first processed results and the second processed results to form combined processed results, and to provide the combined processed results to the data intake and query system.

6. The method of claim 1 , wherein the at least a subset of data is a first subset of data, the processed results are first processed results, and the external data system is a first external data system, the method further comprising:

determining that the set of data includes a second subset of data associated with a second external data system,

wherein defining the query processing scheme, further comprises:

generating a subquery for the second external data system, the subquery for the second external data system identifying the second subset of data and a manner of processing the second subset of data; and

generating instructions for one or more worker nodes to receive and process results of the subquery for the second external data system to form second processed results and to provide the second processed results to the data intake and query system.

7. The method of claim 1 , wherein the data intake and query system and the external data system each independently execute queries other than the query.

8. The method of claim 1 , wherein the data intake and query system and the external data system each independently receive queries other than the query, generate subqueries based on the queries, and execute the subqueries.

9. The method of claim 1 , wherein the data intake and query system and the external data system each include one or more search heads and one or more indexers.

10. The method of claim 1 , wherein determining that the set of data includes at least the subset of data comprises:

parsing the query;

identifying a search parameter in the query associated with a search of an external data source;

identifying the external data system based on said identifying the search parameter; and

determining access information to access the external data system.

11. The method of claim 1 , wherein determining that the set of data includes at least the subset of data comprises:

parsing the query;

identifying a search parameter in the query that includes an identification of the external data system; and

determining access information to access the external data system based on said identification of the external data system.

12. The method of claim 1 , wherein determining that the set of data includes at least the subset of data comprises:

parsing the query;

identifying a search parameter in the query associated with a search of an external data source;

parsing a configuration file based on the search parameter;

identifying the external data system based on said parsing the configuration file; and

determining access information to access the external data system based on said identifying the external data system.

13. The method of claim 1 , further comprising associating a search identifier with the external data system, wherein the one or more worker nodes process results of the subquery based on the search identifier.

14. The method of claim 1 ,

wherein defining the query processing scheme further comprises associating, by the data intake and query system, a first search identifier with the external data system, and

wherein executing the query comprises:

receiving, by the one or more worker nodes, the results of the subquery, wherein the results of the subquery include a second search identifier assigned to the results of the subquery by the external data system;

mapping the first search identifier to the second search identifier; and

processing the results of the subquery based on said mapping.

15. The method of claim 1 , wherein defining the query processing scheme further comprises:

determining a processing capability of the external data system,

wherein the data ingest estimate is based on the processing capability.

16. The method of claim 1 , wherein generating instructions for the one or more worker nodes comprises:

assigning a worker node of the one or more worker nodes to request a version identifier from the external data system;

receiving the version identifier from the worker node,

wherein the data ingest estimate is based on the version identifier.

17. The method of claim 1 , wherein a worker node of the one or more worker nodes is assigned to determine the data ingest estimate for the subquery, wherein generating instructions for the one or more worker nodes comprises:

communicating the subquery to the worker node, wherein the worker node communicates the subquery to the external data system and receives the data ingest estimate from the external data system.

18. The method of claim 1 , wherein a worker node of the one or more worker nodes is assigned to determine the data ingest estimate for the subquery, wherein generating instructions for the one or more worker nodes comprises:

communicating the subquery to the worker node, wherein the worker node parses the subquery, communicates one or more search parameters to the external data system, and receives the data ingest estimate from the external data system.

19. The method of claim 1 , wherein generating instructions for the one or more worker nodes comprises:

determining a quantity of partitions to ingest the results of the subquery; and

generating instructions for the one or more worker nodes based on the quantity of partitions.

20. The method of claim 1 , wherein defining the query processing scheme, further comprises:

obtaining network access information from at least one worker node of the one or more worker nodes, wherein executing the query comprises communicating the network access information to the external data system.

21. The method of claim 1 , wherein the subquery includes instructions for the external data system to distribute the results of the subquery to a plurality of worker nodes of the one or more worker nodes.

22. The method of claim 1 , wherein the subquery includes instructions for the external data system to communicate the results of the subquery to only one worker node of the one or more worker nodes, and wherein defining the query processing scheme further comprises generating instructions for the one worker node to distribute the results of the subquery to a plurality of worker nodes of the one or more worker nodes.

23. The method of claim 1 , wherein executing the query comprises:

communicating the subquery to the one or more worker nodes, wherein at least one worker node of the one or more worker nodes communicates the subquery to the external data system, the external data system processes and executes the subquery, and the one or more worker nodes receive and process the results of the subquery to form the processed results; and

receiving the processed results from the one or more worker nodes.

24. The method of claim 1 , wherein executing the query comprises:

communicating the subquery to the external data system using the one or more worker nodes, wherein the external data system processes and executes the subquery using the one or more worker nodes and the one or more worker nodes receive and process the results of the subquery to form the processed results; and

receiving the processed results from the one or more worker nodes.

25. A computing system of a data intake and query system, the computing system comprising:

memory; and

one or more processing devices coupled to the memory and configured to:

receive, at a data intake and query system, a query identifying a set of data to be processed and a manner of processing the set of data;

determine that the set of data includes at least a subset of data associated with an external data system;

define, by the data intake and query system, a query processing scheme for obtaining and processing the set of data, wherein to define the query processing scheme, the one or more computing devices are configured to:

obtain search configuration data from the external data system,

determine a subquery for the external data system based on the search configuration data, the subquery identifying the at least a subset of data and a manner of processing the at least a subset of data,

determine a data ingest estimate based on an analysis of the subquery, wherein the data ingest estimate includes an estimate of an amount of data to be received from the external data system as a result of the external data system executing the subquery, and

generate, based on the data ingest estimate, instructions for one or more worker nodes to receive and process results of the subquery to form processed results and to provide the processed results to the data intake and query system; and

initiate execution of the query based on the defined query processing scheme, wherein executing the query comprises communicating the generated instructions to the one or more worker nodes.

26. The computing system of claim 25 , wherein to obtain the search configuration data, the one or more processing devices are configured to:

assign a worker node of the one or more worker nodes to request a version identifier from the external data system;

receive the version identifier from the worker node; and

based on the version identifier, request the search configuration data from the external data system.

27. Non-transitory computer readable media comprising computer-executable instructions that, when executed by a computing system of a first data intake and query system, cause the computing system to:

receive, at a data intake and query system, a query identifying a set of data to be processed and a manner of processing the set of data;

determine that the set of data includes at least a subset of data associated with an external data system;

define, by the data intake and query system, a query processing scheme for obtaining and processing the set of data, wherein to define the query processing scheme the computer-executable instructions cause the computing system to:

obtain search configuration data from the external data system,

determine a subquery for the external data system based on the search configuration data, the subquery identifying the at least a subset of data and a manner of processing the at least a subset of data,

determine a data ingest estimate based on an analysis of the subquery, wherein the data ingest estimate includes an estimate of an amount of data to be received from the external data system as a result of the external data system executing the subquery, and

generate, based on the data ingest estimate, instructions for one or more worker nodes to receive and process results of the subquery to form processed results and to provide the processed results to the data intake and query system; and

initiate execution of the query based on the defined query processing scheme, wherein executing the query comprises communicating the generated instructions to the one or more worker nodes.

28. The non-transitory computer readable media of claim 27 , wherein the search configuration data includes instructions for parsing a search parameter of the subquery.

29. A computer-implemented method, comprising:

receiving, at a data intake and query system, a query identifying a set of data to be processed and a manner of processing the set of data;

determining that the set of data includes at least a first subset of data associated with an external data system and a second subset of data associated with the data intake and query system;

defining, by the data intake and query system, a query processing scheme for obtaining and processing the set of data, wherein the defining the query processing scheme comprises:

obtaining search configuration data from the external data system,

determining a subquery for the external data system based on the search configuration data, the subquery identifying the first subset of data and a manner of processing the first subset of data,

determining a first data ingest estimate based on the subquery, wherein the first data ingest estimate includes an estimate of an amount of data to be received from the external data system as a result of the external data system executing the subquery,

generating, based on the first data ingest estimate, instructions for one or more worker nodes to receive and process results of the subquery to form first processed results and to provide the first processed results to the data intake and query system,

generating a subquery for the data intake and query system, the subquery for the data intake and query system identifying the second subset of data and a manner of processing the second subset of data,

determining a second data ingest estimate based on the subquery for the data intake and query system, and

generating, based on the second data ingest estimate, instructions for one or more worker nodes to receive and process results of the subquery for the data intake and query system to form second processed results and to provide the second processed results to the data intake and query system; and

executing the query based on the query processing scheme.

Assignments (3)
CHANGE OF NAME Recorded Jul 22, 2025
From: SPLUNK INC.
To: SPLUNK LLC
Reel/Frame 072170/0599 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2025
From: SPLUNK LLC
To: CISCO TECHNOLOGY, INC.
Reel/Frame 072173/0058 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2019
From: PAL, SOURAV; BHATTACHARJEE, ARINDAM
To: SPLUNK INC.
Reel/Frame 049092/0497 →
Continuity (24)
Continuation In Part 15665159 · Jul 31, 2017
Continuation In Part 15276717 · Sep 26, 2016
Continuation In Part 16051203 · Jul 31, 2018
Continuation In Part 15665148 · Jul 31, 2017
Continuation In Part 15276717 · Sep 26, 2016
Continuation In Part 16051203 · Jul 31, 2018
Continuation In Part 15665187 · Jul 31, 2017
Continuation In Part 15276717 · Sep 26, 2016
Continuation In Part 16051203 · Jul 31, 2018
Continuation In Part 15665248 · Jul 31, 2017
Continuation In Part 15276717 · Sep 26, 2016
Continuation In Part 16051203 · Jul 31, 2018
Continuation In Part 15665197 · Jul 31, 2017
Continuation In Part 15276717 · Sep 26, 2016
Continuation In Part 16051203 · Jul 31, 2018
Continuation In Part 15665279 · Jul 31, 2017
Continuation In Part 15276717 · Sep 26, 2016
Continuation In Part 16051203 · Jul 31, 2018
Continuation In Part 15665302 · Jul 31, 2017
Continuation In Part 15276717 · Sep 26, 2016
Continuation In Part 16051203 · Jul 31, 2018
Continuation In Part 15665339 · Jul 31, 2017
Continuation In Part 15276717 · Jun 26, 2016
Related Publication 20190138640A1 · May 9, 2019
Cited By (20)
US 12,204,536 US 12,204,538 US 12,204,593 US 12,248,484 US 12,265,525 US 12,271,389 US 12,287,790 US 12,353,413 US 12,393,593 US 12,393,631 US 12,430,332 US 12,436,963 US 12,505,246 US 12,530,841 US 12,585,638 US 12,613,864 US 12,639,379 US 12,650,965 US 12,670,152 US 12,711,426