IP Library › Granted Patent US 11,500,871
Granted Patent B1
US 11,500,871 · App. 17/074,100 · Granted Nov 15, 2022

Systems and methods for decoupling search processing language and machine learning analytics from storage of accessed data

Inventors: Chinmay Madhav Kulkarni (Sunnyvale, CA); Lin Ma (Burnaby, CA); Amir Malekpour (Vancouver, CA); Mohan Rajagopalan (Mountain View, CA); John C. Reed (Saratoga, CA); Ram Sriharsha (Oakland, CA)
Assignee: SPLUNK Inc.
G06F16/24549G06F16/2455G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,500,871
App. No.
17/074,100
Filed
Oct 19, 2020
Granted
Nov 15, 2022
Kind
B1
Art Unit
2156
USPC
707/765
Abstract

A computer-implemented method is disclosed that includes operations of receiving a query to be executed, the query including an indication of a data source at which input data is be to obtained, wherein the query is to be executed on the input data, determining a schema of the input data, determining fields of the input data that are required for execution of the query by analyzing a sequence of operators forming the query, determining one or more alterations to the query to improve efficiency of the execution of the query based on the fields of input data required for the execution, and generating an altered query be altering the query in accordance with the one or more alterations. The method may further include converting the query to a directed acyclic graph (DAG) and providing the DAG to a distributed processing engine configured to execute the DAG.

Claims (71)

1. A computer-implemented method, comprising:

receiving a query to be executed, the query including an indication of a data source at which input data is to be obtained, wherein the query is to be executed on the input data;

determining a schema of the input data;

determining fields of the input data that are required for execution of the query by analyzing a sequence of operators forming the query;

determining one or more alterations to the query to improve efficiency of the execution of the query based on the fields of input data required for the execution wherein determining the one or more alterations includes: (i) selecting a particular operator of the sequence of operators to which to apply either a filter or a projection, (ii) determining a subset of fields to be returned by the particular operator that are required by operators subsequent to the particular operator, and (iii) applying the filter or the projection to the particular operator to limit fields returned by the particular operator to the subset of fields; and

generating an altered query by altering the query in accordance with the one or more alterations.

2. The method of claim 1 , further comprising:

converting the query to a directed acyclic graph (DAG); and

providing the DAG to a distributed processing engine configured to execute the DAG in a distributed runtime environment.

3. The method of claim 1 , wherein determining the schema of the input data includes:

performing an initial preliminary read to obtain a sample of the input data; and

detecting one or more fields and corresponding sourcetypes within the sample of the input data, wherein the detected one or more fields and corresponding sourcetypes comprise the schema of the input data.

4. The method of claim 1 , wherein analyzing the sequence of operators forming the query includes:

parsing a syntax of each operator in a reverse order of the sequence, and

identifying each field required for the execution of the query.

5. The method of claim 1 , wherein determining the one or more alterations to the query includes:

determining a projection to be applied to an operator to reduce an amount of the input data read from the data source.

6. The method of claim 1 , wherein determining the one or more alterations to the query includes:

determining a filter to be applied to an operator to reduce an amount of the input data read from the data source.

7. The method of claim 1 , wherein determining the one or more alterations to the query includes:

determining a particular operator to which to apply either a filter or a projection through analyzing the sequence of operators that is as close to a beginning of the query as possible without affecting a result of the execution.

8. The method of claim 1 , wherein generating the altered query includes applying at least one of a filter or a projection to an operator of the query to reduce an amount of the input data read from the data source.

9. The method of claim 1 , wherein the query is provided in accordance with a pipeline command language that includes the sequence of operators in which a set of data or results produced based on execution of the first operator is applied to the second operator.

10. A computing device, comprising:

a processor; and

a non-transitory computer-readable medium having stored thereon instructions that, when executed by the processor, cause the processor to perform operations including:

receiving a query to be executed, the query including an indication of a data source at which input data is to be obtained, wherein the query is to be executed on the input data;

determining a schema of the input data;

determining fields of the input data that are required for execution of the query by analyzing a sequence of operators forming the query;

determining one or more alterations to the query to improve efficiency of the execution of the query based on the fields of input data required for the execution wherein determining the one or more alterations includes: (i) selecting a particular operator of the sequence of operators to which to apply either a filter or a projection, (ii) determining a subset of fields to be returned by the particular operator that are required by operators subsequent to the particular operator, and (iii) applying the filter or the projection to the particular operator to limit fields returned by the particular operator to the subset of fields; and

generating an altered query by altering the query in accordance with the one or more alterations.

11. The computing device of claim 10 , further comprising:

converting the query to a directed acyclic graph (DAG); and

providing the DAG to a distributed processing engine configured to execute the DAG in a distributed runtime environment.

12. The computing device of claim 10 , wherein determining the schema of the input data includes:

performing an initial preliminary read to obtain a sample of the input data; and

detecting one or more fields and corresponding sourcetypes within the sample of the input data, wherein the detected one or more fields and corresponding sourcetypes comprise the schema of the input data.

13. The computing device of claim 10 , wherein analyzing the sequence of operators forming the query includes:

parsing a syntax of each operator in a reverse order of the sequence, and

identifying each field required for the execution of the query.

14. The computing device of claim 10 , wherein determining the one or more alterations to the query includes:

determining a projection to be applied to an operator to reduce an amount of the input data read from the data source.

15. The computing device of claim 10 , wherein determining the one or more alterations to the query includes:

determining a filter to be applied to an operator to reduce an amount of the input data read from the data source.

16. The computing device of claim 10 , wherein determining the one or more alterations to the query includes:

determining a particular operator to which to apply either a filter or a projection through analyzing the sequence of operators that is as close to a beginning of the query as possible without affecting a result of the execution.

17. The computing device of claim 10 , wherein generating the altered query includes applying at least one of a filter or a projection to an operator of the query to reduce an amount of the input data read from the data source.

18. The computing device of claim 10 , wherein the query is provided in accordance with a pipeline command language that includes the sequence of operators in which a set of data or results produced based on execution of the first operator is applied to the second operator.

19. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processor, cause the one or more processors to perform operations including:

receiving a query to be executed, the query including an indication of a data source at which input data is to be obtained, wherein the query is to be executed on the input data;

determining a schema of the input data;

determining fields of the input data that are required for execution of the query by analyzing a sequence of operators forming the query;

determining one or more alterations to the query to improve efficiency of the execution of the query based on the fields of input data required for the execution wherein determining the one or more alterations includes: (i) selecting a particular operator of the sequence of operators to which to apply either a filter or a projection, (ii) determining a subset of fields to be returned by the particular operator that are required by operators subsequent to the particular operator, and (iii) applying the filter or the projection to the particular operator to limit fields returned by the particular operator to the subset of fields; and

generating an altered query by altering the query in accordance with the one or more alterations.

20. The non-transitory computer-readable medium of claim 19 , further comprising:

converting the query to a directed acyclic graph (DAG); and

providing the DAG to a distributed processing engine configured to execute the DAG in a distributed runtime environment.

21. The non-transitory computer-readable medium of claim 19 , wherein determining the schema of the input data includes:

performing an initial preliminary read to obtain a sample of the input data; and

detecting one or more fields and corresponding sourcetypes within the sample of the input data, wherein the detected one or more fields and corresponding sourcetypes comprise the schema of the input data.

22. The non-transitory computer-readable medium of claim 19 , wherein analyzing the sequence of operators forming the query includes:

parsing a syntax of each operator in a reverse order of the sequence, and

identifying each field required for the execution of the query.

23. The non-transitory computer-readable medium of claim 19 , wherein determining the one or more alterations to the query includes:

determining a projection to be applied to an operator to reduce an amount of the input data read from the data source.

24. The non-transitory computer-readable medium of claim 19 , wherein determining the one or more alterations to the query includes:

determining a filter to be applied to an operator to reduce an amount of the input data read from the data source.

25. The non-transitory computer-readable medium of claim 19 , wherein determining the one or more alterations to the query includes:

determining a particular operator to which to apply either a filter or a projection through analyzing the sequence of operators that is as close to a beginning of the query as possible without affecting a result of the execution.

26. The non-transitory computer-readable medium of claim 19 , wherein generating the altered query includes applying at least one of a filter or a projection to an operator of the query to reduce an amount of the input data read from the data source.

27. The non-transitory computer-readable medium of claim 19 , wherein the query is provided in accordance with a pipeline command language that includes the sequence of operators in which a set of data or results produced based on execution of the first operator is applied to the second operator.

Assignments (3)
CHANGE OF NAME Recorded Jul 22, 2025
From: SPLUNK INC.
To: SPLUNK LLC
Reel/Frame 072170/0599 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2025
From: SPLUNK LLC
To: CISCO TECHNOLOGY, INC.
Reel/Frame 072173/0058 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2020
From: KULKARNI, CHINMAY MADHAV; MA, LIN; MALEKPOUR, AMIR; RAJAGOPALAN, MOHAN; REED, JOHN C.; SRIHARSHA, RAM
To: SPLUNK INC.
Reel/Frame 054246/0381 →
Cited By (9)
US 12,210,514 US 12,222,938 US 12,235,851 US 12,367,022 US 12,411,670 US 12,554,712 US 12,625,869 US 12,699,928 US 12,724,593