Externally distributed buckets for execution of queries
A data intake and query system can manage the search of data stored at an external location relative to the data intake and query system using one or more indexers. The data intake and query system can receive data stored at the external location. The data intake and query system can process the data and generate an index using the one or more indexers. The data intake and query system can discard the data and store the index and a location identifier of the external location in one or more buckets. In response to a query, the data intake and query system can identify that at least a subset of the data is responsive to the query using the index and can obtain the at least the subset of the data from the external location using the location identifier.
1 . A method, comprising:
receiving, at a first indexer, from a first external location in a first computing system, a first set of data;
processing, by the first indexer, the first set of data;
generating, by the first indexer, a first index using the first set of data from the first external location based at least in part on processing the first set of data;
storing, by the first indexer, the first index and a location identifier generated to identify the first external location from which the first set of data is obtained, in one or more buckets of a data store; and
executing a search query, wherein executing the search query comprises:
identifying, using the first index generated using the first set of data from the first external location, that at least a first portion of the first set of data from the first external location is responsive to the search query based at least in part on processing the search query;
based on identifying the at least the first portion of the first set of data being responsive to the search query, obtaining, using the location identifier stored by the first indexer and identifying the first external location from which the first set of data is obtained from the first external location, the at least the first portion of the first set of data based at least in part on the search query; and
processing the at least the first portion of the first set of data obtained from the first external location in accordance with the search query to generate one or more results.
2 . The method of claim 1 , wherein the location identifier comprises at least one of a link, a pointer, a reference, or an address.
3 . The method of claim 1 , wherein the first external location comprises a data lake.
4 . The method of claim 1 , further comprising:
discarding the first set of data.
5 . The method of claim 1 , further comprising:
discarding the first set of data from the data store based at least in part on processing the first set of data.
6 . The method of claim 1 , wherein storing the first index and the location identifier in the one or more buckets comprises storing the first index and the location identifier in a first bucket of the one or more buckets, the method further comprising:
storing, in a second bucket of the one or more buckets, a second index and a second location identifier of a second external location in a second computing system, wherein the second index is based at least in part on a second set of data received from the second external location.
7 . The method of claim 1 , wherein storing the first index and the location identifier in the one or more buckets comprises storing the first index and the location identifier in a particular bucket of the one or more buckets, the method further comprising:
receiving, from a second external location in a second computing system, a second set of data; and
adjusting the first index based at least in part on the second set of data such that the first index is associated with the first set of data and the second set of data.
8 . The method of claim 1 , wherein the first set of data is stored at a first location, wherein the first index is stored at a second location, and wherein the first location and the second location are separate, distinct locations.
9 . The method of claim 1 , wherein the first set of data comprises a Parquet file.
10 . The method of claim 1 , wherein, based at least in part on the first index, a second indexer downloads at least a second portion of the first set of data from the first external location and executes one or more searches on the downloaded at least the second portion of the first set of data.
11 . The method of claim 1 , wherein the first index comprises an inverted index, and wherein the inverted index has information regarding at least one of a partition associated with events or an origin of the events, each event of the events comprising a respective portion of the first set of data associated with a respective time stamp, the inverted index comprising a plurality of entries, each entry of the plurality of entries comprising:
a respective token or a respective field-value pair, and
one or more respective event references, each event reference of the one or more respective event references indicative of a respective event that includes the respective token or a field-value corresponding to the respective field-value pair.
12 . The method of claim 1 , wherein the first set of data comprises raw machine data.
13 . The method of claim 1 , wherein the first indexer is an indexer of a data intake and query system, wherein the data intake and query system is distinct and separate from the first computing system.
14 . The method of claim 1 , further comprising:
identifying that the first set of data is stored at the first external location, wherein receiving the first set of data comprises receiving the first set of data in response to identifying that the first set of data is stored at the first external location.
15 . The method of claim 1 , wherein metadata associated with the first set of data is stored at the first external location, and wherein a second indexer utilizes the metadata to execute one or more searches.
16 . The method of claim 1 , wherein the location identifier is generated based on one or more communications from an external data source at the first external location.
17 . A system comprising:
a data store; and
one or more processors configured to:
receive, at a first indexer, from a first external location in a first computing system, a first set of data;
process, by the first indexer, the first set of data;
generate, by the first indexer, a first index using the first set of data from the first external location based at least in part on processing the first set of data;
store, by the first indexer, the first index and a location identifier generated to identify the first external location from which the first set of data is obtained, in one or more buckets of the data store; and
execute a search query, wherein executing the search query comprises:
identifying, using the first index generated using the first set of data from the first external location, that at least a first portion of the first set of data from the first external location is responsive to the search query based at least in part on processing the search query;
based on identifying the at least the first portion of the first set of data being responsive to the search query, obtaining, using the location identifier stored by the first indexer and identifying the first external location from which the first set of data is obtained from the first external location, the at least the first portion of the first set of data based at least in part on the search query; and
process the at least the first portion of the first set of data obtained from the first external location in accordance with the search query to generate one or more results.
18 . The system of claim 17 , wherein the one or more processors are further configured to:
discard the first set of data from the data store based at least in part on processing the first set of data.
19 . The system of claim 17 , wherein to store the first index and the location identifier in the one or more buckets, the one or more processors are further configured to store the first index and the location identifier in a particular bucket of the one or more buckets, wherein the one or more processors are further configured to:
receive, from a second external location in a second computing system, a second set of data; and
adjust the first index based at least in part on the second set of data such that the first index is associated with the first set of data and the second set of data.
20 . Non-transitory computer-readable media including computer-executable instructions that, when executed by a particular computing system, cause the particular computing system to:
receive, at a first indexer, from a first external location in a first computing system, a first set of data;
process, by the first indexer, the first set of data;
generate, by the first indexer, a first index using the first set of data from the first external location based at least in part on processing the first set of data;
store, by the first indexer, the first index and a location identifier generated to identify the first external location from which the first set of data is obtained, in one or more buckets of a data store; and
execute a search query, wherein executing the search query comprises:
identifying, using the first index generated using the first set of data from the first external location, that at least a first portion of the first set of data from the first external location is responsive to the search query based at least in part on processing the search query;
based on identifying the at least the first portion of the first set of data being responsive to the search query, obtaining, using the location identifier stored by the first indexer and identifying the first external location from which the first set of data is obtained from the first external location, the at least the first portion of the first set of data based at least in part on the search query; and
processing the at least the first portion of the first set of data obtained from the first external location in accordance with the search query to generate one or more results.