Search recovery using a shared storage system following a failed search peer
Systems and methods are described for improving data availability and/or resiliency of indexers of a data intake and query system. A data intake and query system can index and search large amounts of data using one or more indexers. An indexer can store a copy of the data it is processing, the results of processing the data, or a copy of the data that the indexer is assigned to search, in the shared storage system. In the event an indexer fails or is otherwise unable to search data that it has been assigned to search, a cluster master can assign one or more second indexers to search the data. The one or more second indexers can download the data from the shared storage system.
1 . A method comprising:
receiving a bucket map identifier from a first search node for execution of at least a portion of a query by the first search node, wherein the bucket map identifier is associated with a group of data and stored in a data structure that associates the bucket map identifier with the group of data;
identifying a set of data identifiers that identify one or more groups of data with which the bucket map identifier and the first search node are associated, the one or more groups of data stored in a shared storage system accessible by a plurality of search nodes;
communicating the set of data identifiers to the first search node to execute the query;
upon communicating the set of data identifiers to the first search node to execute the query, determining that the first search node is not available by monitoring status information from each search node of the plurality of search nodes;
in response to the determining that the first search node is not available, determining that execution of the query has not been completed by the first search node in association with a portion of the one or more groups of data stored in the shared storage system and determining an uncompleted portion of the query;
generating a new query based on the uncompleted portion of the query, the new query including an indication of the portion of the one or more groups of data for which execution of the query was not completed by the first search node prior to the first search node being unavailable; and
assigning one or more second search nodes to obtain the portion of the one or more groups of data from the shared storage system and execute the new query on the portion of the one or more groups of data for which execution of the query was not completed by the first search node.
2 . The method of claim 1 , wherein the bucket map identifier is a first bucket map identifier, wherein a second bucket map identifier is associated with the one or more second search nodes and not including the first search node, and wherein assigning the one or more second search nodes is based on:
receiving the second bucket map identifier and a set of search node identifiers, wherein the set of search node identifiers identifies the one or more second search nodes, wherein a search head is configured to communicate the second bucket map identifier to the one or more second search nodes.
3 . The method of claim 1 , wherein said assigning the one or more second search nodes is based on an availability of the one or more second search nodes.
4 . The method of claim 1 , wherein execution of the query is to be performed by a first set of search nodes comprising the first search node and the one or more second search nodes.
5 . The method of claim 1 , wherein execution of the query is to be performed by a first set of search nodes that does not include at least one of the one or more second search nodes.
6 . The method of claim 1 , wherein determining that the first search node is not available is based on one or more metrics associated with the first search node satisfying a metrics threshold.
7 . The method of claim 1 , wherein determining that the first search node is not available is based on an amount of available processing resources of the first search node satisfying a processing resources threshold.
8 . The method of claim 1 , wherein the portion of the one or more groups of data comprises the one or more groups of data.
9 . The method of claim 1 , wherein assignment to the one or more second search nodes is based on utilization rates.
10 . The method of claim 1 , further comprising disassociating the first search node from the one or more groups of data.
11 . The method of claim 1 , wherein the determination that execution of the query has not been completed is based on obtaining partial search results from the first search node.
12 . The method of claim 1 , further comprising:
identifying a plurality of bucket map identifiers that are associated with the first search node, wherein each of the plurality of bucket map identifiers associate the first search node with one or more data identifiers; and
for each of the plurality of bucket map identifiers, disassociating the one or more data identifiers from the first search node.
13 . The method of claim 1 , further comprising:
identifying a plurality of bucket map identifiers that are associated with the first search node, wherein each of the plurality of bucket map identifiers associate the first search node with one or more data identifiers; and
for each of the plurality of bucket map identifiers:
associating the one or more data identifiers with the one or more second search nodes, and
disassociating the one or more data identifiers from the first search node.
14 . The method of claim 1 , wherein said assigning the one or more second search nodes comprises:
assigning a second search node of the one or more second search nodes to execute a first portion of the new query on the portion of the one or more groups of data; and
assigning a third search node of the one or more second search nodes to execute a second portion of the new query on the portion of the one or more groups of data.
15 . The method of claim 1 , wherein said assigning the one or more second search nodes comprises associating the one or more second search nodes with the bucket map identifier.
16 . The method of claim 1 , wherein said assigning the one or more second search nodes comprises disassociating the first search node from the bucket map identifier.
17 . The method of claim 1 , wherein the one or more groups of data comprise only one group of data.
18 . The method of claim 1 , wherein the bucket map identifier is associated with filter criteria, the filter criteria associated with the query, wherein the query identifies a set of data and manner of processing the set of data, and wherein the one or more groups of data correspond to a subset of the set of data identified by the query.
19 . The method of claim 1 , wherein the bucket map identifier indicates data associated with the set of data identifiers for execution of the query by a first set of search nodes, including the first search node, at a first time period and indicates data associated with a second set of data identifiers for execution of the query by the first set of search nodes at a second time period.
20 . The method of claim 1 , wherein the bucket map identifier is associated with each data identifier of the set of data identifiers.
21 . The method of claim 1 , wherein the one or more groups of data comprise one or more slices of data, wherein the one or more slices of data comprise at least one of raw machine data, structured data, unstructured data, performance metrics data, correlation data, data files, directories of files, data sent over a network, event logs, registries, JSON blobs, XML data, data in a data model, report data, tabular data, messages published to streaming data sources, data exposed in an API, data in a relational database, sensor data, image data, or video data.
22 . The method of claim 1 , wherein the one or more groups of data comprise a bucket, wherein the bucket comprises an inverted index and a plurality of events corresponding to one or more slices of data, the inverted index corresponding to the plurality of events, wherein the one or more slices of data comprise at least one of raw machine data, structured data, unstructured data, performance metrics data, correlation data, data files, directories of files, data sent over a network, event logs, registries, JSON blobs, XML data, data in a data model, report data, tabular data, messages published to streaming data sources, data exposed in an API, data in a relational database, sensor data, image data, or video data.
23 . The method of claim 1 , wherein the one or more groups of data comprise a plurality of buckets, wherein each bucket of the plurality of buckets comprises a particular inverted index and a particular plurality of events corresponding to a particular one or more slices of data, the particular inverted index corresponding to the particular plurality of events, wherein each particular one or more slices of data comprise at least one of raw machine data, structured data, unstructured data, performance metrics data, correlation data, data files, directories of files, data sent over a network, event logs, registries, JSON blobs, XML data, data in a data model, report data, tabular data, messages published to streaming data sources, data exposed in an API, data in a relational database, sensor data, image data, or video data.
24 . The method of claim 1 , wherein the one or more groups of data comprise a plurality of buckets, wherein each data identifier of the set of data identifiers identifies a corresponding bucket of the plurality of buckets.
25 . The method of claim 1 , wherein said determining that the first search node is not available is based on an absence of communications from the first search node.
26 . The method of claim 1 , wherein said determining that the first search node is not available comprises determining at least one of a network failure, an error associated with the first search node, a utilization rate of the first search node, an amount of processing resources used or in use by the first search node, or an amount of memory used or in use by the first search node.
27 . The method of claim 1 , wherein said determining that the first search node is not available is based on at least one of a determination that the first search node did not execute the query on the portion of the one or more groups of data, a determination that the first search node is busy, or a determination that the first search node is failing.
28 . A computing system corresponding to a search head of a data intake and query system, the computing system comprising:
memory; and
one or more processors coupled to the memory and configured to:
receive a bucket map identifier from a first search node for execution of at least a portion of a query by the first search node, wherein the bucket map identifier is associated with a group of data and stored in a data structure that associates the bucket map identifier with the group of data;
identify a set of data identifiers that identify one or more groups of data with which the bucket map identifier and the first search node are associated, the one or more groups of data stored in a shared storage system accessible by a plurality of search nodes;
communicate the set of data identifiers to the first search node to execute the query;
upon communicating the set of data identifiers to the first search node to execute the query, determine that the first search node is not available by monitoring status information from each search node of the plurality of search nodes;
in response to the determining that the first search node is not available, determine that execution of the query has not been completed by the first search node in association with a portion of the one or more groups of data stored in the shared storage system and determining an uncompleted portion of the query;
generate a new query based on the uncompleted portion of the query, the new query including an indication of the portion of the one or more groups of data for which execution of the query was not completed by the first search node prior to the first search node being unavailable; and
assign one or more second search nodes to obtain the portion of the one or more groups of data from the shared storage system and execute the new query on the portion of the one or more groups of data for which execution of the query was not completed by the first search node.
29 . Non-transitory computer readable media comprising computer-executable instructions that, when executed by a computing system corresponding to a search head of a data intake and query system, cause the computing system to:
receive a bucket map identifier from a first search node for execution of at least a portion of a query by the first search node, wherein the bucket map identifier is associated with a group of data and stored in a data structure that associates the bucket map identifier with the group of data;
identify a set of data identifiers that identify one or more groups of data with which the bucket map identifier and the first search node are associated, the one or more groups of data stored in a shared storage system accessible by a plurality of search nodes;
communicate the set of data identifiers to the first search node to execute the query;
upon communicating the set of data identifiers to the first search node to execute the query, determine that the first search node is not available by monitoring status information from each search node of the plurality of search nodes;
in response to the determining that the first search node is not available, determine that execution of the query has not been completed by the first search node in association with a portion of the one or more groups of data stored in the shared storage system and determining an uncompleted portion of the query;
generate a new query based on the uncompleted portion of the query, the new query including an indication of the portion of the one or more groups of data for which execution of the query was not completed by the first search node prior to the first search node being unavailable; and
assign one or more second search nodes to obtain the portion of the one or more groups of data from the shared storage system and execute the new query on the portion of the one or more groups of data for which execution of the query was not completed by the first search node.