Systems and methods for coordinated summarization of indexed data
Provided are systems and methods for concurrent summarization of indexed data. In some embodiments, two or more summary processes can be executed concurrently (e.g., in parallel) by an indexer to generate summaries for respective subsets of indexed data (e.g., partitions or buckets of indexed data) managed by the indexer.
1 . A method of executing a scheduled summary request, said method comprising:
generating a set of time-stamped event records from raw machine data;
indexing and storing the set of time-stamped event records in two or more buckets of event records;
scheduling a summary request for a first bucket of the two or more buckets to be executed by a plurality of indexers;
locking a first summary directory associated with the first bucket;
writing first summary data for the first bucket to the first summary directory, wherein the first summary data comprises at least one summary file that includes a summary representing a statistic associated with data in the first bucket and an indication that an entirety of contents of the first bucket is summarized based on determining a completion of the set of time-stamped event records being stored to the first bucket and summarization of the set of time-stamped event records stored in the first bucket;
unlocking the first summary directory; and
responsive to the indication that the entirety of contents of the first bucket is summarized, writing second summary data for a second bucket to a second summary directory.
2 . The method of claim 1 , wherein the two or more buckets are stored in a memory of an indexer.
3 . The method of claim 2 , further comprising receiving the summary request at the indexer from one or more search heads.
4 . The method of claim 1 , wherein each of the two or more buckets is associated with a timespan and comprises time-stamped event records having a respective timestamp corresponding to a respective time in the timespan.
5 . The method of claim 1 , further comprising:
receiving, from an entity, a request for the second summary data; and
providing, to the entity, the second summary data for the second bucket responsive to the request, wherein the entity is configured to generate a result based at least in part on contents of the second summary data for the second bucket.
6 . The method of claim 1 , wherein the second summary data comprises a first subset of values of fields of events corresponding to a data model, and further comprising:
receiving, from an entity, a request for summary data; and
providing, to the entity, the second summary data, wherein the entity is configured to generate a set of values for the data model determined based at least in part on the first subset of values of fields.
7 . The method of claim 6 , wherein the summary request is scheduled responsive to enabling acceleration of the data model.
8 . The method of claim 1 , wherein the summary request is scheduled responsive to enabling acceleration of a report.
9 . The method of claim 1 , further comprising:
receiving a search request;
generating search results based at least in part on the first summary data for the first bucket; and
providing the search results in response to the search request.
10 . The method of claim 1 , further comprising:
responsive to an indication that the entirety of contents of the first bucket is not summarized, updating first summary data based on processing an un-summarized portion of the first bucket.
11 . The method of claim 1 , further comprising:
determining whether the second summary directory associated with the second bucket is not currently locked by a concurrent process prior to writing second summary data for the second bucket to the second summary directory.
12 . The method of claim 1 , further comprising:
responsive to the indication that the entirety of contents of the first bucket is not summarized, updating the first summary data based on processing an un-summarized portion of the first bucket.
13 . The method of claim 1 , further comprising:
determining whether the second summary directory associated with the second bucket is not currently locked by a concurrent process prior to writing the second summary data for the second bucket to the second summary directory.
14 . One or more non-transitory computer readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform steps of:
generating a set of time-stamped event records from raw machine data;
indexing and storing the set of time-stamped event records in two or more buckets of event records;
scheduling a summary request for a first bucket of the two or more buckets to be executed by a plurality of indexers;
locking a first summary directory associated with the first bucket;
writing first summary data for the first bucket to the first summary directory, wherein the first summary data comprises at least one summary file that includes a summary representing a statistic associated with data in the first bucket and an indication that an entirety of contents of the first bucket is summarized based on determining a completion of the set of time-stamped event records being stored to the first bucket and summarization of the set of time-stamped event records stored in the first bucket;
unlocking the first summary directory; and
responsive to the indication that the entirety of contents of the first bucket is summarized, writing second summary data for a second bucket to a second summary directory.
15 . The one or more non-transitory computer readable media of claim 12 , wherein each of the two or more buckets is associated with a timespan and comprises time-stamped event records having a respective timestamp corresponding to a respective time in the timespan.
16 . The one or more non-transitory computer readable media of claim 12 , wherein the steps further comprise:
receiving, from an entity, a request for the second summary data; and
providing, to the entity, the second summary data for the second bucket responsive to the request, wherein the entity is configured to generate a result based at least in part on contents of the second summary data for the second bucket.
17 . The one or more non-transitory computer readable media of claim 12 , wherein the second summary data comprises a first subset of values of fields of events corresponding to a data model, and further comprising:
receiving, from an entity, a request for summary data; and
providing, to the entity, the second summary data, wherein the entity is configured to generate a set of values for the data model determined based at least in part on the first subset of values of fields.
18 . The one or more non-transitory computer readable media of claim 17 , wherein the summary request is scheduled responsive to enabling acceleration of the data model.
19 . The one or more non-transitory computer readable media of claim 12 , wherein the summary request is scheduled responsive to enabling acceleration of a report.
20 . The one or more non-transitory computer readable media of claim 12 , wherein the steps further comprise:
receiving a search request;
generating search results based at least in part on the first summary data for the first bucket; and
providing the search results in response to the search request.
21 . A system, comprising:
one or more memories storing instructions; and
one or more processors for executing the instructions to:
generating a set of time-stamped event records from raw machine data;
indexing and storing the set of time-stamped event records in two or more buckets of event records;
scheduling a summary request for a first bucket of the two or more buckets to be executed by a plurality of indexers;
locking a first summary directory associated with the first bucket;
writing first summary data for the first bucket to the first summary directory, wherein the first summary data comprises at least one summary file that includes a summary representing a statistic associated with data in the first bucket and an indication that an entirety of contents of the first bucket is summarized based on determining a completion of the set of time-stamped event records being stored to the first bucket and summarization of the set of time-stamped event records stored in the first bucket;
unlocking the first summary directory; and
responsive to the indication that the entirety of contents of the first bucket is summarized, writing second summary data for a second bucket to a second summary directory.
22 . The system of claim 21 , wherein the one or more processors execute the instructions to:
receive, from an entity, a request for the second summary data; and
provide, to the entity, the second summary data for the second bucket responsive to the request, wherein the entity is configured to generate a result based at least in part on contents of the second summary data for the second bucket.