Systems and methods for processing timeseries data
A system is provided for storing, in a database, a plurality of timeseries represented by a plurality of respective documents events in a columnar format. The system further is adapted compress at least one of the values within the plurality of documents. According to some embodiments, the system stores the compressed values as a Simple-8b block and calculates the optimal Simple-8b selector. According to some embodiments, the system is adapted to determine a secondary index based on values within the bucket.
1 . A system comprising:
a database engine configured to store, in a database, a plurality of timeseries events as a plurality of documents within a bucket, the database engine being further configured to:
store, in a columnar format, the plurality of timeseries events represented by the plurality of respective documents;
determine an index associated with the plurality of respective documents within the bucket;
determine a geographically-based index relating to the stored plurality of timeseries events represented by the plurality of documents, wherein the database engine is configured to determine the geographically-based index by:
determining a first geographically-based index comprising information associated with at least a subset of the plurality of timeseries events; and
transforming the first geographically-based index to a second geographically-based index when the information of the first geographically-based index associated with at least the subset is not sufficient to directly index at least the subset of the plurality of timeseries events, by transforming a definition of the first geographically-based index to define the second geographically-based index to indicate a measurement field of the timeseries events to be indexed as a column of points;
generate the determined geographically-based index to index at least the subset of the plurality of timeseries events; and
store the generated geographically-based index in the database.
2 . The system according to claim 1 , wherein the geographically-based index comprises an index determined based on metadata associated with the bucket.
3 . The system according to claim 1 , wherein the geographically-based index comprises an index determined based on timeseries event data.
4 . The system according to claim 1 , wherein the database engine is further configured to sort timeseries events by a distance from a query point.
5 . The system according to claim 1 , wherein the database engine is further configured to filter documents based on a distance from a query point.
6 . The system according to claim 1 , wherein the database engine is further configured to add field data specifying a distance from a query point.
7 . The system according to claim 1 , wherein the database engine is further configured to selectively index documents that fall within a specified boundary.
8 . The system according to claim 1 , wherein the database engine is further configured to process a query using the determined index.
9 . The system according to claim 1 , wherein the index is associated with the plurality of timeseries events based on time values.
10 . The system according to claim 1 , wherein the database engine being further configured to determine a secondary index, wherein the secondary index is an ascending or descending index based on data values associated with a measurement field represented by the plurality of documents.
11 . The system according to claim 1 , wherein the database engine being further configured to determine a secondary index, wherein the secondary index is a compound index based on data values associated with a plurality of measurement fields represented by the plurality of documents.
12 . The system according to claim 1 , wherein the database engine being further configured to determine an index relating to a portion of the stored plurality of timeseries events that meet a specified filter expression.
13 . The system according to claim 1 , wherein the database engine being further configured to determine an index relating to a portion of the stored plurality of timeseries events that meet a specified filter expression.
14 . A system comprising:
storing, by a database engine in a database, a plurality of timeseries events as a plurality of documents within a bucket, the database engine being further configured to:
store, in a columnar format, the plurality of timeseries events represented by the plurality of respective documents;
determine an index associated with the plurality of respective documents within the bucket;
determine a geographically-based index relating to the stored plurality of timeseries events represented by the plurality of documents, wherein the database engine is configured to determine the geographically-based index by:
determining a first geographically-based index comprising information associated with at least a subset of the plurality of timeseries events; and
transforming the first geographically-based index to a second geographically-based index when the information of the first geographically-based index associated with at least the subset is not sufficient to directly index at least the subset of the plurality of timeseries events, by transforming a definition of the first geographically-based index to define the second geographically-based index to indicate a measurement field of the timeseries events to be indexed as a column of points;
generate the determined geographically-based index to index at least the subset of the plurality of timeseries events; and
store the generated geographically-based index in the database.
15 . The method according to claim 14 , wherein the database engine is being further configured to determine a secondary index, wherein the secondary index is at least one of:
an ascending or descending index based on data values associated with a measurement field represented by the plurality of documents,
a compound index based on data values associated with a plurality of measurement fields represented by the plurality of documents, or
an index relating to a portion of the stored plurality of timeseries events that meet a specified filter expression.
16 . The method according to claim 15 , wherein the index is associated with the plurality of timeseries events based on time values.
17 . A non-transitory computer-readable medium containing instructions that, when executed, cause at least one computer hardware processor to perform:
storing, by a database engine in a database, a plurality of timeseries events as a plurality of documents within a bucket, the database engine being further configured to:
store, in a columnar format, the plurality of timeseries events represented by the plurality of respective documents; and
determine an index associated with the plurality of respective documents within the bucket;
determine a geographically-based index relating to the stored plurality of timeseries events represented by the plurality of documents, wherein the database engine is configured to determine the geographically-based index by:
determining a first geographically-based index comprising information associated with at least a subset of the plurality of timeseries events; and
transforming the first geographically-based index to a second geographically-based index when the information of the first geographically-based index associated with at least the subset is not sufficient to directly index at least the subset of the plurality of timeseries events, by transforming a definition of the first geographically-based index to define the second geographically-based index to indicate a measurement field of the timeseries events to be indexed as a column of points;
generate the determined geographically-based index to index at least the subset of the plurality of timeseries events; and
store the generated geographically-based index in the database.
18 . The non-transitory computer-readable medium according to claim 17 , wherein the database engine is being further configured to determine a secondary index, wherein the secondary index is at least one of:
an ascending or descending index based on data values associated with a measurement field represented by the plurality of documents,
a compound index based on data values associated with a plurality of measurement fields represented by the plurality of documents, or
an index relating to a portion of the stored plurality of timeseries events that meet a specified filter expression.
19 . The non-transitory computer-readable medium according to claim 18 , wherein the index is associated with the plurality of timeseries events based on time values.