IP Library Granted Patent US 12,657,175
Granted Patent B2
US 12,657,175 · App. 18/358,238 · Granted Jun 16, 2026

Systems and methods for processing timeseries data

Inventors: Geert Bosch (Brooklyn, NY); Henrik Edin (Portsmouth, NH); Pawel Terlecki (Miami, FL); David Percy (New York, NY); Daniel Larkin-York (Saint Petersburg, FL)
Assignee: MongoDB, Inc.
G06F16/221G06F16/24568G06F16/2477
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,175
App. No.
18/358,238
Granted
Jun 16, 2026
Kind
B2
Abstract

A system is provided for storing, in a database, a plurality of timeseries represented by a plurality of respective documents events in a columnar format. The system further is adapted compress at least one of the values within the plurality of documents. According to some embodiments, the system stores the compressed values as a Simple-8b block and calculates the optimal Simple-8b selector. According to some embodiments, the system is adapted to determine a secondary index based on values within the bucket.

Claims (49)

1 . A system comprising:

a database engine configured to store, in a database, a plurality of timeseries events as a plurality of documents within a bucket, the database engine being further configured to:

store, in a columnar format, the plurality of timeseries events represented by the plurality of respective documents;

determine an index associated with the plurality of respective documents within the bucket;

determine a geographically-based index relating to the stored plurality of timeseries events represented by the plurality of documents, wherein the database engine is configured to determine the geographically-based index by:

determining a first geographically-based index comprising information associated with at least a subset of the plurality of timeseries events; and

transforming the first geographically-based index to a second geographically-based index when the information of the first geographically-based index associated with at least the subset is not sufficient to directly index at least the subset of the plurality of timeseries events, by transforming a definition of the first geographically-based index to define the second geographically-based index to indicate a measurement field of the timeseries events to be indexed as a column of points;

generate the determined geographically-based index to index at least the subset of the plurality of timeseries events; and

store the generated geographically-based index in the database.

2 . The system according to claim 1 , wherein the geographically-based index comprises an index determined based on metadata associated with the bucket.

3 . The system according to claim 1 , wherein the geographically-based index comprises an index determined based on timeseries event data.

4 . The system according to claim 1 , wherein the database engine is further configured to sort timeseries events by a distance from a query point.

5 . The system according to claim 1 , wherein the database engine is further configured to filter documents based on a distance from a query point.

6 . The system according to claim 1 , wherein the database engine is further configured to add field data specifying a distance from a query point.

7 . The system according to claim 1 , wherein the database engine is further configured to selectively index documents that fall within a specified boundary.

8 . The system according to claim 1 , wherein the database engine is further configured to process a query using the determined index.

9 . The system according to claim 1 , wherein the index is associated with the plurality of timeseries events based on time values.

10 . The system according to claim 1 , wherein the database engine being further configured to determine a secondary index, wherein the secondary index is an ascending or descending index based on data values associated with a measurement field represented by the plurality of documents.

11 . The system according to claim 1 , wherein the database engine being further configured to determine a secondary index, wherein the secondary index is a compound index based on data values associated with a plurality of measurement fields represented by the plurality of documents.

12 . The system according to claim 1 , wherein the database engine being further configured to determine an index relating to a portion of the stored plurality of timeseries events that meet a specified filter expression.

13 . The system according to claim 1 , wherein the database engine being further configured to determine an index relating to a portion of the stored plurality of timeseries events that meet a specified filter expression.

14 . A system comprising:

storing, by a database engine in a database, a plurality of timeseries events as a plurality of documents within a bucket, the database engine being further configured to:

store, in a columnar format, the plurality of timeseries events represented by the plurality of respective documents;

determine an index associated with the plurality of respective documents within the bucket;

determine a geographically-based index relating to the stored plurality of timeseries events represented by the plurality of documents, wherein the database engine is configured to determine the geographically-based index by:

determining a first geographically-based index comprising information associated with at least a subset of the plurality of timeseries events; and

transforming the first geographically-based index to a second geographically-based index when the information of the first geographically-based index associated with at least the subset is not sufficient to directly index at least the subset of the plurality of timeseries events, by transforming a definition of the first geographically-based index to define the second geographically-based index to indicate a measurement field of the timeseries events to be indexed as a column of points;

generate the determined geographically-based index to index at least the subset of the plurality of timeseries events; and

store the generated geographically-based index in the database.

15 . The method according to claim 14 , wherein the database engine is being further configured to determine a secondary index, wherein the secondary index is at least one of:

an ascending or descending index based on data values associated with a measurement field represented by the plurality of documents,

a compound index based on data values associated with a plurality of measurement fields represented by the plurality of documents, or

an index relating to a portion of the stored plurality of timeseries events that meet a specified filter expression.

16 . The method according to claim 15 , wherein the index is associated with the plurality of timeseries events based on time values.

17 . A non-transitory computer-readable medium containing instructions that, when executed, cause at least one computer hardware processor to perform:

storing, by a database engine in a database, a plurality of timeseries events as a plurality of documents within a bucket, the database engine being further configured to:

store, in a columnar format, the plurality of timeseries events represented by the plurality of respective documents; and

determine an index associated with the plurality of respective documents within the bucket;

determine a geographically-based index relating to the stored plurality of timeseries events represented by the plurality of documents, wherein the database engine is configured to determine the geographically-based index by:

determining a first geographically-based index comprising information associated with at least a subset of the plurality of timeseries events; and

transforming the first geographically-based index to a second geographically-based index when the information of the first geographically-based index associated with at least the subset is not sufficient to directly index at least the subset of the plurality of timeseries events, by transforming a definition of the first geographically-based index to define the second geographically-based index to indicate a measurement field of the timeseries events to be indexed as a column of points;

generate the determined geographically-based index to index at least the subset of the plurality of timeseries events; and

store the generated geographically-based index in the database.

18 . The non-transitory computer-readable medium according to claim 17 , wherein the database engine is being further configured to determine a secondary index, wherein the secondary index is at least one of:

an ascending or descending index based on data values associated with a measurement field represented by the plurality of documents,

a compound index based on data values associated with a plurality of measurement fields represented by the plurality of documents, or

an index relating to a portion of the stored plurality of timeseries events that meet a specified filter expression.

19 . The non-transitory computer-readable medium according to claim 18 , wherein the index is associated with the plurality of timeseries events based on time values.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2024
From: BOSCH, GEERT; EDIN, HENRIK; TERLECKI, PAWEL; PERCY, DAVID; LARKIN-YORK, DANIEL
To: MONGODB, INC.
Reel/Frame 067042/0901 →
Continuity (4)
Continuation In Part 17858950 · Jul 6, 2022
Provisional Application 63392457 · Jul 26, 2022
Provisional Application 63220332 · Jul 9, 2021
Related Publication 20230367752A1 · Nov 16, 2023
References Cited (24)
US 11010223B2 · Dasgupta et al. · 2021 [cited by applicant]
US 12038926B1 · Pathak · 2024 [cited by examiner]
US 20070255758A1 · Zheng et al. · 2007 [cited by applicant]
US 20120109985A1 · Chandrasekaran · 2012 [cited by applicant]
US 20170262517A1 · Horowitz · 2017 [cited by examiner]
US 20180232459A1 · Park et al. · 2018 [cited by applicant]
US 20180300381A1 · Horowitz et al. · 2018 [cited by applicant]
US 20190087696A1 · Verhoeven et al. · 2019 [cited by applicant]
US 20200372004A1 · Barber · 2020 [cited by examiner]
US 20200387509A1 · Florendo · 2020 [cited by applicant]
US 20210034598A1 · Arye · 2021 [cited by applicant]
US 20210406528A1 · Ramani et al. · 2021 [cited by applicant]
US 20220067980A1 · Kletter · 2022 [cited by applicant]
US 20230037619A1 · Terlecki et al. · 2023 [cited by applicant]
US 20230040530A1 · Terlecki et al. · 2023 [cited by applicant]
US 20230041129A1 · Terlecki et al. · 2023 [cited by applicant]
US 20230367781A1 · Bosch et al. · 2023 [cited by applicant]
US 20230367801A1 · Bosch et al. · 2023 [cited by applicant]
CN 111626623A · 2020 [cited by examiner]
CN 112307177A · 2021 [cited by examiner]
CN 113518081A · 2021 [cited by applicant]
Jeong et al. A data management infrastructure for bridge monitoring. Proceedings vol. 9435, Sensors and Smart Structures Technologies for Civil, Mechanical and Aerospace Systems 2015; 94350P (2015). pp. 1-16 (Year: 2015… [cited by applicant]
MongoDB Manual, https://www.mongodb.com/docs/v4.4 pp. 1-122 (Year: 2020). [cited by applicant]
Qi M. Digital Forensics and NoSQL Databases, 2014 11th International Conference on Fuzzy Systems and Knowledge Discovery pp. 1-6 (Year: 2014). [cited by applicant]