IP Library › Granted Patent US 12,346,325
Granted Patent B2
US 12,346,325 · App. 18/317,635 · Granted Jul 1, 2025

Benchmarking JSON document stores

Inventors: Stefano Belloni (Mannheim, DE); Nils Roerup (Heidelberg, DE); Marco Patrick Schroeder (Heidelberg, DE); Daniel Ritter (Heidelberg, DE)
Assignee: SAP SE
G06F16/24549G06F16/213G06F16/24537G06F16/24542G06F16/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,346,325
App. No.
18/317,635
Granted
Jul 1, 2025
Kind
B2
Abstract

A method may include generating, based at least on a schema configuration of one or more document stores, a plurality of data in a JavaScript Object Notation (JSON) format for storage at the document stores. One or more queries may be generated to match the plurality of data stored at the document stores. The one or more queries may be distributed for execution at the document stores by a scalable quantity of concurrently operating worker nodes. One or more performance metrics for the execution of the one or more queries at the one or more document stores may be generated. Performance improvements at the document stores may be applied based on the performance metrics. Related systems and computer program products are also provided.

Claims (33)

1. A system, comprising:

at least one data processor; and

at least one memory storing instructions which, when executed by the at least one data processor, cause operations comprising:

generating, based at least on a schema configuration of one or more document stores, a plurality of data in a JavaScript Object Notation (JSON) format for storage at the one or more document stores;

generating, based at least on the schema configuration, a query configuration, and a feature matrix of the one or more document stores, one or more queries to match the plurality of data stored at the one or more document stores, wherein the query configuration includes a first object describing an element used in a project clause of a query, and wherein a function selected randomly from the feature matrix is applied to the projected element;

distributing the one or more queries for execution at the one or more document stores by a scalable quantity of concurrently operating worker nodes;

generating one or more performance metrics during the execution of the one or more queries at the one or more document stores; and

applying, based at least on the one or more performance metrics, one or more performance improvements during subsequent query processing at the one or more document stores, wherein the one or more performance improvements include avoiding unnecessary unnesting of JSON objects.

2. The system of claim 1 , wherein the one or more queries are distributed to a first client associated with a first document store for distribution to the scalable quantity of concurrently operating worker nodes associated with a second client.

3. The system of claim 1 , wherein the one or more queries are further distributed to a second client associated with a second document store for distribution to the scalable quantity of concurrently operating worker nodes associated with the second client.

4. The system of claim 1 , wherein the schema configuration includes a first object describing a fixed portion of a document comprising the plurality of data by at least specifying one or more elements required to appear in the document.

5. The system of claim 4 , wherein the schema configuration further includes a second object describing one or more variations in the document, and wherein the one or more variations include variations in a quantity of nested objects, a quantity of keys associated with each nested object, and/or a size of a value associated with each key.

6. The system of claim 1 , wherein the operations further comprise generating random sub-documents in order to increase a size and complexity of a single document.

7. The system of claim 6 , wherein the random sub-documents include variations in a quantity of nested objects, a quantity of keys inside each object, and a size of a value associated with each key.

8. The system of claim 1 , wherein the query configuration further includes a second object describing a where clause of the query that includes a conjunctive combination or a disjunctive combination of multiple predicates with one or more user specified or randomly selected filters.

9. The system of claim 8 , wherein the query configuration further includes a third object describing the element and a nesting depth of the element.

10. The system of claim 1 , wherein the one or more queries include at least one structured query language (SQL) query.

11. The system of claim 1 , wherein the operations further comprise:

generating a single index or a compound index for the execution of the one or more queries at the one or more document stores.

12. The system of claim 1 , wherein the plurality of data includes custom data and/or existing data associated with the one or more document stores.

13. The system of claim 1 , wherein the one or more queries include custom queries and/or existing queries associated with the one or more document stores.

14. The system of claim 1 , wherein the one or more document stores include a JSON document store and/or a relational database with a JSON extension.

15. The system of claim 1 , wherein the one or more performance metrics include a first performance metric of the one or more queries being executed at the one or more document stores by a first quantity of concurrently operating worker nodes and a second performance metric of the one or more queries being executed at the one or more document stores by a second quantity of concurrently operating worker nodes.

16. The system of claim 1 , wherein the plurality of data include documents having various nesting levels, and wherein the one or more performance metrics include a first performance metric of the one or more queries being executed on documents having a first quantity of nesting levels and a second performance metric of the one or more queries being executed on documents having a second quantity of nesting levels.

17. The system of claim 1 , wherein the one or more performance metrics include a first performance metric of the one or more queries being executed with an index, and wherein the one or more metrics further include a second performance metric of the one or more queries being executed without the index.

18. The system of claim 1 , wherein the one or more performance improvements comprise performing a pre-filter operation on a plurality of elements in a JSON array prior to performing an unnest operation on the plurality of elements in the JSON array.

19. The system of claim 1 , wherein the plurality of data and the one or more queries are further generated in accordance with one or more configurable scaling factors associated with data complexity, query complexity, data types, result sizes, and quantity of concurrent users.

20. A computer-implemented method, comprising:

generating, based at least on a schema configuration of one or more document stores, a plurality of data in a JavaScript Object Notation (JSON) format for storage at the one or more document stores;

generating, based at least on the schema configuration, a query configuration, and a feature matrix of the one or more document stores, one or more queries to match the plurality of data stored at the one or more document stores, wherein the query configuration includes a first object describing an element used in a project clause of a query, and wherein a function selected randomly from the feature matrix is applied to the projected element;

distributing the one or more queries for execution at the one or more document stores by a scalable quantity of concurrently operating worker nodes;

generating one or more performance metrics during the execution of the one or more queries at the one or more document stores; and

applying, based at least on the one or more performance metrics, one or more performance improvements during subsequent query processing at the one or more document stores, wherein the one or more performance improvements include avoiding unnecessary unnesting of JSON objects.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2023
From: BELLONI, STEFANO; ROERUP, NILS; SCHROEDER, MARCO PATRICK; RITTER, DANIEL
To: SAP SE
Reel/Frame 063646/0696 →
Continuity (2)
Provisional Application 63350321 · Jun 8, 2022
Related Publication 20230401211A1 · Dec 14, 2023
References Cited (26)
US 11416465B1 · Anwar · 2022 [cited by examiner]
US 20130124467A1 · Naidu · 2013 [cited by examiner]
US 20130166568A1 · Binkert · 2013 [cited by examiner]
US 20160034478A1 · Hernandez-Sherrington · 2016 [cited by examiner]
US 20170300517A1 · Amirsoleymani · 2017 [cited by examiner]
US 20170308555A1 · Hirzel et al. · 2017 [cited by applicant]
US 20210173621A1 · Fender et al. · 2021 [cited by applicant]
GB 2502098A · 2013 [cited by examiner]
Abiteboul. S. et al., “Research Directions for Principles of Data Management,” Dagstuhl Perspectives Workshop 16151, Dagstuhl Afanifestos 7, 1 (2018). 1-29. [cited by applicant]
Bray, T. et al., “The Javascript Object Notation (JSON) Data Interchange Format,” (2014). [cited by applicant]
Chen, Y. et al., “A Study of SQL-on-Hadoop Systems,” In BPOE LNCS, vol. 8807, Springer, 154-166. [cited by applicant]
Cole, R.L. et al., “The mixed workload CH-benCHmark,” Proceedings of the Fourth International Workshop on Testing Database Systems. 2011. [cited by applicant]
Cooper, B.F. et al., “Benchmarking Cloud Serving Systems with YCSB,” In SoCC. ACM, 143-154. [cited by applicant]
Deep, S. et al. “DIAMetrics: Benchmarking Query Engines at Scale.” Proceedings of the VLDB Endowment 13.12. [cited by applicant]
Difallah, D.E. et al., “OLTP-Bench: An Extensible Testbed for Benchmarking Relational Databases,” Proc. VLDB Endow 7, 4 (2013), 277-288. [cited by applicant]
Erling, O. et al., “The LDBC Social Network Benchmark: Interactive Workload,” In SIGMOD. ACM. 619-630. [cited by applicant]
Gray, J. [Ed.] “Database and Transaction Processing Performance Handbook,” In The Benchmark Handbook for Database and Transaction Systems (2nd Edition). Morgan Kaufmann. [cited by applicant]
Ingo, H. et al., “Automated System Performance Testing at MongoDB,” In DBTest@SIGMOD. 3:1-3:6. [cited by applicant]
Jahangiri, S. “Wisconsin Benchmark Data Generator: To JSON and Beyond,” In SIGMOD. ACM, 2887-2889. [cited by applicant]
Kamsky, A. “Adapting TPC-C Benchmark to Measure Performance of Multi-Document Transactions in MongoDB,” Proc. VLDB Endow. 12, 12 (2019). 2254-2262. [cited by applicant]
Read, A.G. “DeWitt clauses: Can we protect purchasers without hurting Microsoft,” Rev. Litig. 25 (2006). 387-421. [cited by applicant]
Rigger, M. et al., “Testing Database Engines via Pivoted Query Synthesis,” (2020), 667-682. [cited by applicant]
Ritter, D. et al., “Bench-marking integration pattern implementations,” In DEBS. ACM, 125-136. [cited by applicant]
Seltenreich, A. et al., “SQLSmith,” 2020. (Available at https://github.com/anse1/sqlsmith). [cited by applicant]
Vogelsgesang, A. et al., “Get Real: How Benchmarks Fail to Represent the Real World,” Proceedings of the Workshop on Testing Database Systems. 2018. [cited by applicant]
Zhong, R. et al., “SQUIRREL: Testing Database Management Systems with Language Validity and Coverage Feedback,” CCS'20: Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security. 2020. [cited by applicant]