IP Library › Granted Patent US 11,727,001
Granted Patent B2
US 11,727,001 · App. 17/497,663 · Granted Aug 15, 2023

Optimized data structures of a relational cache with a learning capability for accelerating query execution by a data system

Inventors: Tomer Shiran (Mountain View, CA); Jacques Nadeau (Santa Clara, CA); Steven Michael Phillips (Mountain View, CA)
Assignee: DREMIO CORPORATION
G06F16/24542G06F16/2228G06F16/24537G06F16/24549G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,727,001
App. No.
17/497,663
Granted
Aug 15, 2023
Kind
B2
Abstract

A method performed by a data system includes automatically learning relationship(s) among datasets based on one or more of a user query or an observation of a data flow through the data system. The method further includes generating an optimized data structure based on the learned relationships among the datasets. The data system then modifies a query plan to obtain query results that satisfy a query by reading the optimized data structure in lieu of reading the datasets.

Claims (75)

1. A computer-readable storage medium, excluding transitory signals and carrying instructions, which, when executed by at least one data processor of a data system, cause the data system to:

automatically determine relationships among a plurality of datasets based on a plurality of virtual datasets,

wherein the plurality of virtual datasets each include a query, a reference to a dataset of the plurality of datasets, and is created by a user of the data system;

generate an optimized data structure based on the relationships among the plurality of datasets,

wherein the optimized data structure is stored in a non-volatile memory of the data system;

receive a first query for a first result that satisfies the first query based on the plurality of datasets;

modify a first query plan for the first query to read the optimized data structure stored in the non-volatile memory in lieu of reading the plurality of datasets;

return the first result that satisfies the first query, wherein the first result is obtained based on the modified first query plan;

receive a second query for a second result that satisfies a second query different from the first query;

modify a second query plan for the second query to read the optimized data structure stored in the non-volatile memory; and

return the second result that satisfies the second query, wherein the second result is obtained based on the modified second query plan.

2. The computer-readable storage medium of claim 1 , wherein to automatically determine relationships among the plurality of datasets comprises causing the data system to:

determine a relationship between two datasets based on a join between the two datasets as indicated in a query of a virtual dataset.

3. The computer-readable storage medium of claim 1 , wherein to automatically determine relationships among the plurality of datasets comprises causing the data system to:

determine a relationship between two datasets based on a stated relationship included in a virtual dataset.

4. The computer-readable storage medium of claim 1 , wherein the data system is further caused to, prior to the query plan being modified:

autonomously decide to generate the optimized data structure,

wherein the decision to generate the optimized data structure is based on a determination that reading the optimized data structure in lieu of reading at least some dataset from a plurality of data sources improves processing of an expected workload, and

wherein the expected workload is based on a quantity or frequency of queries included in the plurality of virtual datasets.

5. A computer-readable storage medium, excluding transitory signals and carrying instructions, which, when executed by at least one data processor of a data system, cause the data system to:

automatically determine relationships among a plurality of datasets based on queries processed by the data system or an observation of a data flow of the data system;

generate an optimized data structure based on the relationships among the plurality of datasets,

wherein the optimized data structure is stored in a non-volatile memory of the data system;

receive a first query for a first result that satisfies the first query based on the plurality of datasets;

modify a first query plan of the first query to read the optimized data structure stored in the non-volatile memory in lieu of reading the plurality of datasets;

return the first result that satisfies the first query, wherein the first result is obtained based on the modified first query plan;

receive a second query for a second result that satisfies a second query different from the first query;

modify a second query plan for the second query to read the optimized data structure stored in the non-volatile memory; and

return the second result that satisfies the second query, wherein the second result is obtained based on the modified second query plan.

6. The computer-readable storage medium of claim 5 , wherein to automatically determine the relationships among the plurality of datasets comprises causing the data system to:

determine the relationships among the plurality of datasets based on a pattern of the queries received by the data system,

wherein the optimized data structure is generated based on the pattern of queries.

7. The computer-readable storage medium of claim 5 , wherein to automatically determine the relationships among the plurality of datasets comprises causing the data system to:

determine the relationships among the plurality of datasets based on the data flow of the data system including a runtime statistic,

wherein the optimized data structure is generated based on the runtime statistic.

8. The computer-readable storage medium of claim 5 , wherein to automatically determine the relationships among the plurality of datasets comprises causing the data system to:

determine the relationships among the plurality of datasets based on the data flow of the data system including a workload distribution,

wherein the optimized data structure is generated based on the workload distribution.

9. The computer-readable storage medium of claim 5 , wherein the optimized data structure is generated based on an aggregation of the plurality of datasets.

10. The computer-readable storage medium of claim 5 , wherein the optimized data structure preserves a quantity of records equal to a quantity of records of one of the plurality of datasets.

11. The computer-readable storage medium of claim 5 , wherein the plurality of datasets includes a fact table and a plurality of dimension tables, and wherein generating the optimized data structure comprises causing the data system to:

preserve a quantity of records of the fact table even though any of the plurality of dimension tables have a quantity of records different from the fact table.

12. The computer-readable storage medium of claim 11 , wherein to automatically determine the relationships among the plurality of datasets comprises causing the data system to:

autonomously determining a fact-dimension relationship among the plurality of datasets.

13. The computer-readable storage medium of claim 5 , wherein the first query plan is associated with a different combination of datasets compared to the second query plan.

14. The computer-readable storage medium of claim 5 , wherein the data system is further caused to, prior to automatically determining the relationships among the plurality of datasets:

processing the plurality of datasets in response to a plurality of queries.

15. The computer-readable storage medium of claim 5 , wherein the first query refers to at least one of the plurality of datasets that are contained in a plurality of data sources, the data system being further caused to:

obtain the result without reading any of the plurality of datasets from the plurality of data sources.

16. The computer-readable storage medium of claim 5 , wherein the first query refers to at least one of the plurality of datasets that are contained in a plurality of data sources, the data system being further caused to:

obtain the result by reading at least some of the plurality of datasets from the plurality of data sources in addition to reading the optimized data structure.

17. The computer-readable storage medium of claim 5 , wherein the data system is caused to, prior to the query plan being modified:

determine that the first query can be accelerated based on the optimized data structure stored in a cache of the data system.

18. The computer-readable storage medium of claim 5 , wherein to automatically determine the relationships among the plurality of datasets and to generate the optimized data structure occurs at runtime of one or more query executions.

19. The computer-readable storage medium of claim 5 , wherein the plurality of datasets corresponds to a plurality of physical datasets, and wherein to modify the query plan comprises causing the data system to:

expand one or more views of the plurality of physical datasets or of other views so that the query plan refers only to the plurality of physical datasets.

20. The computer-readable storage medium of claim 5 , wherein the data system is further caused to, prior to the query plan being modified:

autonomously decide to generate the optimized data structure,

wherein the decision to generate the optimized data structure is based on a determination that reading the optimized data structure in lieu of reading at least some dataset from a plurality of data sources improves processing of an expected workload.

21. A method performed by a data system comprising:

automatically determining relationships among a plurality of datasets based on queries processed by the data system or an observation of a data flow of the data system;

generating an optimized data structure based on the relationships among the plurality of datasets,

wherein the optimized data structure is stored in a non-volatile memory of the data system;

receiving a first query for a first result that satisfies the first query based on the plurality of datasets;

modifying a first query plan of the first query to read the optimized data structure stored in the non-volatile memory in lieu of reading the plurality of datasets; and

returning the first result that satisfies the first query, wherein the first result is obtained based on the modified first query plan;

receiving a second query for a second result that satisfies a second query different from the first query;

modifying a second query plan for the second query to read the optimized data structure stored in the non-volatile memory; and

returning the second result that satisfies the second query, wherein the second result is obtained based on the modified second query plan.

22. The method of claim 21 , wherein automatically determining the relationships among the plurality of datasets comprises:

determining the relationships among the plurality of datasets based on a runtime statistic indicated in the data flow of the data system,

wherein the optimized data structure is generated based on the runtime statistic.

23. The method of claim 21 , wherein automatically determining the relationships among the plurality of datasets comprises:

determining the relationships among the plurality of datasets based on a workload distribution indicated in the data flow of the data system,

wherein the optimized data structure is generated based on the workload distribution.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2021
From: SHIRAN, TOMER; NADEAU, JACQUES; PHILLIPS, STEVEN MICHAEL
To: DREMIO CORPORATION
Reel/Frame 057744/0653 →
Continuity (3)
Continuation 16392483 · Apr 23, 2019
Provisional Application 62662015 · Apr 24, 2018
Related Publication 20220100762A1 · Mar 31, 2022