IP Library › Granted Patent US 11,822,582
Granted Patent B2
US 11,822,582 · App. 17/896,446 · Granted Nov 21, 2023

Metadata clustering

Inventors: Yi Fang (Kirkland, WA); Varun Ganesh (San Bruno, CA); Xinglian Liu (Redmond, WA); Ryan Michael Thomas Shelly (San Francisco, CA); Jiaqi Yan (Menlo Park, CA); Yizhi Zhu (Bellevue, WA)
Assignee: Snowflake Inc.
G06F16/285
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,822,582
App. No.
17/896,446
Granted
Nov 21, 2023
Kind
B2
Abstract

Embodiments of the present disclosure describe systems, methods, and computer program products for improving query processing of a database. An example method can include: storing table data for a table in a plurality of micro-partitions, each micro-partition comprising a portion of the table data for the table; for each micro-partition of the plurality of micro-partitions, storing metadata for the micro-partition in at least one of a plurality of expression properties; and selecting, by a processing device, a subset of the plurality of expression properties to be grouped into a grouping expression property based at least partially on the metadata of the subset of the plurality of the expression properties. The grouping expression property may include cumulative metadata associated with the metadata of the subset of the plurality of expression properties.

Claims (56)

1. A method comprising:

storing table data for a table in a plurality of micro-partitions of a storage device, each micro-partition comprising a portion of the table data for the table, wherein a first data range of a first portion of the table data of a first micro-partition of the plurality of micro-partitions overlaps a second data range of a second portion of the table data of a second micro-partition of the plurality of micro-partitions;

for each micro-partition of the plurality of micro-partitions, storing metadata for the micro-partition in at least one of a plurality of expression properties; and

selecting, by a processing device, a subset of the plurality of expression properties to be grouped into a grouping expression property based at least partially on the metadata of the subset of the plurality of the expression properties,

wherein each expression property of the subset of the plurality of expression properties is associated with a different micro-partition of the plurality of micro-partitions,

wherein the grouping expression property comprises cumulative metadata associated with the metadata of the micro-partitions associated with the subset of the plurality of expression properties and is stored in immutable storage of a storage device separately from the plurality of micro-partitions storing the table data,

wherein the selecting the subset of the plurality of expression properties to be grouped into the grouping expression property is at least partially based on a portion of the metadata of the subset of the plurality of the expression properties that is associated with a clustering key of the micro-partitions, and

wherein the clustering key comprises at least one of a subset of columns of the table or an expression on the table that identifies portions of the table which are to be used for making decisions about co-locating portions of the table data within a same one of the plurality of micro-partitions.

2. The method of claim 1 , wherein the selecting the subset of the plurality of expression properties is further at least partially based on a time of creation of a micro-partition of the plurality of micro-partitions whose data is stored in the subset of the plurality of the expression properties.

3. The method of claim 1 , further comprising:

calculating an average depth value and an average overlap value for the subset of the plurality of expression properties of the grouping expression property;

calculating a grouping level depth value and a grouping level overlap value for the grouping expression property; and

reorganizing the subset of the plurality of expression properties of the grouping expression property responsive to a comparison of the average depth value, the average overlap value, the grouping level depth value, and the grouping level overlap value to respective thresholds.

4. The method of claim 1 , wherein the selecting the subset of the plurality of expression properties is at least partially based on determining whether a range of the metadata of the plurality of expression properties overlaps with a range of the cumulative metadata of the grouping expression property.

5. The method of claim 1 , further comprising:

creating a new micro-partition to be included in the plurality of micro-partitions of a storage device, the new micro-partition comprising data values of the table data;

creating a new expression property to be included in the plurality of expression properties and associated with the new micro-partition, the new expression property comprising metadata describing the data values of the new micro-partition; and

including the new expression property in the subset of the plurality of expression properties to be grouped into the grouping expression property at least partially based on the metadata of the new expression property.

6. A system comprising:

a memory; and

a processing device, operatively coupled to the memory, to:

store table data for a table in a plurality of micro-partitions of a storage device, each micro-partition comprising a portion of the table data for the table, wherein a first data range of a first portion of the table data of a first micro-partition of the plurality of micro-partitions overlaps a second data range of a second portion of the table data of a second micro-partition of the plurality of micro-partitions;

for each micro-partition of the plurality of micro-partitions, store metadata for the micro-partition in at least one of a plurality of expression properties; and

select a subset of the plurality of expression properties to be grouped into a grouping expression property based at least partially on the metadata of the subset of the plurality of the expression properties,

wherein each expression property of the subset of the plurality of expression properties is associated with a different micro-partition of the plurality of micro-partitions,

wherein the grouping expression property comprises cumulative metadata associated with the metadata of the micro-partitions associated with the subset of the plurality of expression properties and is stored in immutable storage of a storage device separately from the plurality of micro-partitions storing the table data,

wherein the processing device is to select the subset of the plurality of expression properties to be grouped into the grouping expression property at least partially based on a portion of the metadata of the subset of the plurality of the expression properties that is associated with a clustering key of the micro-partitions, and

wherein the clustering key comprises at least one of a subset of columns of the table or an expression on the table that identifies portions of the table which are to be used for making decisions about co-locating portions of the table data within a same one of the plurality of micro-partitions.

7. The system of claim 6 , wherein the processing device is further to select the subset of the plurality of expression properties at least partially based on a time of creation of a micro-partition of the plurality of micro-partitions whose data is stored in the subset of the plurality of the expression properties.

8. The system of claim 6 , wherein the processing device is further to:

calculate an average depth value and an average overlap value for the subset of the plurality of expression properties of the grouping expression property;

calculate a grouping level depth value and a grouping level overlap value for the grouping expression property; and

reorganize the subset of the plurality of expression properties of the grouping expression property responsive to a comparison of the average depth value, the average overlap value, the grouping level depth value, and the grouping level overlap value to respective thresholds.

9. The system of claim 6 , wherein the processing device is further to select the subset of the plurality of expression properties at least partially based on determining whether a range of the metadata of the plurality of expression properties overlaps with a range of the cumulative metadata of the grouping expression property.

10. The system of claim 6 , wherein the processing device is further to:

create a new micro-partition to be included in the plurality of micro-partitions of a storage device, the new micro-partition comprising data values of the table data;

create a new expression property to be included in the plurality of expression properties and associated with the new micro-partition, the new expression property comprising metadata describing the data values of the new micro-partition; and

include the new expression property in the subset of the plurality of expression properties to be grouped into the grouping expression property at least partially based on the metadata of the new expression property.

11. A non-transitory computer-readable storage medium including instructions that, when executed by a processing device, cause the processing device to:

store table data for a table in a plurality of micro-partitions of a storage device, each micro-partition comprising a portion of the table data for the table, wherein a first data range of a first portion of the table data of a first micro-partition of the plurality of micro-partitions overlaps a second data range of a second portion of the table data of a second micro-partition of the plurality of micro-partitions;

for each micro-partition of the plurality of micro-partitions, store metadata for the micro-partition in at least one of a plurality of expression properties; and

select, by the processing device, a subset of the plurality of expression properties to be grouped into a grouping expression property based at least partially on the metadata of the subset of the plurality of the expression properties,

wherein each expression property of the subset of the plurality of expression properties is associated with a different micro-partition of the plurality of micro-partitions,

wherein the grouping expression property comprises cumulative metadata associated with the metadata of the micro-partitions associated with the subset of the plurality of expression properties and is stored in immutable storage of a storage device separately from the plurality of micro-partitions storing the table data,

wherein the processing device is to select the subset of the plurality of expression properties to be grouped into the grouping expression property at least partially based on a portion of the metadata of the subset of the plurality of the expression properties that is associated with a clustering key of the micro-partitions, and

wherein the clustering key comprises at least one of a subset of columns of the table or an expression on the table that identifies portions of the table which are to be used for making decisions about co-locating portions of the table data within a same one of the plurality of micro-partitions.

12. The non-transitory computer-readable storage medium of claim 11 , wherein the processing device is further to select the subset of the plurality of expression properties at least partially based on a time of creation of a micro-partition of the plurality of micro-partitions whose data is stored in the subset of the plurality of the expression properties.

13. The non-transitory computer-readable storage medium of claim 11 , wherein the processing device is further to:

calculate an average depth value and an average overlap value for the subset of the plurality of expression properties of the grouping expression property;

calculate a grouping level depth value and a grouping level overlap value for the grouping expression property; and

reorganize the subset of the plurality of expression properties of the grouping expression property responsive to a comparison of the average depth value, the average overlap value, the grouping level depth value, and the grouping level overlap value to respective thresholds.

14. The non-transitory computer-readable storage medium of claim 11 , wherein the processing device is further to select the subset of the plurality of expression properties at least partially based on determining whether a range of the metadata of the plurality of expression properties overlaps with a range of the cumulative metadata of the grouping expression property.

15. The non-transitory computer-readable storage medium of claim 11 , wherein the processing device is further to:

create a new micro-partition to be included in the plurality of micro-partitions of a storage device, the new micro-partition comprising data values of the table data;

create a new expression property to be included in the plurality of expression properties and associated with the new micro-partition, the new expression property comprising metadata describing the data values of the new micro-partition; and

include the new expression property in the subset of the plurality of expression properties to be grouped into the grouping expression property at least partially based on the metadata of the new expression property.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2022
From: FANG, YI; GANESH, VARUN; LIU, XINGLIAN; SHELLY, RYAN MICHAEL THOMAS; YAN, JIAQI; ZHU, YIZHI
To: SNOWFLAKE INC.
Reel/Frame 060912/0411 →
Continuity (2)
Provisional Application 63301157 · Jan 20, 2022
Related Publication 20230229676A1 · Jul 20, 2023