IP Library › Granted Patent US 12,468,673
Granted Patent B2
US 12,468,673 · App. 18/499,762 · Granted Nov 11, 2025

Schema evolution support in hybrid transactional/analytical processing (HTAP) workloads

Inventors: Cristian Diaconu (Kirkland, WA); Chen Luo (San Mateo, CA); Corbin McElhanney (San Mateo, CA); Wumengjian Zhu (Cupertino, CA)
Assignee: Snowflake Inc.
G06F16/213G06F16/24552G06F16/24573
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,673
App. No.
18/499,762
Granted
Nov 11, 2025
Kind
B2
Abstract

The subject technology receives a request to perform a table scan operation of a table. The subject technology determines that the table is being accessed for an initial time. The subject technology populates a columnar cache with data of the table provided by the table scan operation. The subject technology determines a set of schema versions of a set of rows from the data of the table. The subject technology determines schema information of each schema from the set of schema versions. The subject technology generates a result rowset and a second rowset comprising a union of columns that have appeared at least once in each row. The subject technology performs deserialization of rows from the result rowset and the second rowset. The subject technology provides the rows from the result rowset and the second rowset to write to a file in a particular format.

Claims (86)

1 . A system comprising:

at least one hardware processor; and

a memory storing instructions that cause the at least one hardware processor to perform operations comprising:

receiving a request to perform a table scan operation of a table;

determining that the table is being accessed for an initial time;

populating a columnar cache with data of the table provided by the table scan operation;

determining a set of schema versions of a set of rows from the data of the table;

determining, using at least a schema cache, schema information of each schema from the set of schema versions, the schema cache being provided by an execution node;

generating a result rowset and a second rowset comprising a union of columns that have appeared at least once in each row;

performing deserialization of rows from the result rowset and the second rowset;

providing the rows from the result rowset and the second rowset to write to a file in a particular format for storing in the columnar cache;

generating a particular schema version based on a union of columns that have appeared at least once in a set of rows, the particular schema version being utilized as part of a write operation, the write operation comprising writing data from at least one blob file to a particular file in a different format for storing in the columnar cache; and

storing the particular schema version in the schema cache provided by the execution node.

2 . The system of claim 1 , wherein the operations further comprise:

receiving a query including a statement to perform a read operation on a particular table;

determining a set of column ordinals referenced by the query; and

performing a scan of the columnar cache based on the set of column ordinals to determine at least one requested column.

3 . The system of claim 2 , wherein the operations further comprise:

determining that the columnar cache does not store the at least one requested column;

filling the at least one requested column with a current default value; and

providing the at least one requested column with the current default value in a particular result rowset from executing the query.

4 . The system of claim 2 , wherein the operations further comprise:

determining that the at least one requested column is defined as not storing a null value;

determining that the columnar cache does not store the at least one requested column; and

providing an indication of an error as a result of the query.

5 . The system of claim 2 , wherein the operations further comprise:

determining that the at least one requested column is defined as not storing a null value;

determining that the columnar cache includes a null value stored in the at least one requested column; and

providing an indication of an error as a result of the query.

6 . The system of claim 1 , wherein the operations further comprise:

receiving a particular query, the particular query including a query range for processing the query;

sending a request to a key-value store for blob metadata and a set of recent writes for the query range;

receiving the blob metadata, the blob metadata including information related to at least one blob file; and

determining that the at least one blob file includes a set of rows with multiple schema versions.

7 . The system of claim 1 , wherein the particular format comprises a Parquet file.

8 . The system of claim 1 , wherein the columnar cache is stored locally in an execution node, the execution node executing the table scan operation of the table, the table scan operation being performed in connection with a particular query received by the execution node for execution.

9 . The system of claim 1 , wherein the operations further comprise:

determining, by querying a metadata database, the schema information of each schema from the set of schema versions.

10 . The system of claim 1 , wherein the second rowset includes a set of dropped columns based on previous schema information of at least one previous schema version.

11 . A method comprising:

receiving a request to perform a table scan operation of a table;

determining that the table is being accessed for an initial time;

populating a columnar cache with data of the table provided by the table scan operation;

determining a set of schema versions of a set of rows from the data of the table;

determining, using at least a schema cache, schema information of each schema from the set of schema versions, the schema cache being provided by an execution node;

generating a result rowset and a second rowset comprising a union of columns that have appeared at least once in each row;

performing deserialization of rows from the result rowset and the second rowset;

providing the rows from the result rowset and the second rowset to write to a file in a particular format for storing in the columnar cache;

generating a particular schema version based on a union of columns that have appeared at least once in a set of rows, the particular schema version being utilized as part of a write operation, the write operation comprising writing data from at least one blob file to a particular file in a different format for storing in the columnar cache; and

storing the particular schema version in the schema cache provided by the execution node.

12 . The method of claim 11 , further comprising:

receiving a query including a statement to perform a read operation on a particular table;

determining a set of column ordinals referenced by the query; and

performing a scan of the columnar cache based on the set of column ordinals to determine at least one requested column.

13 . The method of claim 12 , further comprising:

determining that the columnar cache does not store the at least one requested column;

filling the at least one requested column with a current default value; and

providing the at least one requested column with the current default value in a particular result rowset from executing the query.

14 . The method of claim 12 , further comprising:

determining that the at least one requested column is defined as not storing a null value;

determining that the columnar cache does not store the at least one requested column; and

providing an indication of an error as a result of the query.

15 . The method of claim 12 , further comprising:

determining that the at least one requested column is defined as not storing a null value;

determining that the columnar cache includes a null value stored in the at least one requested column; and

providing an indication of an error as a result of the query.

16 . The method of claim 11 , further comprising:

receiving a particular query, the particular query including a query range for processing the query;

sending a request to a key-value store for blob metadata and a set of recent writes for the query range;

receiving the blob metadata, the blob metadata including information related to at least one blob file; and

determining that the at least one blob file includes a set of rows with multiple schema versions.

17 . The method of claim 11 , wherein the particular format comprises a Parquet file.

18 . The method of claim 11 , wherein the columnar cache is stored locally in an execution node, the execution node executing the table scan operation of the table, the table scan operation being performed in connection with a particular query received by the execution node for execution.

19 . The method of claim 11 , further comprising:

determining, by querying a metadata database, the schema information of each schema from the set of schema versions.

20 . A non-transitory computer-storage medium comprising instructions that, when executed by one or more processors of a machine, configure the machine to perform operations comprising:

receiving a request to perform a table scan operation of a table;

determining that the table is being accessed for an initial time;

populating a columnar cache with data of the table provided by the table scan operation;

determining a set of schema versions of a set of rows from the data of the table;

determining, using at least a schema cache, schema information of each schema from the set of schema versions, the schema cache being provided by an execution node;

generating a result rowset and a second rowset comprising a union of columns that have appeared at least once in each row;

performing deserialization of rows from the result rowset and the second rowset;

providing the rows from the result rowset and the second rowset to write to a file in a particular format for storing in the columnar cache;

generating a particular schema version based on a union of columns that have appeared at least once in a set of rows, the particular schema version being utilized as part of a write operation, the write operation comprising writing data from at least one blob file to a particular file in a different format for storing in the columnar cache; and

storing the particular schema version in the schema cache provided by the execution node.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2024
From: DIACONU, CRISTIAN; LUO, CHEN; MCELHANNEY, CORBIN; ZHU, WUMENGJIAN
To: SNOWFLAKE INC.
Reel/Frame 066031/0382 →
Continuity (2)
Continuation In Part 18455229 · Aug 24, 2023
Related Publication 20250068605A1 · Feb 27, 2025
References Cited (19)
US 8396886B1 · Tsimelzon · 2013 [cited by examiner]
US 20090187610A1 · Guo · 2009 [cited by applicant]
US 20130018903A1 · Taranov · 2013 [cited by applicant]
US 20150350316A1 · Calder et al. · 2015 [cited by applicant]
US 20210286806A1 · Ahmadi · 2021 [cited by examiner]
US 20220318223A1 · Ahluwalia et al. · 2022 [cited by applicant]
US 20220382758A1 · Schreter · 2022 [cited by applicant]
US 20230141891A1 · Huang · 2023 [cited by examiner]
US 20230141902A1 · Ma · 2023 [cited by examiner]
US 20230259521A1 · Haelen · 2023 [cited by examiner]
US 20230336592A1 · Narayanaswamy et al. · 2023 [cited by applicant]
US 20240111743A1 · Lewis · 2024 [cited by examiner]
US 20240168929A1 · Sigoure · 2024 [cited by examiner]
US 20240378186A1 · Paulraj · 2024 [cited by examiner]
“U.S. Appl. No. 18/455,229, Non Final Office Action mailed Nov. 16, 2023”, 30 pages. [cited by applicant]
“U.S. Appl. No. 18/455,229, Response filed Jan. 30, 2024 to Non Final Office Action mailed Nov. 16, 2023”, 12 pgs. [cited by applicant]
“U.S. Appl. No. 18/455,229, Final Office Action mailed Mar. 5, 2024”, 29 pgs. [cited by applicant]
“U.S. Appl. No. 18/455,229, Response filed May 6, 2024 to Final Office Action mailed Mar. 5, 2024”, 12 pgs. [cited by applicant]
“U.S. Appl. No. 18/455,229, Notice of Allowance mailed Jun. 3, 2024”, 8 pgs. [cited by applicant]