IP Library › Granted Patent US 11,893,029
Granted Patent B2
US 11,893,029 · App. 18/049,325 · Granted Feb 6, 2024

Real-time streaming data ingestion into database tables

Inventors: Tyler Arthur Akidau (Seattle, WA); Istvan Cseri (Seattle, WA); Tyler Jones (Redwood City, CA); Daniel E. Sotolongo (Seattle, WA); Zhuo Zhang (Kirkland, WA)
Assignee: Snowflake Inc.
G06F16/24568G06F16/2219G06F16/2456G06F16/24544G06F16/258
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,893,029
App. No.
18/049,325
Granted
Feb 6, 2024
Kind
B2
Abstract

A streaming ingest platform can improve latency and expense issues related to uploading data into a cloud data system. The streaming ingest platform can organize the data to be ingested into per-table chunks and per-account blobs. This data may be committed and may be made available for query processing before it is ingested into the target source tables. This significantly improves latency issues. The streaming ingest platform can also accommodate uploading data from various sources with different processing and communication capabilities, such as Internet of Things (IOT) devices.

Claims (65)

1. A method comprising:

receiving a registration request for a per-account group of files, the registration request including identification information for the per-account group of files and a storage location of the per-account group of files;

accessing the per-account group of files based on the registration request, the per-account group of files being stored in a first format;

deduping and validating the per-account group of files using sequencing information included in the per-account group of files;

committing the per-account group of files stored in the storage location and making data in the per-account group of files in the first format accessible for query processing before the data is ingested into one or more source tables;

generating a hybrid table for query processing, the hybrid table including the committed data in the first format and data from the one or more source tables;

receiving a query prior to the data being ingested;

executing the query using the hybrid table to generate results for the query including:

converting the committed data from the first format into a common format;

converting the data from the one or more source tables into the common format;

joining the committed data in the common format and the data from the one or more source tables in the common format to generate joined data; and

executing the query based on the joined data; and

ingesting the data into the one or more source tables in a second format.

2. The method of claim 1 , wherein the data is organized into per-table sets, data in each set belonging to a single source table.

3. The method of claim 2 , wherein per-table sets are organized into the per-account group, data in the per-account group belonging to a single account.

4. The method of claim 1 , wherein ordering of the received data is maintained based on the sequencing information.

5. The method of claim 1 , further comprising:

writing the data to a metadata store.

6. The method of claim 1 , further comprising:

retrieving expression properties of the committed data; and

pruning the committed data based on the expression properties and the query.

7. A machine-storage medium embodying instructions that, when executed by a machine, cause the machine to perform operations comprising:

receiving a registration request for a per-account group of files, the registration request including identification information for the per-account group of files and a storage location of the per-account group of files;

accessing the per-account group of files based on the registration request, the per-account group of files being stored in a first format;

deduping and validating the per-account group of files using sequencing information included in the per-account group of files;

committing the per-account group of files stored in the storage location and making data in the per-account group of files in the first format accessible for query processing before the data is ingested into one or more source tables;

generating a hybrid table for query processing, the hybrid table including the committed data in the first format and data from the one or more source tables;

receiving a query prior to the data being ingested;

executing the query using the hybrid table to generate results for the query including:

converting the committed data from the first format into a common format;

converting the data from the one or more source tables into the common format;

joining the committed data in the common format and the data from the one or more source tables in the common format to generate joined data; and

executing the query based on the joined data; and

ingesting the data into the one or more source tables in a second format.

8. The machine-storage medium of claim 7 , wherein the data is organized into per-table sets, data in each set belonging to a single source table.

9. The machine-storage medium of claim 8 , wherein per-table sets are organized into the per-account group, data in the per-account group belonging to a single account.

10. The machine-storage medium of claim 7 , wherein ordering of the received data is maintained based on the sequencing information.

11. The machine-storage medium of claim 7 , further comprising:

writing the data to a metadata store.

12. The machine-storage medium of claim 7 , further comprising:

retrieving expression properties of the committed data; and

pruning the committed data based on the expression properties and the query.

13. A system comprising:

at least one hardware processor; and

at least one memory storing instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform operations comprising:

receiving a registration request for a per-account group of files, the registration request including identification information for the per-account group of files and a storage location of the per-account group of files;

accessing the per-account group of files based on the registration request, the per-account group of files being stored in a first format;

deduping and validating the per-account group of files using sequencing information included in the per-account group of files;

committing the per-account group of files stored in the storage location and making data in the per-account group of files in the first format accessible for query processing before the data is ingested into one or more source tables;

generating a hybrid table for query processing, the hybrid table including the committed data in the first format and data from the one or more source tables;

receiving a query prior to the data being ingested;

executing the query using the hybrid table to generate results for the query including:

converting the committed data from the first format into a common format;

converting the data from the one or more source tables into the common format;

joining the committed data in the common format and the data from the one or more source tables in the common format to generate joined data; and

executing the query based on the joined data; and

ingesting the data into the one or more source tables in a second format.

14. The system of claim 13 , wherein the data is organized into per-table sets, data in each set belonging to a single source table.

15. The system of claim 14 , wherein per-table sets are organized into the per-account group, data in the per-account group belonging to a single account.

16. The system of claim 13 , wherein ordering of the received data is maintained based on the sequencing information.

17. The system of claim 13 , the operations further comprising:

writing the data to a metadata store.

18. The system of claim 13 , the operations further comprising:

retrieving expression properties of the committed data; and

pruning the committed data based on the expression properties and the query.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2022
From: AKIDAU, TYLER ARTHUR; CSERI, ISTVAN; JONES, TYLER; SOTOLONGO, DANIEL E.; ZHANG, ZHUO
To: SNOWFLAKE INC.
Reel/Frame 061525/0068 →
Continuity (4)
Continuation 17647500 · Jan 10, 2022
Continuation 17386258 · Jul 27, 2021
Continuation 17226423 · Apr 9, 2021
Related Publication 20230070152A1 · Mar 9, 2023
Cited By (1)
US 12,399,900