IP Library › Granted Patent US 11,461,335
Granted Patent B1
US 11,461,335 · App. 17/455,594 · Granted Oct 4, 2022

Optimized processing of data in different formats

Inventors: Tyler Arthur Akidau (Seattle, WA); Thierry Cruanes (San Mateo, CA); Istvan Cseri (Seattle, WA); Benoit Dageville (San Mateo, CA); Tyler Jones (Redwood City, CA); Dinesh Chandrakant Kulkarni (Sammamish, WA)
Assignee: Snowflake Inc.
G06F16/24568G06F16/24544
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,461,335
App. No.
17/455,594
Filed
Nov 18, 2021
Granted
Oct 4, 2022
Kind
B1
Examiner
JAMI, HARES
Art Unit
2162
USPC
707/714
Abstract

Hybrid tables can be used in different use-case scenarios. Hybrid tables provide a flexible mechanism to support files and data in different formats while providing access to the different types of data as part of one table. This flexibility can allow the use of hybrid tables in data lake or other similar environments.

Claims (62)

1. A method comprising:

storing data for one or more source tables in a hybrid table, including a first set of data in a first format received from an external cloud storage location and a second set of data in a second format received from a network-based data warehouse system, the first format being a raw format and the second format being a formatted format used by the network-based data warehouse system;

classifying a first subset of the first set of data in the first format as high-value data and classifying a second subset of the first set of data as low-value data;

ingesting a copy of the high-value data from the external cloud storage location into the network-based data warehouse system in the second format, wherein the first subset of the first set of data in the first format is maintained and not deleted in the external cloud storage location in response to ingesting the copy of the high-value data; and

updating metadata of the hybrid table based on ingesting the copy of the high-value data.

2. The method of claim 1 , further comprising:

receiving a query involving the first subset; and

executing the query using the ingested copy of the first subset in the second format.

3. The method of claim 1 , further comprising:

receiving a query involving the second subset and the second set;

converting the second subset from the first format into a common format;

converting the second set from the second format into the common format;

joining the second subset in the common format and the second set in the common format to generate joined data;

executing the query based on the joined data.

4. The method of claim 1 , wherein the classifying is performed based on query patterns.

5. The method of claim 4 , wherein the classifying is performed based on scan statistics.

6. The method of claim 4 , wherein the classifying is performed based on metadata received from a client.

7. The method of claim 4 , further comprising:

re-classifying the first subset from high-value data to low-value data; and

in response to re-classifying, deleting the ingested copy of the first subset in the second format.

8. A machine-storage medium embodying instructions that, when executed by a machine, cause the machine to perform operations comprising:

storing data for one or more source tables in a hybrid table, including a first set of data in a first format received from an external cloud storage location and a second set of data in a second format received from a network-based data warehouse system, the first format being a raw format and the second format being a formatted format used by the network-based data warehouse system;

classifying a first subset of the first set of data in the first format as high-value data and classifying a second subset of the first set of data as low-value data;

ingesting a copy of the high-value data from the external cloud storage location into the network-based data warehouse system in the second format, wherein the first subset of the first set of data in the first format is maintained and not deleted in the external cloud storage location in response to ingesting the copy of the high-value data; and

updating metadata of the hybrid table based on ingesting the copy of the high-value data.

9. The machine-storage medium of claim 8 , further comprising:

receiving a query involving the first subset; and

executing the query using the ingested copy of the first subset in the second format.

10. The machine-storage medium of claim 8 , further comprising:

receiving a query involving the second subset and the second set;

converting the second subset from the first format into a common format;

converting the second set from the second format into the common format;

joining the second subset in the common format and the second set in the common format to generate joined data;

executing the query based on the joined data.

11. The machine-storage medium of claim 8 , wherein the classifying is performed based on query patterns.

12. The machine-storage medium of claim 8 , wherein the classifying is performed based on scan statistics.

13. The machine-storage medium of claim 8 , wherein the classifying is performed based on metadata received from a client.

14. The machine-storage medium of claim 8 , further comprising:

re-classifying the first subset from high-value data to low-value data; and

in response to re-classifying, deleting the ingested copy of the first subset in the second format.

15. A system comprising:

at least one hardware processor; and

at least one memory storing instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform operations comprising:

storing data for one or more source tables in a hybrid table, including a first set of data in a first format received from an external cloud storage location and a second set of data in a second format received from a network-based data warehouse system, the first format being a raw format and the second format being a formatted format used by the network-based data warehouse system;

classifying a first subset of the first set of data in the first format as high-value data and classifying a second subset of the first set of data as low-value data;

ingesting a copy of the high-value data from the external cloud storage location into the network-based data warehouse system in the second format, wherein the first subset of the first set of data in the first format is maintained and not deleted in the external cloud storage location in response to ingesting the copy of the high-value data; and

updating metadata of the hybrid table based on ingesting the copy of the high-value data.

16. The system of claim 15 , the operations further comprising:

receiving a query involving the first subset; and

executing the query using the ingested copy of the first subset in the second format.

17. The system of claim 15 , the operations further comprising:

receiving a query involving the second subset and the second set;

converting the second subset from the first format into a common format;

converting the second set from the second format into the common format;

joining the second subset in the common format and the second set in the common format to generate joined data;

executing the query based on the joined data.

18. The system of claim 15 , wherein the classifying is performed based on query patterns.

19. The system of claim 15 , wherein the classifying is performed based on scan statistics.

20. The system of claim 15 , wherein the classifying is performed based on metadata received from a client.

21. The system of claim 15 , the operations further comprising:

re-classifying the first subset from high-value data to low-value data; and

in response to re-classifying, deleting the ingested copy of the first subset in the second format.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2022
From: AKIDAU, TYLER ARTHUR; CRUANES, THIERRY; CSERI, ISTVAN; DAGEVILLE, BENOIT; JONES, TYLER; KULKARNI, DINESH CHANDRAKANT
To: SNOWFLAKE INC.
Reel/Frame 058776/0579 →
Continuity (2)
Continuation In Part 17386258 · Jul 27, 2021
Continuation 17226423 · Apr 9, 2021
Cited By (2)
US 12,399,900 US 12,592,950