IP Library Granted Patent US 10,019,454
Granted Patent B2
US 10,019,454 · App. 15/582,126 · Granted Jul 10, 2018

Data management systems and methods

Inventors: Benoit Dageville (Foster City, CA); Thierry Cruanes (San Mateo, CA); Marcin Zukowski (San Mateo, CA)
Assignee: SNOWFLAKE COMPUTING INC.
G06F17/30106G06F9/5083G06F17/30445G06F17/30463
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,019,454
App. No.
15/582,126
Granted
Jul 10, 2018
Kind
B2
Abstract

Example data management systems and methods are described. In one implementation, a method identifies multiple files to process based on a received query and identifies multiple execution nodes available to process the multiple files. The method initially creates multiple scansets, each including a portion of the multiple files, and assigns each scanset to one of the execution nodes based on a file assignment model. The multiple scansets are processed by the multiple execution nodes. If the method determines that a particular execution node has finished processing all files in its assigned scanset, an unprocessed file is reassigned from another execution node to the particular execution node.

Claims (26)

1. A method comprising:

receiving a query directed to a database;

identifying a plurality of files within the database to process in order to generate a response to the query;

identifying a plurality of execution nodes available to process the plurality of files;

creating a plurality of scansets and assigning each scanset to a different node of the plurality of execution nodes based on a file assignment model;

processing, by the plurality of execution nodes, the plurality of scansets in parallel;

determining, during the processing, that a first execution node has finished processing all files in its assigned scanset; and

responding to the determining by identifying an unprocessed file within one of the plurality of scansets assigned to a second execution node;

assigning the unprocessed file to the first execution node to be processed by the first execution node; and

generating the response to the query.

2. The method of claim 1 , further comprising arranging two or more files in each scanset of the plurality of scansets based on the size of each file.

3. The method of claim 1 , further comprising arranging two or more files in each scanset of the plurality of scansets to prioritize files cached by the assigned execution node of the plurality of execution nodes.

4. The method of claim 1 , wherein the file assignment model uses a consistent hashing model.

5. The method of claim 1 , wherein assigning the unprocessed file to the first execution node comprises removing the unprocessed file from the scanset that was assigned to the second execution node.

6. The method of claim 1 , wherein identifying the unprocessed file comprises identifying the unprocessed file as having already been cached by the first execution node.

7. The method of claim 1 , wherein identifying the unprocessed file comprises identifying the unprocessed file based on a file stealing model.

8. The method of claim 7 , wherein the file stealing model uses consistent hashing at different ownership levels.

9. The method of claim 8 , wherein the different ownership levels determine an order in which files are processed by each of the plurality of execution nodes.

10. The method of claim 1 , wherein the first execution node initiates retrieval of the unprocessed file from a remote storage device.

11. The method of claim 10 , further comprising:

concluding that the second execution node has become available to process the unprocessed file, the second execution node has cached the unprocessed file, and the first execution node has not finished retrieving the unprocessed file from the remote storage device; and

instructing, in response to the concluding, the first section node to stop processing the unprocessed file.

12. The method of claim 11 , wherein the instructing further comprises instructing the second execution node to process the unprocessed file.

13. The method of claim 1 , wherein the file assignment model uses consistent hashing at different ownership levels.

14. The method of claim 7 , wherein both the file assignment model and the file stealing model use consistent hashing at different ownership levels.

15. The method of claim 1 , wherein each scanset of the plurality of scansets includes a different portion of the plurality of files and each file of the plurality of files is found somewhere within the plurality of scansets.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE EXECUTION DATE TO APRIL 1,2019, THAT WAS INCORRECTLY RECOREDED AS MARCH 14,2019 PREVIOUSLY RECORDED AT REEL: 049127 FRAME: 0027. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 11, 2021
From: SNOWFLAKE COMPUTING, INC.
To: SNOWFLAKE INC.
Reel/Frame 057160/0204 →
CHANGE OF NAME Recorded Apr 11, 2019
From: SNOWFLAKE COMPUTING, INC.
To: SNOWFLAKE INC.
Reel/Frame 049127/0027 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2017
From: DAGEVILLE, BENOIT; CRUANES, THIERRY; ZUKOWSKI, MARCIN
To: SNOWFLAKE COMPUTING INC.
Reel/Frame 042186/0108 →
Continuity (3)
Continuation 14518873 · Oct 20, 2014
Provisional Application 61941986 · Feb 19, 2014
Related Publication 20170235750A1 · Aug 17, 2017