IP Library › Granted Patent US 11,615,092
Granted Patent B1
US 11,615,092 · App. 17/515,232 · Granted Mar 28, 2023

Lightweight database pipeline scheduler

Inventors: Sebastian Breß (Berlin, DE); Moritz Eyssen (Berlin, DE); Max Heimel (Berlin, DE); Max Jendruk (Berlin, DE)
Assignee: Snowflake Inc.
G06F16/24542G06F9/4881G06F16/24532G06F16/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,615,092
App. No.
17/515,232
Granted
Mar 28, 2023
Kind
B1
Abstract

A database scheduler system can be implemented on a distributed database system. The system schedules operations in a lightweight approach that reduces idling and increases parallel processing of database operations for a query on data of the database. The system performs restarts of individual operators or fragments of a query without restarting the entire query.

Claims (65)

1. A method comprising:

generating a query plan for a query on a distributed database, the query plan comprising a plurality of database operations;

identifying a plurality of contingent database operations from the plurality of database operations of the query plan, the plurality of contingent database operations configured to generate query data for the query based on being executed at a specific position in the query plan in relation to one or more other operators of the plurality of database operations;

scheduling the plurality of contingent database operations for execution using a scheduler of an execution node of the distributed database;

setting, by the scheduler, the scheduled contingent database operations to execute at the specific position in the query plan;

setting, by the scheduler, the remaining database operations for execution by any available thread in the execution node as threads that are processing the query plan become available in the execution node;

processing, using the execution node, of the distributed database, the query plan according to the scheduled contingent database operations; and

storing, by the distributed database, data generated in processing the query plan.

2. The method of claim 1 , wherein the plurality of contingent database operations are scheduled without scheduling remaining database operations of the plurality of database operations that are excluded from the plurality of contingent database operations.

3. The method of claim 1 , wherein the plurality of contingent database operations comprises a beginning leaf operation that generates query data based on being executed at a beginning of the query plan.

4. The method of claim 3 , wherein the plurality of contingent database operations comprises an end operation that generates query data based on being executed at an end of the query plan.

5. The method of claim 3 , wherein the plurality of contingent database operations comprises a dependent operation that generates query data using data that is generated by completion of a previous database operation upon which the dependent operation depends.

6. The method of claim 1 , further comprising:

listing, by the execution node, one or more pipelines as ready for processing by the execution node, a pipeline of the one or more pipelines comprising an additional plurality of database operations of the query plan; and

executing the one or more pipelines while an existing pipeline that comprises the plurality of database operations is executing on the execution node.

7. The method of claim 6 , wherein the distributed database comprises a limit for a quantity of new pipelines that can be executed while the existing pipeline is executing on the execution node.

8. The method of claim 7 , wherein the one or more pipelines comprises a plurality of pipelines.

9. The method of claim 8 , wherein the plurality of pipelines that are executed based on being listed as ready is less than the limit of the quantity of new pipelines that can be executed while the existing pipeline is executing on the execution node.

10. The method of claim 1 , further comprising:

restarting one or more of the plurality of database operations without restarting other database operations in the query plan.

11. The method of claim 10 , wherein:

a restart instruction is issued by the execution node to restart the one or more of the plurality of database operations; and

wherein the other database operations in the query plan complete processing based on the restart instruction.

12. The method of claim 11 , wherein the restart instruction is received, followed by completion of processing of the other database operations, and further followed by restarting the one or more of the plurality of database operations according to the restart instruction.

13. A system comprising:

one or more processors of a machine; and

at least one memory storing instructions that, when executed by the one or more processors, cause the machine to perform operations comprising:

generating a query plan for a query on a distributed database, the query plan comprising a plurality of database operations;

identifying a plurality of contingent database operations from the plurality of database operations of the query plan, the plurality of contingent database operations configured to generate query data for the query based on being executed at a specific position in the query plan in relation to one or more other operators of the plurality of database operations;

scheduling the plurality of contingent database operations for execution using a scheduler of an execution node of the distributed database;

setting, by the scheduler, the scheduled contingent database operations to execute at the specific position in the query plan;

setting, by the scheduler, the remaining database operations for execution by any available thread in the execution node as threads that are processing the query plan become available in the execution node;

processing, using the execution node, of the distributed database, the query plan according to the scheduled contingent database operations; and

storing, by the distributed database, data generated in processing the query plan.

14. The system of claim 13 , wherein the plurality of contingent database operations comprises a beginning leaf operation that generates query data based on being executed at a beginning of the query plan.

15. The system of claim 14 , wherein the plurality of contingent database operations comprises an end leaf operation that generates query data based on being executed at an end of the query plan.

16. The system of claim 14 , wherein the plurality of contingent database operations comprises a dependent operation that generates query data using data that is generated by completion of a previous database operation upon which the dependent operation depends.

17. The system of claim 13 , the operations further comprising:

listing, by the execution node, one or more pipelines as ready for processing by the execution node, a pipeline of the one or more pipelines comprising an additional plurality of database operations of the query plan; and

executing the one or more pipelines while an existing pipeline that comprises the plurality of database operations is executing on the execution node.

18. The system of claim 17 , wherein the distributed database comprises a limit for a quantity of new pipelines that can be executed while the existing pipeline is executing on the execution node.

19. The system of claim 18 , wherein the one or more pipelines comprises a plurality of pipelines.

20. The system of claim 19 , wherein the plurality of pipelines that are executed based on being listed as ready is less than the limit of the quantity of new pipelines that can be executed while the existing pipeline is executing on the execution node.

21. The system of claim 13 , the operations further comprising:

restarting one or more of the plurality of database operations without restarting other database operations in the query plan.

22. The system of claim 21 , wherein:

a restart instruction is issued by the execution node to restart the one or more of the plurality of database operations; and

wherein the other database operations in the query plan complete processing based on the restart instruction.

23. The system of claim 22 , wherein the restart instruction is received, followed by completion of processing of the other database operations, and further followed by restarting the one or more of the plurality of database operations according to the restart instruction.

24. A machine storage medium embodying instructions that, when executed by a machine, cause the machine to perform operations comprising:

generating a query plan for a query on a distributed database, the query plan comprising a plurality of database operations;

identifying a plurality of contingent database operations from the plurality of database operations of the query plan, the plurality of contingent database operations configured to generate query data for the query based on being executed at a specific position in the query plan in relation to one or more other operators of the plurality of database operations;

scheduling the plurality of contingent database operations for execution using a scheduler of an execution node of the distributed database;

setting, by the scheduler, the scheduled contingent database operations to execute at the specific position in the query plan;

setting, by the scheduler, the remaining database operations for execution by any available thread in the execution node as threads that are processing the query plan become available in the execution node;

processing, using the execution node, of the distributed database, the query plan according the scheduled contingent database operations; and

storing, by the distributed database, data generated in processing the query plan.

25. The machine storage medium of claim 24 , wherein the plurality of contingent database operations comprises a beginning leaf operation that generates query data based on being executed at a beginning of the query plan.

26. The machine storage medium of claim 25 , wherein the plurality of contingent database operations comprises an end operation that generates query data based on being executed at an end of the query plan.

27. The machine storage medium of claim 25 , wherein the plurality of contingent database operations comprises a dependent operation that generate query data using data that is generated by completion of a previous database operation upon which the dependent operation depends.

28. The machine storage medium of claim 24 , the operations further comprising:

listing, by the execution node, one or more pipelines as ready for processing by the execution node, a pipeline of the one or more pipelines comprising an additional plurality of database operations of the query plan; and

executing the one or more pipelines while an existing pipeline that comprises the plurality of database operations is executing on the execution node.

29. The machine storage medium of claim 28 , wherein the distributed database comprises a limit for a quantity of new pipelines that can be executed while the existing pipeline is executing on the execution node.

30. The machine storage medium of claim 29 , wherein the one or more pipelines comprises a plurality of pipelines.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2021
From: BRESS, SEBASTIAN; EYSSEN, MORITZ; HEIMEL, MAX; JENDRUK, MAX
To: SNOWFLAKE INC.
Reel/Frame 058252/0289 →
Cited By (3)
US 12,399,896 US 12,488,051 US 12,511,289