IP Library Granted Patent US 9,769,032
Granted Patent B1
US 9,769,032 · App. 14/663,748 · Granted Sep 19, 2017

Cluster instance management system

Inventors: Ali Ghodsi (Berkeley, CA); Ion Stoica (Piedmont, CA); Matei Zaharia (Boston, MA)
Assignee: Databricks Inc.
H04L41/50H04L41/0654H04L43/0805H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,769,032
App. No.
14/663,748
Granted
Sep 19, 2017
Kind
B1
Abstract

A system for cluster management comprises a status monitor and an instance replacement manager. The status monitor is for monitoring status of an instance of a set of instances on a cluster provider. The instance replacement manager is for determining a replacement strategy for the instance in the event the instance does not respond. The replacement strategy for the instance is based at least in part on a management criteria for on-demand instances and spot instances on the cluster provider.

Claims (43)

1. A system for cluster management, comprising:

a processor; and

a memory coupled with the processor, wherein the memory is configured to provide the processor with instructions which when executed cause the processor to:

monitor status of an instance of a set of instances on a cluster provider; and

determine a replacement strategy for the instance in the event the instance does not respond, wherein the replacement strategy for the instance is based at least in part on a management criteria for on-demand instances and spot instances on the cluster provider, wherein the determining of the replacement strategy comprises to:

determine whether it is acceptable to not replace the instance immediately based on the management criteria; and

in response to a determination that it is acceptable to not replace the instance immediately based on the management criteria:

monitor availability of a new spot instance; and

in response to a determination that the new spot instance is available, replace the instance that does not respond with the new spot instance.

2. A system as in claim 1 , wherein the processor is further configured to receive the management criteria.

3. A system as in claim 2 , wherein the management criteria comprise one or more of the following: a maximum number of on-demand instances or a minimum number of on-demand instances.

4. A system as in claim 2 , wherein the management criteria comprise one or more of the following: a maximum number of total instances or a minimum number of total instances.

5. A system as in claim 2 , wherein the management criteria comprise a budget limit.

6. A system as in claim 2 , wherein the management criteria comprise a failure criterion.

7. A system as in claim 2 , wherein the management criteria comprise a reserve criterion.

8. A system as in claim 2 , wherein the management criteria comprise a reserve poolsize.

9. A system as in claim 1 , wherein the processor is further configured to determine a new set of instances to request based at least in part on the management criteria.

10. A system as in claim 9 , wherein the processor is further configured to:

determine availability of the new set of instances on the cluster provider; and

in the event there is availability, indicate to create the new set of instances.

11. A system as in claim 9 , wherein the processor is further configured to determine a request strategy based at least in part on a cluster availability and on the management criteria.

12. A system as in claim 1 , wherein the monitoring of the status comprising to monitor the status of the instance periodically.

13. A system as in claim 1 , wherein the monitoring of the status comprising to monitor the status of each instance of the set of instances.

14. A system as in claim 1 , wherein the monitoring of the status comprising to monitor the status of each spot instance of the set of instances.

15. A system as in claim 1 , wherein the replacement strategy comprises replacing the instance with a spot instance from a reserve pool.

16. A system as in claim 1 , wherein the replacement strategy comprises replacing the instance with an on-demand instance in the event that this is allowed under the management criteria.

17. A system as in claim 1 , wherein the replacement strategy comprises stopping the set of instances.

18. A system as in claim 17 , wherein the replacement strategy comprises saving the state of the set of instances.

19. A system as in claim 17 , wherein the replacement strategy comprises maintaining one on-demand instance of the set of instances.

20. A method for cluster management, comprising:

monitoring status of an instance of a set of instances on a cluster provider; and

determining, using a processor, a replacement strategy for the instance in the event the instance does not respond, wherein the replacement strategy for the instance is based at least in part on a management criteria for on-demand instances and spot instances on the cluster provider, wherein the determining of the replacement strategy comprises:

determining whether it is acceptable to not replace the instance immediately based on the management criteria; and

in response to a determination that it is acceptable to not replace the instance immediately based on the management criteria:

monitoring availability of a new spot instance; and

in response to a determination that the new spot instance is available, replacing the instance that does not respond with the new spot instance.

21. A computer program product for cluster management, the computer program product being embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

monitoring status of an instance of a set of instances on a cluster provider; and

determining a replacement strategy for the instance in the event the instance does not respond, wherein the replacement strategy for the instance is based at least in part on a management criteria for on-demand instances and spot instances on the cluster provider, wherein the determining of the replacement strategy comprises:

determining whether it is acceptable to not replace the instance immediately based on the management criteria; and

in response to a determination that it is acceptable to not replace the instance immediately based on the management criteria:

monitoring availability of a new spot instance; and

in response to a determination that the new spot instance is available, replacing the instance that does not respond with the new spot instance.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME FROM DATABRICKS INC. TO DATABRICKS, INC. AND ASSIGNEE STREET ADDRESS FROM 160 SPEAR STREET, SUITE 1300 TO 160 SPEAR STREET, 15TH FLOOR PREVIOUSLY RECORDED ON REEL 35847 FRAME 637. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 21, 2025
From: GHODSI, ALI; STOICA, ION; ZAHARIA, MATEI
To: DATABRICKS, INC.
Reel/Frame 070288/0560 →
SECURITY INTEREST Recorded Jan 6, 2025
From: DATABRICKS, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 069825/0419 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2015
From: GHODSI, ALI; STOICA, ION; ZAHARIA, MATEI
To: DATABRICKS INC.
Reel/Frame 035847/0637 →