IP Library Granted Patent US 10,361,928
Granted Patent B2
US 10,361,928 · App. 15/682,397 · Granted Jul 23, 2019

Cluster instance management system

Inventors: Ali Ghodsi (Berkeley, CA); Ion Stoica (Piedmont, CA); Matei Zaharia (Boston, MA)
Assignee: Databricks Inc.
H04L41/5051G06F11/30H04L41/5096H04L43/0817
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,361,928
App. No.
15/682,397
Granted
Jul 23, 2019
Kind
B2
Abstract

A system for cluster management comprises a status monitor and an instance replacement manager. The status monitor is for monitoring status of an instance of a set of instances on a cluster provider. The instance replacement manager is for determining a replacement strategy for the instance in the event the instance does not respond. The replacement strategy for the instance is based at least in part on a management criteria for on-demand instances and spot instances on the cluster provider.

Claims (44)

1. A system for cluster management, comprising:

a processor; and

a memory coupled with the processor, wherein the memory is configured to provide the processor with instructions which when executed cause the processor to:

monitor status of an instance of a set of instances on a cluster provider; and

determine a replacement strategy for the instance in response to a determination the instance does not respond, wherein the replacement strategy for the instance is based at least in part on a management criteria for on-demand instances and spot instances on the cluster provider, and wherein the determining of the replacement strategy comprises to:

in response to a determination that the set of instances is to be stopped:

determine whether to maintain a master on-demand instance; and

in response to a determination to maintain the master on-demand instance, stop the set of instances except the master on-demand instance.

2. A system as in claim 1 , wherein the processor is further configured to:

receive the management criteria.

3. A system as in claim 2 , wherein the management criteria comprises one or more of the following: a maximum number of on-demand instances or a minimum number of on-demand instances.

4. A system as in claim 2 , wherein the management criteria comprises one or more of the following: a maximum number of total instances or a minimum number of total instances.

5. A system as in claim 2 , wherein the management criteria comprises a budget limit.

6. A system as in claim 2 , wherein the management criteria comprises a failure criterion.

7. A system as in claim 2 , wherein the management criteria comprises a reserve criterion.

8. A system as in claim 2 , wherein the management criteria comprises a reserve poolsize.

9. A system as in claim 1 , wherein the processor is further configured to:

determine a new set of instances to request based at least in part on the management criteria.

10. A system as in claim 9 , wherein the processor is further configured to:

determine availability of the new set of instances on the cluster provider; and

in response to a determination that there is availability, indicate to create the new set of instances.

11. A system as in claim 9 , wherein the determining of the replacement strategy comprises to:

determine a request strategy based at least in part on a cluster availability and on the management criteria.

12. A system as in claim 1 , wherein the monitoring of the status of the instance includes to monitor the status of the instance periodically.

13. A system as in claim 1 , wherein the monitoring of the status of the instance includes to monitor a status of each instance of the set of instances.

14. A system as in claim 1 , wherein the monitoring of the status of the instance includes to monitor a status of each spot instance of the set of instances.

15. A system as in claim 1 , wherein the replacement strategy comprises replacing the instance with a spot instance from a reserve pool.

16. A system as in claim 1 , wherein the replacement strategy comprises replacing the instance with an on-demand instance in response to a determination that this is allowed under the management criteria.

17. A system as in claim 1 , wherein the replacement strategy comprises monitoring for a replacement instance.

18. A system as in claim 1 , wherein the replacement strategy comprises stopping the set of instances.

19. A system as in claim 18 , wherein the replacement strategy comprises saving the state of the set of instances.

20. A system as in claim 18 , wherein the replacement strategy comprises maintaining one on-demand instance of the set of instances.

21. A method for cluster management, comprising:

monitoring status of an instance of a set of instances on a cluster provider; and

determining, using a processor, a replacement strategy for the instance in response to a determination the instance does not respond, wherein the replacement strategy for the instance is based at least in part on a management criteria for on-demand instances and spot instances on the cluster provider, and wherein the determining of the replacement strategy comprises:

in response to a determination that the set of instances is to be stopped:

determining whether to maintain a master on-demand instance; and

in response to a determination to maintain the master on-demand instance, stopping the set of instances except the master on-demand instance.

22. A computer program product for cluster management, the computer program product being embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

monitoring status of an instance of a set of instances on a cluster provider; and

determining, a replacement strategy for the instance in response to a determination the instance does not respond, wherein the replacement strategy for the instance is based at least in part on a management criteria for on-demand instances and spot instances on the cluster provider, and wherein the determining of the replacement strategy comprises:

in response to a determination that the set of instances is to be stopped:

determining whether to maintain a master on-demand instance; and

in response to a determination to maintain the master on-demand instance, stopping the set of instances except the master on-demand instance.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2025
From: GHODSI, ALI; STOICA, ION; ZAHARIA, MATEI
To: DATABRICKS, INC.
Reel/Frame 070338/0894 →
SECURITY INTEREST Recorded Jan 6, 2025
From: DATABRICKS, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 069825/0419 →
Continuity (2)
Continuation 14663748 · Mar 20, 2015
Related Publication 20180048536A1 · Feb 15, 2018
Cited By (2)
US 12,292,870 US 12,608,366