IP Library Granted Patent US 11,144,355
Granted Patent B2
US 11,144,355 · App. 16/751,851 · Granted Oct 12, 2021

System and method of providing system jobs within a compute environment

Inventor: David B. Jackson (Spanish Fork, UT)
Assignee: III Holdings 12, LLC
G06F9/5011G06F9/4843G06F9/542G06F8/61
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,144,355
App. No.
16/751,851
Granted
Oct 12, 2021
Kind
B2
Abstract

The disclosure relates to systems, methods and computer-readable media for using system jobs for performing actions outside the constraints of batch compute jobs submitted to a compute environment such as a cluster or a grid. The method for modifying a compute environment from a system job disclosure associating a system job to a queuable object, triggering the system job based on an event and performing arbitrary actions on resources outside of compute nodes in the compute environment. The queuable objects include objects such as batch compute jobs or job reservations. The events that trigger the system job may be time driven, such as ten minutes prior to completion of the batch compute job, or dependent on other actions associated with other system jobs. The system jobs may be utilized also to perform rolling maintenance on a node by node basis.

Claims (37)

1. A method comprising:

receiving a submission of a system job associated with a node in a compute environment;

automatically performing, at a predetermined time, an update to the node, wherein the update comprises one or more of reinstalling software or updating software on the node;

determining whether the update to the node was successful;

in response to the update to the node being successful, terminating the system job leaving the node available for use in the compute environment;

in response to the update to the node not being successful, communicating an unsuccessful status report and creating a reservation for the node; and

performing an update to at least one additional node in the compute environment, wherein the updates to the node and the at least one additional node are performed on a node-by-node basis independently.

2. The method of claim 1 , wherein the update includes updating an operating system.

3. The method of claim 1 , wherein the predetermined time is set to be one of an earliest possible time, a scheduled time, or an earliest possible time after a predetermined period of time.

4. The method of claim 1 , wherein the system job is set to be submitted at both a grid level and a cluster level within the compute environment.

5. The method of claim 1 , further comprising associating the system job to a queueable object.

6. The method of claim 5 , wherein performance of the system job is based on a time offset associated with one of a beginning of or a completion of the queueable object.

7. The method of claim 1 , wherein communicating an unsuccessful status report comprises communicating a message to an administrator.

8. A system comprising:

a processor; and

a computer-readable storage medium storing instructions which, when executed by the processor, cause the processor to perform operations comprising:

receiving a submission of a system job associated with a node in a compute environment;

automatically performing, at a predetermined time an update to the node, wherein the update comprises one or more of installing software or updating software on the node;

determining whether the update to the node was successful;

in response to the update to the node being successful, terminating the system job leaving the node available for use in the compute environment;

in response to the update to the node not being successful, communicating an unsuccessful status report and creating a reservation for the node; and

perform an update to at least one additional node in the compute environment, wherein the updates to the node and the at least one additional node are performed on a node-by-node basis independently.

9. The system of claim 8 , wherein the predetermined time is set to be one of an earliest possible time, a scheduled time, or an earliest possible time after a predetermined period of time.

10. The system of claim 8 , wherein the system job is set to be submitted at both a grid level and a cluster level within the compute environment.

11. The system of claim 8 , wherein the computer-readable storage medium further stores instructions which, when executed by the processor, cause the processor to associate the system job to a queueable object.

12. The system of claim 11 , wherein performance of the system job is based on a time offset associated with one of a beginning of or a completion of the queueable object.

13. A non-transitory computer-readable storage medium storing instructions which, when executed by a processor, cause the processor to perform operations comprising:

receiving a submission of a system job associated with a node in a compute environment;

automatically performing, at a predetermined time, an update to the node, wherein the update comprises one or more of installing software or updating software on the node;

determining whether the update to the node was successful;

in response to the update to the node being successful, terminating the system job leaving the node available for use in the compute environment; and

in response to the update to the node not being successful, communicating an unsuccessful status report and creating a reservation for the node; and

performing an update to at least one additional node in the compute environment, wherein the updates to the node and the at least one additional node are performed on a node-by-node basis independently.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the predetermined time operation is set to be one of an earliest possible time, a scheduled time or an earliest possible time after a predetermined period of time.

15. The non-transitory computer-readable storage medium of claim 13 , wherein the system job is set to be submitted at both a grid level and a cluster level within the compute environment.

16. The non-transitory computer-readable storage medium of claim 13 , wherein the processor is further configured to associate the system job to a queueable object.

17. The non-transitory computer-readable storage medium of claim 16 , wherein performance of the system job is based on a time offset associated with one of a beginning of or a completion of the queueable object.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2021
From: JACKSON, DAVID B.
To: CLUSTER RESOURCES
Reel/Frame 057527/0278 →
CONFIRMATORY ASSIGNMENT Recorded Sep 9, 2021
From: JACKSON, DAVID B.
To: CLUSTER RESOURCES, INC.
Reel/Frame 057527/0284 →
CHANGE OF NAME Recorded Sep 9, 2021
From: CLUSTER RESOURCES, INC.
To: ADAPTIVE COMPUTING ENTERPRISES, INC.
Reel/Frame 057527/0303 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2021
From: ADAPTIVE COMPUTING ENTERPRISES, INC.
To: III HOLDINGS 12, LLC
Reel/Frame 057527/0358 →
Continuity (6)
Continuation 15437135 · Feb 20, 2017
Continuation 14872645 · Oct 1, 2015
Continuation 13621987 · Sep 18, 2012
Continuation 11718867
Provisional Application 60625894 · Nov 8, 2004
Related Publication 20200233711A1 · Jul 23, 2020