IP Library Granted Patent US 10,540,202
Granted Patent B1
US 10,540,202 · App. 15/719,324 · Granted Jan 21, 2020

Transient sharing of available SAN compute capability

Inventors: Stephen Smaldone (Woodstock, CT); Ian Wigmore (Westborough, MA); Steven Chalmer (Redwood City, CA); Jonathan Krasner (Coventry, RI); Chakib Ouarraoui (Watertown, MA)
Assignee: EMC IP Holding Company LLC
G06F9/4881G06F9/485G06F9/4856G06F9/5027G06F2209/5019
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,540,202
App. No.
15/719,324
Granted
Jan 21, 2020
Kind
B1
Abstract

Embodiments are described for a executing a processing job using one or more nodes of a storage area network using computing resources on the SAN that are predicted to be idle. A predictive model is generated by monitoring idle states of resources of nodes of the SAN and using machine learning to build the predictive model. A scheduler executes jobs on one or more nodes of the SAN with sufficient predicted idle resources to process the job, in accordance with resource requirements and job attributes in a manifest of the job. If a job cannot be completed during a window of time that the necessary resources are predicted to be idle, or if one or more resources become unavailable, the job can be paused and resumed, migrated to another node, or restarted at a later time when the required resources to complete the job are predicted to be idle.

Claims (73)

1. A computer-implemented comprising:

receiving, by a scheduler, a processing job and a manifest for the processing job, the manifest including an estimate of one or more resources required to perform the processing job;

determining, by a scheduler on a first node in a multi-node network, at least one second node in the multi-node network, on which to execute the processing job using a plurality of computing resources predicted to be idle to execute the processing job on the at least one second node, wherein the determining is based at least in part upon whether a predictive model of idle times of computing resources in the multi-node network predicts that there are idle computing resources on the at least one second node having a magnitude that meets or exceeds an estimated requirements to perform the processing job; and

in response to determining that an actual state of the computing resources predicted to be idle on the at least one second node meets or exceeds a predicted state of the computing resources on the at least one second node, executing the processing job on the at least one second node.

2. The method of claim 1 , wherein the multi-node network comprises a storage area network (SAN) and nodes of the multi-node network comprise a plurality of host computers and at least one storage appliance, and the predictive model includes predictions for resources of the at least one storage appliance.

3. The method of claim 1 , further comprising updating the predictive model of idle times of computing resources on a plurality of nodes of the multi-node network by updating an estimate of one or more resources required to perform the processing job with actual execution resources used to perform the processing job.

4. The method of claim 1 , further comprising:

in response to determining that the actual state of computing resources predicted to be idle on the at least one second node does not meet the computing resources predicted to be idle for executing the processing job on the at least one second node:

determining at least one third node on which to execute the processing job;

executing the processing job on the at the least one third node.

5. The method of claim 1 , further comprising:

in response to determining, during the execution of the processing job on the at least one second node, that at least some of the computing resources predicted to be idle on the at least one second node are no longer idle or are no longer available to the processing job:

determining that the processing job is pausible;

pausing the processing job until the computing resources required to complete the processing job on the at least one second node are predicted to be idle for a predicted remaining execution time of the processing job;

resuming execution of the processing job on the at least one second node.

6. The method of claim 1 , further comprising:

in response to determining, during the execution of the processing job on the at least one second node, that at least some of the computing resources predicted to be idle on the at least one second node are no longer idle or are no longer available to the processing job:

determining that the processing job is restartable;

determining a second predicted time at which the resources for executing the processing job are predicted to be idle on the at least one second node;

restarting execution of the processing job on the at least one second node.

7. The method of claim 1 , further comprising:

in response to determining that at least some of the computing resources predicted to be idle on the at least one second node are not idle or are no longer available to the processing job:

migrating the processing job to a fourth node in the multi-node network that has the plurality of resources for executing the processing job predicted to be idle;

executing the processing job on the at least one fourth node.

8. A non-transitory computer-readable medium programmed with executable instructions that, when executed by a processing system having at least one hardware processor, perform operations comprising:

receiving a processing job and a manifest for the processing job, the manifest including an estimate of one or more resources required to perform the processing job;

determining, by a scheduler on a first node in a multi-node network, at least one second node the a multi-node network, on which to execute the processing job using a plurality of computing resources predicted to be idle to execute the processing job on the at least one second node, wherein the determining is based at least in part upon whether a predictive model of idle times of computing resources in the multi-node network predicts that there are idle computing resources on the at least one second node having a magnitude that meets or exceeds estimated requirements to perform the processing job; and

in response to determining that an actual state of the computing resources predicted to be idle on the at least one second node meets or exceeds a predicted state of the computing resources on the at least one second node, executing the processing job on the at least one second node.

9. The medium of claim 8 , wherein the multi-node network comprises a storage area network (SAN) and nodes of the multi-node network comprise a plurality of host computers and at least one storage appliance, and the predictive model includes predictions for resources of the at least one storage appliance.

10. The medium of claim 8 , wherein the operations further comprise updating the predictive model of idle times of computing resources on a plurality of nodes of the multi-node network by updating an estimate of one or more resources required to perform the processing job with actual execution resources used to perform the processing job.

11. The medium of claim 8 , wherein operations further comprise:

in response to determining that an actual state of computing resources predicted to be idle on the at least one second node does not meet the computing resources predicted to be idle for executing the processing job on the at least one second node:

determining at least one third node on which to execute the processing job;

executing the processing job on the at the least one third node.

12. The medium of claim 8 , the operations further comprising:

in response to determining, during the execution of the processing job on the at least one second node, that at least some of the computing resources predicted to be idle on the at least one second node are no longer idle or are no longer available to the processing job:

determining that the processing job is pausible;

pausing the processing job until the computing resources required to complete the processing job on the at least one second node are predicted to be idle for a predicted remaining execution time of the processing job;

resuming execution of the processing job on the at least one second node.

13. The medium of claim 8 , the operations further comprising:

in response to determining, during the execution of the processing job on the at least one second node, that at least some of the computing resources predicted to be idle on the at least one second node are no longer idle or are no longer available to the processing job:

determining that the processing job is restartable;

determining a second predicted time at which the resources for executing the processing job are predicted to be idle on the at least one second node;

restarting execution of the processing job on the at least one second node.

14. The medium of claim 8 , the operations further comprising:

in response to determining that at least some of the computing resources predicted to be idle on the at least one second node are not idle or are no longer available to the processing job:

migrating the processing job to a fourth node in the multi-node network that has the plurality of resources for executing the processing job predicted to be idle;

executing the processing job on the at least one fourth node.

15. A system comprising:

a processing system having at least one hardware processor, the processing system coupled to a memory programmed with executable instructions that, when executed by the processing system, perform operations comprising:

receiving a processing job and a manifest for the processing job, the manifest including an estimate of one or more resources required to perform the processing job;

determining, by a scheduler on a first node in a multi-node network, at least one second node in the multi-node network, on which to execute the processing job using a plurality of computing resources predicted to be idle to execute the processing job on the at least one second node, wherein the determining is based at least in part upon whether a predictive model of idle times of computing resources in the multi-node network predicts that there are idle computing resources on at least one second node having a magnitude that meets or exceeds estimated requirements to perform the processing job; and

in response to determining that an actual state of the computing resources predicted to be idle on the at least one second node meets or exceeds a predicted state of the computing resources on the at least one second node, executing the processing job on the at least one second node.

16. The system of claim 15 , wherein the multi-node network comprises a storage area network (SAN) and nodes of the multi-node network comprise a plurality of host computers and at least one storage appliance, and the predictive model includes predictions for resources of the at least one storage appliance.

17. The system of claim 15 , wherein the operations further comprise updating the predictive model of idle times of computing resources on a plurality of nodes of the multi-node network by updating an estimate of one or more resources required to perform the processing job with actual execution resources used to perform the processing job.

18. The system of claim 15 , wherein the operations further comprise:

in response to determining that an actual state of computing resources predicted to be idle on the at least one second node does not meet the computing resources predicted to be idle for executing the processing job on the at least one second node:

determining at least one third node on which to execute the processing job;

executing the processing job on the at the least one third node.

19. The system of claim 15 , the operations further comprising:

in response to determining, during the execution of the processing job on the at least one second node, that at least some of the computing resources predicted to be idle on the at least one second node are no longer idle or are no longer available to the processing job:

determining that the processing job is pausible;

pausing the processing job until the computing resources required to complete the processing job on the at least one second node are predicted to be idle for a predicted remaining execution time of the processing job;

resuming execution of the processing job on the at least one second node.

20. The system of claim 15 , the operations further comprising:

in response to determining, during the execution of the processing job on the at least one second node, that at least some of the computing resources predicted to be idle on the at least one second node are no longer idle or are no longer available to the processing job:

determining that the processing job is restartable;

determining a second predicted time at which the resources for executing the processing job are predicted to be idle on the at least one second node;

restarting execution of the processing job on the at least one second node.

21. The system of claim 15 , the operations further comprising:

in response to determining that at least some of the computing resources predicted to be idle on the at least one second node are not idle or are no longer available to the processing job:

migrating the processing job to a fourth node in the multi-node network that has the plurality of resources for executing the processing job predicted to be idle;

executing the processing job on the at least one fourth node.

Assignments (7)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (044535/0109) Recorded May 20, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO WYSE TECHNOLOGY L.L.C.)
Reel/Frame 060753/0414 →
RELEASE OF SECURITY INTEREST AT REEL 044535 FRAME 0001 Recorded Nov 2, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.
Reel/Frame 058298/0475 →
SECURITY AGREEMENT Recorded Mar 21, 2019
From: CREDANT TECHNOLOGIES, INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 049452/0223 →
PATENT SECURITY AGREEMENT (NOTES) Recorded Nov 29, 2017
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 044535/0109 →
PATENT SECURITY AGREEMENT (CREDIT) Recorded Nov 29, 2017
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 044535/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2017
From: SMALDONE, STEPHEN; WIGMORE, IAN; CHALMER, STEVEN; KRASNER, JONATHAN; OUARRAOUI, CHAKIB
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 043743/0628 →
Cited By (11)
US 12,197,940 US 12,360,808 US 12,399,749 US 12,401,627 US 12,411,710 US 12,412,108 US 12,450,088 US 12,517,754 US 12,613,738 US 12,675,339 US 12,681,759