IP Library Patent Application 13632200
Patent Application
App. No. 13/632,200

WORKLOAD MANAGEMENT CONSIDERING HARDWARE RELIABILITY

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
13/632,200
Abstract

A method identifies uptime for each of a plurality of components within a cluster of nodes, and determines a reliability level for each of the plurality of components, where the reliability level of each component is determined by comparing the identified uptime for the component with mean-time-between-failure data for components of the same component type. The method also determines a priority level and a job type for a job to be scheduled. Then, at least one target component type is selected in consideration of the job type, and a target reliability level for the at least one target component type is selected in consideration of the priority level. The job is then scheduled on one of the nodes that includes a component of the at least one target component type having the target reliability level.

Claims (34)

1 . A method, comprising:

identifying uptime for each of a plurality of components within a cluster of nodes;

determining a reliability level for each of the plurality of components, wherein the reliability level of each component is determined by comparing the identified uptime for the component with mean-time-between-failure data for components of the same component type;

determining a priority level and a job type for a job to be scheduled;

selecting at least one target component type in consideration of the job type determined for the job;

selecting a target reliability level for the at least one target component type in consideration of the priority level determined for the job; and

scheduling the job on one of the nodes that includes a component of the at least one target component type having the target reliability level.

2 . The method of claim 1 , wherein identifying uptime for each of a plurality of components within a cluster of nodes, includes reading vital product data for each of the plurality of components.

3 . The method of claim 2 , wherein vital product data for each component includes the component uptime and a component type.

4 . The method of claim 2 , wherein each node in the cluster includes a management controller that reads the vital product data and makes the vital product data available to a cluster management node.

5 . The method of claim 4 , wherein the cluster management node provides the vital product data to a provisioning manager that is responsible for scheduling the job.

6 . The method of claim 1 , wherein the target reliability level for the at least one target component type is selected in direct relation to the priority level determined for the job.

7 . The method of claim 1 , wherein the reliability level of each component is determined by comparing the identified uptime for the component with mean-time-between-failure data for components of the same component type and component manufacturer.

8 . The method of claim 7 , wherein the mean-time-between-failure data for the component type includes one or more reliability level designated by a range of component uptime.

9 . The method of claim 1 , wherein a plurality of predetermined job types each have a predetermined priority level.

10 . The method of claim 1 , wherein a plurality of predetermined job types each have a predetermined priority level and at least one predetermined component type.

11 . The method of claim 10 , wherein the at least one predetermined component type is a component type on which the job will place the highest workload.

12 . The method of claim 1 , wherein the target reliability level of the at least one target component type is selected in direct relation to the priority level determined for the job.

13 . The method of claim 1 , wherein the at least one target component type is selected from a processing device, a memory device, a data storage device, and a data communication device.

14 . The method of claim 1 , wherein the reliability level determined for each component is increased to reflect the presence of a redundant component within the same node.

15 . The method of claim 1 , further comprising:

identifying the cost of each of the plurality of components; and

scheduling the job on one of the nodes that includes a component of the at least one target component type having the target reliability level and a cost in proportion to the priority of the job.

16 . A computer program product including computer usable program code embodied on a tangible computer usable storage medium, the computer program product including:

computer usable program code for identifying uptime for each of a plurality of components within a cluster of nodes;

computer usable program code for determining a reliability level for each of the plurality of components, wherein the reliability level of each component is determined by comparing the identified uptime for the component with mean-time-between-failure data for components of the same component type;

computer usable program code for determining a priority level and a job type for a job to be scheduled;

computer usable program code for selecting at least one target component type in consideration of the job type determined for the job;

computer usable program code for selecting a target reliability level for the at least one target component type in consideration of the priority level determined for the job; and

computer usable program code for scheduling the job on one of the nodes that includes a component of the at least one target component type having the target reliability level.

17 . The computer program product of claim 16 , wherein the computer usable program code for identifying uptime for each of a plurality of components within a cluster of nodes, includes computer usable program code for reading vital product data for each of the plurality of components.

18 . The computer program product of claim 17 , wherein vital product data for each component includes the component uptime and a component type.

19 . The computer program product of claim 17 , wherein each node in the cluster includes a management controller that reads the vital product data and makes the vital product data available to a cluster management node.

20 . The computer program product of claim 19 , wherein the cluster management node provides the vital product data to a provisioning manager that is responsible for scheduling the job.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2014
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: LENOVO ENTERPRISE SOLUTIONS (SINGAPORE) PTE. LTD.
Reel/Frame 034194/0111 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2012
From: ALSHINNAWI, SHAREEF F.; CUDAK, GARY D.; SUFFERN, EDWARD S.; WEBER, J. MARK
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 029053/0279 →