IP Library Granted Patent US 11,023,281
Granted Patent B2
US 11,023,281 · App. 15/964,424 · Granted Jun 1, 2021

Parallel processing apparatus to allocate job using execution scale, job management method to allocate job using execution scale, and recording medium recording job management program to allocate job using execution scale

Inventors: Ryosuke Kokubo (Kannami, JP); Tsuyoshi Hashimoto (Kawasaki, JP)
Assignee: FUJITSU LIMITED
G06F9/5038G06F9/4887G06F9/5066G06F9/5077
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,023,281
App. No.
15/964,424
Granted
Jun 1, 2021
Kind
B2
Abstract

A parallel processing apparatus includes a memory and a processor. The memory stores a program and the processor is coupled to the memory. The processor calculates, based on a number of nodes to be used in execution of respective jobs that are waiting to be executed and a scheduled execution time period for execution of the respective jobs, an execution scale of the respective jobs and allocates the respective jobs to an area in which a number of problem nodes that have a high failure possibility is small from among a plurality of areas into which a region in which a plurality of nodes are disposed is partitioned and divided. The allocation of the jobs is performed in descending order of the execution scale beginning with the job whose execution scale is the largest.

Claims (30)

1. A parallel processing apparatus, comprising:

a memory that stores a program; and

a processor coupled to the memory, the processor:

calculates, based on a number of nodes to be used in execution of respective jobs that are waiting to be executed and a scheduled execution time period for execution of the respective jobs, an execution scale of the respective jobs;

sorts the jobs in descending order of the execution scale to acquire a specific order of the jobs;

acquires a plurality of areas by dividing a node area in which a plurality of nodes are arranged, some of the plurality of areas including one or more nodes with a failure possibility; and

starts an allocation of the respective jobs from an area in which a number of the one or more nodes with the failure possibility is small from among the plurality of areas, the allocation of the jobs being performed the specific order of the jobs beginning with the job whose execution scale is the largest,

wherein the processor selects, when the allocation of the respective jobs is to be performed, an area that does not include the one or more nodes with the failure possibility and performs the allocation of the respective jobs to the selected area, and

wherein the processor selects, when selection of the area that does not include the one or more nodes with the failure possibility results in failure in regard to all of the plurality of areas, an area such that the number of the one or more nodes becomes minimum and performs the allocation of the job to the selected area.

2. The parallel processing apparatus according to claim 1 , wherein the plurality of nodes form a torus network.

3. The parallel processing apparatus according to claim 1 , wherein each of the one or more nodes with the failure possibility is a node whose hardware failure is foreseen from system logs recorded individually in the plurality of nodes.

4. A job management method, comprising:

calculating, by a computer, based on a number of nodes to be used in execution of respective jobs that are waiting to be executed and a scheduled execution time period for execution of the respective jobs, an execution scale of the respective jobs;

sorting the jobs in descending order of the execution scale to acquire a specific order of the jobs;

acquiring a plurality of areas by dividing a node area in which a plurality of nodes are arranged, some of the plurality of areas including one or more nodes with a failure possibility;

starting an allocation of the respective jobs from an area in which a number of the one or more nodes with the failure possibility is small from among the plurality of areas, the allocation of the jobs being performed in the specific order of the jobs beginning with the job whose execution scale is the largest;

selecting, when the allocation of the respective jobs is to be performed, an area that does not include the one or more nodes with the failure possibility; and performing the allocation of the respective jobs to the selected area; and

selecting, when selection of the area that does not include the one or more nodes with the failure possibility results in failure in regard to all of the plurality of areas, an area such that the number of the one or more nodes becomes minimum and performs the allocation of the job to the selected area.

5. The job management method according to claim 4 , wherein the plurality of nodes form a torus network.

6. The job management method according to claim 4 , wherein each of the one or more nodes with the failure possibility is a node whose hardware failure is foreseen from system logs recorded individually in the plurality of nodes.

7. A non-transitory computer-readable recording medium recording a job management program which causes a computer to perform operations, the operations comprising:

calculating, based on a number of nodes to be used in execution of respective jobs that are waiting to be executed and a scheduled execution time period for execution of the respective jobs, an execution scale of the respective, jobs;

sorting the jobs in descending order of the execution scale to acquire a specific order of the jobs;

acquiring a plurality of areas by dividing a node area in which a plurality of nodes are arranged, some of the plurality of areas including one or more nodes with a failure possibility; and

starting an allocation of the respective jobs from an area in which a number of the one or more nodes with the failure possibility is small from among the plurality of areas, the allocation of the jobs being performed in the specific order of the jobs beginning with the job whose execution scale is the largest;

selecting, when the allocation of the respective jobs is to be performed, an area that does not include the one or more nodes with the failure possibility;

performing the allocation of the respective jobs to the selected area; and

selecting, when selection of the area that does not include the one or more nodes with the failure possibility results in failure in regard to all of the plurality of areas, an area such that the number of the one or more nodes becomes minimum and performs the allocation of the job to the selected area.

8. The non-transitory computer-readable recording medium to claim 7 , wherein the plurality of nodes form a torus network.

9. The non-transitory computer-readable recording medium according to claim 7 , wherein each of the one or more nodes with the failure possibility is a node whose hardware failure is foreseen from system logs recorded individually in the plurality of nodes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2018
From: KOKUBO, RYOSUKE; HASHIMOTO, TSUYOSHI
To: FUJITSU LIMITED
Reel/Frame 045656/0099 →
Priority Claims (1)
JP JP2017-095200 · May 12, 2017 · national
Continuity (1)
Related Publication 20180329752A1 · Nov 15, 2018
Cited By (1)
US 12,650,877