IP Library Granted Patent US 12,333,343
Granted Patent B2
US 12,333,343 · App. 17/456,210 · Granted Jun 17, 2025

Avoidance of workload duplication among split-clusters

Inventors: Hai Hui Wang (Xian, CN); Shan Gao (Beijing, CN); Yang Gao (Xian, CN)
Assignee: International Business Machines Corporation
G06F9/505G06F9/542
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,343
App. No.
17/456,210
Granted
Jun 17, 2025
Kind
B2
Abstract

A computer implemented method avoids workload duplication in a cluster environment. The computer identifies a state change among a set of cluster resources in a cluster of nodes. Responsive to identifying the state change, the computer predicts resource requirements for a queued workload. The computer determines a pre-assignment of the queued workload to a sub-cluster according to the resource requirements that were predicted for the queued workload. The computer marks the queued workload to indicate the pre-assignment to the sub-cluster.

Claims (64)

1. A method for avoiding workload duplication in a cluster environment, the method comprising:

identifying, by a computer system, a state change among a set of cluster resources in a cluster of nodes;

responsive to identifying the state change, predicting, by the computer system, resource requirements for a queued workload;

determining, by the computer system, a pre-assignment of the queued workload to a sub-cluster according to the resource requirements that were predicted for the queued workload;

marking, by the computer system, the queued workload to indicate the pre-assignment to the sub-cluster; and

scheduling, by the sub-cluster of the computer system, the queued workload for execution on the sub-cluster, wherein the sub-cluster does not schedule queued workloads that are pre-assigned to other sub-clusters.

2. The method of claim 1 , wherein the steps of predicting resource requirements, determining the pre-assignment, and marking the queued workload are provided as a set of dispatch policies that are automatically enabled when the state change is identified among the set of cluster resources.

3. The method of claim 2 wherein the set of dispatch policies comprises custom and built-in policies that provide configurable rules for distinguishing the sub-cluster from other sub-cluster s.

4. The method of claim 1 , further comprising:

responsive to a failure of a network connection between the cluster of nodes, scheduling, by the computer system, the queued workload according to the pre-assignment.

5. The method of claim 4 , wherein the step of scheduling the queued workload further comprises:

identifying, by the sub-cluster of the computer system, the pre-assignment of the queued workload to the sub-cluster.

6. The method of claim 1 , wherein predicting the resource requirements further comprises:

modeling, by the computer system, the queued workload on a set of machine learning models trained on the resource requirements of a completed workload; and

predicting, by the computer system, the resource requirements of the queued workload from the set of machine learning models.

7. The method of claim 6 , wherein modeling the queued workload further comprises:

generating, by the computer system, metadata about the resource requirements of the completed workload;

collecting, by the computer system, the metadata into a training data set; and

training, by the computer system, the set of machine learning models from the training data set.

8. The method of claim 7 , wherein the metadata comprises:

a pending time of the completed workload;

a running time of the completed workload; and

a resource usage of the completed workload.

9. A computer system comprising:

a number of storage devices that store program instructions; and

a number of processor units in communication with the number of storage devices, wherein the number of processor units executes program instructions to:

identify a state change among a set of cluster resources in a cluster of nodes;

predict, in response to identifying the state change, resource requirements for a queued workload;

determine a pre-assignment of the queued workload to a sub-cluster according to the resource requirements that were predicted for the queued workload;

mark the queued workload to indicate the pre-assignment to the sub-cluster; and

scheduling, by the sub-cluster of the computer system, the queued workload for execution on the sub-cluster, wherein the sub-cluster does not schedule queued workloads that are pre-assigned to other sub-clusters.

10. The computer system of claim 9 , wherein the steps of predicting resource requirements, determining the pre-assignment, and marking the queued workload are provided as a set of dispatch policies that are automatically enabled when the state change is identified among the set of cluster resources.

11. The computer system of claim 10 , wherein the set of dispatch policies comprises custom and built-in policies that provide configurable rules for distinguishing the sub-cluster from other sub-clusters.

12. The computer system of claim 9 , wherein the processor further executes the program instructions to:

schedule, responsive to a failure of a network connection between the cluster of nodes, the queued workload according to the pre-assignment.

13. The computer system of claim 12 , wherein in scheduling the queued workload, the processor further executes the program instructions:

to identify the pre-assignment of the queued workload to the sub-cluster.

14. The computer system of claim 9 , wherein in predicting the resource requirements, the processor further executes the program instructions to:

model the queued workload on a set of machine learning models trained on the resource requirements of a completed workload; and

predict the resource requirements of the queued workload from the set of machine learning models.

15. The computer system of claim 14 , wherein in modeling the queued workload, the processor further executes the program instructions to:

generate metadata about the resource requirements of the completed workload;

collect the metadata into a training data set; and

train the set of machine learning models from the training data set.

16. The computer system of claim 15 , wherein the metadata comprises:

a pending time of the completed workload;

a running time of the completed workload; and

a resource usage of the completed workload.

17. A computer program product for avoiding workload duplication in a cluster, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer system to cause the computer system to perform a method of:

identifying a state change among a set of cluster resources in a cluster of nodes;

responsive to identifying the state change, predicting resource requirements for a queued workload;

determining a pre-assignment of the queued workload to a sub-cluster according to the resource requirements that were predicted for the queued workload;

marking the queued workload to indicate the pre-assignment to the sub-cluster; and

scheduling, by the sub-cluster of the computer system, the queued workload for execution on the sub-cluster, wherein the sub-cluster does not schedule queued workloads that are pre-assigned to other sub-clusters.

18. The computer program product of claim 17 , further comprising:

responsive to a failure of a network connection between the cluster of nodes, scheduling the queued workload according to the pre-assignment, wherein scheduling the queued workload further comprises:

identifying, by the sub-cluster, the pre-assignment of the queued workload to the sub-cluster.

19. The computer program product of claim 17 , wherein predicting the resource requirements further comprises:

modeling the queued workload on a set of machine learning models trained on the resource requirements of a completed workload; and

predicting the resource requirements of the queued workload from the set of machine learning models.

20. The computer program product of claim 19 , wherein modeling the queued workload further comprises:

generating metadata about the resource requirements of the completed workload;

collecting the metadata into a training data set; and

training the set of machine learning models from the training data set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 23, 2021
From: WANG, HAI HUI; GAO, SHAN; GAO, YANG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058193/0683 →
Continuity (1)
Related Publication 20230161633A1 · May 25, 2023
References Cited (23)
US 7953843B2 · Cherkasova · 2011 [cited by examiner]
US 8108715B1 · Agarwal et al. · 2012 [cited by applicant]
US 9460183B2 · Dalton · 2016 [cited by applicant]
US 10320703B2 · Gahlot · 2019 [cited by examiner]
US 10560315B2 · Yuan · 2020 [cited by applicant]
US 10956230B2 · Gopalan · 2021 [cited by examiner]
US 11169854B2 · Gururaj · 2021 [cited by examiner]
US 11593583B2 · Kallanagoudar · 2023 [cited by examiner]
US 12045667B2 · Wang et al. · 2024 [cited by applicant]
US 20130227359A1 · Butterworth · 2013 [cited by examiner]
US 20140130057A1 · Fu et al. · 2014 [cited by applicant]
US 20160283270A1 · Amaral et al. · 2016 [cited by applicant]
US 20180300174A1 · Karanasos et al. · 2018 [cited by applicant]
US 20200250002A1 · Gururaj et al. · 2020 [cited by applicant]
US 20200410284A1 · Kallanagoudar et al. · 2020 [cited by applicant]
CN 102308559A · 2012 [cited by applicant]
CN 102394914A · 2012 [cited by applicant]
CN 104158707A · 2014 [cited by applicant]
CN 112463390A · 2021 [cited by applicant]
WO 2017211042A1 · 2017 [cited by applicant]
WO 2023093354A1 · 2023 [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing,” Computer Security Division, National Institute of Standards and Technology, Jan. 2011, 7 pages. [cited by applicant]
PCT International Search Report and Written Opinion, dated Nov. 28, 2022, regarding Application No. PCT/CN2022/125396, 9 pages. [cited by applicant]