Workload management among candidate distributed computing environments based on governance policies
Aspects of the present disclosure provide systems, methods, and computer-readable storage media that support workload management. A computing device may generate recommendation information for processing of the workload, the recommendation information associated with a first execution environment selected from the set of candidate execution environments based on execution information that indicates, for each candidate execution environment of the set of candidate execution environments, a carbon intensity value associated with the candidate execution environment. The computing device may output a first indicator that indicates the recommendation information.
1 . A method for workload management performed by one or more processors, the method comprising:
training, by the one or more processors, a Named Entity Recognition (NER) model on a corpus of governance policies to identify entities including scheduling entities and provisioning entities,
identifying, by the one or more processors, a governance policy associated with a workload within distributed computing environments, wherein the distributed computing environments include one or more computing devices in cloud computing environments, and
wherein the distributed computing environment allocates computing resources to the workload for execution;
identifying, by the one or more processors based on the governance policy, a set of candidate execution environments for processing the workload;
identifying, by the one or more processors, allowed locations based on the governance policy and identifying the set of candidate execution environments based on the allowed locations,
wherein the workload includes a batch job, the set of candidate execution environments, each candidate execution environment of the set of candidate execution environments is a different cloud computing environment, and the set of candidate execution environments includes multiple candidate execution environments;
generating, by the one or more processors, recommendation information for processing of the workload, the recommendation information associated with a first execution environment selected from the set of candidate execution environments based on execution information that indicates, for each candidate execution environment of the set of candidate execution environments, a carbon intensity value associated with the candidate execution environment;
outputting, by the one or more processors, a first indicator that indicates the recommendation information; and
transferring, by the one or more processors, the workload from a storage location of the workload to the first execution environment.
2 . The method of claim 1 , further comprising:
determining the execution information for the set of candidate execution environments;
for each candidate execution environment of the set of candidate execution environments, determining the execution information associated with the candidate execution environment, the execution information associated with the candidate execution environment indicates:
a predicated carbon intensity value for the candidate execution environment,
a transfer energy costs associated with transferring the workload to the candidate execution environment, the transfer energy cost based on a location of the candidate execution environment, a distance between the first execution environment and the candidate execution environment, network information, or a combination thereof;
a carbon intensity average value based on a duration associated with a duration of processing the workload at the candidate execution environment, a power usage effectiveness value of the candidate execution environment, or a combination thereof; or
a combination thereof; and
selecting, based on the execution information, the first execution environment from the set of candidate execution environments, the first execution environment further selected based on scheduling information associated with the set of candidate execution environments, historic workload scheduling data, or a combination thereof.
3 . The method of claim 2 , wherein:
for each candidate execution environment of the set of candidate execution environments, determining the execution information associated with the candidate execution environment includes determining the carbon intensity average value for each time of multiple processing start times; and
the recommendation information is generated based on:
multiple workloads that includes the workload,
the set of candidate execution environments,
scheduling information associated with each candidate execution environment of the set of candidate execution environments,
for each candidate execution environment of the set of candidate execution environments:
the carbon intensity value of the candidate execution environment,
an infrastructure cost of the candidate execution environment,
a provisioning time of the candidate execution environment,
a service level agreement or latency associated with the candidate execution environment, or
a combination thereof,
a cost threshold or an energy efficiency threshold,
a weather prediction, or
a combination thereof.
4 . The method of claim 1 , further comprising:
receiving a second indicator that indicates an acceptance of the first execution environment for processing of the workload; and
based on the second indicator:
generating a schedule of the workload to indicate the processing of the workload at the first execution environment at a first time.
5 . The method of claim 4 , further comprising:
provisioning the first execution environment for the processing of the workload; and
initiating the processing of the workload by the first execution environment.
6 . The method of claim 1 , further comprising:
receiving scheduling information that indicates that the workload that is scheduled for processing at a second execution environment at a second time;
selecting the first execution environment from the set of candidate execution environments;
transferring the workload from the second execution environment to the first execution environment; and
initiating processing of the workload by the first execution environment.
7 . A system for workload management, the system comprising:
a memory; and
one or more processors communicatively coupled to the memory, the one or more processors configured to:
train a Named Entity Recognition (NER) model on a corpus of governance policies to identify entities including scheduling entities and provisioning entities,
identify a governance policy associated with a workload within distributed computing environments, wherein the distributed computing environments include one or more computing devices in cloud computing environments, and
wherein the distributed computing environment allocates computing resources to the workload for execution;
identify based on the governance policy, a set of candidate execution environments for processing the workload;
identify allowed locations based on the governance policy and identify the set of candidate execution environments based on the allowed locations,
wherein the workload includes a batch job, the set of candidate execution environments, each candidate execution environment of the set of candidate execution environments is a different cloud computing environment, and the set of candidate execution environments includes multiple candidate execution environments;
generate recommendation information for processing of the workload, the recommendation information associated with a first execution environment selected from the set of candidate execution environments based on execution information that indicates, for each candidate execution environment of the set of candidate execution environments, a carbon intensity value associated with the candidate execution environment;
output a first indicator that indicates the recommendation information; and
transfer the workload from a storage location of the workload to the first execution environment.
8 . The system of claim 7 , wherein the one or more processors configured to, for each candidate execution environment of the set of candidate execution environments, receive a predicated carbon intensity value for the candidate execution environment.
9 . The system of claim 7 , wherein the one or more processors configured to, for each candidate execution environment of the set of candidate execution environments, determine a transfer energy costs associated with transferring the workload to the candidate execution environment, the transfer energy cost based on a location of the candidate execution environment, a distance between the first execution environment and the candidate execution environment, network information, or a combination thereof.
10 . The system of claim 7 , wherein the one or more processors configured to, for each candidate execution environment of the set of candidate execution environments, determine a carbon intensity average value based on a duration associated with a duration of processing the workload at the candidate execution environment, a power usage effectiveness value of the candidate execution environment, or a combination thereof.
11 . The system of claim 10 , wherein, to determine the execution information for each candidate execution environment of the set of execution environments, the one or more processors is further configured to determine the carbon intensity average value for each time of multiple processing start times.
12 . The system of claim 7 , wherein the one or more processors configured to select, based on the execution information, the first execution environment from the set of candidate execution environments, the first execution environment further selected based on scheduling information associated with the set of candidate execution environments and historic workload scheduling data.
13 . The system of claim 7 , wherein the recommendation information is generated based on, for each candidate execution environment of the set of candidate execution environments:
the carbon intensity value of the candidate execution environment,
an infrastructure cost of the candidate execution environment,
a provisioning time of the candidate execution environment, and
a service level agreement or latency associated with the candidate execution environment.
14 . A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations for application portfolio management, the operations comprising:
training a Named Entity Recognition (NER) model on a corpus of governance policies to identify entities including scheduling entities and provisioning entities,
identifying a governance policy associated with a workload within distributed computing environments, wherein the distributed computing environments include one or more computing devices in cloud computing environments, and
wherein the distributed computing environment allocates computing resources to the workload for execution;
identifying based on the governance policy, a set of candidate execution environments for processing the workload;
identifying allowed locations based on the governance policy and identifying the set of candidate execution environments based on the allowed locations,
wherein the workload includes a batch job, the set of candidate execution environments, each the candidate execution environment of the set of candidate execution environments is a different cloud computing environment, and the set of candidate execution environments includes multiple candidate execution environments;
generating recommendation information for processing of the workload, the recommendation information associated with a first execution environment selected from the set of candidate execution environments based on execution information that indicates, for each candidate execution environment of the set of candidate execution environments, a carbon intensity value associated with the candidate execution environment;
outputting a first indicator that indicates the recommendation information; and
transferring the workload from a storage location of the workload to the first execution environment.
15 . The non-transitory computer-readable storage medium of claim 14 , the operations further comprising:
determining the execution information for the set of candidate execution environments;
for each candidate execution environment of the set of candidate execution environments, determining the execution information associated with the candidate execution environment, the execution information associated with the candidate execution environment indicates:
a predicated carbon intensity value for the candidate execution environment,
a transfer energy costs associated with transferring the workload to the candidate execution environment, the transfer energy cost based on a location of the candidate execution environment, a distance between the first execution environment and the candidate execution environment, network information, or a combination thereof; and
a carbon intensity average value based on a duration associated with a duration of processing the workload at the candidate execution environment, a power usage effectiveness value of the candidate execution environment, or a combination thereof.
16 . The non-transitory computer-readable storage medium of claim 14 , wherein the recommendation information is generated based on:
multiple workloads that includes the workload,
the set of candidate execution environments,
scheduling information associated with each candidate execution environment of the set of candidate execution environments,
for each candidate execution environment of the set of candidate execution environments:
the carbon intensity value of the candidate execution environment,
an infrastructure cost of the candidate execution environment,
a provisioning time of the candidate execution environment,
a service level agreement or latency associated with the candidate execution environment, or
a combination thereof,
a cost threshold or an energy efficiency threshold, and
a weather prediction.
17 . The non-transitory computer-readable storage medium of claim 14 , the operations further comprising:
receiving a second indicator that indicates an acceptance of the first execution environment for processing of the workload;
based on the second indicator:
generating a schedule of the workload to indicate the processing of the workload at the first execution environment at a first time;
provisioning the first execution environment for the processing of the workload; and
initiating the processing of the workload by the first execution environment.
18 . The non-transitory computer-readable storage medium of claim 14 , the operations further comprising:
receiving scheduling information that indicates that the workload that is scheduled for processing at a second execution environment at a second time;
selecting the first execution environment from the set of candidate execution environments;
transferring the workload from the second execution environment to the first execution environment; and
initiating processing of the workload by the first execution environment.