Method and system for managing resource utilization based on reinforcement learning
The present teaching relates to managing computing resources. In one example, information about resource utilization on a computing node is received from the computing node. Available resource on the computing node is determined based on the information. A model generated in accordance with reinforcement learning based on simulated training data is obtained. An adjusted available resource is generated based on the available resource and the model with respect to the computing node. The adjusted available resource is sent to a scheduler for scheduling one or more jobs to be executed on the computing node based on the adjusted available resource.
1 . A method for managing computing resources, the method comprising:
running, by a computing node, a first job;
sending, to a resource management engine, container information about a container in the computing node for running the first job, wherein the container information indicates a maximum usage of the container, a minimum usage of the container, an average usage of the container, and a current usage of the container;
determining current available computing resources on the computing node based on the container information and historical information indicating past available computing resources on the computing node; and
receiving, from a job scheduler distinct from the resource management engine, a second job to be launched on the computing node, wherein the second job is determined by the job scheduler based on an adjusted available computing resources for the computing node to be assigned with a new job, the adjusted available computing resources are determined by the resource management engine according to a machine-trained model for adjusting the current available computing resources to become the adjusted available computing resources of the computing node in assigning a new job to the computing node to handle, and the job scheduler that assigns jobs has no knowledge that the current available computing resources have been adjusted by the resource management engine to become the adjusted available computing sources.
2 . The method of claim 1 , wherein the container comprises a reserved space on the computing node for running the first job.
3 . The method of claim 1 , further comprising:
generating a new container to run the second job.
4 . The method of claim 1 , wherein the machine-trained model is trained based on training data generated via simulation based on raw training data associated with records of previous resource adjustments with respect to the computing node.
5 . The method of claim 1 , wherein the machine-trained model is trained based on an aggressiveness score with values indicative of aggressiveness of the machine-trained model in controlling adjusting the available computing resource of the computing node in assigning a job to the computing node to handle.
6 . The method of claim 5 , wherein the aggressiveness indicates whether the machine-trained model causes loss of jobs at the computing node or waste of computing resources at the computing node.
7 . The method of claim 1 , wherein the machine-trained model is trained by maximizing a score associated with a fitness function of the machine-trained model.
8 . A non-transitory, computer-readable medium having information recorded thereon for managing computing resources, when read by at least one processor, effectuate operations comprising:
running, by a computing node, a first job;
sending, to a resource management, container information about a container in the computing node for running the first job, wherein the container information indicates a maximum usage of the container, a minimum usage of the container, an average usage of the container, and a current usage of the container;
determining current available computing resources on the computing node based on the container information and historical information indicating past available computing resources on the computing node; and
receiving, from a job scheduler distinct from the resource management engine, a second job to be launched on the computing node, wherein the second job is determined by the job scheduler based on an adjusted available computing resources for the computing node to be assigned with a new job, the adjusted available computing resources are determined by the resource management engine according to a machine-trained model for adjusting the current available computing resources to become the adjusted adjusting available computing resources of the computing node in assigning a new job to the computing node to handle, and the job scheduler that assigns jobs has no knowledge that the current available computing resources have been adjusted by the resource management engine to become the adjusted available computing sources.
9 . The medium of claim 8 , wherein the container comprises a reserved space on the computing node for running the first job.
10 . The medium of claim 8 , wherein the operations further comprise:
generating a new container to run the second job.
11 . The medium of claim 8 , wherein the machine-trained model is trained based on training data generated via simulation based on raw training data associated with records of previous resource adjustments with respect to the computing node.
12 . The medium of claim 8 , wherein the machine-trained model is trained based on an aggressiveness score with values indicative of aggressiveness of the machine-trained model in controlling adjusting the available computing resource of the computing node in assigning a job to the computing node to handle.
13 . The medium of claim 12 , wherein the aggressiveness indicates whether the machine-trained model causes loss of jobs at the computing node or waste of computing resources at the computing node.
14 . The medium of claim 8 , wherein the machine-trained model is trained by maximizing a score associated with a fitness function of the machine-trained model.
15 . A system for managing computing resources, the system comprising:
memory storing computer program instructions; and
one or more processors that, in response to executing the computer program instructions, effectuate operations comprising:
running, by a computing node, a first job;
sending, to a resource management engine, container information about a container in the computing node for running the first job, wherein the container information indicates a maximum usage of the container, a minimum usage of the container, an average usage of the container, and a current usage of the container;
determining current available computing resources on the computing node based on the container information and historical information indicating past available computing resources on the computing node; and
receiving, from a job scheduler distinct from the resource management engine, a second job to be launched on the computing node, wherein the second job is determined by the job scheduler based on an adjusted available computing resources for the computing node to be assigned with a new job, the adjusted available computing resources are determined by the resource management engine according to a machine-trained model for adjusting the current available computing resources to become the adjusted adjusting available computing resources of the computing node in assigning a new job to the computing node to handle, and the job scheduler that assigns jobs has no knowledge that the current available computing resources have been adjusted by the resource management engine to become the adjusted available computing sources.
16 . The system of claim 15 , wherein the container comprises a reserved space on the computing node for running the first job.
17 . The system of claim 15 , wherein the operations further comprise:
generating a new container to run the second job.
18 . The system of claim 15 , wherein the machine-trained model is trained based on training data generated via simulation based on raw training data associated with records of previous resource adjustments with respect to the computing node.
19 . The system of claim 15 , wherein the machine-trained model is trained based on an aggressiveness score with values indicative of aggressiveness of the machine-trained model in controlling adjusting the available computing resource of the computing node in assigning a job to the computing node to handle.
20 . The system of claim 19 , wherein the aggressiveness indicates whether the machine-trained model causes loss of jobs at the computing node or waste of computing resources at the computing node.