Resource management in a virtual machine cluster
Managing resources in a VM cluster; allocation of resources among competing VMs. Assigning VMs to real devices, responsive to needs for resources: processor usage, disk I/O, network I/O, memory space, disk space. Assigning VMs responsive to needs for cluster activity: network response latency, QoS. Using predictive models of resource usage by VMs. Transferring VMs, improving utilization. Failing-soft onto alternative VM resources. Providing a resource buffer for collective VM resource demand. Transferring VMs to a cloud service that charges for resources.
1 . A method, including steps of
identifying a category of application programs into which a selected virtual machine fits;
identifying a current workplace usage associated with a distributed fault-tolerant multiprocessor cluster, to provide a preferred set of real devices on which to execute the selected virtual machine;
transferring one or more virtual machines, including the selected virtual machine, between real devices in the distributed fault-tolerant multiprocessor cluster, in response to a set of time-changing resource requirements;
measuring use of a plurality of different resources by the selected virtual machine;
providing a model of future actual resource usage by the selected virtual machine in response to the steps of measuring, wherein the steps of identifying a category are responsive to the model of future resource usage by the selected virtual machine; and
predicting said set of time-changing resource requirements in response to the current workplace usage and a result of the steps of identifying a category;
wherein the step of transferring includes steps of
halting a particular one virtual machine;
reallocating a new set of real devices to said particular one virtual machine; and
restarting said particular one virtual machine;
wherein said particular one virtual machine does not know the step of transferring has occurred.
2 . The method as in claim 1 , wherein the step of transferring one or more virtual machines is responsive to the set of time-changing resource requirements.
3 . The method as in claim 1 , further comprising predicting said set of time-changing resource requirements in response to past behavior of one or more said virtual machines.
4 . A distributed fault-tolerant multiprocessor cluster including one or more real devices disposed to execute one or more virtual machines, wherein the one or more real devices are disposed to collectively perform steps of
identifying a category of application programs into which a selected one or more of the virtual machines fits;
identifying a current workplace usage associated with the distributed fault-tolerant multiprocessor cluster, to provide a preferred set of the one or more real devices on which to execute the selected one or more of the virtual machines;
transferring one or more of the virtual machines, including the selected one or more virtual machines, between the one or more real devices in the distributed fault-tolerant multiprocessor cluster, in response to a set of time-changing resource requirements;
measuring use of a plurality of different resources by the selected one or more virtual machines;
providing a model of future actual resource usage by the selected virtual machine in response to the steps of measuring, wherein the steps of identifying a category are responsive to the model of future resource usage by the selected virtual machine; and
predicting said set of time-changing resource requirements in response to the current workplace usage and a result of the steps of identifying a category;
wherein the step of transferring includes steps of
halting a particular one virtual machine;
reallocating a new set of real devices to said particular one virtual machine; and
restarting said particular one virtual machine;
wherein said particular one virtual machine does not know the step of transferring has occurred.
5 . A distributed fault-tolerant multiprocessor cluster as in claim 4 , wherein the step of transferring one or more virtual machines is responsive to the set of time-changing resource requirements.
6 . A distributed fault-tolerant multiprocessor cluster as in claim 4 , wherein the steps further comprise predicting said set of time-changing resource requirements in response to past behavior of virtual machines.
7 . A method of operating a distributed fault-tolerant multiprocessor cluster, including steps of
measuring use of resources by one or more virtual machines in the distributed fault-tolerant multiprocessor cluster, the use of resources disposed to vary over a time duration;
generating at least one model of resource usage by each such one or more virtual machines, the model having a time-varying component; and
allocating resources to the one or more virtual machines in response to a time-varying prediction by the at least one model;
wherein the step of allocating resources includes steps of transferring one or more virtual machines between a first set of re-sources and a second set of resources available at the distributed fault-tolerant multiprocessor cluster, wherein the steps of transferring include steps of
halting a particular one virtual machine;
reallocating a new set of resources to said particular one virtual machine; and
restarting said particular one virtual machine;
wherein said particular one virtual machine does not know the step of transferring has occurred.
8 . A method as in claim 7 , wherein the model distinguishes between daily work and background applications.
9 . A method as in claim 7 , wherein the steps of generating include steps of determining a category into which a selected virtual machine fits.
10 . A method as in claim 7 , wherein the model includes a time of day during which the virtual machine is to execute, for which the time-varying prediction is made.
11 . A distributed fault-tolerant multiprocessor cluster including one or more real devices disposed to execute one or more virtual machines, wherein the one or more real devices perform at least steps of
measuring use of resources by the one or more virtual machines over a time duration;
generating at least one model of resource usage by each such one or more virtual machines; and
allocating resources to the one or more virtual machines in response to the at least one model;
wherein the step of allocating resources includes steps of transferring one or more virtual machines between a first set of re-sources and a second set of resources available at the distributed fault-tolerant multiprocessor cluster, wherein the step of transferring includes steps of
halting a particular one virtual machine;
reallocating a new set of resources to said particular one virtual machine; and
restarting said particular one virtual machine;
wherein said particular one virtual machine does not know the step of transferring has occurred.
12 . A distributed fault-tolerant multiprocessor cluster as in claim 11 , wherein the model distinguishes between daily work and background applications.
13 . A distributed fault-tolerant multiprocessor cluster as in claim 11 , wherein the steps of generating includes steps of determining a category into which a selected virtual machine fits.
14 . A distributed fault-tolerant multiprocessor cluster as in claim 11 , wherein the category includes a time of day.