Intelligent scheduling of backups
Normal virtual machine operation is observed to automatically determine patterns of resource utilization. Backup activities are then scheduled, taking into account these utilization patterns. For example, if a normally scheduled backup would occur during a busy period, it may be rescheduled to a less busy period. As another example, backups made by made opportunistically during less busy periods even if not required by the normal backup schedule, in order to alleviate backup demands during more busy periods.
1. A method for scheduling a backup of a set of virtual machines, the method comprising:
monitoring resource utilization metrics of the set of virtual machines;
training a model based on the monitored resource utilization metrics;
using the trained model, determining a plurality of opportunistic windows of reduced resource utilization for the set of virtual machines and a plurality of blackout windows of increased resource utilization for the set of virtual machines;
identifying, based at least in part on a service schedule for the set of virtual machines, that the backup for the set of virtual machines would occur at least partially during a blackout window of the plurality of blackout windows determined via the trained model;
identifying an opportunistic window, of the plurality of opportunistic windows determined via the trained model, as a substitute for the blackout window; and
scheduling, based at least in part on identifying that the backup would occur at least partially during the blackout window, the backup of the set of virtual machines to occur during the opportunistic window identified as the substitute for the blackout window.
2. The method of claim 1 , wherein monitoring resource utilization metrics by the set of virtual machines includes running a continuous analytic process to observe and learn resource utilization patterns of a hypervisor managing the set of virtual machines.
3. The method of claim 2 , wherein the resource utilization patterns identify cyclical gaps in the resource utilization to identify and predict the plurality of opportunistic windows of reduced resource utilization.
4. The method of claim 3 , wherein incremental snapshots of the set of virtual machines are scheduled to be performed during one or more of the plurality of opportunistic windows of reduced resource utilization.
5. The method of claim 2 , wherein the resource utilization patterns identify cyclical gaps in the resource utilization to identify and predict the plurality of blackout windows of high resource contention.
6. The method of claim 5 , wherein scheduling the backup further comprises:
deferring the backup past the blackout window while still meeting a maintenance service schedule.
7. The method of claim 1 , further comprising updating the model as new resource utilization metrics are collected.
8. A system for scheduling a backup of a set of virtual machines, the system comprising:
processors; and
a memory storing instructions that, when executed by at least one processor among the processors, cause the system to perform operations comprising, at least:
monitoring resource utilization metrics of the set of virtual machines;
training a model based on the monitored resource utilization metrics;
using the trained model, determining a plurality of opportunistic windows of reduced resource utilization for the set of virtual machines and a plurality of blackout windows of increased resource utilization for the set of virtual machines;
identifying, based at least in part on a service schedule for the set of virtual machines, that the backup for the set of virtual machines would occur at least partially during a blackout window of the plurality of blackout windows determined via the trained model;
identifying an opportunistic window, of the plurality of opportunistic windows determined via the trained model, as a substitute for the blackout window; and
scheduling, based at least in part on identifying that the backup would occur at least partially during the blackout window, the backup of the set of virtual machines to occur during the opportunistic window identified as the substitute for the blackout window.
9. The system of claim 8 , wherein monitoring resource utilization metrics by the set of virtual machines includes running a continuous analytic process to observe and learn resource utilization patterns of a hypervisor managing the set of virtual machines.
10. The system of claim 9 , wherein the resource utilization patterns identify cyclical gaps in the resource utilization to identify and predict the plurality of opportunistic windows of reduced resource utilization.
11. The system of claim 10 , wherein the operations further comprise scheduling incremental snapshots of the set of virtual machines for performance during one or more of the plurality of opportunistic windows of reduced resource utilization.
12. The system of claim 9 , wherein the resource utilization patterns identify cyclical gaps in the resource utilization to identify and predict the plurality of blackout windows of high resource contention.
13. The system of claim 12 , wherein the operations to schedule the backup further comprise:
deferring the backup past the blackout window while still meeting a maintenance service schedule.
14. The system of claim 8 , wherein the operations further comprise updating the model as new resource utilization metrics are collected.
15. A non-transitory machine-readable medium including instructions which, when read by a machine, cause the machine to perform operations including, at least:
monitoring resource utilization metrics of a set of virtual machines;
training a model based on the monitored resource utilization metrics;
using the trained model, determining a plurality of opportunistic windows of reduced resource utilization for the set of virtual machines and a plurality of blackout windows of increased resource utilization for the set of virtual machines;
identifying, based at least in part on a service schedule for the set of virtual machines, that a backup for the set of virtual machines would occur at least partially during a blackout window of the plurality of blackout windows determined via the trained model:
identifying an opportunistic window, of the plurality of opportunistic windows determined via the trained model, as a substitute for the blackout window; and
scheduling, based at least in part on identifying that the backup would occur at least partially during the blackout window, the backup of the set of virtual machines to occur during the opportunistic window identified as the substitute for the blackout window.
16. The medium of claim 15 , wherein monitoring resource utilization metrics by the set of virtual machines includes running a continuous analytic process to observe and learn resource utilization patterns of a hypervisor managing the set of virtual machines.
17. The medium of claim 16 , wherein the resource utilization patterns identify cyclical gaps in the resource utilization to identify and predict the plurality of opportunistic windows of reduced resource utilization.
18. The medium of claim 17 , wherein the operations further comprise scheduling incremental snapshots of the set of virtual machines for performance during one or more of the plurality of opportunistic windows of reduced resource utilization.
19. The medium of claim 16 , wherein the resource utilization patterns identify cyclical gaps in the resource utilization to identify and predict the plurality of blackout windows of high resource contention.
20. The medium of claim 19 , wherein the operations to schedule the backup further comprise:
deferring the backup past the blackout window while still meeting a maintenance service schedule.
21. The medium of claim 15 , wherein the operations further comprise updating the model as new resource utilization metrics are collected.