Managing an application's resource stack
Managing an application's resource stack, including: detecting, in dependence upon one or more storage system metrics, an occurrence of a storage system performance anomaly; and responsive to detecting the storage system performance anomaly, identifying, in dependence upon codified relationships between one or more storage system metrics and one or more elements in the application stack that are external to the storage system, a root cause of the storage system performance anomaly.
1. A method implemented by a computing device, the method comprising:
predicting, based on one or more storage system metrics, an occurrence of a performance anomaly affecting a performance of a storage system;
responsive to the prediction, automatically identifying, by the computing device, based on codified relationships between one or more storage system metrics and external metrics from one or more execution environments of an application stack that is external to the storage system, a root cause of the predicted performance anomaly; and
based on the identified root cause, automatically causing, by the computing device, the execution environment external to the storage system to redirect execution of operations to an other storage system based on the determination of the root cause.
2. The method of claim 1 wherein identifying, in dependence upon codified relationships between one or more storage system metrics and one or more elements in the application stack that are external to the storage system, a root cause of the storage system performance anomaly further comprises retrieving one or more upstack performance metrics from the one or more elements in the application stack that are external to the storage system.
3. The method of claim 1 wherein initiating remedial actions associated with the identified root cause further comprises migrating a workload.
4. The method of claim 3 wherein migrating the workload further comprises migrating an application that is executing on a virtual machine that is supported by a first host to a second host that is coupled to a second storage system.
5. The method of claim 1 further comprising:
receiving one or more storage system metrics from a plurality of storage systems;
receiving metrics associated with one or more elements in the application stack that are external to the storage system; and
identifying relationships between one or more storage system metrics, metrics associated with and one or more elements in the application stack that are external to the storage system, and a root cause.
6. The method of claim 1 wherein the identifying of the root cause of the storage system performance anomaly uses one or more machine learning models.
7. An apparatus comprising a computer processor, a computer memory operatively coupled to the computer processor, the computer memory having disposed within it computer program instructions that, when executed by the computer processor, cause the apparatus to carry out the steps of:
predicting, based on one or more storage system metrics, an occurrence of a performance anomaly affecting a performance of a storage system;
responsive to the prediction, automatically identifying, by the computing device, based on codified relationships between one or more storage system metrics and external metrics from one or more execution environments of an application stack that is external to the storage system, a root cause of the predicted performance anomaly; and
based on the identified root cause, automatically causing, by the computing device, the execution environment external to the storage system to redirect execution of operations to an other storage system based on the determination of the root cause.
8. The apparatus of claim 7 wherein identifying, in dependence upon codified relationships between one or more storage system metrics and one or more elements in the application stack that are external to the storage system, a root cause of the storage system performance anomaly further comprises retrieving one or more upstack performance metrics from the one or more elements in the application stack that are external to the storage system.
9. The apparatus of claim 8 wherein initiating remedial actions associated with the identified root cause further comprises migrating a workload.
10. The apparatus of claim 9 wherein migrating the workload further comprises migrating an application that is executing on a virtual machine that is supported by a first host to a second host that is coupled to a second storage system.
11. The apparatus of claim 7 further comprising computer program instructions that, when executed by the computer processor, cause the apparatus to carry out the steps of:
receiving one or more storage system metrics from a plurality of storage systems;
receiving metrics associated with one or more elements in the application stack that are external to the storage system; and
identifying relationships between one or more storage system metrics, metrics associated with and one or more elements in the application stack that are external to the storage system, and a root cause.
12. The apparatus of claim 7 wherein the identifying of the root cause of the storage system performance anomaly uses one or more machine learning models.
13. A non-transitory computer readable medium comprising computer program instructions that, when executed, cause a computer to carry out the steps of:
predicting, based on one or more storage system metrics, an occurrence of a performance anomaly affecting a performance of a storage system;
responsive to the prediction, automatically identifying, by the computing device, based on codified relationships between one or more storage system metrics and external metrics from one or more execution environments of an application stack that is external to the storage system, a root cause of the predicted performance anomaly; and
based on the identified root cause, automatically causing, by the computing device, the execution environment external to the storage system to redirect execution of operations to an other storage system based on the determination of the root cause.
14. The non-transitory computer readable medium of claim 13 wherein identifying, in dependence upon codified relationships between one or more storage system metrics and one or more elements in the application stack that are external to the storage system, a root cause of the storage system performance anomaly further comprises retrieving one or more upstack performance metrics from the one or more elements in the application stack that are external to the storage system.
15. The non-transitory computer readable medium of claim 13 wherein initiating remedial actions associated with the identified root cause further comprises migrating a workload.
16. The non-transitory computer readable medium of claim 15 wherein migrating the workload further comprises migrating an application that is executing on a virtual machine that is supported by a first host to a second host that is coupled to a second storage system.
17. The non-transitory computer readable medium of claim 13 wherein the identifying of the root cause of the storage system performance anomaly uses one or more machine learning models.