Reactive semi-analytical orchestration for reproducing failures in container-based data protection systems
A method for managing a data protection system includes obtaining, by an interaction manager, a failure notification for a data protection workload executing on the data protection system, wherein the data protection system comprises a plurality of services, and wherein the data protection workload is executed by the plurality of services communicating with each other using application programming interface (API) calls, based on the failure notification, identifying a portion of the plurality of services used for the data protection workload, performing a recreation of the data protection workload to identify a failure point associated with the failure notification, and performing a remediation action of the data protection workload based on the recreation.
1 . A method for managing a data protection system, the method comprising:
obtaining, by an interaction manager, a failure notification for a data protection workload executing on the data protection system,
wherein the data protection system comprises a plurality of services, and
wherein the data protection workload is executed by the plurality of services communicating with each other using application programming interface (API) calls;
based on the failure notification, identifying a portion of the plurality of services used for the data protection workload;
performing a recreation of the data protection workload to identify a failure point associated with the failure notification; and
performing a remediation action of the data protection workload based on the recreation.
2 . The method of claim 1 , wherein the recreation is performed by simulating the portion of the plurality of services to obtain a simulated environment and simulating the API calls in the simulated environment to identify which of the API calls indicates the failure point.
3 . The method of claim 1 , wherein the remediation action comprises one of: restarting the data protection workload in the data protection system, replacing at least one service of the portion of the plurality of services based on the failure point, issuing a notification to an administrator of the failure point, and modifying a sequence of the API calls.
4 . The method of claim 1 , wherein performing the remediation action comprises applying the failure point to a failure prediction engine, and wherein the failure prediction engine determines the remediation action based on the failure point.
5 . The method of claim 4 , further comprising:
updating the failure prediction engine based on the remediation action by updating an action-reward associated with the remediation action and applying a reinforcement learning on the failure prediction engine based on the action-reward.
6 . The method of claim 1 , wherein the plurality of services comprises: an asset backup service, a pillar service, an asset restoration service, a scheduler, and a user interface communicating with a client device.
7 . The method of claim 1 , wherein the failure point is based on one of a list consisting of: one of the API calls being out of order, an overloading of one of the plurality of services, and a detected latency issue.
8 . A non-transitory computer readable medium comprising computer readable program code, which when executed by a computer processor enables the computer processor to perform a method for managing a data protection system, the method comprising:
obtaining, by an interaction manager, a failure notification for a data protection workload executing on the data protection system,
wherein the data protection system comprises a plurality of services, and
wherein the data protection workload is executed by the plurality of services communicating with each other using application programming interface (API) calls;
based on the failure notification, identifying a portion of the plurality of services used for the data protection workload;
performing a recreation of the data protection workload to identify a failure point associated with the failure notification; and
performing a remediation action of the data protection workload based on the recreation.
9 . The non-transitory computer readable medium of claim 8 , wherein the recreation is performed by simulating the portion of the plurality of services to obtain a simulated environment and simulating the API calls in the simulated environment to identify which of the API calls indicates the failure point.
10 . The non-transitory computer readable medium of claim 8 , wherein the remediation action comprises one of: restarting the data protection workload in the data protection system, replacing at least one service of the portion of the plurality of services based on the failure point, issuing a notification to an administrator of the failure point, and modifying a sequence of the API calls.
11 . The non-transitory computer readable medium of claim 8 , wherein performing the remediation action comprises applying the failure point to a failure prediction engine, and wherein the failure prediction engine determines the remediation action based on the failure point.
12 . The non-transitory computer readable medium of claim 11 , further comprising:
updating the failure prediction engine based on the remediation action by updating an action-reward associated with the remediation action and applying a reinforcement learning on the failure prediction engine based on the action-reward.
13 . The non-transitory computer readable medium of claim 8 , wherein the plurality of services comprises: an asset backup service, a pillar service, an asset restoration service, a scheduler, and a user interface communicating with a client device.
14 . The non-transitory computer readable medium of claim 8 , wherein the failure point is based on one of a list consisting of: one of the API calls being out of order, an overloading of one of the plurality of services, and a detected latency issue.
15 . A system comprising:
an interaction manager, operating on a processor; and
memory comprising instructions, which when executed by the processor, perform a method comprising:
obtaining a failure notification for a data protection workload executing on a data protection system,
wherein the data protection system comprises a plurality of services, and
wherein the data protection workload is executed by the plurality of services communicating with each other using application programming interface (API) calls;
based on the failure notification, identifying a portion of the plurality of services used for the data protection workload;
performing a recreation of the data protection workload to identify a failure point associated with the failure notification,
wherein the recreation is performed by simulating the portion of the plurality of services to obtain a simulated environment and simulating the API calls in the simulated environment to identify which of the API calls indicates the failure point; and
performing a remediation action of the data protection workload based on the recreation.
16 . The system of claim 15 , wherein the remediation action comprises one of: restarting the data protection workload in the data protection system, replacing at least one service of the portion of the plurality of services based on the failure point, issuing a notification to an administrator of the failure point, and modifying a sequence of the API calls.
17 . The system of claim 15 , wherein performing the remediation action comprises applying the failure point to a failure prediction engine, and wherein the failure prediction engine determines the remediation action based on the failure point.
18 . The system of claim 17 , further comprising:
updating the failure prediction engine based on the remediation action by updating an action-reward associated with the remediation action and applying a reinforcement learning on the failure prediction engine based on the action-reward.
19 . The system of claim 15 , wherein the plurality of services comprises: an asset backup service, a pillar service, an asset restoration service, a scheduler, and a user interface communicating with a client device.
20 . The system of claim 15 , wherein the failure point is based on one of a list consisting of: one of the API calls being out of order, an overloading of one of the plurality of services, and a detected latency issue.