Intelligent management of cloud services
There provided is a method for managing cloud service, wherein the method include detecting, at client nodes, anomalies based on a predetermined set of rules; determining, by the client nodes, whether to upload one or more anomalies of the detected anomalies to a topology aggregation service, in response to a decision not to upload the one or more anomalies to the topology aggregation service, automatically selecting an action from a plurality of available actions, by the client nodes, wherein the action is selected based at least in part on the service level associated with the entity that is experiencing the one or more anomalies, and undertaking the selected action; and in response to a decision to upload the one or more anomalies to the topology aggregation service, transmitting the one or more anomalies to the topology aggregation service.
1 . A method for managing a cloud service, comprising:
detecting, at client nodes associated with the cloud service, anomalies in the cloud service based on a predetermined set of rules;
determining, by the client nodes, whether to upload one or more anomalies of the detected anomalies to a topology aggregation service configured to collect information about a topology of the cloud service and provide information regarding a structure and an operation of the cloud service based on the information about the topology of the cloud service, wherein the determining is based at least in part on a severity level associated with the one or more anomalies and/or a service level associated with an entity that is experiencing the one or more anomalies;
in response to a decision not to upload the one or more anomalies to the topology aggregation service,
generating, by the client nodes, a list of available actions that is ordered based on an expected effectiveness of each action in mitigating an impact of the one or more anomalies on the cloud service, the available actions comprising one or more of aborting a task associated with the one or more anomalies, interrupting a service associated with the one or more anomalies, retrying an operation associated with the one or more anomalies, and isolating a problem area associated with the one or more anomalies, and
executing, by the client nodes, an action selected from the list of available actions to mitigate the impact of the one or more anomalies on the cloud service, wherein the action is selected based at least in part on the service level associated with the entity that is experiencing the one or more anomalies; and
in response to a decision to upload the one or more anomalies to the topology aggregation service,
transmitting, by the client nodes, the one or more anomalies to the topology aggregation service.
2 . The method of claim 1 , wherein the action is executed by the client nodes to mitigate the impact of the one or more anomalies on the cloud service without uploading the one or more anomalies to the topology aggregation service.
3 . The method of claim 1 , further comprising:
transmitting the uploaded one or more anomalies from the topology aggregation service to a conductor;
prioritizing, by the conductor, the uploaded one or more anomalies based on one or more of severity levels, contagious levels, historical data, computational power associated with the uploaded one or more anomalies;
generating, by an auto-remediate service, proposed solutions to the uploaded one or more anomalies; and
propagating the proposed solutions to the client nodes.
4 . The method of claim 3 , wherein the auto-remediate service further comprises an auto-adaptive thresholds operation, the auto-adaptive thresholds operation detecting, in real-time or near real-time, whether one or more of the operations corresponding to the proposed solutions is available.
5 . The method of claim 1 , wherein the list of available actions is ordered by the client nodes based further on a potential disruption caused by each action to the cloud service.
6 . The method of claim 1 , wherein the topology aggregation service comprises an indexed assembly triplet, the indexed assembly triplet comprising an error severity level, an original service associated with the error, and a potential operation list associated with the anomaly.
7 . The method of claim 1 , wherein the determining of whether to upload the one or more anomalies is based at least in part on a quality or cost associated with the service level.
8 . A system, comprising:
a programmable processor; and
a non-transient machine-readable medium storing instructions that, when executed by the processor, cause the at least one programmable processor to perform operations comprising:
detecting, at client nodes associated with a cloud service, anomalies in the cloud service based on a predetermined set of rules;
determining, by the client nodes, whether to upload one or more anomalies of the detected anomalies to a topology aggregation service configured to collect information about a topology of the cloud service and provide information regarding a structure and an operation of the cloud service based on the information about the topology of the cloud service, wherein the determining is based at least in part on a severity level associated with the one or more anomalies and/or a service level associated with an entity that is experiencing the one or more anomalies;
in response to a decision not to upload the one or more anomalies to the topology aggregation service,
generating, by the client nodes, a list of available actions that is ordered based on an expected effectiveness of each action in mitigating an impact of the one or more anomalies on the cloud service, the available actions comprising one or more of aborting a task associated with the one or more anomalies, interrupting a service associated with the one or more anomalies, retrying an operation associated with the one or more anomalies, and isolating a problem area associated with the one or more anomalies, and
executing, by the client nodes, an action selected from the list of available actions to mitigate the impact of the one or more anomalies on the cloud service, wherein the action is selected based at least in part on the service level associated with the entity that is experiencing the one or more anomalies; and
in response to a decision to upload the one or more anomalies to the topology aggregation service,
transmitting, by the client nodes, the one or more anomalies to the topology aggregation service.
9 . The system of claim 8 , wherein the action is executed by the client nodes to mitigate the impact of the one or more anomalies on the cloud service without uploading the one or more anomalies to the topology aggregation service.
10 . The system of claim 8 , wherein the operations further comprising:
transmitting the uploaded one or more anomalies from the topology aggregation service to a conductor;
prioritizing, by the conductor, the uploaded one or more anomalies based on one or more of severity levels, contagious levels, historical data, computational power associated with the uploaded one or more anomalies;
generating, by an auto-remediate service, proposed solutions to the uploaded one or more anomalies; and
propagating the proposed solutions to the client nodes.
11 . The system of claim 10 , wherein the auto-remediate service further comprises an auto-adaptive thresholds operation, the auto-adaptive thresholds operation detecting, in real-time or near real-time, whether one or more of the operations corresponding to the proposed solutions is available.
12 . The system of claim 8 , wherein the list of available actions is ordered by the client nodes based further on a potential disruption caused by each action to the cloud service.
13 . The system of claim 8 , wherein the topology aggregation service comprises an indexed assembly triplet, the indexed assembly triplet comprising an error severity level, an original service associated with the error, and a potential operation list associated with the anomaly.
14 . The system of claim 8 , wherein the determining of whether to upload the one or more anomalies is based at least in part on a quality or cost associated with the service level.
15 . A non-transitory computer-readable medium storing instructions, which when executed by at least one data processor, result in operations comprising:
detecting, at client nodes associated with a cloud service, anomalies in the cloud service based on a predetermined set of rules;
determining, by the client nodes, whether to upload one or more anomalies of the detected anomalies to a topology aggregation service configured to collect information about a topology of the cloud service and provide information regarding a structure and an operation of the cloud service based on the information about the topology of the cloud service, wherein the determining is based at least in part on a severity level associated with the one or more anomalies and/or a service level associated with an entity that is experiencing the one or more anomalies;
in response to a decision not to upload the one or more anomalies to the topology aggregation service,
generating, by the client nodes, a list of available actions that is ordered based on an expected effectiveness of each action in mitigating an impact of the one or more anomalies on the cloud service, the available actions comprising one or more of aborting a task associated with the one or more anomalies, interrupting a service associated with the one or more anomalies, retrying an operation associated with the one or more anomalies, and isolating a problem area associated with the one or more anomalies, and
executing, by the client nodes, an action selected from the list of available actions to mitigate the impact of the one or more anomalies on the cloud service, wherein the action is selected based at least in part on the service level associated with the entity that is experiencing the one or more anomalies; and
in response to a decision to upload the one or more anomalies to the topology aggregation service,
transmitting, by the client nodes, the one or more anomalies to the topology aggregation service.
16 . The non-transitory computer-readable medium of claim 15 , wherein the action is executed by the client nodes to mitigate the impact of the one or more anomalies on the cloud service without uploading the one or more anomalies to the topology aggregation service.
17 . The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:
transmitting the uploaded one or more anomalies from the topology aggregation service to a conductor;
prioritizing, by the conductor, the uploaded one or more anomalies based on one or more of severity levels, contagious levels, historical data, computational power associated with the uploaded one or more anomalies;
generating, by an auto-remediate service, proposed solutions to the uploaded one or more anomalies; and
propagating the proposed solutions to the client nodes.
18 . The non-transitory computer-readable medium of claim 17 , wherein the auto-remediate service further comprises an auto-adaptive thresholds operation, the auto-adaptive thresholds operation detecting, in real-time or near real-time, whether one or more of the operations corresponding to the proposed solutions is available.
19 . The non-transitory computer-readable medium of claim 15 , wherein the list of available actions is ordered by the client nodes based further on a potential disruption caused by each action to the cloud service.
20 . The non-transitory computer-readable medium of claim 15 , wherein the topology aggregation service comprises an indexed assembly triplet, the indexed assembly triplet comprising an error severity level, an original service associated with the error, and a potential operation list associated with the anomaly.