Hardware management based on failure prediction in a multi-tiered architecture
Techniques are disclosed for hardware management in an information processing system. For example, a method obtains, at an edge platform, one or more failure prediction values for one or more device types of the edge platform, wherein a failure prediction value for a device type represents a likelihood of failure associated with the device type. The method, at the edge platform, computes one or more health indicator values for the one or more device types based on the one or more failure prediction values, computes one or more behavior indicator values for the one or more device types based on the one or more health indicator values. The method causes, in response to the one or more behavior indicator values for the one or more device types, determination of one or more proactive actions to be initiated prior to a failure of one or more devices of the edge platform.
1 . A method comprising:
obtaining, at an edge platform, one or more failure prediction values for one or more device types of the edge platform, wherein a failure prediction value for a device type represents a likelihood of failure associated with the device type;
computing, at the edge platform, one or more health indicator values for the one or more device types based on the one or more failure prediction values;
computing, at the edge platform, one or more behavior indicator values for the one or more device types based on the one or more health indicator values;
causing, in response to the one or more behavior indicator values for the one or more device types, one or more proactive actions to be initiated prior to a failure of one or more devices of the edge platform, wherein the one or more proactive actions comprise one of updating and upgrading one or more hardware resources of the one or more devices;
recomputing, at the edge platform, the one or more behavior indicator values based on a set of monitored parameters of one or more devices of the edge platform; and
sending, from the edge platform to a centralized backend device from which the one or more failure prediction values originate, at least a portion of the set of monitored parameters to enable the centralized backend device to update at least a portion of the one or more failure prediction values;
wherein the set of monitored parameters are received at the edge platform as part of a monitoring policy from the centralized backend device; and
wherein the obtaining, computing, causing, recomputing and sending steps are executed by a processing device operatively coupled to a memory.
2 . The method of claim 1 wherein the one or more health indicator values are computed using one or more probability distribution functions.
3 . The method of claim 2 wherein the one or more probability distribution functions comprise a probability mass function.
4 . The method of claim 3 wherein the one or more probability distribution functions further comprise a cumulative distribution function.
5 . The method of claim 4 wherein a health indicator value for a device type is computed via a logical addition of a computation result of the probability mass function and a computation result of the cumulative distribution function.
6 . The method of claim 2 wherein at least one of the one or more probability distribution functions comprises a Poisson distribution function.
7 . The method of claim 1 wherein the one or more failure prediction values comprise mean coefficient values representative of respective failure predictions for the one or more device types.
8 . The method of claim 1 further comprises initiating a registration, by the edge platform, with a content delivery network device in a network of distributed content delivery network devices that are connected to a centralized backend device from which the one or more failure prediction values originate.
9 . The method of claim 1 wherein the steps of the method are executed by an edge analyzer client installed at the edge platform.
10 . An apparatus comprising:
a processing device operatively coupled to a memory and configured:
to obtain, at an edge platform, one or more failure prediction values for one or more device types of the edge platform, wherein a failure prediction value for a device type represents a likelihood of failure associated with the device type;
to compute, at the edge platform, one or more health indicator values for the one or more device types based on the one or more failure prediction values;
to compute, at the edge platform, one or more behavior indicator values for the one or more device types based on the one or more health indicator values;
to cause, in response to the one or more behavior indicator values for the one or more device types, one or more proactive actions to be initiated prior to a failure of one or more devices of the edge platform, wherein the one or more proactive actions comprise one of updating and upgrading one or more hardware resources of the one or more devices;
to recompute, at the edge platform, the one or more behavior indicator values based on a set of monitored parameters of one or more devices of the edge platform and
to send, from the edge platform to a centralized backend device from which the one or more failure prediction values originate, at least a portion of the set of monitored parameters to enable the centralized backend device to update at least a portion of the one or more failure prediction values;
wherein the set of monitored parameters are received at the edge platform as part of a monitoring policy from the centralized backend device.
11 . The apparatus of claim 10 wherein the one or more health indicator values are computed using one or more probability distribution functions.
12 . The apparatus of claim 11 wherein the one or more probability distribution functions comprise a probability mass function.
13 . The apparatus of claim 12 wherein the one or more probability distribution functions further comprise a cumulative distribution function.
14 . The apparatus of claim 13 wherein a health indicator value for a device type is computed via a logical addition of a computation result of the probability mass function and a computation result of the cumulative distribution function.
15 . The apparatus of claim 11 wherein at least one of the one or more probability distribution functions comprises a Poisson distribution function.
16 . An article of manufacture comprising a non-transitory processor-readable storage medium having stored therein program code of one or more software programs, wherein the program code when executed by at least one processing device causes said at least one processing device to perform steps of:
obtaining, at an edge platform, one or more failure prediction values for one or more device types of the edge platform, wherein a failure prediction value for a device type represents a likelihood of failure associated with the device type;
computing, at the edge platform, one or more health indicator values for the one or more device types based on the one or more failure prediction values;
computing, at the edge platform, one or more behavior indicator values for the one or more device types based on the one or more health indicator values;
causing, in response to the one or more behavior indicator values for the one or more device types, one or more proactive actions to be initiated prior to a failure of one or more devices of the edge platform, wherein the one or more proactive actions comprise one of updating and upgrading one or more hardware resources of the one or more devices;
recomputing, at the edge platform, the one or more behavior indicator values based on a set of monitored parameters of one or more devices of the edge platform; and
sending, from the edge platform to a centralized backend device from which the one or more failure prediction values originate, at least a portion of the set of monitored parameters to enable the centralized backend device to update at least a portion of the one or more failure prediction values;
wherein the set of monitored parameters are received at the edge platform as part of a monitoring policy from the centralized backend device.
17 . The article of manufacture of claim 16 wherein the one or more health indicator values are computed using one or more probability distribution functions.
18 . The article of manufacture of claim 17 wherein the one or more probability distribution functions comprise a probability mass function.
19 . The article of manufacture of claim 18 wherein the one or more probability distribution functions further comprise a cumulative distribution function.
20 . The article of manufacture of claim 19 wherein a health indicator value for a device type is computed via a logical addition of a computation result of the probability mass function and a computation result of the cumulative distribution function.