Thermal aware predictive failure analysis
Predictive Failure Analysis (PFA) techniques that are thermal aware can enable the prevention of uncorrectable errors without premature replacement of memory resources with redundant memory resources. In one example, a management controller (such as Baseboard Management Controller (BMC)) can monitor the rate of corrected errors. When the BMC detects that there are too many corrected errors occurring within certain time duration, the BMC can check the temperature and airflow rate of memory modules. Based on temperature data, the BMC can boost the fans and verify the reduction in memory corrected errors. If after multiple retries the rate of corrected error remains same, the BMC can enable memory resource replacement techniques such as SDDC or ADDDC or sparing on the failing memory module.
1 . A management controller comprising:
an interface to separately couple with a memory controller and with a memory module, the memory module located in a thermal zone; and
logic to increase air flow to the thermal zone in response to a determination that:
a corrected error count for the memory module is lower than a first threshold at which memory resource replacement is triggered,
a rate of corrected errors for the memory module is greater than a second threshold, and
a fluctuation of a temperature of the memory module is greater than a third threshold.
2 . The management controller of claim 1 , wherein:
the logic is to trigger replacement of memory resources with redundant memory resources in response to a determination that the corrected error count is greater than the first threshold.
3 . The management controller of claim 2 , wherein:
the replacement of the memory resources includes: rank sparing or the replacement of the memory resources with the redundant memory resources.
4 . The management controller of claim 1 , wherein:
the logic is to determine that the rate of the corrected errors is greater than the second threshold when a number of corrected errors within a predetermined time period is greater than a predetermined value that is lower than the first threshold.
5 . The management controller of claim 1 , wherein:
the logic to increase the air flow to the thermal zone is to:
increase a speed of a fan directing air flow to the thermal zone.
6 . The management controller of claim 1 , wherein:
the logic is to:
continue monitoring the rate of corrected errors after the increase in air flow, and
reduce air flow to the thermal zone in response to a determination that the rate of corrected errors is below the second threshold.
7 . The management controller of claim 1 , wherein:
the logic is to:
read a temperature sensor of the memory module multiple times within a period of time to determine whether the fluctuation of the temperature is greater than the third threshold.
8 . The management controller of claim 7 , wherein:
the logic is to determine the fluctuation of the temperature is greater than the third threshold when a difference between a minimum temperature and a maximum temperature in the period of time exceeds the third threshold or when a difference between the minimum or maximum temperature in the period of time and an average temperature exceeds the third threshold.
9 . The management controller of claim 7 , wherein:
the logic is to read the temperature sensor of the memory module via a direct link between the management controller and the memory module.
10 . The management controller of claim 1 , wherein:
the logic is to increase air flow to the thermal zone further in response to a determination that the air flow to the thermal zone is below a fourth threshold.
11 . A system comprising:
a memory controller to couple with a memory modules, the memory module located in a thermal zone from among a plurality of thermal zones; and
management controller logic coupled with the memory controller and the memory module, the management control logic to:
increase air flow to the thermal zone in response to a determination that:
a corrected error count for the memory module is lower than a first threshold,
a rate of corrected errors for the memory module is greater than a second threshold, and
a fluctuation of a temperature of the memory module is greater than a third threshold.
12 . The system of claim 11 , wherein:
the memory controller is included in a processor.
13 . The system of claim 11 , further comprising one or more of:
the memory module; and
one or more fans to cause the increase in air flow to the thermal zone.
14 . The system of claim 11 , wherein:
the management controller logic is to trigger replacement of memory resources with redundant memory resources in response to a determination that the corrected error count is greater than the first threshold; and
the management controller logic is to determine that the rate of the corrected errors is greater than the second threshold when a number of corrected errors within a predetermined time period is greater than a predetermined value that is lower than the first threshold.
15 . The system of claim 11 , wherein:
the management controller logic comprises a baseboard management controller.
16 . The system of claim 11 , wherein:
the management controller logic is to:
continue monitoring the rate of corrected errors after the increase in air flow, and
reduce air flow to the thermal zone in response to a determination that the rate of corrected errors is below the second threshold.
17 . A non-transitory machine-readable medium having instructions stored thereon that when executed by a management controller separately coupled with a memory controller and a memory module cause the management controller to:
monitor a rate of corrected errors for the memory module; and
increase air flow to a thermal zone via which the memory modules are located, wherein the air flow is increased to the thermal zone in response to a determination that:
a corrected error count for the memory module is lower than a first threshold,
a rate of corrected errors for the memory module is greater than a second threshold, and
a fluctuation of a temperature of the memory module is greater than a third threshold.
18 . The non-transitory machine-readable medium of claim 17 , wherein:
replacement of memory resources with redundant memory resources is triggered in response to a determination that the corrected error count is greater than the first threshold; and
the rate of the corrected errors is greater than the second threshold when a number of corrected errors within a predetermined time period is greater than a predetermined value that is lower than the first threshold.
19 . The non-transitory machine-readable medium of claim 17 , wherein:
to increase the air flow to the thermal zone includes increasing a speed of a fan directing air flow to the thermal zone.
20 . The non-transitory machine-readable medium of claim 17 , wherein the instructions further cause the management controller to:
continue to monitor the rate of corrected errors after the increase in air flow; and
reduce the air flow to the thermal zone in response to a determination that the rate of corrected errors is below the second threshold.