METHODS AND APPARATUS FOR MAINTAINING THE COOLING SYSTEMS OF DISTRIBUTED COMPUTE SYSTEMS
Methods and apparatus for maintaining the cooling systems of distributed compute systems are disclosed. An example apparatus disclosed herein includes memory, machine-readable instructions, and processor circuitry to execute the machine-readable instructions to determine a health of a server, determine a threshold based on a workload service level agreement associated with the server, and in response to determining the health does not a satisfy the threshold, throttle a workload on the server.
1 . An apparatus comprising:
memory;
machine-readable instructions; and
processor circuitry to execute the machine-readable instructions to:
determine a health of a server;
determine a threshold based on a workload service level agreement associated with the server; and
in response to determining the health does not a satisfy the threshold, throttle a workload on the server.
2 . The apparatus of claim 1 , wherein the server is disposed in a marine environment, a deep Earth environment, or a high altitude environment.
3 . The apparatus of claim 1 , wherein the server is a first server, the threshold is a first threshold, and the processor circuitry is to execute the machine-readable instructions to:
determine a second threshold based on the workload service level agreement; and
migrate the workload to a second server different than the first server.
4 . The apparatus of claim 3 , wherein the first server includes a liquid-based cooling system and the second server includes an air-based cooling system.
5 . The apparatus of claim 1 , wherein the server includes a cooling system and the processor circuitry is to execute the machine-readable instructions to determine the health of the server by:
determining a first component health of a compute component of the server;
determining a second component health of a cooling component of the cooling system; and
determining the health of the server based on the first component health and the second component health.
6 . The apparatus of claim 5 , in response to determining the health of the server does not satisfy the threshold, the processor circuitry is to execute the machine-readable instructions to:
determine a first maintenance window of the compute component based on the first component health;
determine a second maintenance window of the cooling component based on the second component health; and
schedule a maintenance period of the server, the maintenance period scheduled to be within the first maintenance window and the second maintenance window.
7 . The apparatus of claim 5 , wherein the processor circuitry is to execute the machine-readable instructions to determine the health of the server as a lesser one of the first component health and the second component health.
8 . A method comprising:
determining a health of a server;
determining a threshold based on a workload service level agreement associated with the server; and
in response to determining the health does not a satisfy the threshold, throttling a workload on the server.
9 . The method of claim 8 , wherein the server is located at of a marine environment, a deep Earth environment, or a high altitude environment.
10 . The method of claim 8 , wherein the server is a first server, the threshold is a first threshold, and the method further includes:
determining a second threshold based on the workload service level agreement; and
migrating the workload to a second server different than the first server.
11 . The method of claim 10 , wherein the first server includes a liquid-based cooling system and the second server includes an air-based cooling system.
12 . The method of claim 8 , wherein the server includes a cooling system and the determination the health of the server includes:
determining a first component health of a compute component of the server;
determining a second component health of a cooling component of the cooling system; and
determining the health of the server based on the first component health and the second component health.
13 . The method of claim 12 , further including, in response to determining the health of the server does not satisfy the threshold:
determining a first maintenance window of the compute component based on the first component health;
determining a second maintenance window of the cooling component based on the second component health; and
scheduling a maintenance period of the server, the maintenance period scheduled to be within the first maintenance window and the second maintenance window.
14 . A non-transitory computer readable storage medium comprising machine-readable instructions, which when executed, caused one or more processors to:
determine a health of a server including a cooling system;
determine a threshold based on a workload service level agreement associated with the server; and
in response to determining the health does not a satisfy the threshold, throttle a workload on the server.
15 . The non-transitory computer readable storage medium of claim 14 , wherein the server is disposed in at least one of a marine environment, a deep Earth environment, or a high altitude environment.
16 . The non-transitory computer readable storage medium of claim 14 , wherein the server is a first server, the threshold is a first threshold, and the instructions, when executed, caused the one or more processors to execute the machine-readable instructions to:
determine a second threshold based on the workload service level agreement; and
migrate the workload to a second server different than the first server.
17 . The non-transitory computer readable storage medium of claim 16 , wherein the first server includes a liquid-based cooling system and the second server includes an air-based cooling system.
18 . The non-transitory computer readable storage medium of claim 14 , wherein the instructions, when executed, caused the one or more processors to execute the machine-readable instructions to:
determine a first component health of a compute component of the server;
determine a second component health of a cooling component of the cooling system; and
determine the health of the server based on the first component health and the second component health.
19 . The non-transitory computer readable storage medium of claim 18 , wherein the instructions, when executed, caused the one or more processors to execute the machine-readable instructions to, in response to determining the health does not satisfy the threshold:
determine a first maintenance window of the compute component based on the first component health;
determine a second maintenance window of the cooling component based on the second component health; and
schedule a maintenance period of the server, the maintenance period scheduled to be within the first maintenance window and the second maintenance window.
20 . The non-transitory computer readable storage medium of claim 18 , wherein the instructions, when executed, caused the one or more processors to determine the health of the server as a lesser one of the first component health and the second component health.