Methods and apparatus for maintaining the cooling systems of distributed compute systems
Methods and apparatus for maintaining the cooling systems of distributed compute systems are disclosed. An example apparatus disclosed herein includes memory, machine-readable instructions, and processor circuitry to execute the machine-readable instructions to determine a health of a server, determine a threshold based on a workload service level agreement associated with the server, and in response to determining the health does not a satisfy the threshold, throttle a workload on the server.
1 . An apparatus comprising:
memory;
machine-readable instructions; and
at least one processor circuit to execute the machine-readable instructions to:
determine a first component health of a compute component of a server;
determine a second component health of a cooling component of a cooling system for the server;
determine a health of the server based on the first component health and the second component health;
determine a threshold based on a workload service level agreement associated with the server; and
in response to determining the health of the server does not a satisfy the threshold:
throttle a workload on the server;
determine a first maintenance window of the compute component based on the first component health;
determine a second maintenance window of the cooling component based on the second component health; and
schedule a maintenance period for the server, the maintenance period scheduled to be within the first maintenance window and the second maintenance window.
2 . The apparatus of claim 1 , wherein the server is disposed in a marine environment, a deep Earth environment, or a high altitude environment.
3 . The apparatus of claim 1 , wherein the server is a first server, the threshold is a first threshold, and one or more of the at least one processor circuit is to execute the machine-readable instructions to:
determine a second threshold based on the workload service level agreement; and
migrate the workload to a second server different than the first server.
4 . The apparatus of claim 3 , wherein the first server includes a liquid-based cooling system and the second server includes an air-based cooling system.
5 . The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to execute the machine-readable instructions to determine the health of the server as a lesser one of the first component health and the second component health.
6 . A method comprising:
determining a first component health of a compute component of a server;
determining a second component health of a cooling component of a cooling system for the server;
determining a health of the server based on the first component health and the second component health;
determining a threshold based on a workload service level agreement associated with the server; and
in response to determining the health of the server does not a satisfy the threshold:
throttling a workload on the server;
determining a first maintenance window of the compute component based on the first component health;
determining a second maintenance window of the cooling component based on the second component health; and
scheduling, by executing instructions with at least one processor circuit, a maintenance period for the server, the maintenance period scheduled to be within the first maintenance window and the second maintenance window.
7 . The method of claim 6 , wherein the server is located at of in a marine environment, a deep Earth environment, or a high altitude environment.
8 . The method of claim 6 , wherein the server is a first server, the threshold is a first threshold, and the method further includes:
determining a second threshold based on the workload service level agreement; and
migrating the workload to a second server different than the first server.
9 . The method of claim 8 , wherein the first server includes a liquid-based cooling system and the second server includes an air-based cooling system.
10 . The method of claim 6 , wherein the server is located in a deep sea environment.
11 . The method of claim 6 , wherein the server is located in a deep Earth environment.
12 . The method of claim 6 , wherein the server is located in a high altitude environment.
13 . A non-transitory computer readable storage medium comprising machine-readable instructions to cause one or more processor circuits to:
determine a first component health of a compute component of a server;
determine a second component health of a cooling component of a cooling system for the server;
determine a health of the server including a cooling system based on the first component health and the second component health;
determine a threshold based on a workload service level agreement associated with the server; and
in response to determining the health of the server does not a satisfy the threshold:
throttle a workload on the server;
determine a first maintenance window of the compute component based on the first component health;
determine a second maintenance window of the cooling component based on the second component health; and
schedule a maintenance period for the server, the maintenance period scheduled to be within the first maintenance window and the second maintenance window.
14 . The non-transitory computer readable storage medium of claim 13 , wherein the server is disposed in at least one of a marine environment, a deep Earth environment, or a high altitude environment.
15 . The non-transitory computer readable storage medium of claim 13 , wherein the server is a first server, the threshold is a first threshold, and the instructions are to cause at least one of the one or more processor circuits to execute the machine-readable instructions to:
determine a second threshold based on the workload service level agreement; and
migrate the workload to a second server different than the first server.
16 . The non-transitory computer readable storage medium of claim 15 , wherein the first server includes a liquid-based cooling system and the second server includes an air-based cooling system.
17 . The non-transitory computer readable storage medium of claim 13 , wherein the instructions are to cause at least one of the one or more processor circuits to determine the health of the server as a lesser one of the first component health and the second component health.