IP Library Granted Patent US 12,613,785
Granted Patent B2
US 12,613,785 · App. 18/329,627 · Granted Apr 28, 2026

System wear leveling

Inventors: Gary Smerdon (Los Gatos, CA); Isaac R. Nassi (Los Gatos, CA); David P. Reed (Los Gatos, CA)
Assignee: Hewlett Packard Enterprise Development LP
G06F11/2041G06F2201/81
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,613,785
App. No.
18/329,627
Granted
Apr 28, 2026
Kind
B2
Abstract

A technique includes receiving information relating to wear of computer nodes. Based on the information relating to the wear, the technique includes ranking the computer nodes according to respective expected remaining lifetimes of the computer nodes. The technique includes, responsive to an event that corresponds to at least one of a spare computer node being added to the system or a given computer node being removed from the system, reconfiguring the system based on the ranking.

Claims (65)

1 . A non-transitory machine-readable storage medium comprising instructions that upon execution cause a system of a plurality of computer nodes to:

execute a plurality of hyper-kernels on respective computer nodes of the plurality of computer nodes;

responsive to the execution of the plurality of hyper-kernels, provide a guest operating system distributed across the plurality of computer nodes;

receive information relating to wear of the computer nodes;

based on the information relating to the wear, rank the computer nodes according to respective expected remaining lifetimes of the computer nodes; and

responsive to an event corresponding to at least one of a spare computer node being added to the system or a given computer node of the plurality of computer nodes being removed from the system, reconfigure the system based on the ranking.

2 . The non-transitory machine-readable storage medium of claim 1 , wherein:

the event corresponds to a maintenance service in which the given computer node is removed from the system for maintenance;

the reconfiguring comprises removing the given computer node from the system; and

the instructions upon execution further cause the system to select the given computer node for removal from the system responsive to a determination that the respective expected remaining lifetime of the given computer node is less than at least one other respective expected remaining lifetime of the expected remaining lifetimes.

3 . The non-transitory machine-readable storage medium of claim 2 , wherein:

the instructions upon execution further cause the system to use a model to estimate the expected remaining lifetimes, and responsive to an expected remaining lifetime determined from a maintenance inspection of the given computer node after the given computer node was removed, then update the model.

4 . The non-transitory machine-readable storage medium of claim 1 , wherein:

the event corresponds to a maintenance service in which the given computer node is removed from the system for maintenance;

the reconfiguring comprises adding the spare computer node to the system; and

the instructions upon execution further cause the system to select the spare computer node from among a plurality of candidate spare nodes responsive to a determination that a respective remaining lifetime of the spare computer node is greater than at least one other respective expected remaining lifetime of another spare computer node of the plurality of candidate spare computer nodes.

5 . The non-transitory machine-readable storage medium of claim 1 , wherein the system comprises a distributed system, and the instructions upon execution further cause the system to:

assess an application workload provided by the plurality of computer nodes;

determine that resources provided by the plurality of computer nodes to support the application workload are undersubscribed;

responsive to the determination, initiate the event to upscale the resources; and

in association with the upscaling, select the spare computer node from among a plurality of candidate spare computer nodes responsive to a determination that a respective remaining lifetime of the spare computer node is greater than at least one other respective expected remaining lifetime of another spare computer node of the plurality of candidate spare computer nodes.

6 . The non-transitory machine-readable storage medium of claim 1 , wherein the system comprises a distributed system, and the instructions upon execution further cause the system to:

assess an application workload provided by the plurality of computer nodes;

determine that resources provided by the plurality of computer nodes to support the application workload are oversubscribed;

responsive to the determination, initiate the event to downscale the resources; and

in association with the downscaling, to select the given computer node for removal from the system responsive to a determination that the respective expected remaining lifetime of the given computer node is less than at least one other respective expected remaining lifetime of the expected remaining lifetimes.

7 . The non-transitory machine-readable storage medium of claim 6 , wherein the instructions upon execution further cause the system to:

select the given computer node responsive to a determination that the respective expected remaining lifetime of the given computer node is the least expected remaining lifetime of the expected remaining lifetimes.

8 . The non-transitory machine-readable storage medium of claim 1 , wherein the information relating to the wear comprises information relating to at least one of errors attributable to physical memories of the plurality of computer nodes or ages of the physical memories.

9 . The non-transitory machine-readable storage medium of claim 1 , wherein the information relating to the wear comprises information relating to at least one of temperatures of physical processors of the plurality of computer nodes, clock rates of the physical processors, or ages of the physical processors.

10 . The non-transitory machine-readable storage medium of claim 1 , wherein the information relating to the wear comprises information relating to at least one of errors attributable physical network interfaces of the plurality of computer nodes or ages of the physical network interfaces.

11 . The non-transitory machine-readable storage medium of claim 1 , wherein the information relating to the wear comprises information relating to at least one of errors attributable to physical storage devices associated with the plurality of computer nodes or ages of the physical storage devices.

12 . The non-transitory machine-readable storage medium of claim 1 , wherein the information relating to the wear is from a program that reads data indicative of the wear from at least one of hardware registers of the plurality of computer nodes or system event logs of the plurality of computer nodes.

13 . The non-transitory machine-readable storage medium of claim 1 , wherein the information relating to the wear is provided by a baseboard management controller.

14 . The non-transitory machine-readable storage medium of claim 1 , wherein the instructions, upon execution, further cause the system to:

execute a single virtual machine across the plurality of computer nodes.

15 . The non-transitory machine-readable storage medium of claim 1 , wherein the system comprises a distributed system, a hypervisor is distributed among the computer nodes, and reconfiguring the system is by the distributed hypervisor.

16 . A distributed system comprising:

a plurality of computer nodes;

wherein the plurality of computer nodes comprises instructions executable to:

execute a plurality of hyper-kernels on respective computer nodes of the plurality of computer nodes;

responsive to the execution of the plurality of hyper-kernels, provide a guest operating system distributed across the plurality of computer nodes;

receive an indication of an event corresponding to changing a number of the plurality of computer nodes; and

responsive to the event, identify a particular computer node associated with effecting the change from a pool of candidate computer nodes based on respective wear levels associated with the candidate computer nodes.

17 . The distributed system of claim 16 , wherein:

the event comprises adding the particular computer node to the plurality of computer nodes;

the pool of candidate computer nodes comprises a plurality of spare computer nodes; and

the computer node wear-leveling instructions are further executable to:

receive wear-leveling information relating to the plurality of spare computer nodes; and

identify a spare computer node of the plurality of spare computer nodes as the particular computer node based on the wear-leveling information.

18 . The distributed system of claim 16 , wherein:

the event comprises removing the particular computer node from the plurality of computer nodes;

the plurality of computer nodes corresponds to the pool of candidate computer nodes; and

the computer node wear-leveling instructions are further executable to:

receive wear-leveling information relating to the plurality of computer nodes; and

identify, by the wear-leveling engine, a particular computer node associated with effecting the change from a pool of candidate computer nodes based on the ranking.

19 . A method of a distributed system comprising a plurality of computer nodes, comprising:

executing, by the plurality of computer nodes, a plurality of hyper-kernels;

providing, by the plurality of hyper-kernels, a guest operating system distributed across the plurality of computer nodes;

receiving, by a node wear-leveling engine, information relating to wear of the computer nodes;

based on the information relating to the wear, ranking, by the node wear-leveling engine, the computer nodes according to respective expected remaining lifetimes of the computer nodes; and

responsive to an event corresponding to changing a number of the computer nodes, identifying, by the node wear-leveling engine, a particular computer node associated with effecting the change from a pool of candidate computer nodes based on the ranking.

20 . The method of claim 19 , wherein the identifying further comprises:

selecting a list of the computer nodes based on the ranking, wherein the list includes the particular computer node; and

randomly or pseudo-randomly selecting the particular computer node from the list.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2023
From: SMERDON, GARY; NASSI, ISAAC R.; REED, DAVID P.
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 063862/0241 →
Continuity (2)
Provisional Application 63345691 · May 25, 2022
Related Publication 20250086076A1 · Mar 13, 2025
References Cited (12)
US 10082965B1 · Tamilarasan · 2018 [cited by examiner]
US 11119872B1 · Jacobson · 2021 [cited by examiner]
US 11252068B1 · Kamen · 2022 [cited by examiner]
US 20140068153A1 · Gu et al. · 2014 [cited by applicant]
US 20170199681A1 · Jain et al. · 2017 [cited by applicant]
US 20170286176A1 · Artman et al. · 2017 [cited by applicant]
US 20200177451A1 · Easterling et al. · 2020 [cited by applicant]
US 20210216398A1 · Davis et al. · 2021 [cited by applicant]
US 20220276905A1 · Manousakis · 2022 [cited by examiner]
US 20220398175A1 · Ukawa · 2022 [cited by examiner]
US 20230267040A1 · Davis · 2023 [cited by examiner]
International Search Report and Written Opinion received for PCT Patent Application No. PCT/US2023/014025, mailed on Jun. 19, 2023, 10 pages. [cited by applicant]