IP Library › Granted Patent US 12,737,244
Granted Patent B2
US 12,737,244 · App. 18/938,516 · Granted Sep 15, 2026

Proactive operating system memory unit replacement based on namespace health

Inventors: Asimuddin Kazi (Naperville, IL); Alejandro Reyther (Fox River Grove, IL); Edward James Weisbrod (Arlington Heights, IL); Adam Gray (Chicago, IL)
Assignee: International Business Machines Corporation
G06F11/004G06F3/0616G06F3/0653G06F3/067G06F2201/81
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,244
App. No.
18/938,516
Granted
Sep 15, 2026
Kind
B2
Abstract

A computer-implemented method includes receiving diagnostic data about a plurality of operating system (OS) memory units of a distributed computing system. The method further includes defining a failure state for predicting whether at least one of the plurality of OS memory units will experience a failure. The method further includes predicting whether the at least one of the plurality of OS memory units will experience the failure based at least in part on the failure state and the diagnostic data about the plurality of OS memory units. The method further includes causing the at least one of the plurality of OS memory units predicted to experience the failure to be replaced with at least one replacement OS memory unit without impacting a namespace availability of the distributed computing system.

Claims (39)

1 . A computer-implemented method comprising:

receiving diagnostic data about a plurality of operating system (OS) memory units of a distributed computing system, the diagnostic data comprises key-value pairs of metrics;

converting, for each of the key-value pairs of metrics, a value of a key-value pair into an array of numbers, quantizing the array into a vector using k-means clustering, classifying the vector, and storing, when classified within an unhealthy cluster, a distance to a centroid of the unhealthy cluster within the value of a key-value store;

defining a failure state for predicting whether at least one of the plurality of OS memory units will experience a failure;

predicting whether the at least one of the plurality of OS memory units will experience the failure based at least in part on the failure state and the diagnostic data about the plurality of OS memory units, wherein the predicting comprises determining that the at least one of the plurality of OS memory units is likely to fail responsive to the distance to the centroid of the unhealthy cluster being below a threshold distance; and

causing the at least one of the plurality of OS memory units predicted to experience the failure to be replaced with at least one replacement OS memory unit without impacting a namespace availability of the distributed computing system.

2 . The computer-implemented method of claim 1 , wherein receiving the diagnostic data about the plurality of OS memory units comprises collecting the diagnostic data about the plurality of OS memory units.

3 . The computer-implemented method of claim 2 , wherein the collecting is performed on a periodic basis.

4 . The computer-implemented method of claim 1 , the diagnostic data comprises self-monitoring, analysis, and reporting technology (SMART) data.

5 . The computer-implemented method of claim 1 , wherein causing the at least one of the plurality of OS memory units predicted to experience the failure to be replaced prevents the distributed computing system from entering a hang state.

6 . The computer-implemented method of claim 1 , wherein causing the at least one of the plurality of OS memory units predicted to experience the failure to be replaced prevents the distributed computing system from experiencing performance degradation.

7 . The computer-implemented method of claim 1 , wherein the failure state comprises a threshold, wherein the at least one of the plurality of OS memory units is predicted to experience the failure responsive to a metric about the at least one of the plurality of OS memory units exceeding the threshold.

8 . The computer-implemented method of claim 1 , further comprising, responsive to predicting that the at least one of the plurality of OS memory units will experience the failure, backing up data and logs stored on the at least one of the plurality of OS memory units predicted to experience the failure to at least one of a local storage vault and a management vault and synchronizing the data to the at least one replacement OS memory unit.

9 . A system comprising:

a memory comprising computer readable instructions; and

a processing device for executing the computer readable instructions, the computer readable instructions controlling the processing device to perform operations comprising:

receiving diagnostic data about a plurality of operating system (OS) memory units of a distributed computing system, the diagnostic data comprises key-value pairs of metrics;

converting, for each of the key-value pairs of metrics, a value of a key-value pair into an array of numbers, quantizing the array into a vector using k-means clustering, classifying the vector, and storing, when classified within an unhealthy cluster, a distance to a centroid of the unhealthy cluster within the value of a key-value store;

defining a failure state for predicting whether at least one of the plurality of OS memory units will experience a failure;

predicting whether the at least one of the plurality of OS memory units will experience the failure based at least in part on the failure state and the diagnostic data about the plurality of OS memory units, wherein the predicting comprises determining that the at least one of the plurality of OS memory units is likely to fail responsive to the distance to the centroid of the unhealthy cluster being below a threshold distance; and

causing the at least one of the plurality of OS memory units predicted to experience the failure to be replaced with at least one replacement OS memory unit without impacting a namespace availability of the distributed computing system.

10 . The system of claim 9 , wherein receiving the diagnostic data about the plurality of OS memory units comprises collecting the diagnostic data about the plurality of OS memory units.

11 . The system of claim 10 , wherein the collecting is performed on a periodic basis.

12 . The system of claim 9 , the diagnostic data comprises self-monitoring, analysis, and reporting technology (SMART) data.

13 . The system of claim 9 , wherein causing the at least one of the plurality of OS memory units predicted to experience the failure to be replaced prevents the distributed computing system from entering a hang state.

14 . The system of claim 9 , wherein causing the at least one of the plurality of OS memory units predicted to experience the failure to be replaced prevents the distributed computing system from experiencing performance degradation.

15 . The system of claim 9 , wherein the failure state comprises a threshold, wherein the at least one of the plurality of OS memory units is predicted to experience the failure responsive to a metric about the at least one of the plurality of OS memory units exceeding the threshold.

16 . The system of claim 9 , wherein the operations further comprise, responsive to predicting that the at least one of the plurality of OS memory units will experience the failure, backing up data and logs stored on the at least one of the plurality of OS memory units predicted to experience the failure to at least one of a local storage vault and a management vault and synchronizing the data to the at least one replacement OS memory unit.

17 . A computer program product comprising:

a set of one or more computer-readable storage media;

program instructions, collectively stored in the set of one or more storage media, for causing a processor set to perform the following computer operations:

receiving diagnostic data about a plurality of operating system (OS) memory units of a distributed computing system, the diagnostic data comprises key-value pairs of metrics;

converting, for each of the key-value pairs of metrics, a value of a key-value pair into an array of numbers, quantizing the array into a vector using k-means clustering, classifying the vector, and storing, when classified within an unhealthy cluster, a distance to a centroid of the unhealthy cluster within the value of a key-value store;

defining a failure state for predicting whether at least one of the plurality of OS memory units will experience a failure;

predicting whether the at least one of the plurality of OS memory units will experience the failure based at least in part on the failure state and the diagnostic data about the plurality of OS memory units, wherein the predicting comprises determining that the at least one of the plurality of OS memory units is likely to fail responsive to the distance to the centroid of the unhealthy cluster being below a threshold distance; and

causing the at least one of the plurality of OS memory units predicted to experience the failure to be replaced with at least one replacement OS memory unit without impacting a namespace availability of the distributed computing system.

18 . The computer program product of claim 17 , wherein receiving the diagnostic data about the plurality of OS memory units comprises collecting the diagnostic data about the plurality of OS memory units.

19 . The computer program product of claim 18 , wherein the collecting is performed on a periodic basis.

20 . The computer program product of claim 17 , the diagnostic data comprises self-monitoring, analysis, and reporting technology (SMART) data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2024
From: KAZI, ASIMUDDIN; REYTHER, ALEJANDRO; WEISBROD, EDWARD JAMES; GRAY, ADAM
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 069154/0596 →
Continuity (1)
Related Publication 20260127057A1 · May 7, 2026
References Cited (92)
US 5761411A · Teague et al. · 1998 [cited by applicant]
US 7584219B2 · Zybura et al. · 2009 [cited by applicant]
US 7987167B1 · Kazar et al. · 2011 [cited by applicant]
US 8108483B2 · Aust et al. · 2012 [cited by applicant]
US 8151352B1 · Novitchi · 2012 [cited by applicant]
US 8447780B1 · Plantenberg et al. · 2013 [cited by applicant]
US 9142322B2 · Baranwal et al. · 2015 [cited by applicant]
US 9171160B2 · Vincent et al. · 2015 [cited by applicant]
US 9535774B2 · Cher et al. · 2017 [cited by applicant]
US 9686293B2 · Golshan et al. · 2017 [cited by applicant]
US 10275361B2 · Ish et al. · 2019 [cited by applicant]
US 10884648B2 · Guo et al. · 2021 [cited by applicant]
US 10891068B2 · Guo et al. · 2021 [cited by applicant]
US 11226860B1 · Baptist et al. · 2022 [cited by applicant]
US 11314442B2 · Kazi et al. · 2022 [cited by applicant]
US 11412041B2 · Tamborski et al. · 2022 [cited by applicant]
US 11588892B1 · Khadiwala et al. · 2023 [cited by applicant]
US 11973823B1 · Ghorpade et al. · 2024 [cited by applicant]
US 20070006048A1 · Zimmer et al. · 2007 [cited by applicant]
US 20110145838A1 · de Melo · 2011 [cited by examiner]
US 20130326284A1 · Losh et al. · 2013 [cited by applicant]
US 20140032962A1 · Deshpande · 2014 [cited by examiner]
US 20150074469A1 · Cher et al. · 2015 [cited by applicant]
US 20150127919A1 · Baldwin · 2015 [cited by examiner]
US 20150243373A1 · Chun · 2015 [cited by examiner]
US 20160019983A1 · Chung · 2016 [cited by examiner]
US 20160072889A1 · Jung et al. · 2016 [cited by applicant]
US 20170004309A1 · Pavlyushchik et al. · 2017 [cited by applicant]
US 20170031959A1 · Zayas et al. · 2017 [cited by applicant]
US 20170090824A1 · Tamborski · 2017 [cited by applicant]
US 20170093978A1 · Cilfone et al. · 2017 [cited by applicant]
US 20180025025A1 · Davis et al. · 2018 [cited by applicant]
US 20180239697A1 · Huang et al. · 2018 [cited by applicant]
US 20180314441A1 · Suryanarayana et al. · 2018 [cited by applicant]
US 20190065524A1 · Johnson et al. · 2019 [cited by applicant]
US 20190146675A1 · Subramanian et al. · 2019 [cited by applicant]
US 20190146907A1 · Frolikov · 2019 [cited by applicant]
US 20190188589A1 · Ponnuru · 2019 [cited by examiner]
US 20190278498A1 · Dedrick · 2019 [cited by applicant]
US 20190281114A1 · Cocagne · 2019 [cited by applicant]
US 20190394272A1 · Tamborski et al. · 2019 [cited by applicant]
US 20200019447A1 · Tamborski et al. · 2020 [cited by applicant]
US 20200050365A1 · Tamborski et al. · 2020 [cited by applicant]
US 20200104056A1 · Benisty et al. · 2020 [cited by applicant]
US 20200192799A1 · Johns et al. · 2020 [cited by applicant]
US 20200278896A1 · Kumari · 2020 [cited by examiner]
US 20200409559A1 · Sharon et al. · 2020 [cited by applicant]
US 20210173582A1 · Kazi et al. · 2021 [cited by applicant]
US 20210216227A1 · Kazi et al. · 2021 [cited by applicant]
US 20210349782A1 · Ki et al. · 2021 [cited by applicant]
US 20210367834A1 · Palavalli et al. · 2021 [cited by applicant]
US 20220236916A1 · Esaka et al. · 2022 [cited by applicant]
US 20230195577A1 · Darnell et al. · 2023 [cited by applicant]
US 20230401120A1 · Gim · 2023 [cited by examiner]
US 20240220227A1 · Kerr et al. · 2024 [cited by applicant]
US 20240361939A1 · Brown et al. · 2024 [cited by applicant]
US 20240377948A1 · Zhang et al. · 2024 [cited by applicant]
US 20250077475A1 · Qi et al. · 2025 [cited by applicant]
“Intel® Optane™ DC SSD Series”, Intel, Feb. 2012, 27 pages, https://ark.intel.com/content/www/us/en/ark/products/series/213706/intel-optane-dc-ssd-series.html. [cited by applicant]
“Intel® Solid-State Drive 520 Series”, Intel, Sep. 9, 2024, 5 pages, https://www.intel.com/content/dam/www/public/us/en/documents/product-specifications/ssd-520-specification.pdf. [cited by applicant]
“SMART Attribute Details”, Kingston, 2015, 9 pages, https://media.kingston.com/support/downloads/MKP_306_SMART_attribute.pdf. [cited by applicant]
“SSD Failures: Common Causes and Main Bad Symptoms”, Multi-cloud backup Solutions, Oct. 27, 2020, 9 pages, https://www.salvagedata.com/ssd-failures-common-causes-and-main-symptom/. [cited by applicant]
“Tabular Classification”, Hugging Face, retrieved from web dated Dec. 20, 2024, 4 pages. [cited by applicant]
“What is supervised learning?”, IBM, retrieved from web dated Dec. 20, 2024, 9 pages. [cited by applicant]
Diamantopoulos et al. “WannaLaugh: A Configurable Ransomware Emulator—Learning to Mimic Malicious Storage Traces”, arXiv:2403.07540v2 [cs.CR], Jun. 12, 2024, 22 pages, https://arxiv.org/abs/2403.07540v1#. [cited by applicant]
Gagulic et al. “Ransomware Detection with Machine Learning in Storage Systems”, University of Zurich Department of Informatics (IFI) Binzmühlestrasse 14, CH-8050 Zürich, Switzerland, Feb. 13, 2023, 114 pages, https://fi… [cited by applicant]
Kruegel Christopher. “Full System Emulation: Achieving Successful Automated Dynamic Analysis of Evasive Malware”, Lastline, 2014, 7 pages. [cited by applicant]
Lucia Theo. “An Ultimate Guide to Hard Drive Problems, Solutions and Tips”, The Wayback Machine, Jan. 13, 2021, 21 pages, https://web.archive.org/web/20210122072236/https://recoverit.wondershare.com/computer-problem/com… [cited by applicant]
Rout Sidhartha Sankar. “Reliability Aware Intelligent Memory Management (RAIMM)”, Indraprastha Institute of Information Technology Delhi, New Delhi, 2014, 54 pages, https://repository.iiitd.edu.in/jspui/handle/123456789… [cited by applicant]
Sharma Natasha. “K-Means Clustering Explained”, Neptune Blog, Apr. 15, 2024, 25 pages. [cited by applicant]
Sharma Pulkit. “An Introduction to K-Means Clustering”, Machine Learning, Dec. 18, 2024, 43 pages. [cited by applicant]
Taylor et al. “Sensor-based Ransomware Detection”, Future Technologies Conference (FTC), Nov. 29-30, 2017, 8 pages, https://s2.smu.edu/~mitch/ftp_dir/pubs/ftc17.pdf. [cited by applicant]
Wijaya Cornellius Yudha. “LLMs Implementation for Tabular Classifications”, Trying out ML Tabular Classification Task with LLM, Nov. 16, 2023, 9 pages. [cited by applicant]
Zhang et al. “Towards Foundation Models for Learning on Tabular Data”, arXiv:2310.07338 [cs.LG], Oct. 11, 2023, 22 pages. [cited by applicant]
Authors et al.: Disclosed Without Attribution, IP.com No. IPCOM000252417D, “Method and System for Classifying Memory Devices and Slices to Create Optimal Storage Decisions in Distributed Storage Network (DSN)”, Jan. 9, … [cited by applicant]
Authors et al.: Disclosed Without Attribution, IP.com No. IPCOM000263306D, “Dispersed Storage Namespace Health Based Device Prioritization and Recovery Algorithm”, Aug. 17, 2020, 7 pages. [cited by applicant]
Authors et al.: Disclosed Without Attribution, IP.com No. IPCOM000263385D, “Methodology of Hinting the NVME Host about the NVME Queues Optimized for NVME Namespace Operation”, Aug. 26, 2020, 7 pages. [cited by applicant]
Authors et al.: Seagate Technology, LLC, IP.com No. IPCOM000268073D, “Staged NVMe in a Primary/Secondary Relationship”, Dec. 21, 2021, 4 pages. [cited by applicant]
Han et al. “ZNS+: Advanced Zoned Namespace Interface for Supporting In-Storage Zone Compaction”, Proceedings of the 15th USENIX Symposium on Operating Systems Design and Implementation, Jul. 14-16, 2021, 17 pages. [cited by applicant]
Kazi, et al. “Namespace Range-Based Memory Device Compaction,” U.S. Appl. No. 19/078,385, filed Mar. 13, 2025, 29 pages. [cited by applicant]
Kazi, et al. “Upgrade Orchestration of a Storage System Based on Namespace Range Gaps,” U.S. Appl. No. 18/938,502, filed Nov. 6, 2024, 40 pages. [cited by applicant]
Min et al. “eZNS: Elastic Zoned Namespace for Enhanced Performance Isolation and Device Utilization”, ACM Transactions on Storage, Jun. 6, 2024, pp. 1-41, vol. 20, Issue 3. [cited by applicant]
Tamborski, et al. “Identifying and Visualizing Namespace Range Gaps,” U.S. Appl. No. 18/938,511, filed Nov. 6, 2024, 41 pages. [cited by applicant]
Tamborski, et al. “Processing Namespace Range Information by a Global Coordinator,” U.S. Appl. No. 18/938,523, filed Nov. 6, 2024, 27 pages. [cited by applicant]
Wang David. “Application Optimization with Flexible Zone Namespace Configurations in QLC-Based SSDs”, Silicon Motion, 2023, 16 pages. [cited by applicant]
United States Notice of Allowance dated Apr. 27, 2026, 48 pages, in U.S. Appl. No. 19/078,385. [cited by applicant]
United States Non-Final Rejection dated Feb. 9, 2026, 45 pages, in U.S. Appl. No. 18/938,502. [cited by applicant]
United States Notice of Allowance dated Mar. 10, 2026, 17 pages, in U.S. Appl. No. 18/938,511. [cited by applicant]
United States Non-Final Rejection dated Dec. 17, 2025, 33 pages, in U.S. Appl. No. 18/938,511. [cited by applicant]
Rahi Rohit, “Object Storage”, Oct. 2019, 18 pages, https://www.oracle.com/a/ocom/docs/cloud/object-storage-100.pdf. [cited by applicant]
SNIA, “The SNIA Dictionary”, Mar. 2022, 02 pages, https://www.snia.org/sites/default/files/dictionary/SNIADictionary.pdf. [cited by applicant]
United States Non-Final Rejection dated Oct. 27, 2025, 15 pages, in U.S. Appl. No. 18/938,523. [cited by applicant]