IP Library › Granted Patent US 12,585,556
Granted Patent B2
US 12,585,556 · App. 17/937,537 · Granted Mar 24, 2026

Active component driven computational server reliability and failure prevention system

Inventors: Madhana Sunder (Meridian, ID); Christopher Muzzy (Burlington, VT); James Mansfield Crafts (Warren, VT); Noah Singer (White Plains, NY)
Assignee: International Business Machines Corporation
G06F11/2025G06F11/0793G06F11/2035G06F11/3006G06F2201/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,585,556
App. No.
17/937,537
Granted
Mar 24, 2026
Kind
B2
Abstract

An approach for managing and minimized failure of one or more devices in a computerized cluster and/or vehicle infrastructure is disclosed. The proactive approach for mitigating such black swan events as it relate to hardware failures (e.g., servers, network, vehicle systems/architecture, etc.). The approach would proactively inactivate “suspect” components (i.e., components that are completely functional in multiple systems) based on component vintage data from systems where components have failed or malfunctioned. A dedicated service is actively updating suspect components and their respective vintages spread across various systems. Furthermore, the approach backups components using different vintages are effectively utilized to avoid complete system failure.

Claims (71)

1 . A computer-implemented method for managing and minimized failure of one or more devices in a computerized cluster, the computer-implemented method comprising:

monitoring a health status of the one or more devices associated with the computerized cluster;

determining whether the one or more devices is malfunctioning;

in responsive to determining that the one or more devices is malfunctioning, determining whether there are backup components of the one or more devices;

in responsive to determining that there are backup components, determining whether the backup components have a different vintage information than the vintage information of the one or more malfunctioning devices;

in responsive to determining that the backup components does not have a different vintage than the one or more malfunctioning devices, activating the backup components; and

in responsive to determining that there are no backup components, managing the one or more malfunctioning devices.

2 . The computer-implemented method of claim 1 , wherein monitoring the health status further comprising:

receiving device data from the one or more devices, wherein device data comprises of connectivity status and operational status; and

updating a device database based on the received data, by changing suspect identifier associated with the one or more devices from a suspect designation to a non-suspect designation and vice versa.

3 . The computer-implemented method of claim 1 , wherein determining whether there are backup components of the one or more devices further comprising:

searching the device database;

identifying whether the backup components resides in the computerized cluster; and

responsive to identifying that the backup components does reside in the computerized cluster, identifying if the backup components are active or inactive based on connectivity.

4 . The computer-implemented method of claim 1 , wherein determining whether the backup components have a different vintage than from the one or more malfunctioning devices further comprising:

retrieving a list from the device database containing the one or more devices; and

comparing the vintage information of the one or more malfunctioning devices against the vintage information of the one or more backup components.

5 . The computer-implemented method of claim 1 further comprising:

in responsive to determining that the backup components is the same vintage as the one or more malfunctioning devices, activating the backup components, wherein the backup components can share the workload.

6 . The computer-implemented method of claim 1 , wherein updating the device database further comprising:

updating the health status in the device database by changing the suspect identifier associated with the one or more devices from a suspect designation to a non-suspect designation and vice versa.

7 . The computer-implemented method of claim 1 further comprising:

selecting the one or more devices with the suspect identifier with the suspect designation for failure analysis and test work based on active component vintage data from a field.

8 . A computer program product for managing and minimized failure of one or more devices in a computerized cluster further comprising:

one or more computer readable storage media having computer-readable program instructions stored on the one or more computer readable storage media, said program instructions executes a computer-implemented method comprising the steps of:

monitoring a health status of the one or more devices associated with the computerized cluster;

determining whether the one or more devices is malfunctioning;

in responsive to determining that the one or more devices is malfunctioning, determining whether there are backup components of the one or more devices;

in responsive to determining that there are backup components, determining whether the backup components have a different vintage information than the vintage information of the one or more malfunctioning devices;

in responsive to determining that the backup components does not have a different vintage than the one or more malfunctioning devices, activating the backup components; and

in responsive to determining that there are no backup components, managing the one or more malfunctioning devices.

9 . The computer program product of claim 8 , wherein monitoring the health status further comprising:

receiving device data from the one or more devices, wherein device data comprises of connectivity status and operational status; and

updating a device database based on the received data, by changing suspect identifier associated with the one or more devices from a suspect designation to a non-suspect designation and vice versa.

10 . The computer program product of claim 8 , wherein determining whether there are backup components of the one or more devices further comprising:

searching the device database;

identifying whether the backup components resides in the computerized cluster; and

responsive to identifying that the backup components does reside in the computerized cluster, identifying if the backup components are active or inactive based on connectivity.

11 . The computer program product of claim 8 , wherein determining whether the backup components have a different vintage than from the one or more malfunctioning devices further comprising:

retrieving a list from the device database containing the one or more devices; and

comparing the vintage information of the one or more malfunctioning devices against the vintage information of the one or more backup components.

12 . The computer program product of claim 8 , further comprising:

in responsive to determining that the backup components is the same vintage as the one or more malfunctioning devices, activating the backup components, wherein the backup components can share the workload.

13 . The computer program product of claim 8 , wherein updating the device database further comprising:

updating the health status in the device database by changing the suspect identifier associated with the one or more devices from a suspect designation to a non-suspect designation and vice versa.

14 . The computer program product of claim 8 , further comprising:

selecting the one or more devices with the suspect identifier with the suspect designation for failure analysis and test work based on active component vintage data from a field.

15 . A computer system for managing and minimized failure of one or more devices in a computerized cluster, the computer system comprising:

one or more computer processors;

one or more computer readable storage media;

one or more computer readable storage media having computer-readable program instructions stored on the one or more computer readable storage media for execution by at least one of the one or more computer processors, said program instructions executes a computer-implemented method comprising the steps of:

monitoring a health status of the one or more devices associated with the computerized cluster;

determining whether the one or more devices is malfunctioning;

in responsive to determining that the one or more devices is malfunctioning, determining whether there are backup components of the one or more devices;

in responsive to determining that there are backup components, determining whether the backup components have a different vintage information than the vintage information of the one or more malfunctioning devices;

in responsive to determining that the backup components does not have a different vintage than the one or more malfunctioning devices, activating the backup components; and

in responsive to determining that there are no backup components, managing the one or more malfunctioning devices.

16 . The computer system of claim 15 , wherein monitoring the health status further comprising:

receiving device data from the one or more devices, wherein device data comprises of connectivity status and operational status; and

updating a device database based on the received data, by changing suspect identifier associated with the one or more devices from a suspect designation to a non-suspect designation and vice versa.

17 . The computer system of claim 15 , wherein determining whether there are backup components of the one or more devices further comprising:

searching the device database;

identifying whether the backup components resides in the computerized cluster; and

responsive to identifying that the backup components does reside in the computerized cluster, identifying if the backup components are active or inactive based on connectivity.

18 . The computer system of claim 15 , wherein determining whether the backup components have a different vintage than from the one or more malfunctioning devices further comprising:

retrieving a list from the device database containing the one or more devices; and

comparing the vintage information of the one or more malfunctioning devices against the vintage information of the one or more backup components.

19 . The computer system of claim 15 , further comprising:

in responsive to determining that the backup components is the same vintage as the one or more malfunctioning devices, activating the backup components, wherein the backup components can share the workload.

20 . The computer system of claim 15 , wherein updating the device database further comprising:

updating the health status in the device database by changing the suspect identifier associated with the one or more devices from a suspect designation to a non-suspect designation and vice versa.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2022
From: SUNDER, MADHANA; MUZZY, CHRISTOPHER; CRAFTS, JAMES MANSFIELD; SINGER, NOAH
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061288/0490 →
Continuity (1)
Related Publication 20240111643A1 · Apr 4, 2024
References Cited (15)
US 7233989B2 · Srivastava · 2007 [cited by applicant]
US 7401263B2 · Dubois, Jr. · 2008 [cited by applicant]
US 7849368B2 · Srivastava · 2010 [cited by applicant]
US 8830031B2 · Atsutomo · 2014 [cited by applicant]
US 9600370B2 · Chiu · 2017 [cited by applicant]
US 10831619B2 · Park · 2020 [cited by applicant]
US 11176029B2 · Salame · 2021 [cited by applicant]
US 20210056001A1 · Dwarampudi · 2021 [cited by applicant]
US 20210334168A1 · Kanp · 2021 [cited by applicant]
CN 109582422A · 2019 [cited by applicant]
JP 2010165148A · 2010 [cited by examiner]
“Component Health Record Evaluation Framework,” An IP.com Prior Art Database Technical Disclosure, Authors et. al.: Disclosed Anonymously, IP.com No. IPCOM000222504D, IP.com Electronic Publication Date: Oct. 11, 2012, 6… [cited by applicant]
“Hardware Performance Monitoring & Management,” SolarWinds, Publish Date: Mar. 4, 2013, 14 pages, <https://www.solarwinds.com/resources/tech-tip/hardware-performance-monitoring-and-management>. [cited by applicant]
“Method and System for Providing Emergency Backup Services for IoT Devices During an Outage by Learning Critical and/or Important IoT Devices and Services,” An IP.com Prior Art Database Technical Disclosure, Authors et.… [cited by applicant]
Liu et al., “Active Message Oriented Adaptation Middleware for Collaborative Applications in Heterogeneous Environments,” IEEE, ICC 2008 Proceedings, pp. 1866-1870. [cited by applicant]