IP Library Granted Patent US 12,572,401
Granted Patent B2
US 12,572,401 · App. 18/018,653 · Granted Mar 10, 2026

Predictively addressing hardware component failures

Inventors: Sree Nandan Atur (Newark, CA); Ravi Kumar Alluboyina (Santa Clara, CA)
Assignee: Rakuten Symphony, Inc.
G06F11/008G06F11/0709G06F11/0712G06F11/0793
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,401
App. No.
18/018,653
Granted
Mar 10, 2026
Kind
B2
Abstract

The present invention extends to methods, systems, and computer program products for predictively addressing hardware component failures. Network packets can be received over time at a platform. Metrics derived from platform hardware components and derived from one or more workloads utilizing the platform hardware components can be monitored. Model training data can be formulated from the metrics. A health check model can be trained using the model training data. The health check model can be executed to compute a probability that a monitored platform hardware component is on a path to failure. It can be determined that the probability exceeds a threshold. A workload can be relocated from a pod containing the monitored platform hardware component to another pod. Additional network packets can be received over time at the platform. The workload can process data contained in the additional network packets at the other pod.

Claims (48)

1 . A computer implemented method comprising:

receiving network packets over time at a platform;

monitoring metrics derived from platform hardware components and derived from one or more workloads utilizing the platform hardware components and processing data contained in the network packets;

formulating model training data from the metrics;

training a health check model using the model training data;

automating execution of the health check model computing a probability that a monitored platform hardware component is on a path to failure;

determining that the probability exceeds a threshold;

relocating a workload from a pod containing the monitored platform hardware component to another pod;

receiving additional network packets over time at the platform; and

processing data contained in the additional network packets using the workload at the other pod.

2 . The method of claim 1 , wherein training the health check model comprises training a Recurrent Neural Network (RNN).

3 . The method of claim 1 , wherein training the health check model comprises training a Long Short-Term Memory (LSTM) model.

4 . The method of claim 1 , wherein computing the probability that the monitored platform hardware component is on the path to failure comprises detecting a probability that a network interface card is on a path to failure.

5 . The method of claim 1 , wherein computing the probability that the monitored platform hardware component is on the path to failure comprises detecting a probability that a storage device is on a path to failure.

6 . The method of claim 1 , wherein computing the probability that the monitored platform hardware component is on the path to failure comprises detecting a probability that one of: a processor unit, system memory or a field programmable gate array (FPGA) is on a path to failure.

7 . The method of claim 1 , wherein relocating the workload to the other pod comprises relocating the workload from one pod at a node to another pod at the node.

8 . The method of claim 1 , wherein relocating the workload to the other pod comprises relocating the workload from a pod at a current node to another pod at a destination node.

9 . The method of claim 8 , wherein relocating the workload from the pod at the current node to the other pod at the destination node comprises:

backing up a workload state of the workload from the pod at the current node;

deleting the workload from the pod at the current node;

selecting a destination pod;

importing the workload state to the destination pod; and

restoring the workload at the destination pod.

10 . A computer system comprising:

a processor; and

system memory coupled to the processor and storing instructions configured to cause the processor to:

receive network packets over time at a platform;

monitor metrics derived from platform hardware components and derived from one or more workloads utilizing the platform hardware components and processing data contained in the network packets;

formulate model training data from the metrics;

train a health check model using the model training data;

automate execution of the health check model computing a probability that a monitored platform hardware component is on a path to failure;

determine that the probability exceeds a threshold;

relocate a workload from a pod containing the monitored platform hardware component to another pod;

receive additional network packets over time at the platform; and

process data contained in the additional network packets using the workload at the other pod.

11 . The computer system of claim 10 , wherein the instructions configured to train the health check model comprise instructions configured to train a Recurrent Neural Network (RNN).

12 . The computer system of claim 10 , wherein the instructions configured to train the health check model comprise instructions configured to train a Long Short-Term Memory (LSTM) model.

13 . The computer system of claim 10 , wherein the instructions configured to compute the probability that the monitored platform hardware component is on the path to failure comprise instructions configured to detect a probability that a network interface card is on a path to failure.

14 . The computer system of claim 10 , wherein the instructions configured to compute the probability that the monitored platform hardware component is on the path to failure comprise instructions configured to detect a probability that a storage device is on a path to failure.

15 . The computer system of claim 10 , wherein the instructions configured to compute the probability that the monitored platform hardware component is on the path to failure comprise instructions configured to detect a probability that one of: a processor unit, a system memory, or a field programmable gate array (FPGA) is on a path to failure.

16 . The computer system of claim 10 , wherein the instructions configured to relocate the workload to the other pod comprise instructions configured to relocate the workload from one pod at a node to another pod at the node.

17 . The computer system of claim 10 , wherein the instructions configured to relocate the workload to the other pod comprise instructions configured to relocate the workload from a pod at a current node to another pod at a destination node.

18 . The computer system of claim 17 , wherein the instructions configured to relocate the workload from the pod at the current node to the other pod at the destination node comprise instructions configured to:

back up a workload state of the workload from the pod at the current node;

delete the workload from the pod at the current node;

select a destination pod;

import the workload state to the destination pod; and

restore the workload at the destination pod.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2024
From: ROBIN SYSTEMS, INC.
To: RAKUTEN SYMPHONY, INC.
Reel/Frame 068193/0367 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2023
From: ATUR, SREE NANDAN; ALLUBOYINA, RAVI KUMAR
To: ROBIN SYSTEMS, INC
Reel/Frame 062534/0335 →
Continuity (1)
Related Publication 20250307042A1 · Oct 2, 2025
References Cited (16)
US 7577090B2 · Xu · 2009 [cited by applicant]
US 9300548B2 · Asthana · 2016 [cited by applicant]
US 9483338B2 · Bhalla · 2016 [cited by applicant]
US 9628340B2 · Blair · 2017 [cited by applicant]
US 9645899B1 · Felstaine · 2017 [cited by applicant]
US 9774522B2 · Vasseur · 2017 [cited by applicant]
US 10048996B1 · Bell · 2018 [cited by applicant]
US 11212184B2 · Strom · 2021 [cited by applicant]
US 11709741B1 · Uppal · 2023 [cited by examiner]
US 20070186251A1 · Horowitz · 2007 [cited by applicant]
US 20110191462A1 · Smith · 2011 [cited by applicant]
US 20210096893A1 · Kottomtharayil et al. · 2021 [cited by applicant]
US 20210320936A1 · Wright et al. · 2021 [cited by applicant]
US 20220012112A1 · Wouhaybi · 2022 [cited by examiner]
CN 107069676B · 2019 [cited by applicant]
WO WO2017152763A1 · 2017 [cited by applicant]