IP Library Granted Patent US 12,468,549
Granted Patent B2
US 12,468,549 · App. 18/209,012 · Granted Nov 11, 2025

Automated system for restarting large scale cluster supercomputers

Inventors: Elvis Nyamwange (Little Elm, TX); Sailesh Vezzu (Hillsborough, NJ); Amer Ali (Jersey City, NJ); Rahul Shashidhar Phadnis (Charlotte, NC); Rahul Yaksh (Austin, TX); Hari Vuppala (Concord, NC); Pratap Dande (Saint Johns, FL); Brian Neal Jacobson (Los Angeles, CA); Erik Dahl (Newark, DE)
Assignee: BANK OF AMERICA CORPORATION
G06F9/4416G06F9/442
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,549
App. No.
18/209,012
Granted
Nov 11, 2025
Kind
B2
Abstract

Systems, computer program products, and methods are described herein an automated system for restarting large scale cluster supercomputers. The present disclosure is configured to receive a request to reboot a supercomputer cluster, wherein the request comprises a sequence of reboot instructions; determine, using a data integrity engine, whether a current state of the supercomputer cluster meets reboot requirements, wherein the reboot requirements are associated with a core logic of the data integrity engine; and execute the sequence of reboot instructions in an instance where the current state of the supercomputer cluster meets the reboot requirements.

Claims (52)

1. A system an automated system for restarting large scale cluster supercomputers, the system comprising:

a processing device;

a non-transitory storage device containing instructions when executed by the processing device, causes the processing device to:

receive a request to reboot a supercomputer cluster, wherein the request comprises a sequence of reboot instructions;

determine, using a data integrity engine, whether a current state of the supercomputer cluster meets reboot requirements, wherein the reboot requirements are associated with a core logic of the data integrity engine, wherein the data integrity engine is a holochain application; and

execute the sequence of reboot instructions in an instance where the current state of the supercomputer cluster meets the reboot requirements, wherein executing further comprises retrieving information associated with computational tasks currently being processed by the supercomputer cluster, wherein the information associated with the computational tasks comprises at least a task priority, a task criticality, a task duration, and computational resource requirement.

2. The system of claim 1 , wherein executing the instructions further causes the processing device to:

receive, from a user input device, the sequence of reboot instructions; and

store the sequence of reboot instructions in a reboot sequence repository.

3. The system of claim 2 , wherein executing the instructions further causes the processing device to:

receive, from the user input device, the request to reboot the supercomputer cluster; and

retrieve, from the reboot sequence repository, the sequence of reboot instructions in response to receiving the request.

4. The system of claim 1 , wherein executing the instructions further causes the processing device to:

determine, from the current state of the supercomputer cluster, computational tasks currently being processed by the supercomputer cluster.

5. The system of claim 1 , wherein executing the instructions further causes the processing device to:

determine that the request is a valid signed entry; and

transmit the request and the current state of the supercomputer cluster to a distributed hash table (DHT) associated with the data integrity engine.

6. The system of claim 5 , wherein executing the instructions further causes the processing device to:

receive, from a plurality of peer computing nodes, an indication that the current state of the supercomputer cluster meets the reboot requirements;

determine that a total number of the plurality of peer computing nodes indicating that the current state of the supercomputer cluster meets the reboot requirements is greater than a predetermined threshold; and

determine a confirmation that the current state of the supercomputer cluster meets the reboot requirements in an instance in which the total number of the plurality of peer computing nodes is greater than the predetermined threshold.

7. A computer program product an automated system for restarting large scale cluster supercomputers, the computer program product comprising a non-transitory computer-readable medium comprising code causing an apparatus to:

receive a request to reboot a supercomputer cluster, wherein the request comprises a sequence of reboot instructions;

determine, using a data integrity engine, whether a current state of the supercomputer cluster meets reboot requirements, wherein the reboot requirements are associated with a core logic of the data integrity engine, wherein the data integrity engine is a holochain application; and

execute the sequence of reboot instructions in an instance where the current state of the supercomputer cluster meets the reboot requirements, wherein executing further comprises retrieving information associated with computational tasks currently being processed by the supercomputer cluster, wherein the information associated with the computational tasks comprises at least a task priority, a task criticality, a task duration, and computational resource requirement.

8. The computer program product of claim 7 , wherein the code further causes the apparatus to:

receive, from a user input device, the sequence of reboot instructions; and

store the sequence of reboot instructions in a reboot sequence repository.

9. The computer program product of claim 8 , wherein the code further causes the apparatus to:

receive, from the user input device, the request to reboot the supercomputer cluster; and

retrieve, from the reboot sequence repository, the sequence of reboot instructions in response to receiving the request.

10. The computer program product of claim 7 , the code further causes the apparatus to:

determine, from the current state of the supercomputer cluster, computational tasks currently being processed by the supercomputer cluster.

11. The computer program product of claim 7 , wherein the code further causes the apparatus to:

determine that the request is a valid signed entry; and

transmit the request and the current state of the supercomputer cluster to a distributed hash table (DHT) associated with the data integrity engine.

12. The computer program product of claim 11 , wherein the code further causes the apparatus to:

receive, from a plurality of peer computing nodes, an indication that the current state of the supercomputer cluster meets the reboot requirements;

determine that a total number of the plurality of peer computing nodes indicating that the current state of the supercomputer cluster meets the reboot requirements is greater than a predetermined threshold; and

determine a confirmation that the current state of the supercomputer cluster meets the reboot requirements in an instance in which the total number of the plurality of peer computing nodes is greater than the predetermined threshold.

13. A method an automated system for restarting large scale cluster supercomputers, the method comprising:

receiving a request to reboot a supercomputer cluster, wherein the request comprises a sequence of reboot instructions;

determining, using a data integrity engine, whether a current state of the supercomputer cluster meets reboot requirements, wherein the reboot requirements are associated with a core logic of the data integrity engine, wherein the data integrity engine is a holochain application; and

executing the sequence of reboot instructions in an instance where the current state of the supercomputer cluster meets the reboot requirements, wherein executing further comprises retrieving information associated with computational tasks currently being processed by the supercomputer cluster, wherein the information associated with the computational tasks comprises at least a task priority, a task criticality, a task duration, and computational resource requirement.

14. The method of claim 13 , wherein the method further comprises:

receiving, from a user input device, the sequence of reboot instructions; and

storing the sequence of reboot instructions in a reboot sequence repository.

15. The method of claim 14 , wherein the method further comprises:

receiving, from the user input device, the request to reboot the supercomputer cluster; and

retrieving, from the reboot sequence repository, the sequence of reboot instructions in response to receiving the request.

16. The method of claim 13 , wherein the method further comprises:

determining, from the current state of the supercomputer cluster, computational tasks currently being processed by the supercomputer cluster.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2023
From: NYAMWANGE, ELVIS; VEZZU, SAILESH; ALI, AMER; PHADNIS, RAHUL SHASHIDHAR; YAKSH, RAHUL; VUPPALA, HARI; DANDE, PRATAP; JACOBSON, BRIAN NEAL; DAHL, ERIK
To: BANK OF AMERICA CORPORATION
Reel/Frame 063988/0001 →
Continuity (1)
Related Publication 20240419454A1 · Dec 19, 2024
References Cited (25)
US 7765435B2 · Taylor · 2010 [cited by applicant]
US 7984282B2 · George · 2011 [cited by applicant]
US 8307243B2 · Archer · 2012 [cited by applicant]
US 9021299B2 · Douros · 2015 [cited by applicant]
US 9274902B1 · Morley · 2016 [cited by examiner]
US 9971713B2 · Asaad · 2018 [cited by applicant]
US 10025639B2 · Kozloski · 2018 [cited by applicant]
US 10423428B2 · Georges · 2019 [cited by applicant]
US 10599544B2 · Zhang · 2020 [cited by applicant]
US 10606681B2 · Boenisch · 2020 [cited by applicant]
US 10860367B2 · Mani · 2020 [cited by examiner]
US 10924368B1 · Kaddoura · 2021 [cited by applicant]
US 11270193B2 · Modha · 2022 [cited by applicant]
US 11327796B2 · Quintin · 2022 [cited by applicant]
US 11334392B2 · Quintin · 2022 [cited by applicant]
US 11860754B2 · Murray · 2024 [cited by applicant]
US 20030005068A1 · Nickel · 2003 [cited by applicant]
US 20100251259A1 · Howard · 2010 [cited by applicant]
US 20130080482A1 · Berkowitz · 2013 [cited by applicant]
US 20140278340A1 · Berkowitz · 2014 [cited by applicant]
US 20140372586A1 · Tannenbaum · 2014 [cited by applicant]
US 20190220285A1 · Ali · 2019 [cited by examiner]
US 20190332481A1 · Kulick · 2019 [cited by examiner]
US 20200389521A1 · Brock · 2020 [cited by examiner]
US 20210357270A1 · Karnawat · 2021 [cited by examiner]