IP Library Granted Patent US 11,847,029
Granted Patent B1
US 11,847,029 · App. 17/394,381 · Granted Dec 19, 2023

Log-based rollback-recovery

Inventors: Srinidhi Varadarajan (Blacksburg, VA); Joseph F. Ruscio (Blacksburg, VA)
Assignee: International Business Machines Corporation
G06F11/1458G06F11/1438
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,847,029
App. No.
17/394,381
Granted
Dec 19, 2023
Kind
B1
Abstract

Log-Based Rollback Recovery for system failures. The system includes a storage medium, and a component configured to transition through a series of states. The component is further configured to record in the storage medium the state of the component every time the component communicates with another component in the system, the system being configured to recover the most recent state recorded in the storage medium following a failure of the component.

Claims (49)

1. A system for improved recovery by avoiding replaying of non-deterministic events, the system comprising:

a hardware component configured to:

perform a process resulting in a transition of the hardware component through a series of states;

a spare component that performs one or more operations of the hardware component upon a failure of the hardware component; and

a recovery manager comprising a processor that when executing instructions stored in an associated memory is configured to:

identify the failure of the hardware component due to an occurrence of a non-deterministic event internal to a processing node comprising the hardware component that has occurred during a deterministic interval that is a period between a communication of a most recent state of the hardware component and a communication of a previous state of the hardware component,

identify a non-deterministic event external to the processing node,

wherein, in response to the identification of the failure, the recovery manager is further configured to:

recover the most recent state of the hardware component from a storage medium, wherein the most recent state includes information associated with the non-deterministic event external to the processing node, and

load the recovered most recent state into the spare component without the non-deterministic event internal to the processing node being replayed and communicated to other components external to the processing node, and

wherein after the loading, the spare component is configured to continue to perform the process based on the recovered most recent state.

2. The system of claim 1 , wherein the process has multiple threads.

3. The system of claim 2 , wherein at least two threads of the multiple threads share a common state as the hardware component transitions through the series of states.

4. The system of claim 3 , wherein the common state comprises:

an access by the at least two threads to a common resource.

5. The system of claim 1 , wherein the hardware component is further configured to:

perform multiple processes in parallel, the processes resulting in the hardware component transitioning through the series of states.

6. The system of claim 5 , wherein at least one process of the multiple processes comprises multiple threads.

7. The system of claim 5 , wherein at least two processes of the multiple processes share a common state as the hardware component transitions through the series of states.

8. The system of claim 7 , wherein the common state comprises:

an access by the at least two processes to a common resource.

9. A non-transitory computer-readable media comprising instructions that when executed by one or more processors cause the one or more processors to perform:

executing, by a hardware component, a process resulting in a transition of the hardware component through a series of states; and

identifying, by a recovery manager, a failure of the hardware component due to an occurrence of a non-deterministic event internal to a processing node comprising the hardware component that has occurred during a deterministic interval that is a period between a communication of a most recent state of the hardware component and a communication of a previous state of the hardware component;

identifying, by the recovery manager, a non-deterministic event external to the processing node;

wherein, in response to the identifying a failure, the instructions further cause the one or more processor to perform:

recovering, by the recovery manager, the most recent state of the component from a storage medium, wherein the most recent state includes information associated with the non-deterministic event external to the processing node, and

loading, by the recovery manager, the most recent state in a spare component that performs one or more operations upon a failure of the hardware component to continue performing the process using the recovered most recent state without replaying the non-deterministic event and communicating the non-deterministic event to other components external to the processing node.

10. The computer-readable media of claim 9 , wherein the process has multiple threads.

11. The computer-readable media of claim 10 , wherein at least two threads of the multiple threads share a common state as the hardware component transitions through the series of states.

12. The computer-readable media of claim 11 , wherein the common state comprises:

an access by the at least two threads to a common resource.

13. The computer-readable media of claim 9 , wherein the instructions further cause the one or more processors to perform:

executing, by the hardware component, configured to perform multiple processes in parallel, the processes resulting in the hardware component transitioning through the series of states.

14. The computer-readable media of claim 13 , wherein at least one process of the multiple processes comprises multiple threads.

15. The computer-readable media of claim 14 , wherein at least two processes of the multiple processes share a common state as the hardware component transitions through the series of states.

16. A method for improved recovery by avoiding replaying of non-deterministic events, the method comprising:

executing, by a hardware component, a process resulting in a transition of the hardware component through a series of states; and

identifying, by a recovery manager, a failure of the hardware component due to an occurrence of a non-deterministic event internal to a processing node comprising the hardware component that has occurred during a deterministic interval that is a period between a communication of a most recent state of the hardware component and a communication of a previous state of the hardware component;

identifying, by the recovery manager, a non-deterministic event external to the processing node;

wherein, in response to the identifying a failure, the method further comprises:

recovering, by the recovery manager, the most recent state of the component from a storage medium, wherein the most recent state includes information associated with the non-deterministic event external to the processing node, and

loading, by the recovery manager, the most recent state in a spare component that performs one or more operations upon a failure of the hardware component to continue performing the process using the recovered most recent state without replaying the non-deterministic event and communicating the non-deterministic event to other components external to the processing node.

17. The method of claim 16 , wherein the process has multiple threads.

18. The method of claim 17 , wherein at least two threads of the multiple threads share a common state as the hardware component transitions through the series of states.

19. The method of claim 18 , wherein the common state comprises:

an access by the at least two threads to a common resource.

20. The method of claim 18 , further comprising:

performing, by the hardware component, multiple processes in parallel, the processes resulting in the component transitioning through the series of states.

Assignments (5)
CORRECTIVE ASSIGNMENT TO CORRECT THE EFFECTIVE DATE OF THE PATENT ASSIGNMENT AGREEMENT DATED NOVEMBER 30, 2021 PREVIOUSLY RECORDED AT REEL: 058426 FRAME: 0791. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jan 14, 2022
From: OPEN INVENTION NETWORK LLC
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058736/0436 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2021
From: VARADARAJAN, SRINIDHI; RUSCIO, JOSEPH F.
To: EVERGRID, INC.
Reel/Frame 058467/0932 →
CHANGE OF NAME Recorded Dec 23, 2021
From: EVERGRID, INC.
To: LIBRATO, INC.
Reel/Frame 058585/0267 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2021
From: LIBRATO, INC.
To: OPEN INVENTION NETWORK LLC
Reel/Frame 058925/0793 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2021
From: OPEN INVENTION NETWORK LLC
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058426/0791 →
Continuity (5)
Continuation 16430934 · Jun 4, 2019
Continuation 14152806 · Jan 10, 2014
Continuation 12894877 · Sep 30, 2010
Continuation 11424350 · Jun 15, 2006
Provisional Application 60760026 · Jan 18, 2006