IP Library Granted Patent US 8,812,907
Granted Patent B1
US 8,812,907 · App. 13/186,087 · Granted Aug 19, 2014

Fault tolerant computing systems using checkpoints

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,812,907
App. No.
13/186,087
Granted
Aug 19, 2014
Kind
B1
Abstract

A computer system configured to provide fault tolerance includes a first host system and a second host system. The first host system is programmed to monitor a number of portions of memory of the first host system that have been modified by a guest running on the first host system and, upon determining that the number of portions exceeds a threshold level, determine that a checkpoint needs to be created. Upon determining that the checkpoint needs to be created, operation of the guest is paused and checkpoint data is generated. After generating the checkpoint data, operation of the guest is resumed while the checkpoint data is transmitted to the second host system.

Claims (43)

1. A computer system configured to provide fault tolerance, the computer system comprising a first host system and a second host system, wherein the first host system comprises a first processor and a memory and wherein the first processor is programmed to:

monitor a number of portions of memory of the first host system that have been modified by a guest running on the first host system and, upon determining that the number of portions exceeds a threshold level, determine that a checkpoint needs to be created, and upon determining that the number of portions does not exceed the threshold level, determine that no checkpoint needs to be created and return to monitoring the number of portions of memory that have been modified;

upon determining that the checkpoint needs to be created, pause operation of the guest and generate checkpoint data; and

after generating the checkpoint data, resume operation of the guest and transmit the checkpoint data to the second host system;

wherein operation of the guest is resumed while the checkpoint data is being transmitted to the second host system, and

wherein the threshold level is a predetermined number of portions of memory.

2. The computer system of claim 1 , wherein the first host system comprises multiple processors generating the checkpoint data.

3. The computer system of claim 1 , wherein the checkpoint data includes data corresponding to all portions of memory of the first host system that have been modified since a previous checkpoint was generated.

4. The computer system of claim 3 , wherein the checkpoint data also includes data representing an operating state of the first host system.

5. The computer system of claim 1 , wherein the first host system is further programmed to determine that a checkpoint needs to be created based on network I/O activity of the guest running on the first host system.

6. The computer system of claim 5 , wherein the first processor of the first host system is further programmed to determine that a checkpoint needs to be created when the duration of a time period since a last previous checkpoint was created exceeds a specified level.

7. The computer system of claim 1 , wherein the first processor of the first host system is programmed to monitor a number of portions of memory of the first host system that have been modified by a guest running on the first host system by:

setting permissions for all pages of memory to be read only such that a fault is generated when a page is accessed for modification; and

in response to a fault generated by attempted modification of a page of memory set to be read only:

adding the page to a list of pages that have been modified;

setting the permissions for the page to be read/write; and

allowing the modification to proceed.

8. The computer system of claim 7 , wherein, in response to a fault generated by attempted modification of a page of memory set to be read only, the first processor of the first host system is further programmed to add an entry corresponding to the page to a list that is used in setting permissions for all pages of memory to be read only when a checkpoint is generated.

9. A computer system configured to provide fault tolerance, the computer system comprising a first host system and a second host system, wherein the first processor of the first host system is programmed to:

monitor network I/O activity by a guest running on the first host system and, upon determining that a threshold level of network I/O activity has occurred, determine that a checkpoint needs to be created, and upon determining that the number of portions does not exceed the threshold level, determine that no checkpoint needs to be created and return to monitoring the number of portions of memory that have been modified;

upon determining that the checkpoint needs to be created, pause operation of the guest and generate checkpoint data; and

after generating the checkpoint data, resume operation of the guest and transmit the checkpoint data to the second host system;

wherein operation of the guest is resumed while the checkpoint data is being transmitted to the second host system.

10. A method of implementing a fault tolerant computer system using a first host system comprising a first processor and a memory and a second host system, the method comprising, at the first host system:

monitoring, using the first processor, a number of portions of memory of the first host system that have been modified by a guest running on the first host system and, upon determining that the number of portions exceeds a threshold level, determining that a checkpoint needs to be created, and upon determining that the number of portions does not exceed the threshold level, determine that no checkpoint needs to be created and return to monitoring the number of portions of memory that have been modified;

upon determining that the checkpoint needs to be created, pausing operation of the guest and generating checkpoint data; and

after generating the checkpoint data, resuming operation of the guest and transmitting the checkpoint data to the second host system;

wherein operation of the guest is resumed while the checkpoint data is being transmitted to the second host system, and

wherein the threshold level is a predetermined number of portions of memory.

11. The method of claim 10 , wherein the first host system comprises multiple processors generating the checkpoint data.

12. The method of claim 10 , wherein the checkpoint data includes data corresponding to all portions of memory of the first host system that have been modified since a previous checkpoint was generated.

13. The method of claim 12 , wherein the checkpoint data also includes data representing an operating state of the first host system.

14. The method of claim 10 , further comprising determining that a checkpoint needs to be created based on network I/O activity of the guest running on the first host system.

15. The method of claim 14 , further comprising determining that a checkpoint needs to be created when the duration of a time period since a last previous checkpoint was created exceeds a specified level.

16. The method of claim 10 , wherein monitoring a number of portions of memory by the first processor of the first host system that have been modified by a guest running on the first host system comprises:

setting permissions for all pages of memory to be read only such that a fault is generated when a page is accessed for modification; and

in response to a fault generated by attempted modification of a page of memory set to be read only:

adding the page to a list of pages that have been modified;

setting the permissions for the page to be read/write; and

allowing the modification to proceed.

17. The method of claim 16 , further comprising, in response to a fault generated by attempted modification of a page of memory set to be read only, adding an entry corresponding to the page to a list that is used in setting permissions for all pages of memory to be read only when a checkpoint is generated.

18. The system of claim 7 wherein the permission of the page is reset using a reverse page table entry list.

19. The method of claim 16 further comprising the step of resetting the permission of the page using a reverse page table entry list.

Assignments (8)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (057254/0557) Recorded Aug 29, 2022
From: CERBERUS BUSINESS FINANCE AGENCY, LLC
To: STRATUS TECHNOLOGIES IRELAND LIMITED; STRATUS TECHNOLOGIES BERMUDA LTD.
Reel/Frame 061354/0599 →
GRANT OF SECURITY INTEREST IN PATENT RIGHTS Recorded Jun 9, 2021
From: STRATUS TECHNOLOGIES IRELAND LIMITED; STRATUS TECHNOLOGIES BERMUDA LTD.
To: CERBERUS BUSINESS FINANCE AGENCY, LLC, AS COLLATERAL AGENT
Reel/Frame 057254/0557 →
SECURITY INTEREST Recorded Apr 3, 2020
From: STRATUS TECHNOLOGIES IRELAND LIMITED
To: TRUST BANK (AS SUCCESSOR BY MERGER TO SUNTRUST BANK)
Reel/Frame 052316/0371 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2020
From: STRATUS TECHNOLOGIES BERMUDA LTD.
To: STRATUS TECHNOLOGIES IRELAND LTD.
Reel/Frame 052210/0411 →
SECURITY INTEREST Recorded Apr 28, 2014
From: STRATUS TECHNOLOGIES BERMUDA LTD.
To: SUNTRUST BANK
Reel/Frame 032776/0595 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2012
From: CITRIX SYSTEMS, INC.
To: STRATUS TECHNOLOGIES BERMUDA LTD.
Reel/Frame 029518/0502 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2011
From: MARATHON TECHNOLOGIES CORPORATION
To: CITRIX SYSTEMS, INC.
Reel/Frame 026975/0827 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2011
From: BISSETT, THOMAS D.; LEVEILLE, PAUL A.; LIN, TED M.; MELNICK, JERRY; PAGAN, ANGEL L.; TREMBLAY, GLENN A.
To: MARATHON TECHNOLOGIES CORPORATION
Reel/Frame 026747/0400 →