IP Library Granted Patent US 10,095,590
Granted Patent B2
US 10,095,590 · App. 15/147,083 · Granted Oct 9, 2018

Controlling the operating state of a fault-tolerant computer system

Inventors: Thomas D Bissett (Shirley, MA); Stephen J Wark (Shrewsbury, MA); Paul A Leveille (Grafton, MA); James D McCollum (Princeton, MA); Angel L Pagan (Holden, MA)
G06F11/1484G06F9/45558G06F11/301G06F11/302G06F11/3051G06F11/3495G06F2009/45591G06F2201/815
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,095,590
App. No.
15/147,083
Granted
Oct 9, 2018
Kind
B2
Abstract

A fault tolerant computer system having two virtual machines (VMs), each running on a separate host device, is connected over a network to one or more I/O devices. The system operates to monitor the health of one or more operational characteristics associated with each VM, and in the event that the health of both virtual machines dictates that one or the other of the VMs should be downgraded, but the system is not able to determine which VM should be downgraded and there is an imbalance in a monitored system operational characteristic, the system can defer downgrading one VM for a selected period of time during which the operational characteristic that is in imbalance is monitored. If the imbalance is resolved, the downgrade is cancelled, if an operational fault is confirmed prior to the expiration of the deferral period or if the deferral period expires, then one host is downgraded.

Claims (15)

1. A method of controlling the operational state of a first or a second virtual machine in a fault tolerant computer system, comprising:

monitoring, by a first availability manager running on the first virtual machine associated with a first host computer, a current health of a plurality of an immediately detectable and delayed detection operational characteristics that are associated with the first virtual machine;

monitoring, by a second availability manager running on the second virtual machine associated with a second host computer, a current health of a plurality of an immediately detectable and delayed detection operational characteristics that are associated with the second virtual machine, wherein the fault tolerant computer system is comprised of the first and the second host computers;

examining, by the first and second availability managers, the health of at least one of the monitored immediately detectible operational characteristics on each of the corresponding first and the second virtual machines, and if it is determined after the examination of the immediately detectible operational characteristics of the corresponding first and second virtual machines that the health of each of the first and second virtual machines is indistinguishably poor; then

examining, by the first and the second availability managers, the delayed detection operational characteristics associated with the corresponding first and second virtual machines and if the examination of the delayed detection operational characteristic associated with the first and the second virtual machines indicates that at least one of the plurality of the monitored delayed detection operational characteristics associated with the first virtual machine is in poor health which causes the first virtual machine to be in poor health relative to the health of the second virtual machine, then;

overriding, by the first availability manager, an operational state downgrade of the first virtual machine during a pre-specified deferral period of time while the health of the monitored delayed detection operational characteristic associated with the first virtual machine continues to be monitored.

2. The method of claim 1 , further comprising cancelling the operational state downgrade if during the pre-specified deferral period it is determined that the monitored delayed detection operational characteristic associated with the first virtual machine returns to good health.

3. The method of claim 1 , further comprising downgrading the operational state of the first virtual machine if during the pre-selected deferral period the delayed detection operational characteristic associated with the first virtual machine does not return to normal health.

4. The method of claim 1 , wherein the operational state of the first and second virtual machines is one of an active and on-line state, an on-line state, or off-line state.

5. The method of claim 1 , wherein the immediately detectable operational characteristic is any operational characteristic that can be used to qualify the operational state of either the first or the second virtual machine and that is detected by the first of the second availability managers prior to the expiration of a pre-specified period of time.

6. The method of claim 5 , wherein the pre-specified period of time is a time value that is less than the time it takes to detect that communication is lost between either the first of the second virtual machines and functionality that they rely upon to provide fault tolerant services.

7. The method of claim 1 , wherein the immediately detectable operational characteristic is any one of a loss of connectivity between the first and the second virtual machine, a loss of connectivity between either the first or the second virtual machine and another computational device, and the loss of connectivity between either the first and the second virtual machine and a network.

8. The method of claim 1 , wherein the delayed detection operational characteristic is any operational characteristic that can be used to qualify the operational state of either the first or the second virtual machine and that is detected by the fault tolerant computer subsequent to the expiration of a pre-specified period of time.

9. The method of claim 8 , wherein the pre-specified period of time is a time value that is greater than the time it takes to detect that communication is lost between the virtual machine and functionality that it relies upon to provide fault tolerant service.

10. The method of claim 1 , wherein the delayed detectable operational characteristic is the loss of connectivity between the first or the second virtual machine and a virtual container.

Assignments (8)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (057254/0557) Recorded Aug 29, 2022
From: CERBERUS BUSINESS FINANCE AGENCY, LLC
To: STRATUS TECHNOLOGIES IRELAND LIMITED; STRATUS TECHNOLOGIES BERMUDA LTD.
Reel/Frame 061354/0599 →
GRANT OF SECURITY INTEREST IN PATENT RIGHTS Recorded Jun 9, 2021
From: STRATUS TECHNOLOGIES IRELAND LIMITED; STRATUS TECHNOLOGIES BERMUDA LTD.
To: CERBERUS BUSINESS FINANCE AGENCY, LLC, AS COLLATERAL AGENT
Reel/Frame 057254/0557 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 17, 2020
From: STRATUS TECHNOLOGIES BERMUDA LTD
To: STRATUS TECHNOLOGIES IRELAND LTD
Reel/Frame 052960/0896 →
SECURITY INTEREST Recorded Apr 3, 2020
From: STRATUS TECHNOLOGIES IRELAND LIMITED
To: TRUST BANK (AS SUCCESSOR BY MERGER TO SUNTRUST BANK)
Reel/Frame 052316/0371 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2020
From: STRATUS TECHNOLOGIES BERMUDA LTD.
To: STRATUS TECHNOLOGIES IRELAND LTD.
Reel/Frame 052210/0411 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2017
From: STRATUS TECHNOLOGIES INC
To: STRATUS TECHNOLOGIES BERMUDA LTD
Reel/Frame 043248/0444 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2016
From: LEVEILLE, PAUL A; MCCOLLUM, JAMES D; BISSETT, THOMAS D
To: STRATUS TECHNOLOGIES, INC.
Reel/Frame 038703/0097 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2016
From: PAGAN, ANGEL L; WARK, STEPHEN J; LEVEILLE, PAUL A; MCCOLLUM, JAMES D
To: STRATUS TECHNOLOGIES, INC.
Reel/Frame 038543/0620 →
Continuity (2)
Provisional Application 62157826 · May 6, 2015
Related Publication 20160328302A1 · Nov 10, 2016