IP Library › Granted Patent US 12,711,030
Granted Patent B2
US 12,711,030 · App. 18/399,466 · Granted Aug 18, 2026

Method to detect boot failure of peer compute node by determining if watchdog timers exceeds threshold of normal boot operation

Inventor: Derek Shute (Ashland, MA)
Assignee: STRATUS TECHNOLOGIES IRELAND LTD.
G06F11/203G06F11/2023G06F11/2028G06F11/2038G06F11/2048G06F11/2094G06F11/1417
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,711,030
App. No.
18/399,466
Filed
Dec 28, 2023
Granted
Aug 18, 2026
Kind
B2
Art Unit
2184
USPC
713/2
Abstract

In part, in one aspect, the disclosure relates to a method of monitoring an active compute node in a fault tolerant system. The method may include monitoring a boot process of a first compute node by interrogating one or more watchdog timers running on the first compute node during the boot process; determining a boot process outcome for the first compute node, wherein the boot process outcome is a failure if the one or more of the watchdog timers exceeds a threshold time indicative of normal boot operation; performing a first action if the outcome is a failure; and performing a second action if the outcome is not a failure.

Claims (30)

1 . A method of monitoring an active compute node in a fault tolerant system comprising:

monitoring a boot process of a first compute node, using a second compute node of the fault tolerant system, by interrogating one or more watchdog timers running on the first compute node during the boot process, wherein the first compute node is the active compute node;

determining a boot process outcome for the first compute node, wherein the boot process outcome is a failure if the one or more of the watchdog timers exceeds a threshold time indicative of normal boot operation;

performing a first action if the outcome is a failure; and

performing a second action if the outcome is not a failure, wherein the second action is initiating a migration process such that the second compute node can take over for the first compute node with the second compute node becoming a new active compute node.

2 . The method of claim 1 , wherein determining the boot process outcome is performed using the second compute node of the fault tolerant system.

3 . The method of claim 2 , wherein the first compute node and the second compute node are connected using a communication interface, wherein watchdog timer values are received at the second compute node from the first compute node using the communication interface.

4 . The method of claim 3 , wherein the communication interface is an Intelligent Platform Management Interface (IPMI), wherein the first compute node comprises a first BMC, wherein the second compute node comprises a second BMC, wherein the second BMC performs the interrogation of the one or more watchdog times.

5 . The method of claim 1 , wherein the first action is rebooting the first compute node.

6 . The method of claim 1 , wherein the first action is generating an alert or a log indicative of a problem with the first computer node.

7 . The method of claim 1 , wherein the second action is reprovisioning a plurality of peripheral devices from the first compute node to the second compute node.

8 . The method of claim 1 , wherein the boot process comprises a Unified Extensible Firmware Interface (UEFI) control phase and an operating system (OS) control phase.

9 . The method of claim 8 , wherein the one or more watchdog timers, comprises a first watchdog timer and a second watchdog timer that are started during the UEFI control phase of the boot process.

10 . The method of claim 9 , wherein the one or more watchdog timers further comprises a third watchdog timer that is started during the OS control phase of the boot process.

11 . The method of claim 10 , wherein in response to all of the first, second, and third watchdog timers not exceeding their respective threshold values, the second action is for the first compute node to remain the active node.

12 . The system of claim 1 , wherein the fault tolerant system comprises the first computer node and the second compute node.

13 . A fault tolerant computer system comprising:

a first compute node comprising a Fault Resilient Boot (FRB) hardware device, wherein the FRB hardware device is configured to generate a set of watchdog timers, wherein each watchdog timer in the set corresponds to a phase or an event in a boot process;

a communications interface;

a management processor; and

a second compute node, wherein the second compute node is in communication with the FRB hardware device using the communications interface,

wherein the second compute node comprises the management processor,

the management processor configured to read each watchdog timer in the set of watchdog timers,

the management processor configured to determine if each watchdog timer exceeds a threshold time for its respective phase of the boot process, wherein not exceeding the threshold time for each respective timer corresponds to the normal operation of the compute node during the corresponding phase of the boot process.

14 . The system of claim 13 , wherein the communications interface is an intelligent platform management bus or an IPMI.

15 . The system of claim 14 , wherein the management processor is a BMC.

16 . The system of claim 15 , wherein the second compute node is configured to take control of the first compute node in response to one or more of the threshold times for each respective watchdog timer being exceeded.

17 . The system of claim 13 wherein the second compute node is configured to reboot the first compute node in response to one or more of the threshold times for each respective watchdog timer being exceeded.

18 . The system of claim 13 wherein the second compute node is configured to initiate a transfer of processor state and memory information from the first compute node to the second compute node in response to one or more of the threshold times for each respective watchdog timer being exceeded.

19 . The system of claim 13 wherein the second compute node is configured to generate a report or an error log in response to one or more of the threshold times for each respective watchdog timer being exceeded.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2025
From: SHUTE, DEREK
To: STRATUS TECHNOLOGIES IRELAND LTD.
Reel/Frame 072716/0587 →
Continuity (2)
Provisional Application 63545153 · Oct 20, 2023
Related Publication 20250130899A1 · Apr 24, 2025
References Cited (75)
US 6355991B1 · Goff et al. · 2002 [cited by applicant]
US 6633996B1 · Suffin et al. · 2003 [cited by applicant]
US 6687851B1 · Somers et al. · 2004 [cited by applicant]
US 6691225B1 · Suffin · 2004 [cited by applicant]
US 6691257B1 · Suffin · 2004 [cited by applicant]
US 6708283B1 · Nevin et al. · 2004 [cited by applicant]
US 6718474B1 · Somers et al. · 2004 [cited by applicant]
US 6766413B2 · Newman · 2004 [cited by applicant]
US 6766479B2 · Edwards · 2004 [cited by applicant]
US 6802022B1 · Olson · 2004 [cited by applicant]
US 6813721B1 · Tetreault et al. · 2004 [cited by applicant]
US 6842823B1 · Olson · 2005 [cited by applicant]
US 6862689B2 · Bergsten et al. · 2005 [cited by applicant]
US 6874102B2 · Doody et al. · 2005 [cited by applicant]
US 6886171B2 · MacLeod · 2005 [cited by applicant]
US 6928583B2 · Griffin et al. · 2005 [cited by applicant]
US 6948010B2 · Somers et al. · 2005 [cited by applicant]
US 6970892B2 · Green et al. · 2005 [cited by applicant]
US 6971043B2 · McLoughlin et al. · 2005 [cited by applicant]
US 6996750B2 · Tetreault · 2006 [cited by applicant]
US 7065672B2 · Long et al. · 2006 [cited by applicant]
US 7496786B2 · Graham et al. · 2009 [cited by applicant]
US 7496787B2 · Edwards et al. · 2009 [cited by applicant]
US 7669073B2 · Graham et al. · 2010 [cited by applicant]
US 7904906B2 · Puthukattukaran et al. · 2011 [cited by applicant]
US 7958076B2 · Bergsten et al. · 2011 [cited by applicant]
US 8117495B2 · Graham · 2012 [cited by applicant]
US 8161311B2 · Wiebe · 2012 [cited by applicant]
US 8234521B2 · Graham et al. · 2012 [cited by applicant]
US 8271416B2 · Al-Biek et al. · 2012 [cited by applicant]
US 8312318B2 · Graham et al. · 2012 [cited by applicant]
US 8381012B2 · Wiebe · 2013 [cited by applicant]
US 8812907B1 · Bissett et al. · 2014 [cited by applicant]
US 9251002B2 · Manchek et al. · 2016 [cited by applicant]
US 9588844B2 · Bissett et al. · 2017 [cited by applicant]
US 9652338B2 · Bissett et al. · 2017 [cited by applicant]
US 9760442B2 · Bissett et al. · 2017 [cited by applicant]
US 10216598B2 · Haid et al. · 2019 [cited by applicant]
US 10360117B2 · Haid et al. · 2019 [cited by applicant]
US 11288143B2 · Horvath et al. · 2022 [cited by applicant]
US 11550655B2 · Yu · 2023 [cited by examiner]
US 11586504B2 · Lin · 2023 [cited by examiner]
US 11586514B2 · Pawlowski et al. · 2023 [cited by applicant]
US 20010042202A1 · Horrath et al. · 2001 [cited by applicant]
US 20020016935A1 · Bergsten et al. · 2002 [cited by applicant]
US 20020070717A1 · Pellegrino · 2002 [cited by applicant]
US 20030046670A1 · Marlow · 2003 [cited by applicant]
US 20030095366A1 · Pellegrino · 2003 [cited by applicant]
US 20060222125A1 · Edwards et al. · 2006 [cited by applicant]
US 20060222126A1 · Edwards et al. · 2006 [cited by applicant]
US 20060259815A1 · Graham et al. · 2006 [cited by applicant]
US 20060274508A1 · LaRiviere et al. · 2006 [cited by applicant]
US 20070011499A1 · Begsten et al. · 2007 [cited by applicant]
US 20070028144A1 · Graham et al. · 2007 [cited by applicant]
US 20070038891A1 · Graham · 2007 [cited by applicant]
US 20070106873A1 · Lally et al. · 2007 [cited by applicant]
US 20070174484A1 · Lussier et al. · 2007 [cited by applicant]
US 20090249129A1 · Femia · 2009 [cited by applicant]
US 20150205688A1 · Haid et al. · 2015 [cited by applicant]
US 20150263983A1 · Brennan et al. · 2015 [cited by applicant]
US 20170324609A1 · Hong et al. · 2017 [cited by applicant]
US 20180046480A1 · Dong et al. · 2018 [cited by applicant]
US 20180088962A1 · Balakrishnan · 2018 [cited by examiner]
US 20180143885A1 · Dong et al. · 2018 [cited by applicant]
US 20210034447A1 · Horvath et al. · 2021 [cited by applicant]
US 20210034464A1 · Dailey et al. · 2021 [cited by applicant]
US 20210034465A1 · Haid et al. · 2021 [cited by applicant]
US 20210034483A1 · Haid · 2021 [cited by applicant]
US 20210034523A1 · Dailey · 2021 [cited by applicant]
US 20210037092A1 · Cao · 2021 [cited by applicant]
US 20230185681A1 · Pawlowski et al. · 2023 [cited by applicant]
US 20230291721A1 · Robinson · 2023 [cited by applicant]
US 20240176739A1 · Alden et al. · 2024 [cited by applicant]
Dong et al., “COLO: COarse-grain LOck-stepping Virtual Machine for Non-stop Service”, SoCC'13, Oct. 1-3, 2013, Santa Clara, California, USA, ACM 978-1-4503-2428-1; 16 pages. [cited by applicant]
Dong et al., “COLO: COarse-grain LOck-stepping Virtual Machine for Non-stop Service”, https://www.linux-kvm.org/images/1/1d/Kvm-forum-2013-COLO.pdf; 24 pages. [cited by applicant]