IP Library › Granted Patent US 12,277,040
Granted Patent B2
US 12,277,040 · App. 18/330,651 · Granted Apr 15, 2025

In-place recovery of fatal system errors at virtualization hosts

Inventors: Binit Ranjan Mishra (Kenmore, WA); Mukhtar Ahmed (Everett, WA); Christina Marianne Curlette (Redmond, WA); Steven Adrian West (Redmond, WA); Gaurav Jagtiani (Kirkland, WA); Naga Kiran Govindaraju (Medina, WA); James George Cavalaris (Bothell, WA); Drew Douglas Cross (Bothell, WA); Jason Stewart Wohlgemuth (Seattle, WA); James Anthony Schwartz, Jr. (Seattle, WA); Jennifer Marie Bourlier (Seattle, WA); Sri Harsha Kanukuntla (Morganville, NJ); Emma Sutherland Boyd (Richmond, VA); Scott Chao-Chueh Lee (Bellevue, WA); Vijaybalaji Madhanagopal (Redmond, WA); Terence Kwok Tak Chan (Redmond, WA); Yuri Dotsenko (Redmond, WA); Peter Hanpeng Jiang (Kirkland, WA); Aacer Hatem Daken (Renton, WA); Emily Nicole Wilson (Seattle, WA); Emily Cara Clemens (Snohomish, WA); Cody Dean Hartwig (Seattle, WA); Raz Meir Aloni (Seattle, WA); Sharon Scarlet Tang (Kirkland, WA); Minsang Kim (Bellevue, WA); Shen Wang (Sammamish, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F11/1471G06F11/0772G06F11/1441
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,040
App. No.
18/330,651
Granted
Apr 15, 2025
Kind
B2
Abstract

In-place recovery of fatal system errors at virtualization hosts. A device identifies an occurrence of a fatal system error in the first instance of a host operating system (OS) executing in a computer system. The device determines to perform an in-place recovery for the fatal system error. The device performs the in-place recovery, including pausing the execution of a virtual machine (VM) by the first instance of the host OS, preserving a state of the VM within system memory of the computer system, and resuming the execution of the VM by a second instance of the host OS executing in the computer system based on the state of the VM that is preserved within the system memory of the computer system.

Claims (61)

1. A method implemented in a computer system that includes a processor system, comprising:

identifying an occurrence of a fatal system error in a first instance of a host operating system (OS) executing in the computer system;

determining to perform an in-place recovery for the fatal system error, wherein the in-place recovery preserves a virtual machine (VM) in the computer system while transitioning from the first instance of the host OS to a second instance of the host OS executing at the computer system; and

performing the in-place recovery, including,

pausing execution of the VM by the first instance of the host OS;

preserving a state of the VM within a system memory of the computer system;

preserving a virtualization stack state from a first virtualization stack in the first instance of the host OS;

restoring the virtualization stack state into a second virtualization stack in the second instance of the host OS; and

resuming the execution of the VM by the second instance of the host OS using the state of the VM that is preserved within the system memory of the computer system.

2. The method of claim 1 , wherein the method further comprises holding normal handling of the fatal system error by the first instance of the OS.

3. The method of claim 1 , wherein preserving the state of the VM within the system memory of the computer system comprises preserving the state of the VM during a shutdown of the first instance of the host OS.

4. The method of claim 1 , wherein the second instance of the host OS replaces the first instance of the host OS within a root partition.

5. The method of claim 4 , wherein performing the in-place recovery further includes initiating a soft reboot of the first instance of the host OS to create the second instance of the host OS within the root partition.

6. The method of claim 1 , wherein the first instance of the host OS resides in a first partition and the second instance of the host OS resides in a second partition that is different from the first partition.

7. The method of claim 6 , wherein the second instance of the host OS is a standby host OS.

8. The method of claim 6 , wherein performing the in-place recovery further includes starting the second instance of the host OS within the second partition.

9. The method of claim 1 , wherein preserving the virtualization stack state comprises preserving the virtualization stack state within the system memory of the computer system.

10. The method of claim 1 , wherein determining to perform the in-place recovery for the fatal system error includes sending a request to a control plane.

11. The method of claim 1 , wherein determining to perform the in-place recovery for the fatal system error includes analyzing at least one of,

a fatal system error type,

a first result of a first prior in-place recovery for the fatal system error type,

a second result of a second prior in-place recovery at the computer system,

an interruption tolerance of a workload executing on the VM, or

a downtime threshold of a service level agreement (SLA) associated with the VM.

12. The method of claim 1 , wherein,

pausing execution of the VM by the first instance of the host OS comprises pausing execution of each VM of a plurality of VMs; and

preserving the state of the VM within the system memory of the computer system comprises preserving the state of each VM of the plurality of VMs.

13. The method of claim 12 , wherein resuming the execution of the VM by the second instance of the host OS comprises resuming the execution of each VM of the plurality of VMs.

14. The method of claim 12 , wherein resuming the execution of the VM by the second instance of the host OS comprises resuming the execution of a subset of the plurality of VMs.

15. A computer system comprising:

a processor system;

a system memory; and

a computer storage media that stores computer-executable instructions that are executable by the processor system to at least:

identify an occurrence of a fatal system error in a first instance of a host operating system (OS) executing in the computer system; and

perform an in-place recovery for the fatal system error, wherein the in-place recovery preserves a virtual machine (VM) in the computer system while transitioning from the first instance of the host OS to a second instance of the host OS executing at the computer system, including,

pausing execution of the VM by the first instance of the host OS;

preserving a state of the VM within the system memory of the computer system;

preserving a virtualization stack state from a first virtualization stack in the first instance of the host OS;

restoring the virtualization stack state into a second virtualization stack in the second instance of the host OS; and

resuming the execution of the VM by the second instance of the host OS using the state of the VM that is preserved within the system memory of the computer system.

16. The computer system of claim 15 , wherein the second instance of the host OS,

replaces the first instance of the host OS within a root partition; or

resides in a second partition that is different from a first partition in which the first instance of the host OS resides.

17. The computer system of claim 15 , the computer-executable instructions also executable by the processor system to determine to perform the in-place recovery for the fatal system error based on at least one of,

a fatal system error type,

a first result of a first prior in-place recovery for the fatal system error type,

a second result of a second prior in-place recovery at the computer system,

an interruption tolerance of a workload executing on the VM, or

a downtime threshold of a service level agreement (SLA) associated with the VM.

18. A method implemented in a computer system that includes a processor system comprising:

identifying an occurrence of a fatal system error in a first instance of a host operating system (OS) executing in the computer system;

holding normal handling of the fatal system error by the first instance of the OS;

determining not to perform an in-place recovery for the fatal system error, including analyzing at least one of,

a fatal system error type,

a first result of a first prior in-place recovery for the fatal system error type,

a second result of a second prior in-place recovery at the computer system,

an interruption tolerance of a workload executing on a virtual machine (VM) executing at the computer system, or

a downtime threshold of a service level agreement (SLA) associated with the VM; and

resuming the normal handling of the fatal system error by the first instance of the OS.

19. The method of claim 18 , wherein determining not to perform the in-place recovery for the fatal system error includes sending a request to a control plane.

20. The computer system of claim 15 , wherein preserving the virtualization stack state comprises preserving the virtualization stack state within the system memory of the computer system.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2023
From: AHMED, MUKHTAR; CURLETTE, CHRISTINA MARIANNE; GOVINDARAJU, NAGA KIRAN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065105/0707 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2023
From: MISHRA, BINIT RANJAN; WEST, STEVEN ADRIAN; JAGTIANI, GAURAV; CAVALARIS, JAMES GEORGE; CROSS, DREW DOUGLAS; WOHLGEMUTH, JASON STEWART; SCHWARTZ, JAMES ANTHONY, JR.; BOURLIER, JENNIFER MARIE; KANUKUNTLA, SRI HARSHA; BOYD, EMMA SUTHERLAND; LEE, SCOTT CHAO-CHUEH; MADHANAGOPAL, VIJAYBALAJI; CHAN, TERENCE KWOK TAK; DOTSENKO, YURI; JIANG, PETER HANPENG; DAKEN, AACER HATEM; WILSON, EMILY NICOLE; CLEMENS, EMILY CARA; HARTWIG, CODY DEAN; ALONI, RAZ MEIR; TANG, SHARON SCARLET; KIM, MINSANG; WANG, SHEN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 064641/0712 →
Continuity (2)
Provisional Application 63494205 · Apr 4, 2023
Related Publication 20240338282A1 · Oct 10, 2024
References Cited (5)
US 10860412B2 · Noe · 2020 [cited by examiner]
US 11307921B2 · Noe · 2022 [cited by examiner]
US 20160110210A1 · Vecera · 2016 [cited by examiner]
Russinovich, Mark, “Improving Azure Virtual Machines Resiliency with Project Tardigrade”, Retrieved From: https://azure.microsoft.com/en-us/blog/improving-azure-virtual-machines-resiliency-with-project-tardigrade/, Aug.… [cited by applicant]
Russinovich, Mark, “Inside Azure Datacenter Architecture with Mark Russinovich—BRK3060”, Retrieved From: https://www.youtube.com/watch?v=S2zguwKvlQk, May 9, 2019, 3 Pages. [cited by applicant]