IP Library › Granted Patent US 12,117,895
Granted Patent B2
US 12,117,895 · App. 17/709,947 · Granted Oct 15, 2024

Memory error recovery using write instruction signaling

Inventor: Jue Wang (Redmond, WA)
Assignee: Google LLC
G06F11/0793G06F11/0712G06F11/0772G06F11/1484G06F2201/805
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,117,895
App. No.
17/709,947
Granted
Oct 15, 2024
Kind
B2
Abstract

A system and method for balancing data storage among a plurality of groups of computing devices, each group comprising one or more respective computing devices, each group having an available storage capacity. The method may involve, for each group of computing devices, determining an amount of used storage at the group of computing devices exceeding a predefined first threshold value that is less than the available storage capacity and calculating a storage cost based on the determined amount of used storage exceeding the predefined first threshold value, determining a total storage cost of the plurality of groups of computing devices based on a sum of the calculated storage costs, determining a transfer of one or more projects between the groups of computing devices that reduces the total storage and directing the plurality of groups of computing devices to execute the determined transfer.

Claims (44)

1. A method for memory error recovery comprising:

receiving, by a monitoring agent, an indication of a memory error generated in response to a write instruction at a virtual machine (VM) of a computing system; and

transmitting, by the monitoring agent, an instruction to a scheduler of the computing system to initiate migration of the VM in response to the memory error generated in response to the write instruction.

2. The method of claim 1 , wherein the indication of the memory error is a corrected machine check interrupt (CMCI) signal.

3. The method of claim 2 , further comprising:

determining, by the monitoring agent, that the CMCI signal is associated with an uncorrectable error, wherein transmitting the instruction to the scheduler is in response to the determination that the CMCI signal is associated with the uncorrectable error.

4. The method of claim 3 , wherein the monitoring agent determines that the CMCI signal is associated with the uncorrectable error and transmits the instruction to the scheduler on an order of milliseconds.

5. The method of claim 1 , wherein the monitoring agent transmits the instruction to the scheduler prior to a read instruction being executed at the VM.

6. The method of claim 1 , further comprising migrating, by one or more processors, the VM from a source machine to a target machine according to a migration instruction from the scheduler.

7. The method of claim 6 , wherein migrating the VM from the source machine to the target machine comprises:

copying memory associated with the source machine to the target machine;

detecting, during the copying, the memory error; and

injecting a software recoverable action optional (SRAO) machine check exception (MCE) into the copied memory at a memory page containing the memory error, whereby the memory page containing the memory error is isolated.

8. The method of claim 7 , wherein detecting the memory error and injecting the SRAO MCE are performed by a live migration pre-copy thread.

9. The method of claim 7 , wherein the SRAO MCE is injected to a single virtual processor core of the computing system.

10. The method of claim 6 , wherein migrating the VM from the source machine to the target machine comprises:

copying memory associated with the source machine to the target machine; and

determining whether a memory page containing the memory error is in use by one or more applications; and

in response to determining that the memory page is in use, setting a page fault, such that an attempt by the one or more applications to access the memory page avoids an MCE.

11. The method of claim 10 , further comprising, in response to determining that the memory page is not in use, unmapping the memory page, such that the memory page is invisible to the one or more applications.

12. A system for memory error recovery comprising

one or more processors; and

memory in communication with the one or more processors, wherein the memory contains instructions configured to cause the one or more processors to:

perform write error monitoring of data being written to a VM of the system; and

perform migration of the VM from a source machine to a target machine in response to an uncorrected error with no action (UCNA) detected by the write error monitoring, wherein the UCNA is a memory error.

13. A system of claim 12 , wherein the error monitoring includes:

receiving CMCI signaling; and

interpreting the CMCI signaling as the UNCA.

14. The system of claim 12 , wherein the write_error monitoring occurs on an order of milliseconds.

15. The system of claim 12 , wherein the instructions are configured to cause the one or more processors to perform migration of the VM by transmitting an instruction to a scheduler.

16. The system of claim 12 , wherein the instructions are further configured to cause the one or more processors to:

copy memory associated with the source machine to the target machine;

detect, during a read operation of the copying, the memory error; and

inject an SRAO MCE into the copied memory at a memory page containing the memory error.

17. The system of claim 16 , wherein the instructions are configured to cause the one or more processors to inject the SRAO MCE into a single virtual processor core of the target machine.

18. The system of claim 12 , wherein the instructions are further configured to cause the one or more processors to:

copy the memory error to a memory page of the target machine; and

either (i) set a page fault, such that an attempt by one or more applications to access the memory page avoids an MCE; or (ii) unmap the memory page, such that the memory page is invisible to the one or more applications.

19. The system of claim 18 , wherein the instructions are further configured to cause the one or more processors to:

determine whether a memory page to which the memory error is copied is in use by one or more applications; and

in response to a determination that the memory page is in use, set the page fault, such that the attempt by the one or more applications to access the memory page avoids the MCE.

20. The system of claim 18 , wherein the instructions are further configured to cause the one or more processors to:

determine whether a memory page to which the memory error is copied is in use by one or more applications; and

in response to a determination that the memory page is not in use, unmap the memory page, such that the memory page is invisible to the one or more applications.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2022
From: WANG, JUE
To: GOOGLE LLC
Reel/Frame 059468/0090 →
Continuity (1)
Related Publication 20230315561A1 · Oct 5, 2023