IP Library Granted Patent US 12,242,339
Granted Patent B2
US 12,242,339 · App. 17/462,151 · Granted Mar 4, 2025

Memory error processing method and apparatus

Inventors: Zhong Li (Hangzhou, CN); Jia Lou (Hangzhou, CN); Dongshu Zhou (Hangzhou, CN)
Assignee: XFUSION DIGITAL TECHNOLOGIES CO., LTD
G06F11/1024G06F9/4411G06F11/0727G06F11/0793G06F11/2023G06F9/4403
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,339
App. No.
17/462,151
Granted
Mar 4, 2025
Kind
B2
Abstract

In a memory error processing method, a processor of a computer apparatus obtains from a basic input/output system (BIOS) first error description information that describes a type of a first error that has occurred in a first memory page. Based on the first error description information, the processing device identifies the type of the first error to be a first type, wherein an error of the first type is a corrected error and is not a mirror scrub success error. The processor then determines that a number of errors of the first type that occurred in the first memory page has reached a threshold. In response to the determining, the processing device takes the first memory page offline.

Claims (36)

1. A memory error processing method performed by a computer apparatus, comprising:

obtaining first error description information from a basic input/output system (BIOS), wherein the first error description information describes a type of a first error that has occurred in a first memory page;

identifying, based on the first error description information, that the type of the first error is a first type, wherein an error of the first type is a correctable patrol error, a correctable read/write error, a correctable sparing error, or a mirror scrub failover error that is correctable;

determining that a number of errors of the first type that occurred in the first memory page has reached a threshold; and

in response to the determining that the number of errors of the first type that occurred in the first memory page has reached the threshold, taking the first memory page offline,

wherein, when an error occurred in the first memory page is a mirror scrub success error, the first memory page is not taken offline, the mirror scrub success error being a correctable error,

wherein, after the first memory page is taken offline, an accumulated quantity of times that a correctable error of a non-mirror scrub success error type occurs in the first memory page is cleared.

2. The method according to claim 1 , further comprising:

obtaining second error description information from the BIOS, wherein the second error description information describes a type of a second error that has occurred in a second memory page;

identifying, based on the second error description information, that the type of the second error is a second type, wherein an error of the second type is an uncorrectable error and is not a burst fatal error; and

taking the second memory page offline in response to identifying that the type of the second error is the second type.

3. The method according to claim 2 , wherein the second error is an uncorrectable no action (UCNA) error, a software recoverable action optional (SRAO) error, or a software recoverable action required (SRAR) error.

4. The method according to claim 3 , wherein the second memory page is used by an application, and taking the second memory page offline comprises taking the second memory page offline without closing the application.

5. The method according to claim 2 , wherein the second error is an uncorrectable patrol error.

6. The method according to claim 2 , wherein, when the error occurred in the first memory page is an uncorrectable error that is of a type other than the second type, the first memory page is not taken offline.

7. The method according to claim 1 , wherein, when the error occurred in the first memory page is a burst fatal error, the first memory page is not taken offline.

8. The method according to claim 1 , wherein, when the number of errors of the first type that occurred in the first memory page has not reached the threshold, the first memory page is not taken offline.

9. A computer apparatus, comprising:

at least one processor,

a basic input/output system (BIOS), and

a memory comprising a first memory page, wherein the memory is configured to store computer-executable instructions that, when executed by the at least one processor, cause the computer apparatus to:

obtain first error description information from the BIOS, wherein the first error description information describes a type of a first error that has occurred in the first memory page;

based on the first error description information, identify that the type of the first error is a first type, wherein an error of the first type is a correctable patrol error, a correctable read/write error, a correctable sparing error, or a mirror scrub failover error that is correctable;

determine that a number of errors of the first type that occurred in the first memory page has reached a threshold; and

in response to determining that the number of errors of the first type has reached the threshold, take the first memory page offline,

wherein, when an error occurred in the first memory page is a mirror scrub success error, the first memory page is not taken offline, the mirror scrub success error being a correctable error,

wherein, after the first memory page is taken offline, an accumulated quantity of times that a correctable error of a non-mirror scrub success error type occurs in the first memory page is cleared.

10. The apparatus according to claim 9 , wherein the computer-executable instructions, when executed by the at least one processor, further cause the computer apparatus to:

obtain second error description information from the BIOS, wherein the second error description information describes a type of a second error that has occurred in a second memory page;

based on the second error description information, identify that the type of the second error is a second type, wherein an error of the second type is an uncorrectable error and is not a burst fatal error; and

take the second memory page offline in response to identifying that the type of the second error is the second type.

11. The apparatus according to claim 10 , wherein the second error is an uncorrectable no action (UCNA) error, a software recoverable action optional (SRAO) error, or a software recoverable action required (SRAR) error.

12. The apparatus according to claim 11 , wherein the second memory page is used by an application, and taking the second memory page offline comprises taking the second memory page offline without closing the application.

13. The apparatus according to claim 10 , wherein the second error is an uncorrectable patrol error.

14. The apparatus according to claim 10 , wherein, when the error occurred in the first memory page is an uncorrectable error that is of a type other than the second type, the first memory page is not taken offline.

15. The apparatus according to claim 9 , wherein, when the error occurred in the first memory page is a burst fatal error, the first memory page is not taken offline.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2022
From: HUAWEI TECHNOLOGIES CO., LTD.
To: XFUSION DIGITAL TECHNOLOGIES CO., LTD.
Reel/Frame 058682/0312 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 2, 2021
From: LI, ZHONG; LOU, JIA; ZHOU, DONGSHU
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 058263/0141 →
Priority Claims (1)
CN 201910157218.0 · Mar 1, 2019 · national
Continuity (2)
Continuation PCTCN2020072925 · Jan 19, 2020
Related Publication 20210389956A1 · Dec 16, 2021
References Cited (9)
US 20140331015A1 · Prasad · 2014 [cited by examiner]
US 20150178142A1 · Raj · 2015 [cited by examiner]
US 20170103780A1 · Nakata · 2017 [cited by examiner]
US 20190188092A1 · Prasad · 2019 [cited by examiner]
US 20220050603A1 · Zhou · 2022 [cited by examiner]
US 20230185659A1 · Bao · 2023 [cited by examiner]
Google Patents/Scholar search—text refined (Year: 2023). [cited by examiner]
Reliability, Availability, and Serviceability (RAS), intel, 2020 https://www.intel.com/content/www/us/en/developer/articles/technical/pmem-RAS.html#:˜:text=Patrol%20scrub%20(also%20known%20as,the%20background%20during%2… [cited by examiner]
Google Scholar/Patents search—text refined (Year: 2024). [cited by examiner]