IP Library Granted Patent US 10,019,312
Granted Patent B2
US 10,019,312 · App. 15/583,019 · Granted Jul 10, 2018

Error monitoring of a memory device containing embedded error correction

Inventors: Michael B. Healy (Cortlandt Manor, NY); Hillery C. Hunter (Chappaqua, NY); Charles A. Kilmer (Essex Junction, VT); Kyu-hyoun Kim (Chappaqua, NY); Warren E. Maule (Cedar Park, TX)
Assignee: International Business Machines Corporation
G06F11/1044H03M13/1515H03M13/19
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,019,312
App. No.
15/583,019
Granted
Jul 10, 2018
Kind
B2
Abstract

Embodiments of the present disclosure provide an approach for monitoring the health and predicting the failure of dynamic random-access memory (DRAM) devices with embedded error-correcting code (ECC). Additional registers are embedded on the DRAM device to store information about the DRAM, such as the number and location of soft errors detected by the device. When the DRAM device detects a soft error, it will update the information stored in the additional registers. A controller compares the information stored in the additional registers to associated thresholds. In some embodiments, after comparing the information to the associated thresholds, the controller may determine whether to schedule a repair action. In other embodiments, the controller may determine whether to alert the memory controller that the DRAM may be failing.

Claims (40)

1. A dynamic random-access memory (DRAM) device having embedded error-correcting code (ECC), the DRAM device comprising:

a DRAM array;

a first register to store an error count;

a first register bank to store a set of error addresses;

a second register bank to store a plurality of decision parameters;

an ECC controller, wherein the ECC controller is configured to perform error detection using an ECC, increment the error count whenever an error is detected, and write an error address in an available register in the first register bank; and

a failure detection unit to predict failure of the DRAM device and to detect failing rows or columns of the DRAM device.

2. The DRAM device of claim 1 , further comprising a second register to store a multi-bit error count.

3. The DRAM device of claim 1 , further comprising a second register to store an uncorrectable error flag.

4. A method for logging and correcting dynamic random-access memory (DRAM) errors, the method comprising:

detecting an error in a word in a DRAM device using an error-correcting code;

incrementing, in response to detecting the error, an error count stored in a first register;

saving, in response to detecting the error, an error address corresponding to a location of the error in an available register in a first register bank;

determining a first error count at a first time;

determining a second error count at a second time, the second time being subsequent to the first time;

determining a number of new errors by comparing the second error count to the first error count;

determining whether the number of new errors is greater than a new error threshold; and

scheduling, in response to determining that the number of new errors exceeds the new error threshold, a repair action.

5. The method of claim 4 , further comprising:

determining whether the error is uncorrectable; and

setting, in response to the error being uncorrectable, an uncorrectable error flag.

6. A method for logging and correcting dynamic random-access memory (DRAM) errors, the method comprising:

detecting an error in a word in a DRAM device using an error-correcting code;

incrementing, in response to detecting the error, an error count stored in a first register;

saving, in response to detecting the error, an error address corresponding to a location of the error in an available register in a first register bank;

determining whether the error is in a first memory bank; and

incrementing, in response to the error being detected in the first memory bank, a first bank-specific error count, wherein the first bank-specific error count stores a number of errors in the first memory bank and is stored in a second register bank.

7. A method for predicting failure in a DRAM device, the method comprising:

receiving a memory information about the DRAM device;

processing, using a set of decision parameters, the memory information to determine an error indicator;

determining whether the error indicator exceeds an associated error threshold; and

alerting, in response to the error indicator exceeding the associated error threshold, a controller, wherein alerting the controller comprises raising a dedicated pin.

8. The method of claim 7 , wherein the error indicator is an error count.

9. The method of claim 7 , wherein the error indicator is an error rate.

10. The method of claim 7 , wherein the error indicator is an error acceleration.

11. A method for predicting failure in a DRAM device, the method comprising:

receiving a memory information about the DRAM device;

processing, using a set of decision parameters, the memory information to determine an error indicator;

determining whether the error indicator exceeds an associated error threshold; and

alerting, in response to the error indicator exceeding the associated error threshold, a controller, wherein alerting the controller comprises sending a predefined read data pattern to the controller.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2017
From: HEALY, MICHAEL B.; HUNTER, HILLERY C.; KILMER, CHARLES A.; KIM, KYU-HYOUN; MAULE, WARREN E.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 042192/0840 →
Continuity (3)
Continuation 15401744 · Jan 9, 2017
Continuation 14611351 · Feb 2, 2015
Related Publication 20170235632A1 · Aug 17, 2017
Cited By (2)
US 12,205,661 US 12,650,898