IP Library › Granted Patent US 12,277,029
Granted Patent B2
US 12,277,029 · App. 18/308,470 · Granted Apr 15, 2025

Systems and methods for predictive memory maintenance

Inventors: Tim Breitenbach (Mitgenfeld, DE); Patrick Jahnke (Leimen, DE)
Assignee: SAP SE
G06F11/1068G06F11/076G06F11/0772
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,029
App. No.
18/308,470
Filed
Apr 27, 2023
Granted
Apr 15, 2025
Kind
B2
Art Unit
2112
USPC
714/764
Abstract

Embodiments of the present disclosure include techniques for predictive memory maintenance. In one embodiment, locations of correctable errors in a memory are observed. A machine learning (ML) system may be trained with patterns of correctable errors that result in uncorrectable errors. A trained ML monitors correctable errors to predict when memory requires maintenance. In another embodiment, error rates from multiple memories are monitored to predict memory channel and other upstream device failures.

Claims (38)

1. A method of managing memory errors comprising:

receiving memory error locations;

processing a plurality of values corresponding to the memory error locations in a machine learning system configured to recognize patterns of correctable error locations resulting in uncorrectable errors; and

generating an alert when the machine learning system recognizes a pattern of the error locations corresponding to an uncorrectable error.

2. The method of claim 1 , wherein the memory error locations correspond to correctable memory errors.

3. The method of claim 1 , wherein the plurality of values is received in the machine learning system as a two-dimensional array of values, and wherein positions in the two-dimensional array correspond to the memory error locations.

4. The method of claim 3 , wherein the memory error locations are processed in predefined frames, and wherein the predefined frames are a subset of the two-dimensional array of values.

5. The method of claim 1 , wherein the memory error locations are received over a time period, and wherein the plurality of values aggregates the received memory error locations over the time period.

6. The method of claim 1 , wherein the plurality of values is a plurality of binary values.

7. The method of claim 1 , wherein the plurality of values is a plurality of values between a maximum value and a minimum value.

8. The method of claim 7 , wherein a first value of the plurality of values corresponds to a number of times a particular error occurs at a particular memory error location.

9. The method of claim 7 , wherein a first value of the plurality of values corresponds to an error rate.

10. The method of claim 7 , wherein a first value of the plurality of values corresponds to a number of bits read between errors for a particular memory error location.

11. The method of claim 1 , wherein the plurality of values corresponding to the memory error locations are a first plurality of values, the method further comprising processing a second plurality of values corresponding to the memory error locations in a machine learning system configured to recognize patterns of correctable error locations resulting in uncorrectable errors, wherein the first plurality of values corresponds to aggregated memory error locations over a first time period and the second plurality of values corresponds to aggregated memory error locations over a second time period.

12. The method of claim 1 , wherein the plurality of memory error locations are a first plurality of memory error locations and the plurality of values corresponding to the memory error locations are a first plurality of values, the method further comprising;

storing a second plurality of memory error locations;

converting the second plurality of memory error locations into a second plurality of values;

detecting uncorrectable errors for at least a portion of the second plurality of memory error locations; and

training the machine learning system using the second plurality of values and the detected uncorrectable errors to recognize patterns of correctable error locations resulting in uncorrectable errors.

13. The method of claim 12 , further comprising associating values corresponding to an aggregated plurality of memory error locations with an uncorrectable error value when an uncorrectable error occurs subsequent to a plurality of correctable errors in the aggregated plurality of memory error locations.

14. The method of claim 1 , wherein the machine learning system is a convolutional neural network.

15. The method of claim 1 , wherein the memory is a dynamic random access memory on a dual in-line memory module comprising a plurality of dynamic random access memories.

16. The method of claim 1 , wherein the memory error locations comprise first memory error data from a plurality of random access memories coupled to a memory controller over a memory channel, the method further comprising:

determining a memory error rate from the first memory error data;

comparing the memory error rate to a threshold; and

eliminating the first memory error data from the memory error locations when the memory error rate is above the threshold.

17. A system for managing memory errors comprising:

at least one processor;

at least one non-transitory computer readable medium storing computer executable instructions that, when executed by the at least one processor, cause the system to perform actions comprising:

receiving memory error locations;

processing a plurality of values corresponding to the memory error locations in a machine learning system configured to recognize patterns of correctable error locations resulting in uncorrectable errors; and

generating an alert when the machine learning system recognizes a pattern of the error locations corresponding to an uncorrectable error.

18. The system of claim 17 , wherein the memory is a dynamic random access memory on a dual in-line memory module comprising a plurality of dynamic random access memories.

19. A non-transitory computer-readable medium storing computer-executable instructions that, when executed by at least one processor, perform a method of managing memory errors, the method comprising:

receiving memory error locations;

processing a plurality of values corresponding to the memory error locations in a machine learning system configured to recognize patterns of correctable error locations resulting in uncorrectable errors; and

generating an alert when the machine learning system recognizes a pattern of the error locations corresponding to an uncorrectable error.

20. The non-transitory computer-readable medium of claim 19 , wherein the memory is a dynamic random access memory on a dual in-line memory module comprising a plurality of dynamic random access memories.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2023
From: BREITENBACH, TIM; JAHNKE, PATRICK
To: SAP SE
Reel/Frame 063477/0660 →
Continuity (1)
Related Publication 20240362113A1 · Oct 31, 2024
References Cited (46)
US 6009545A · Tsutsui · 1999 [cited by applicant]
US 6067259A · Handa · 2000 [cited by applicant]
US 6108253A · Ohta · 2000 [cited by applicant]
US 10599504B1 · BeSerra · 2020 [cited by examiner]
US 10887331B2 · Nomura · 2021 [cited by applicant]
US 11537468B1 · Ghosh · 2022 [cited by applicant]
US 20020099996A1 · Demura · 2002 [cited by applicant]
US 20040083410A1 · Cutter · 2004 [cited by applicant]
US 20040098644A1 · Wuu · 2004 [cited by applicant]
US 20040243891A1 · Ohta · 2004 [cited by applicant]
US 20050160311A1 · Hartwell et al. · 2005 [cited by applicant]
US 20070247937A1 · Moriyama · 2007 [cited by applicant]
US 20090006886A1 · O'Connor et al. · 2009 [cited by applicant]
US 20090052330A1 · Matsunaga et al. · 2009 [cited by applicant]
US 20090144579A1 · Swanson · 2009 [cited by examiner]
US 20090319822A1 · Coronado et al. · 2009 [cited by applicant]
US 20110320864A1 · Gower et al. · 2011 [cited by applicant]
US 20130036325A1 · Matsui et al. · 2013 [cited by applicant]
US 20140359367A1 · Friman et al. · 2014 [cited by applicant]
US 20180046541A1 · Niu et al. · 2018 [cited by applicant]
US 20190004896A1 · Tokoyoda · 2019 [cited by applicant]
US 20200066353A1 · Pletka · 2020 [cited by applicant]
US 20200135293A1 · Raychaudhuri et al. · 2020 [cited by applicant]
US 20200183616A1 · Liikanen · 2020 [cited by applicant]
US 20200273499A1 · Jaenisch · 2020 [cited by applicant]
US 20210133022A1 · Chao · 2021 [cited by examiner]
US 20210318922A1 · Roberts · 2021 [cited by applicant]
US 20210357287A1 · Kim · 2021 [cited by examiner]
US 20220050603A1 · Zhou et al. · 2022 [cited by applicant]
US 20220156160A1 · Wang · 2022 [cited by examiner]
US 20220284980A1 · Chiang · 2022 [cited by applicant]
US 20230098902A1 · Desai et al. · 2023 [cited by applicant]
US 20230207045A1 · Harry-Hak-Lay et al. · 2023 [cited by applicant]
US 20230342245A1 · Zhang · 2023 [cited by applicant]
US 20240012726A1 · Thacker · 2024 [cited by applicant]
US 20240362101A1 · Breitenbach · 2024 [cited by examiner]
Du, Xiaoming et al., “Predicting Uncorrectable Memory Errors from the Correctable Error History: No Free Predictors in the Field”, The International Symposium on Memory Systems, 2021, 8 pages. [cited by applicant]
Du, Xiaoming et al., “Fault-Aware Prediction-Guided Page Offlining for Uncorrectable Memory Error Prevention”, 2021 IEEE 39th International Conference on Computer Design (ICCD), 2021, 8 pages. [cited by applicant]
Du, Xiaoming et al., “Predicting Uncorrectable Memory Errors for Proactive Replacement: An Empirical Study on Large-Scale Field Data”, 2020 16th European Dependable Computing Conference (EDCC). 2020, 6 pages. [cited by applicant]
Du, Yuyang et al., “A Rising Tide Lifts All Boats: How Memory Error Prediction and Prevention Can Help with Virtualized System Longevity”, Sixth Workshop on Hot Topics in System Dependability (HotDep 10), 2010, 6 pages. [cited by applicant]
Giurgiu, Iloana et al., “Predicting DRAM Reliability in the Field with Machine Learning”, Proceedings of the 18th Acm/IFIP/USENIX Middleware Conference: Industrial Track, 2017, 7 pages. [cited by applicant]
Meza, Justin et al., “Revisiting Memory Errors in Large-Scale Production Data Centers: Analysis and Modeling of New Trends from the Field”, 2015 45th Annual IEEE/IFIP International Conference on Dependable Systems and N… [cited by applicant]
Schroeder, Bianca et al., “DRAM Errors in the Wild: A Large-Scale Field Study”, ACM Sigmetrics Performance Evaluation Review, 2009, 12 pages. [cited by applicant]
Second Generation Intel Xeon Scalable Processors Datasheet, vol. Two: Registers, Section 3.4, pp. 41-48, Apr. 2019, 60 pages. [cited by applicant]
Intel 64 and IA-32 Architectures Software Developer's Manual, vol. 3B: System Programming Guide, Part 2, Intel Machine Check architecture described in Chapters 15-16, pp. 43-112, Sep. 2016, 582 pages. [cited by applicant]
Intel Xeon Processor E7 Family: Reliability, Availability, and Serviceability, Advanced data integrity and resiliency support for mission-critical deployments, 2011, 16 pages. [cited by applicant]