IP Library › Granted Patent US 12,517,798
Granted Patent B2
US 12,517,798 · App. 16/825,276 · Granted Jan 6, 2026

Techniques for memory error isolation

Inventors: Jonathon Stuart Ramsay Evans (Santa Clara, CA); Naveen Cherukuri (San Jose, CA); Jerome Francis Duluk, Jr. (Palo Alto, CA); Shailendra Singh (Fremont, CA); Vaibhav Vyas (Santa Clara, CA); Wishwesh Gandhi (Sunnyvale, CA); Arvind Gopalakrishnan (Fremont, CA); Manas Mandal (Palo Alto, CA)
Assignee: NVIDIA Corporation
G06F11/2094G06T1/20G06F2201/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,517,798
App. No.
16/825,276
Granted
Jan 6, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to detect memory errors and isolate or migrate partitions on a parallel processing unit using an application programming interface to facilitate parallel computing, such as CUDA. In at least one embodiment, interrupts are intercepted and processed on a graphics processing unit indicating a memory error for one or more partitions, and a policy is applied to isolate that memory error from other partitions.

Claims (49)

1 . One or more processors, comprising:

circuitry to:

cause one or more graphics processing unit (GPU) slices to be associated with a second one or more storage locations based, at least in part, on whether an error is detected within a first one or more storage locations; and

facilitate migration of the one or more GPU slices based, at least in part, on whether the error is detected within the first one or more storage locations.

2 . The one or more processors of claim 1 , wherein the one or more GPU slices are to be associated with a third one or more storage locations if an error is detected within the second one or more storage locations.

3 . The one or more processors of claim 1 , wherein the first one or more storage locations are indicated to contain the error if the error is detected within the first one or more storage locations.

4 . The one or more processors of claim 1 , wherein the one or more GPU slices are physically isolated computing slices of a GPU.

5 . The one or more processors of claim 1 , wherein if the error is detected within the first one or more storage locations, data values indicating the error are to be set in the first one or more storage locations.

6 . The one or more processors of claim 1 , wherein if the error is detected within the first one or more storage locations, the one or more GPU slices are prevented from accessing the first one or more storage locations containing the error.

7 . The one or more processors of claim 1 , wherein if the error is detected, the one or more GPU slices are to be attributed with the error.

8 . The processer one or more processors of claim 7 , wherein one or more software programs being performed by the one or more GPU slices are to be indicated as being performed by the one or more GPU slices attributed with the error.

9 . The one or more processors of claim 1 , wherein the circuitry is to cause the one or more GPU slices to start communicating with the second one or more storage locations.

10 . The one or more processors of claim 1 , wherein the one or more GPU slices are to collect error information and report said error information to one or more software programs being performed by the one or more GPU slices, the one or more software programs to be used to enforce error isolation policies for the one or more GPU slices.

11 . The one or more processors of claim 1 , wherein each of the one or more GPU slices comprise an individual interrupt table and interrupts generated by each of the one or more GPU slices are to be ignored by others of the one or more GPU slices.

12 . A system comprising:

one or more processors to cause one or more graphics processing unit (GPU) slices to be associated with a second one or more storage locations based, at least in part, on whether an error is detected within a first one or more storage locations; and

facilitate migration of the one or more GPU slices based, at least in part, on whether the error is detected within the first one or more storage locations.

13 . The system of claim 12 , wherein the one or more GPU slices are to be associated with a third one or more storage locations if an error is detected within the second one or more storage locations.

14 . The system of claim 12 , wherein the first one or more storage locations are indicated to contain the error if the error is detected within the first one or more storage locations.

15 . The system of claim 12 , wherein if the error is detected within the first one or more storage locations, the one or more GPU slices are to be prevented from accessing the first one or more storage locations containing the error.

16 . The system of claim 12 , wherein if the error is detected within the first one or more storage locations, data values indicating the error are set in the first one or more storage locations.

17 . The system of claim 12 , wherein if the error is detected, one or more software programs being performed by the one or more GPU slices are to be attributed with the error.

18 . The system of claim 12 , wherein the one or more processors are graphics processing units.

19 . The system of claim 12 , wherein the migration of the one or more GPU slices is facilitated, by an administrator of the one or more processors, from a first processor of the one or more processors to a second processor of the one or more processors based, at least in part, on whether the error is detected on the first processor.

20 . The system of claim 12 , wherein each of the one or more GPU slices on each of the one or more processors comprise an individual interrupt table, and interrupts generated by each of the one or more GPU slices are to be ignored by other GPU slices of the one or more GPU slices.

21 . A non-transitory machine-readable medium having stored thereon instructions which, if performed by one or more processors, cause the one or more processors to at least:

cause one or more graphics processing unit (GPU) slices to be associated with a second one or more storage locations based, at least in part, on whether an error is detected within a first one or more storage locations; and

facilitate migration of the one or more GPU slices based, at least in part, on whether the error is detected within the first one or more storage locations.

22 . The non-transitory machine-readable medium of claim 21 , wherein the one or more GPU slices are to be associated with a third one or more storage locations if an error is detected within the second one or more storage locations.

23 . The non-transitory machine-readable medium of claim 21 , wherein the first one or more storage locations are indicated to contain the error if the error is detected within the first one or more storage locations.

24 . The non-transitory machine-readable medium of claim 21 , wherein if the error is detected within the first one or more storage locations, the one or more GPU slices are to be prevented from accessing the first one or more storage locations containing the error by, at least in part, notifying a memory management unit that the first one or more storage locations are blacklisted.

25 . The non-transitory machine-readable medium of claim 21 , wherein if the error is detected within the first one or more storage locations, the instructions, if performed, are to cause the one or more processors to modify data values in the first one or more storage locations containing the error to indicate that the first one or more storage locations contain an error.

26 . The non-transitory machine-readable medium of claim 21 , wherein if the error is detected, one or more software programs being performed by the one or more GPU slices are to be indicated as being performed by the one or more GPU slices containing the error.

27 . The non-transitory machine-readable medium of claim 21 , wherein the instructions, if performed, further cause the one or more processors to cause the one or more GPU slices to collect information about the error and to provide the information to one or more software programs being performed by the one or more GPU slices, the one or more software programs providing the information to an administrator.

28 . The non-transitory machine-readable medium of claim 21 , wherein the instructions, if performed, further cause the one or more processors to cause the one or more GPU slices to start communicating with the second one or more storage locations.

29 . A method comprising:

causing one or more graphics processing unit (GPU) slices on one or more parallel processing units to be associated with a second one or more storage locations based, at least in part, on whether an error is detected within a first one or more storage locations; and

facilitate migration of the one or more GPU slices based, at least in part, on whether the error is detected within the first one or more storage locations.

30 . The method of claim 29 , wherein the one or more GPU slices are to access a third one or more storage locations if an error is detected within the second one or more storage locations.

31 . The method of claim 29 , wherein the first one or more storage locations are indicated to contain the error if the error is detected within the first one or more storage locations.

32 . The method of claim 29 , wherein the one or more GPU slices are physically isolated computing slices of a parallel processing unit.

33 . The method of claim 29 , wherein if the error is detected within the first one or more storage locations, data values indicating the error are to be set in the first one or more storage locations.

34 . The method of claim 29 , wherein if the error is detected within the first one or more storage locations, the one or more GPU slices are to be prevented from accessing the first one or more storage locations containing the error.

35 . The method of claim 29 , wherein if the error is detected, the one or more GPU slices are to be attributed with the error.

36 . The method of claim 35 , wherein one or more software programs being performed by the one or more GPU slices are to be indicated as being performed by the one or more GPU slices containing the error.

37 . The method of claim 29 , wherein the migration of the one or more GPU slices is from a first parallel processing unit of the one or more parallel processing units to a second parallel processing unit of the one or more parallel processing units based, at least in part, on whether the error is detected in the first one or more storage locations on the first parallel processing unit.

38 . The method of claim 37 , wherein the migration of the one or more GPU slices is by one or more software programs being performed by the one or more GPU slices to perform administration of the one or more parallel processing units.

39 . The method of claim 29 , wherein the one or more GPU slices are to collect information about the error and are to provide the information to one or more software programs being performed by the one or more GPU slices, the one or more software programs providing the information to an administrator.

40 . The method of claim 29 , wherein each of the one or more GPU slices comprise an individual interrupt table and interrupts generated by each of the one or more GPU slices are to be ignored by others of the one or more GPU slices.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2020
From: EVANS, JONATHON STUART RAMSAY; CHERUKURI, NAVEEN; DULUK, JEROME FRANCIS, JR.; SINGH, SHAILENDRA; VYAS, VAIBHAV; GANDHI, WISHWESH; GOPALAKRISHNAN, ARVIND; MANDAL, MANAS
To: NVIDIA CORPORATION
Reel/Frame 052628/0130 →
Continuity (1)
Related Publication 20210294707A1 · Sep 23, 2021
References Cited (25)
US 8156370B2 · Raghunandan · 2012 [cited by applicant]
US 10908987B1 · Pandey · 2021 [cited by examiner]
US 20060020850A1 · Jardine et al. · 2006 [cited by applicant]
US 20090282300A1 · Heyrman et al. · 2009 [cited by applicant]
US 20160342543A1 · Bonzini · 2016 [cited by examiner]
US 20170256018A1 · Gandhi et al. · 2017 [cited by applicant]
US 20180190365A1 · Luck · 2018 [cited by examiner]
US 20200230499A1 · Buser · 2020 [cited by examiner]
US 20200310758A1 · Desoli · 2020 [cited by examiner]
US 20210064429A1 · Stetter, Jr. · 2021 [cited by examiner]
US 20210124614A1 · Gupta · 2021 [cited by examiner]
CN 104662583A · 2015 [cited by applicant]
CN 105579971A · 2016 [cited by applicant]
CN 106095390A · 2016 [cited by applicant]
WO 2013101020A1 · 2013 [cited by applicant]
WO 2014081611A2 · 2014 [cited by applicant]
WO 2019074952A2 · 2019 [cited by applicant]
Wikipedia “CPU” page, retrieved from https://en.wikipedia.org/wiki/Central_processing_unit (Year: 2023). [cited by examiner]
Aggarwal et al., “Configurable Isolation: Building High Availability Systems with Commodity Multi-Core Processors,” ACM Sigarch Computer Architecture News, vol. 35, 2007, 12 pages. [cited by applicant]
United Kingdom Combined Search and Examination Report for Patent Application No. 2103842.7 dated Jan. 5, 2022, 7 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
United Kingdom Search Report for Patent Application No. GB211478.9 dated Apr. 12, 2023, 4 pages. [cited by applicant]
Office Action for Chinese Application No. 202110305077.X, mailed Jul. 19, 2024, 23 pages. [cited by applicant]
Office Action for Chinese Application No. 202110305077.X, mailed Jan. 14, 2025, 17 pages. [cited by applicant]
Decision of Rejection for Chinese Application No. 202110305077.X, mailed Mar. 24, 2025, 12 pages. [cited by applicant]