IP Library › Granted Patent US 12,705,061
Granted Patent B2
US 12,705,061 · App. 19/016,197 · Granted Aug 11, 2026

Application programming interface to indicate accelerator error handlers

Inventors: Karthik Raghavan Ravi (Bangalore, IN); Ashutosh Jain (Bengaluru, IN); Rahul Suresh (Bangalore, IN)
Assignee: NVIDIA Corporation
G06F9/3861G06F9/3836G06F9/3877G06F9/541
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,061
App. No.
19/016,197
Filed
Jan 10, 2025
Granted
Aug 11, 2026
Kind
B2
Art Unit
2114
USPC
714/10
Abstract

Apparatuses, systems, and techniques to execute one or more application programming interfaces (APIs) to perform one or more operations for one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more processors are to perform one or more instructions in response to one or more APIs to indicate one or more functions to be performed in response to one or more errors from one or more accelerators within a heterogeneous processor.

Claims (41)

1 . One or more processors, comprising: circuitry to:

receive an Application Programming Interface (“API”) call, an input parameter of the API indicating error handler program code;

in response to the API call, register the error handler program code to handle one or more asynchronous errors from one or more accelerators, wherein the one or more accelerators and one or more Graphics Processing Units (GPUs) are within a heterogeneous processor; and

cause the error handler program code to be performed in response to one or more asynchronous errors from the one or more accelerators within the heterogeneous processor.

2 . The one or more processors of claim 1 , wherein the input parameter comprises a pointer to a function to be registered as the error handler program code.

3 . The one or more processors of claim 1 , wherein the circuitry is further to register one or more memory buffers as error notification buffers to be polled for error information generated by the one or more accelerators.

4 . The one or more processors of claim 1 , wherein the circuitry is further to:

receive a second Application Programming Interface (“API”) call having a second input parameter indicating one or more memory regions usable to store error information generated by the one or more accelerators; and

in response to the second API call, identify, in the one or more memory regions, error information generated by the one or more accelerators.

5 . The one or more processors of claim 1 , wherein the circuitry is further to:

receive a second Application Programming Interface (“API”) call having a second input parameter indicating one or more memory regions usable to store error information generated by the one or more accelerators; and

in response to the second API call, poll for one or more errors from the one or more accelerators in the one or more memory regions by checking for error information generated by the one or more accelerators in the one or more memory regions.

6 . The one or more processors of claim 1 , wherein the one or more accelerators comprise at least one of: a deep learning accelerator (DLA), a programmable vision accelerator (PVA), or a field-programmable gate array (FPGA).

7 . The one or more processors of claim 1 , wherein the one or more accelerators are packaged on a system-on-chip (SoC) together with at least one of: the one or more processors or the one or more GPUs.

8 . The one or more processors of claim 1 , wherein the one or more processors include a central processing unit (CPU).

9 . The one or more processors of claim 1 , wherein the error handler program code comprises one or more instructions that, if performed, cause the one or more accelerators within the heterogeneous processor to perform one or more computational operations in response to the one or more asynchronous errors.

10 . The one or more processors of claim 1 , wherein the one or more asynchronous errors are to be generated by the one or more accelerators within the heterogeneous processor and one or more other errors are to be generated by one or more other accelerators, and the one or more other errors are to be handled by one or more other error handlers.

11 . A system, comprising:

one or more processors to:

receive an Application Programming Interface (“API”) call, an input parameter of the API indicating error handler program code;

in response to the API call, register the error handler program code to handle one or more asynchronous errors from one or more accelerators, wherein the one or more accelerators and one or more Graphics Processing Units (GPUs) are within a heterogeneous processor; and

cause the error handler program code to be performed in response to one or more asynchronous errors from the one or more accelerators within the heterogeneous processor.

12 . The system of claim 11 , wherein the input parameter comprises a pointer to a function to be registered as the error handler program code.

13 . The system of claim 11 , wherein the one or more processors are further to register one or more memory buffers as error notification buffers to be polled for error information generated by the one or more accelerators.

14 . The system of claim 11 , wherein the one or more processors are further to:

receive a second Application Programming Interface (“API”) call having a second input parameter indicating one or more memory regions usable to store error information generated by the one or more accelerators; and

in response to the second API call, identify, in the one or more memory regions, error information generated by the one or more accelerators.

15 . The system of claim 11 , wherein the one or more processors are further to:

receive a second Application Programming Interface (“API”) call having a second input parameter indicating one or more memory regions usable to store error information generated by the one or more accelerators; and

in response to the second API call, poll for one or more errors from the one or more accelerators in the one or more memory regions by checking for error information generated by the one or more accelerators in the one or more memory regions.

16 . A method, comprising:

receiving an Application Programming Interface (“API”) call, an input parameter of the API indicating error handler program code;

in response to the API call, registering the error handler program code to handle one or more asynchronous errors from one or more accelerators, wherein the one or more accelerators and one or more Graphics Processing Units (GPUs) are within a heterogeneous processor; and

causing the error handler program code to be performed in response to one or more asynchronous errors from the one or more accelerators within the heterogeneous processor.

17 . The method of claim 16 , wherein the input parameter comprises a pointer to a function to be registered as the error handler program code.

18 . The method of claim 16 , further comprising:

registering one or more memory buffers as error notification buffers to be polled for error information generated by the one or more accelerators.

19 . The method of claim 16 , further comprising:

receiving a second Application Programming Interface (“API”) call having a second input parameter indicating one or more memory regions usable to store error information generated by the one or more accelerators; and

in response to the second API call, identifying, in the one or more memory regions, error information generated by the one or more accelerators.

20 . The method of claim 16 , wherein the one or more accelerators comprise at least one of: a deep learning accelerator (DLA), a programmable vision accelerator (PVA), or a field-programmable gate array (FPGA).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2025
From: RAVI, KARTHIK RAGHAVAN; JAIN, ASHUTOSH; SURESH, RAHUL
To: NVIDIA CORPORATION
Reel/Frame 070106/0174 →
Continuity (2)
Continuation 18070148 · Nov 28, 2022
Related Publication 20250342041A1 · Nov 6, 2025
References Cited (46)
US 10120779B1 · Leung · 2018 [cited by examiner]
US 10325343B1 · Zhao · 2019 [cited by examiner]
US 10819680B1 · Santan et al. · 2020 [cited by applicant]
US 10877817B1 · Balle et al. · 2020 [cited by applicant]
US 11132326B1 · Modukuri et al. · 2021 [cited by applicant]
US 20090119677A1 · Stefansson et al. · 2009 [cited by applicant]
US 20090217275A1 · Krishnamurthy et al. · 2009 [cited by applicant]
US 20100161632A1 · Rosen · 2010 [cited by applicant]
US 20130232495A1 · Rossbach et al. · 2013 [cited by applicant]
US 20140146062A1 · Kiel · 2014 [cited by examiner]
US 20140189426A1 · Ben-Kiki et al. · 2014 [cited by applicant]
US 20150007182A1 · Rossbach et al. · 2015 [cited by applicant]
US 20150205629A1 · Kruglick · 2015 [cited by applicant]
US 20160085546A1 · Brower et al. · 2016 [cited by applicant]
US 20170308504A1 · Cook et al. · 2017 [cited by applicant]
US 20180121320A1 · Dolby · 2018 [cited by examiner]
US 20190005606A1 · Yang · 2019 [cited by applicant]
US 20190197655A1 · Sun · 2019 [cited by examiner]
US 20190260818A1 · Ciabarra, Jr. · 2019 [cited by examiner]
US 20200050451A1 · Babich et al. · 2020 [cited by applicant]
US 20200057681A1 · Mamaghani et al. · 2020 [cited by applicant]
US 20200341812A1 · McClure · 2020 [cited by applicant]
US 20200364088A1 · Ashwathnarayan · 2020 [cited by examiner]
US 20200409877A1 · Balle et al. · 2020 [cited by applicant]
US 20210182132A1 · Chen et al. · 2021 [cited by applicant]
US 20210216435A1 · Godefroid · 2021 [cited by examiner]
US 20210294638A1 · Edwards · 2021 [cited by applicant]
US 20210294707A1 · Evans et al. · 2021 [cited by applicant]
US 20210390033A1 · Singhal · 2021 [cited by examiner]
US 20220283851A1 · Lee et al. · 2022 [cited by applicant]
US 20220318604A1 · Xu et al. · 2022 [cited by applicant]
US 20220374277A1 · Kim · 2022 [cited by applicant]
US 20220413921A1 · Hamlin et al. · 2022 [cited by applicant]
Jingwen Leng et al. “Asymmetric Resilience: Exploiting Task-level Idempotency for Transient Error Recovery in Accelerator-based Systems” (Year: 2020). [cited by examiner]
IEEE “IEEE Standard for Floating-Point Arithmetic”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008, 70 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2023/081364, mailed Mar. 20, 2024, filed Nov. 28, 2023, 16 pages. [cited by applicant]
Jung et al., “SnuRHAC: A Runtime for Heterogeneous Accelerator Clusters with CUDS Unified Memory,” ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, 2021, 14 pages. [cited by applicant]
Muthukrishnan et al., “Efficient Multi-GPU Shared Memory via Automatic Optimization of Fine-Grained Transfers,” ACM/IEEE 48th Annual Intemational Symposium on Computer Architecture, 2021, 14 pages. [cited by applicant]
Kevin Hsieh, Accelerating Pointer Chasing in 3D-Stacked Memory: Challenges, Mechanisms, Evaluation. (Year: 2016), 8 pages. [cited by applicant]
Leng et al., “Asymmetric Resilience for Accelerator-Rich Systems”, IEEE Computer Architecture Letters, vol. 18, No. 1, 2019, 4 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2023/081356, mailed Mar. 15, 2024, filed Nov. 28, 2023, 17 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2023/081368, mailed Mar. 5, 2024, filed Nov. 28, 2023, 11 pages. [cited by applicant]
Leng et al. “Asymmetric Resilience: Exploiting Task-level Idempotency for Transient Error Recovery in Accelerator-based Systems,” 2020, 14 pages. [cited by applicant]
Nozal et al., “Exploiting Co-Execution with oneAPI: Heterogeneity from a Modern Perspective,” 2021, 14 pages. [cited by applicant]
Pavlidakis et al., “Arax: A Runtime Framework for Decoupling Applications from Heterogeneous Accelerators,” ACM SIGSAC Conference on Computer and Communications Security, Nov. 7, 2022, 15 pages. [cited by applicant]