IP Library › Granted Patent US 12,675,344
Granted Patent B1
US 12,675,344 · App. 17/978,918 · Granted Jul 7, 2026

Application programming interface to obtain kernel node context information

Inventors: Houston Thompson Hoffman (San Jose, CA); David Anthony Fontaine (Mountain View, CA); Sally Tessa Stevenson (Broomfield, CO)
Assignee: NVIDIA Corporation
G06F9/545
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,675,344
App. No.
17/978,918
Filed
Nov 1, 2022
Granted
Jul 7, 2026
Kind
B1
Examiner
ONAT, UMUT
Art Unit
2194
USPC
719/328
Abstract

Apparatuses, systems, and techniques to execute CUDA programs. In at least one embodiment, an application programming interface is performed to provide one or more indicators of context information corresponding to one or more kernels. The application programming interface can be used to provide indicators of context information for context-free kernels and context.

Claims (37)

1 . One or more processors, comprising: circuitry to:

receive an application programming interface (API) call to provide one or more indicators of context information corresponding to one or more kernels; and

in response to receipt of the API call:

store a data structure into a memory location specified by a parameter of the API call;

extract a value of a field of the data structure; and

in response to determining that the extracted field has a predetermined value, provide one or more indicators of context information of one or more context-free kernels corresponding to one or more nodes of a graph.

2 . The one or more processors of claim 1 , wherein at least one of the one or more kernels is a context-free kernel.

3 . The one or more processors of claim 1 , wherein the one or more kernels are to be indicated by one or more nodes of a graph.

4 . The one or more processors of claim 1 , wherein the context information is based, at least in part, on a context of graphics processing unit (GPU).

5 . The one or more processors of claim 1 , wherein the context information is associated with the one or more kernels at a time when the one or more kernels are performed.

6 . The one or more processors of claim 1 , wherein the context information is associated with the one or more kernels at a time when the one or more kernels are created.

7 . The one or more processors of claim 1 , wherein the one or more kernels include a set of kernel parameters comprising the context information.

8 . The one or more processors of claim 1 , wherein the field of the data structure is a pointer usable to indicate a kernel bound to a context when the kernel is created.

9 . A computer-implemented method comprising:

receiving an application programming interface (API) call to provide one or more indicators of context information corresponding to one or more kernels; and

in response to receipt of the API call:

storing a data structure into a memory location specified by a parameter of the API call;

extracting a value of a field of the data structure; and

in response to determining that the extracted field has a predetermined value, providing one or more indicators of context information of one or more context-free kernels corresponding to one or more nodes of a graph.

10 . The computer-implemented method of claim 9 , wherein the predetermined value is a value that does not correspond to a valid kernel object.

11 . The computer-implemented method of claim 9 , wherein the one or more kernels are to be indicated by the one or more nodes of the graph.

12 . The computer-implemented method of claim 9 , wherein the context information is based, at least in part, on a context of graphics processing unit (GPU).

13 . The computer-implemented method of claim 9 , wherein the context information is associated with the one or more kernels at a time when the one or more kernels are performed.

14 . The computer-implemented method of claim 9 , wherein the context information is associated with the one or more kernels at a time when the API is to be performed.

15 . The computer-implemented method of claim 9 , wherein the API receives, as input, one or more attributes of the one or more kernels usable to perform the one or more kernels.

16 . A computer system comprising:

one or more processors and memory storing executable instructions that, if performed by the one or more processors, are to:

receive an application programming interface (API) call to provide one or more indicators of context information corresponding to one or more kernels; and

in response to receipt of the API call:

store a data structure into a memory location specified by a parameter of the API call;

extract a value of a field of the data structure; and

in response to determining that the extracted field has a predetermined value, provide one or more indicators of context information of one or more context-free kernels corresponding to one or more nodes of a graph.

17 . The computer system of claim 16 , wherein at least one of the one or more kernels is a context-free kernel.

18 . The computer system of claim 16 , wherein the one or more kernels are to be indicated by one or more nodes of a graph.

19 . The computer system of claim 16 , wherein the context information is based, at least in part, on a context of graphics processing unit (GPU).

20 . The computer system of claim 16 , wherein the context information is associated with the one or more kernels at a time when the one or more kernels are performed.

21 . The computer system of claim 16 , wherein the context information is associated with the one or more kernels at a time when the one or more kernels are instantiated.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2022
From: HOFFMAN, HOUSTON THOMPSON; FONTAINE, DAVID ANTHONY; STEVENSON, SALLY TESSA
To: NVIDIA CORPORATION
Reel/Frame 061893/0061 →
References Cited (37)
US 9250956B2 · Munshi et al. · 2016 [cited by applicant]
US 10402223B1 · Santan et al. · 2019 [cited by applicant]
US 11080111B1 · Perelygin et al. · 2021 [cited by applicant]
US 11561826B1 · Nagpal et al. · 2023 [cited by applicant]
US 12243118B2 · Fontaine et al. · 2025 [cited by applicant]
US 20050050553A1 · Hen · 2005 [cited by examiner]
US 20050149726A1 · Joshi · 2005 [cited by examiner]
US 20050257203A1 · Nattinger · 2005 [cited by applicant]
US 20050268173A1 · Kudukoli et al. · 2005 [cited by applicant]
US 20070022147A1 · Bean · 2007 [cited by examiner]
US 20110022817A1 · Gaster et al. · 2011 [cited by applicant]
US 20130132934A1 · Munshi · 2013 [cited by examiner]
US 20150022538A1 · Munshi · 2015 [cited by applicant]
US 20160085527A1 · de Lima Ottoni · 2016 [cited by applicant]
US 20180232403A1 · Bhatti et al. · 2018 [cited by applicant]
US 20200004993A1 · Volos et al. · 2020 [cited by applicant]
US 20200089528A1 · Gutierrez · 2020 [cited by examiner]
US 20210096917A1 · Cerny · 2021 [cited by examiner]
US 20210149719A1 · Jones et al. · 2021 [cited by applicant]
US 20210149734A1 · Gurfinkel et al. · 2021 [cited by applicant]
US 20210248115A1 · Jones et al. · 2021 [cited by applicant]
US 20220100549A1 · Price · 2022 [cited by examiner]
US 20220334845A1 · Foote · 2022 [cited by examiner]
US 20220334851A1 · Zhurba et al. · 2022 [cited by applicant]
WO WO2013178864A1 · 2013 [cited by examiner]
Hoffman, “CUkern Support for Kernel Node Apis,” NVIDIA, Jun. 2, 2022, 13 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
NVIDIA, “CUDA Memcpy3d v2 Struct Reference,” retrieved from https://web.archive.org/web/20220628223336/https://docs.nvidia.com/cuda/cuda-driver-api/structCUDA_MEMCPY3D_V2.html#structCUDA_MEMCPY3D_V2, May 11, 2022, 1 pag… [cited by applicant]
U.S. Appl. No. 17/978,916, filed Nov. 1, 2022. [cited by applicant]
U.S. Appl. No. 17/978,924, filed Nov. 1, 2022. [cited by applicant]
U.S. Appl. No. 17/978,929, filed Nov. 1, 2022. [cited by applicant]
Jain et al., “The OoO VLIW JIT Compiler for GPU Inference”, Jan. 28, 2019, 7 pages. [cited by applicant]
Jeremiah, “Introduction to OpenCL”, Dec. 19, 2011, 3D Game Engine Programming, 31 pages. [cited by applicant]
NVIDIA, “CUDA C++ Best Practices Guide”, Sep. 2021, 100 pages. [cited by applicant]
NVIDIA, “CUDA Driver API”, API Reference Manual, Sep. 2021, 548 pages. [cited by applicant]
Harvey et al., “Swan: A Tool for Porting CUDA Programs to OpenCL”, Jan. 11, 2011, Computer Physics Communications vol. 182, Issue 4, 7 pages. [cited by applicant]
Sedighi et al., “Workload-aware Dynamic GPU Resource Management in Component-based Applications”, Date of Conference: Sep. 26-30, 2022, IEEE, 2022 IEEE International Conference on Cloud Engineering (IC2E), 8 pages. [cited by applicant]