IP Library › Granted Patent US 12,705,060
Granted Patent B1
US 12,705,060 · App. 18/114,711 · Granted Aug 11, 2026

Application programming interface to indicate kernel dependencies

Inventors: David Anthony Fontaine (Mountain View, CA); Steven Arthur Gurfinkel (San Jose, CA)
Assignee: NVIDIA Corporation
G06F9/3838G06F9/3877G06F9/541
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,060
App. No.
18/114,711
Filed
Feb 27, 2023
Granted
Aug 11, 2026
Kind
B1
Art Unit
2183
USPC
712/216
Abstract

Apparatuses, systems, and techniques to perform an application programming interface (API) to add one or more graph nodes to a software graph, wherein the API is to cause a kernel node to be added to a software graph based, at least in part, on a dependency type indicated by the API. In at least one embodiment, one or more nodes are added to a graph in accordance to one or more dependency types.

Claims (28)

1 . One or more processors, comprising:

circuitry to:

receive an application programming interface (API) call comprising one or more input parameters indicating a kernel node of a software graph, the software graph indicating processor-performable operations and one or more dependencies among the processor-performable operations; and

in response to the API call, cause the kernel node to be added to the software graph based, at least in part, on a dependency type indicated by the one or more input parameters, the dependency type indicating one or more constraints on scheduling a first operation corresponding to the kernel node and a second operation corresponding to a graph node.

2 . The one or more processors of claim 1 , wherein the circuitry is to, in response to the API call, cause the kernel node to be added to the software graph, based, at least in part, on the dependency type, one or more edges of the software graph, the one or more edges to or from the kernel node.

3 . The one or more processors of claim 1 , wherein the API call is to further indicate a port identifier of a node associated with the kernel node.

4 . The one or more processors of claim 1 , wherein the circuitry is to cause the kernel node to be added to the software graph by identifying one or more graph nodes of the software graph dependent to or dependent on the kernel node.

5 . The one or more processors of claim 1 , wherein the dependency type includes one or more of a full execution dependency, a launch order dependency, a fast dependent launch, or an anti-deadlock dependency.

6 . The one or more processors of claim 1 , wherein the kernel node is to be performed by one or more graphics processing units (GPUs) based, at least in part, on the dependency type.

7 . The one or more processors of claim 1 , wherein the API call is to receive a set of parameters comprising an identifier of the kernel node and dependency information corresponding to the dependency type.

8 . A computer-implemented method comprising:

receiving an application programming interface (API) call comprising one or more input parameters that indicates a kernel node of a software graph that indicates one or more processor-performable operations, one or more dependencies among two or more processor-performable operations, and one or more dependency types associated with the one or more dependencies; and

in response to the API call, causing the kernel node to be added to the software graph to have the one or more dependencies according to the one or more dependency types, the one or more dependency types indicating one or more constraints on scheduling the two or more processor-performable operations.

9 . The computer-implemented method of claim 8 , wherein the API is to cause the kernel node to be added to the software graph, based, at least in part, on the one or more dependency types, one or more edges of the software graph, the one or more edges to or from the kernel node.

10 . The computer-implemented method of claim 8 , wherein the API is to further indicate a port identifier of a node associated with the kernel node.

11 . The computer-implemented method of claim 8 , wherein the one or more dependency types includes one or more of a full execution dependency, a launch order dependency, or an anti-deadlock dependency.

12 . The computer-implemented method of claim 8 , wherein the kernel node is to be performed by one or more graphics processing units (GPUs) based, at least in part, on the one or more dependency types.

13 . The computer-implemented method of claim 8 , wherein the API is to receive a set of parameters comprising an identifier of the kernel node and dependency information corresponding to the one or more dependency types.

14 . The computer-implemented method of claim 8 , wherein the API is to cause the kernel node to be added to the software graph by identifying one or more graph nodes of the software graph dependent to or dependent on the kernel node.

15 . A computer system comprising:

one or more processors and memory storing executable instructions that, if performed by the one or more processors, are to:

receive an application programming interface (API) call comprising one or more input parameters indicating a kernel node of a software graph, the software graph indicating processor-performable operations and one or more dependencies among the processor-performable operations; and

in response to the API call, cause the kernel node to be added to the software graph, based, at least in part, on a dependency type indicated by the one or more input parameters, the dependency type indicating one or more constraints on scheduling a first operation corresponding to the kernel node and a second operation corresponding to a graph node.

16 . The computer system of claim 15 , wherein the one or more processors are to cause the kernel node to be added to the software graph, based, at least in part, on the dependency type, one or more edges of the software graph, the one or more edges to or from the kernel node.

17 . The computer system of claim 15 , wherein the one or more processors are to further indicate a port identifier of a node associated with the kernel node.

18 . The computer system of claim 15 , wherein the one or more processors are to cause the kernel node to be added to the software graph by identifying one or more graph nodes of the software graph dependent to or dependent on the kernel node.

19 . The computer system of claim 15 , wherein the dependency type includes one or more of a full execution dependency, a launch order dependency, or an anti-deadlock dependency.

20 . The computer system of claim 15 , wherein the API call is to receive a set of parameters comprising an identifier of the kernel node and dependency information corresponding to the dependency type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2023
From: FONTAINE, DAVID ANTHONY; GURFINKEL, STEVEN ARTHUR
To: NVIDIA CORPORATION
Reel/Frame 063003/0821 →
References Cited (114)
US 5689711A · Bardasz et al. · 1997 [cited by applicant]
US 6937969B1 · Vandersteen et al. · 2005 [cited by applicant]
US 7478375B1 · Kersters · 2009 [cited by applicant]
US 8115773B2 · Swift et al. · 2012 [cited by applicant]
US 8181168B1 · Lee et al. · 2012 [cited by applicant]
US 8239404B2 · Zhou et al. · 2012 [cited by applicant]
US 8539516B1 · Wilt et al. · 2013 [cited by applicant]
US 9251225B2 · Stanfill · 2016 [cited by applicant]
US 9372670B1 · Cartey et al. · 2016 [cited by applicant]
US 9411706B1 · van Schaik · 2016 [cited by applicant]
US 9542192B1 · Wilt et al. · 2017 [cited by applicant]
US 9684944B2 · Taylor et al. · 2017 [cited by applicant]
US 10417058B1 · Kesler · 2019 [cited by applicant]
US 10540270B1 · Surkatty et al. · 2020 [cited by applicant]
US 10673712B1 · Gosar et al. · 2020 [cited by applicant]
US 11003423B2 · Mazurskiy · 2021 [cited by applicant]
US 11113030B1 · Monga · 2021 [cited by examiner]
US 11150961B2 · Agarwal et al. · 2021 [cited by applicant]
US 11340873B2 · Cangea et al. · 2022 [cited by applicant]
US 11422797B1 · Zhang et al. · 2022 [cited by applicant]
US 11455152B2 · Zhang · 2022 [cited by applicant]
US 11842221B2 · Glass et al. · 2023 [cited by applicant]
US 11868237B2 · Balasubramanian et al. · 2024 [cited by applicant]
US 12073263B1 · Thompson · 2024 [cited by applicant]
US 12159217B1 · Borkovic · 2024 [cited by applicant]
US 12443462B1 · Fontaine et al. · 2025 [cited by applicant]
US 20040088666A1 · Poznanovic et al. · 2004 [cited by applicant]
US 20050034106A1 · Kornerup et al. · 2005 [cited by applicant]
US 20050155034A1 · Jiang et al. · 2005 [cited by applicant]
US 20050174984A1 · O'Neill · 2005 [cited by applicant]
US 20070174494A1 · Bonwick et al. · 2007 [cited by applicant]
US 20070220031A1 · MacMahon et al. · 2007 [cited by applicant]
US 20080278482A1 · Farmanbar et al. · 2008 [cited by applicant]
US 20090055630A1 · Isshiki et al. · 2009 [cited by applicant]
US 20090102846A1 · Flockermann et al. · 2009 [cited by applicant]
US 20090113396A1 · Rosen et al. · 2009 [cited by applicant]
US 20100079462A1 · Breeds et al. · 2010 [cited by applicant]
US 20100333110A1 · Luo et al. · 2010 [cited by applicant]
US 20120008530A1 · Kulkarni et al. · 2012 [cited by applicant]
US 20120072887A1 · Basak · 2012 [cited by applicant]
US 20120278365A1 · Labat et al. · 2012 [cited by applicant]
US 20130127891A1 · Kim et al. · 2013 [cited by applicant]
US 20130212131A1 · Reddy · 2013 [cited by applicant]
US 20150006644A1 · Anantharam et al. · 2015 [cited by applicant]
US 20150016257A1 · Kumar et al. · 2015 [cited by applicant]
US 20150288595A1 · Suzuki · 2015 [cited by applicant]
US 20160210720A1 · Taylor et al. · 2016 [cited by applicant]
US 20160210724A1 · Taylor et al. · 2016 [cited by applicant]
US 20160253625A1 · Casey · 2016 [cited by applicant]
US 20160307353A1 · Ligenza et al. · 2016 [cited by applicant]
US 20170286526A1 · Bar-Or et al. · 2017 [cited by applicant]
US 20170373946A1 · Lewandowski et al. · 2017 [cited by applicant]
US 20180113713A1 · Cheng et al. · 2018 [cited by applicant]
US 20180136933A1 · Kogan et al. · 2018 [cited by applicant]
US 20180181676A1 · Khandelwal et al. · 2018 [cited by applicant]
US 20180218259A1 · Braz et al. · 2018 [cited by applicant]
US 20190182107A1 · Saxena et al. · 2019 [cited by applicant]
US 20190188055A1 · Hunt et al. · 2019 [cited by applicant]
US 20190327154A1 · Sahoo et al. · 2019 [cited by applicant]
US 20190339966A1 · Moondhra et al. · 2019 [cited by applicant]
US 20190370061A1 · Shah et al. · 2019 [cited by applicant]
US 20190370407A1 · Dickie · 2019 [cited by applicant]
US 20190370927A1 · Frenkel et al. · 2019 [cited by applicant]
US 20200050633A1 · Evans et al. · 2020 [cited by applicant]
US 20200057748A1 · Danilak · 2020 [cited by applicant]
US 20200136891A1 · Mdini et al. · 2020 [cited by applicant]
US 20200310937A1 · Takeda · 2020 [cited by applicant]
US 20200364088A1 · Ashwathnarayan et al. · 2020 [cited by applicant]
US 20200371761A1 · Gupta et al. · 2020 [cited by applicant]
US 20200396075A1 · Visegrady et al. · 2020 [cited by applicant]
US 20200409671A1 · Mazurskiy · 2020 [cited by applicant]
US 20200409709A1 · ChoFleming et al. · 2020 [cited by applicant]
US 20210004263A1 · Moita et al. · 2021 [cited by applicant]
US 20210011849A1 · Simpson et al. · 2021 [cited by applicant]
US 20210037397A1 · Guo et al. · 2021 [cited by applicant]
US 20210089368A1 · Goosen et al. · 2021 [cited by applicant]
US 20210096921A1 · Banerjee et al. · 2021 [cited by applicant]
US 20210133089A1 · Khillar et al. · 2021 [cited by applicant]
US 20210149719A1 · Jones et al. · 2021 [cited by applicant]
US 20210149734A1 · Gurfinkel et al. · 2021 [cited by applicant]
US 20210232579A1 · Schechter et al. · 2021 [cited by applicant]
US 20210248115A1 · Jones · 2021 [cited by examiner]
US 20210311727A1 · Khullar · 2021 [cited by applicant]
US 20210318908A1 · Cavus et al. · 2021 [cited by applicant]
US 20210373974A1 · Agarwal et al. · 2021 [cited by applicant]
US 20220214861A1 · Sohrabizadeh et al. · 2022 [cited by applicant]
US 20220334851A1 · Zhurba et al. · 2022 [cited by applicant]
US 20220334891A1 · Fontaine · 2022 [cited by applicant]
US 20230005096A1 · Gurfinkel et al. · 2023 [cited by applicant]
US 20230005097A1 · Gurfinkel et al. · 2023 [cited by applicant]
US 20230084951A1 · Fontaine et al. · 2023 [cited by applicant]
US 20230108560A1 · Wang · 2023 [cited by applicant]
US 20230118695A1 · Zhang et al. · 2023 [cited by applicant]
US 20230140822A1 · Purnomo et al. · 2023 [cited by applicant]
US 20230185634A1 · Fontaine et al. · 2023 [cited by applicant]
US 20230185635A1 · Vaz · 2023 [cited by applicant]
US 20230244523A1 · Gorantla et al. · 2023 [cited by applicant]
US 20230244549A1 · Fontaine · 2023 [cited by examiner]
US 20230297444A1 · Fernandes et al. · 2023 [cited by applicant]
US 20240118965A1 · Ashrafi et al. · 2024 [cited by applicant]
US 20240168795A1 · Edwards et al. · 2024 [cited by applicant]
US 20240220314A1 · Gasparakis · 2024 [cited by applicant]
US 20240289187A1 · Fontaine et al. · 2024 [cited by applicant]
US 20250021407A1 · Fontaine et al. · 2025 [cited by applicant]
US 20250138795A1 · Dubrovsky et al. · 2025 [cited by applicant]
IEEE “IEEE Standard for Floating-Point Arithmetic”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Sid 754-2008, dated Jun. 12, 2008, 70 pages. [cited by applicant]
Abdolrashidi et al., “Wireframe: Supporting Data-dependent Parallelism through Dependency Graph Execution in GPUs,” ACM, 2017, 12 pages. [cited by applicant]
Zhou et al., “Deadlock Prediction via Generalized Dependency,” ACM, 2022, 12 pages. [cited by applicant]
Gutman et al., “CUDA Graph Usage: CUDA FeatureTesting,” retrieved from <https://web.archive.org/web/20201028074137/https://codingbyexample.com/2020/09/25/cuda-graph-usage/,> 2020, 14 pages. [cited by applicant]
Yu et al., “OpenMP to CUDA Graphs: A Compiler-based Transformation to Enhance the Programmability of NVIDIA Devices,” ACM, 2020, 6 pages. [cited by applicant]
Gray, “Getting Started with CUDA Graph,” retrieved from forums.developer.nvidia.com, Sep. 5, 2019, 9 pages. [cited by applicant]
Jones, “CUDA Graphs Updates,” NVIDIA, Oct. 2022, 35 pages. [cited by applicant]
NVIDIA, “CUDA Runtime API, Reference Manual”, Jan. 2022, <https://docs.nvidia.com/cuda/archive/11.6.0/pdf/CUDA_Runtime_API.pdf>,> Chapter 6.30, 638 pages. [cited by applicant]
Clucas et al., “Ripple: Simplified Large-Scale Computation on Heterogeneous Architectures with Polymorphic Data Layout”, Journal of Parallel and Distributed Computing, Apr. 20, 2021, 18 pages. [cited by applicant]