IP Library › Granted Patent US 12,578,993
Granted Patent B2
US 12,578,993 · App. 17/955,175 · Granted Mar 17, 2026

Application programming interface to share memory between groups of blocks of threads

Inventors: Ze Long (San Jose, CA); Kyrylo Perelygin (Broomfield, CO); Harold Carter Edwards (Campbell, CA); Gokul Ramaswamy Hirisave Chandra Shekhara (Bangalore, IN); Jaydeep Marathe (Kirkland, WA); Ronny Meir Krashinsky (Portola Valley, CA); Girish Bhaskarrao Bharambe (Pune, IN)
Assignee: NVIDIA Corporation
G06F9/4881G06F8/456G06F9/30072G06F9/5044G06F9/505G06F9/522G06F9/544G06F9/545
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,578,993
App. No.
17/955,175
Granted
Mar 17, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to execute CUDA programs. In at least one embodiment, an application programming interface is performed to cause memory to be shared between two or more groups of blocks of threads.

Claims (39)

1 . A processor comprising:

one or more circuits to perform an application programming interface (API) to cause memory to be shared between two or more groups of blocks of threads.

2 . The processor of claim 1 , wherein the two or more groups of blocks of threads are groups within a software kernel.

3 . The processor of claim 1 , wherein the two or more groups of blocks of threads are groups within a grid of blocks of threads.

4 . The processor of claim 1 , wherein the two or more groups of blocks of threads are partitions of a partitioning of blocks of threads.

5 . The processor of claim 1 , wherein the memory is memory of a graphics processing unit (GPU).

6 . The processor of claim 1 , wherein the API is to cause at least one thread of one group of blocks of threads to be able to access a memory location accessible to at least one thread of another group of blocks of threads.

7 . The processor of claim 1 , wherein the API is to cause at least one thread of one group of blocks of threads to provide a memory address of the memory to be shared to the two or more groups of blocks of threads.

8 . The processor of claim 1 , wherein the API is to cause at least one thread of a first group of blocks of threads to calculate a memory address of the memory to be shared based, at least in part, on an identifier of a block of the first group of blocks of threads.

9 . The processor of claim 1 , wherein the API is to cause at least one thread of a first group of blocks of threads to calculate a memory address of the memory to be shared based, at least in part, on an identifier of the first group of blocks of threads.

10 . A computer-implemented method comprising:

performing an application programming interface (API) to cause memory to be shared between two or more groups of blocks of threads.

11 . The computer-implemented method of claim 10 , wherein the two or more groups of blocks of threads are groups within a software kernel.

12 . The computer-implemented method of claim 10 , wherein the two or more groups of blocks of threads are groups within a grid of blocks of threads.

13 . The computer-implemented method of claim 10 , wherein the two or more groups of blocks of threads are partitions of a partitioning of blocks of threads.

14 . The computer-implemented method of claim 10 , wherein the memory is memory of a graphics processing unit (GPU).

15 . The computer-implemented method of claim 10 , wherein the API is to cause at least one thread of one group of blocks of threads to be able to access a memory location accessible to at least one thread of another group of blocks of threads.

16 . The computer-implemented method of claim 10 , wherein the API is to cause at least one thread of one group of blocks of threads to provide a memory address of the memory to be shared to the two or more groups of blocks of threads.

17 . The computer-implemented method of claim 10 , wherein the API is to cause at least one thread of a first group of blocks of threads to calculate a memory address of the memory to be shared based, at least in part, on an identifier of a block of the first group of blocks of threads.

18 . The computer-implemented method of claim 10 , wherein the API is to cause at least one thread of a first group of blocks of threads to calculate a memory address of the memory to be shared based, at least in part, on an identifier of the first group of blocks of threads.

19 . A computer system comprising:

one or more processors and memory storing executable instructions that, if performed by the one or more processors, are to perform an application programming interface (API) to cause memory to be shared between two or more groups of blocks of threads.

20 . The computer system of claim 19 , wherein the two or more groups of blocks of threads are groups within a software kernel.

21 . The computer system of claim 19 , wherein the two or more groups of blocks of threads are groups within a grid of blocks of threads.

22 . The computer system of claim 19 , wherein the two or more groups of blocks of threads are partitions of a partitioning of blocks of threads.

23 . The computer system of claim 19 , wherein the memory is memory of a graphics processing unit (GPU).

24 . The computer system of claim 19 , wherein the API is to cause at least one thread of one group of blocks of threads to be able to access a memory location accessible to at least one thread of another group of blocks of threads.

25 . The computer system of claim 19 , wherein the API is to cause at least one thread of one group of blocks of threads to provide a memory address of the memory to be shared to the two or more groups of blocks of threads.

26 . The computer system of claim 19 , wherein the API is to cause at least one thread of a first group of blocks of threads to calculate a memory address of the memory to be shared based, at least in part, on an identifier of a block of the first group of blocks of threads.

27 . The computer system of claim 19 , wherein the API is to cause at least one thread of a first group of blocks of threads to calculate a memory address of the memory to be shared based, at least in part, on an identifier of the first group of blocks of threads.

28 . A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, are to perform an application programming interface (API) to cause memory to be shared between two or more groups of blocks of threads.

29 . The machine-readable medium of claim 28 , wherein the two or more groups of blocks of threads are groups within a software kernel.

30 . The machine-readable medium of claim 28 , wherein the two or more groups of blocks of threads are groups within a grid of blocks of threads.

31 . The machine-readable medium of claim 28 , wherein the two or more groups of blocks of threads are partitions of a partitioning of blocks of threads.

32 . The machine-readable medium of claim 28 , wherein the memory is memory of a graphics processing unit (GPU).

33 . The machine-readable medium of claim 28 , wherein the API is to cause at least one thread of one group of blocks of threads to be able to access a memory location accessible to at least one thread of another group of blocks of threads.

34 . The machine-readable medium of claim 28 , wherein the API is to cause at least one thread of one group of blocks of threads to provide a memory address of the memory to be shared to the two or more groups of blocks of threads.

35 . The machine-readable medium of claim 28 , wherein the API is to cause at least one thread of a first group of blocks of threads to calculate a memory address of the memory to be shared based, at least in part, on an identifier of a block of the first group of blocks of threads.

36 . The machine-readable medium of claim 28 , wherein the API is to cause at least one thread of a first group of blocks of threads to calculate a memory address of the memory to be shared based, at least in part, on an identifier of the first group of blocks of threads.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2023
From: LONG, ZE; PERELYGIN, KYRYLO; EDWARDS, HAROLD CARTER; HIRISAVE CHANDRA SHEKHARA, GOKUL RAMASWAMY; MARATHE, JAYDEEP; KRASHINSKY, RONNY MEIR; BHARAMBE, GIRISH BHASKARRAO
To: NVIDIA CORPORATION
Reel/Frame 062318/0535 →
Priority Claims (1)
IN 202241043444 · Jul 29, 2022 · national
Continuity (1)
Related Publication 20240036957A1 · Feb 1, 2024
References Cited (55)
US 5745778A · Alfieri · 1998 [cited by applicant]
US 7690003B2 · Fuller · 2010 [cited by applicant]
US 7788468B1 · Nickolls et al. · 2010 [cited by applicant]
US 11080111B1 · Perelygin et al. · 2021 [cited by applicant]
US 11568523B1 · Ligowski et al. · 2023 [cited by applicant]
US 11609921B2 · King et al. · 2023 [cited by applicant]
US 20030212671A1 · Meredith et al. · 2003 [cited by applicant]
US 20040078420A1 · Marrow et al. · 2004 [cited by applicant]
US 20070067606A1 · Lin · 2007 [cited by examiner]
US 20070294666A1 · Papakipos et al. · 2007 [cited by applicant]
US 20080276262A1 · Munshi et al. · 2008 [cited by applicant]
US 20090259997A1 · Grover et al. · 2009 [cited by applicant]
US 20090307699A1 · Munshi · 2009 [cited by examiner]
US 20090307704A1 · Munshi et al. · 2009 [cited by applicant]
US 20110063313A1 · Bolz et al. · 2011 [cited by applicant]
US 20110072249A1 · Nickolls · 2011 [cited by examiner]
US 20110087860A1 · Nickolls et al. · 2011 [cited by applicant]
US 20110285729A1 · Munshi et al. · 2011 [cited by applicant]
US 20120198214A1 · Gadre et al. · 2012 [cited by applicant]
US 20120222051A1 · Kakulamarri et al. · 2012 [cited by applicant]
US 20120254875A1 · Marathe et al. · 2012 [cited by applicant]
US 20130085730A1 · Shaw et al. · 2013 [cited by applicant]
US 20150160970A1 · Nugteren et al. · 2015 [cited by applicant]
US 20150187042A1 · Gupta · 2015 [cited by applicant]
US 20150199787A1 · Pechanec et al. · 2015 [cited by applicant]
US 20160232107A1 · Ros et al. · 2016 [cited by applicant]
US 20160364829A1 · Apodaca et al. · 2016 [cited by applicant]
US 20160371081A1 · Powers et al. · 2016 [cited by applicant]
US 20170024924A1 · Wald et al. · 2017 [cited by applicant]
US 20180033114A1 · Chen et al. · 2018 [cited by applicant]
US 20180307529A1 · Koker et al. · 2018 [cited by applicant]
US 20200004602A1 · Pawlowski et al. · 2020 [cited by applicant]
US 20200250005A1 · Munshi et al. · 2020 [cited by applicant]
US 20200394202A1 · Slesarenko et al. · 2020 [cited by applicant]
US 20210165699A1 · Parravicini et al. · 2021 [cited by applicant]
US 20210287325A1 · Kramer et al. · 2021 [cited by applicant]
US 20220051093A1 · Skaljak · 2022 [cited by applicant]
US 20220108497A1 · Panteleev · 2022 [cited by applicant]
US 20220342721A1 · Shveidel et al. · 2022 [cited by applicant]
US 20220342761A1 · Hukerikar et al. · 2022 [cited by applicant]
US 20230084951A1 · Fontaine et al. · 2023 [cited by applicant]
US 20230086989A1 · Ciolkosz et al. · 2023 [cited by applicant]
US 20230185706A1 · Kini et al. · 2023 [cited by applicant]
US 20230244549A1 · Fontaine et al. · 2023 [cited by applicant]
US 20230305853A1 · Ciolkosz et al. · 2023 [cited by applicant]
US 20230350661A1 · Reed et al. · 2023 [cited by applicant]
US 20240036944A1 · Long et al. · 2024 [cited by applicant]
US 20240036945A1 · Long et al. · 2024 [cited by applicant]
CN 102099789A · 2011 [cited by applicant]
CN 107357661A · 2017 [cited by applicant]
JP 2023070746A · 2023 [cited by applicant]
WO 2009148713A1 · 2009 [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Munshi, Aaftab, “The Open Specification” Khronos OpenCL Working Group, Version: 1.2, Document Revision: 19, Last Revision Date Nov. 14, 2012, 380 pages. [cited by applicant]
Harris et al., “Cooperative Groups: Flexible CUDA Thread Programming”, NVIDIA Technical Blog, Oct. 4, 2017, 15 pages. [cited by applicant]