IP Library › Granted Patent US 12,608,320
Granted Patent B1
US 12,608,320 · App. 17/713,021 · Granted Apr 21, 2026

Application programming interface to allocate memory for shared virtual memory

Inventors: James Christopher Beyer (Marshall, MN); Paul J. Sidenblad (Cupertino, CA); Vyas Venkataraman (Sharon, MA); Chetan Gokhale (Santa Clara, CA); Cory Perry (San Jose, CA); Ying Liang (Palo Alto, CA); Harold Carter Edwards (Campbell, CA)
Assignee: NVIDIA Corporation
G06F12/1009G06F9/3877G06F9/45558G06F12/084G06F2009/45583
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,320
App. No.
17/713,021
Granted
Apr 21, 2026
Kind
B1
Abstract

Apparatuses, systems, and techniques to facilitate memory management. In at least one embodiment, an application programming interface is performed to cause physical memory corresponding to shared virtual memory to be designated for use by a plurality of processors.

Claims (66)

1 . A processor comprising:

one or more circuits to perform an application programming interface (API) to receive a location of one or more physical memory addresses and cause the one or more physical memory addresses to be assigned to one or more shared virtual memory addresses to be used by a plurality of processors.

2 . The processor of claim 1 , wherein the one or more shared virtual memory addresses correspond to an address of multicast memory.

3 . The processor of claim 1 , wherein:

a first processor of the plurality of processors is to be on a first device; and

a second processor of the plurality of processors is to be on a second device.

4 . The processor of claim 1 , wherein the one or more shared virtual memory-addresses are to be shared between the plurality of processors.

5 . The processor of claim 1 , wherein the one or more circuits are to perform a memory manager to coordinate the one or more shared virtual memory addresses between the plurality of processors.

6 . The processor of claim 1 , wherein the one or more shared virtual memory-addresses are to be provided to the API as a memory handle.

7 . The processor of claim 1 , wherein the one or more shared virtual memory addresses are to be allocated by a second API that causes the one or more shared virtual memory-addresses to be allocated for use by the plurality of processors.

8 . The processor of claim 1 , wherein the one or more circuits comprise a switch to route the one or more shared virtual memory addresses to one or more different physical memory addresses located on different processors.

9 . The processor of claim 1 , wherein at least one processor of the plurality of processors is a graphics processing unit (GPU).

10 . The processor of claim 1 , wherein:

a first processor of the plurality of processors is of a first node of a compute cluster;

a second processor of the plurality of processors is of a second node of the compute cluster; and

the first processor is to access memory of the second processor using the one or more shared virtual memory addresses.

11 . The processor of claim 1 , wherein the one or more circuits are to further cause at least one processor of the plurality of processors to designate the one or more physical memory addresses to the one or more shared virtual memory addresses of a shared virtual memory.

12 . A computer-implemented method comprising: performing an application programming interface (API) to receive a location of one or more physical memory addresses and cause the one or more physical memory addresses to be assigned to one or more shared virtual memory addresses to be used by the plurality of processors.

13 . The computer-implemented method of claim 12 , wherein the one or more physical memory addresses comprise an address of multicast memory.

14 . The computer-implemented method of claim 12 , wherein:

a first processor of the plurality of processors is on a first node of a compute cluster; and

a second processor of the plurality of processors is on a second node of the compute cluster.

15 . The computer-implemented method of claim 12 , further comprising:

sharing the one or more shared virtual memory addresses between the plurality of processors using a memory handle.

16 . The computer-implemented method of claim 12 , further comprising:

performing a memory manager to coordinate the one or more shared virtual memory addresses between the plurality of processors.

17 . The computer-implemented method of claim 12 , wherein performing the API comprises:

selecting a processor of the plurality of processors;

allocating a physical memory location in memory of the processor; and

mapping the physical memory location to the one or more shared virtual memory addresses.

18 . The computer-implemented method of claim 12 , wherein at least one processor of the plurality of processors is a graphics processing unit (GPU).

19 . The computer-implemented method of claim 12 , further comprising:

routing the one or more shared virtual memory addresses to one or more different physical memory addresses located on a different device.

20 . The computer-implemented method of claim 12 , further comprising:

performing one or more compute operations using the one or more shared virtual memory addresses.

21 . The computer-implemented method of claim 12 , wherein a set processors of the plurality of processors are to share a first memory handle corresponding to the one or more shared virtual memory addresses.

22 . The computer-implemented method of claim 12 , wherein:

a set of processors of the plurality of processors are to share a first memory handle corresponding to a first virtual memory address of the one or more shared virtual memory addresses; and

a subset of the set of processors are to share a second memory handle corresponding to a second virtual memory address.

23 . A computer system comprising:

one or more processors and memory storing executable instructions that, as a result of being executed by the one or more processors, cause the one or more processors to perform an application programming interface (API) to receive a location of one or more physical memory addresses and cause the one or more physical memory addresses to be assigned to one or more shared virtual memory addresses to be used by a plurality of processors.

24 . The computer system of claim 23 , wherein the one or more shared virtual memory addresses correspond to an address of multicast memory.

25 . The computer system of claim 23 , wherein:

a first processor of the plurality of processors is to be on a first device; and

a second processor of the plurality of processors is to be on a second device.

26 . The computer system of claim 23 , wherein the one or more shared virtual memory addresses are to be shared between the plurality of processors.

27 . The computer system of claim 23 , wherein the one or more processors are to perform a load operation using the one or more shared virtual memory addresses.

28 . The computer system of claim 23 , wherein the one or more processors are to perform a store operation using the one or more shared virtual memory addresses.

29 . The computer system of claim 23 , wherein the one or more processors are to perform an atomic operation using the one or more shared virtual memory addresses.

30 . The computer system of claim 23 , wherein the one or more processors are to perform a load operation with reduction using the one or more shared virtual memory addresses.

31 . The computer system of claim 23 , wherein the one or more processors are to perform an atomic operation with reduction using the one or more shared virtual memory addresses.

32 . A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to perform an application programming interface (API) to receive a location of one or more physical memory addresses and cause the one or more physical memory addresses to be assigned to one or more shared virtual memory addresses to be used by a plurality of processors.

33 . The machine-readable medium of claim 32 ,

wherein the one or more shared virtual memory addresses comprise an address of multicast memory.

34 . The machine-readable medium of claim 32 ,

wherein the one or more shared virtual memory addresses comprise an address of unicast memory.

35 . The machine-readable medium of claim 32 , wherein:

a first processor of the plurality of processors is to be on a first node of a compute cluster; and

a second processor of the plurality of processors is to be on a second node of the compute cluster.

36 . The machine-readable medium of claim 32 , wherein the one or more shared virtual memory addresses are to be shared between the plurality of processors.

37 . The machine-readable medium of claim 32 , wherein the API is to receive a memory handle that at least indicates the one or more shared virtual memory addresses.

38 . The machine-readable medium of claim 35 , wherein the first processor is to access memory of the second processor using the one or more shared virtual memory addresses.

39 . The machine-readable medium of claim 32 , wherein the API is to receive a set of flags that indicate one or more properties of the one or more shared virtual memory addresses.

40 . The machine-readable medium of claim 32 , wherein the API is to return a success indicator to a calling process executing on a processor of the one or more processors.

41 . A processor comprising:

one or more circuits to perform an application programming interface (API) to cause one or more physical memory addresses to be assigned to one or more shared virtual memory addresses to be used by a plurality of processors, wherein a first processor of the plurality of processors is of a first node of a compute cluster, a second processor of the plurality of processors is of a second node of the compute cluster, and the first processor is to access memory of the second processor using the one or more shared virtual memory addresses.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2022
From: BEYER, JAMES CHRISTOPHER; SIDENBLAD, PAUL J.; VENKATARAMAN, VYAS; GOKHALE, CHETAN; PERRY, CORY; LIANG, YING; EDWARDS, HAROLD CARTER
To: NVIDIA CORPORATION
Reel/Frame 059751/0748 →
References Cited (40)
US 5914730A · Santos et al. · 1999 [cited by applicant]
US 8402229B1 · Wilt et al. · 2013 [cited by applicant]
US 8776050B2 · Plouffe et al. · 2014 [cited by applicant]
US 9678775B1 · Grover et al. · 2017 [cited by applicant]
US 11288194B2 · Johns et al. · 2022 [cited by applicant]
US 20050034049A1 · Nemawarkar et al. · 2005 [cited by applicant]
US 20050190190A1 · Diard et al. · 2005 [cited by applicant]
US 20080247409A1 · Choudhury et al. · 2008 [cited by applicant]
US 20080301683A1 · Archer et al. · 2008 [cited by applicant]
US 20090128574A1 · Fujii et al. · 2009 [cited by applicant]
US 20100049821A1 · Oved · 2010 [cited by applicant]
US 20110153957A1 · Gao et al. · 2011 [cited by applicant]
US 20120210095A1 · Nellans et al. · 2012 [cited by applicant]
US 20120222052A1 · Magenheimer et al. · 2012 [cited by applicant]
US 20130024875A1 · Wang et al. · 2013 [cited by applicant]
US 20170262567A1 · Vassiliev · 2017 [cited by applicant]
US 20180285174A1 · Che · 2018 [cited by applicant]
US 20180300931A1 · Vembu et al. · 2018 [cited by applicant]
US 20180308202A1 · Appu et al. · 2018 [cited by applicant]
US 20190243780A1 · Gopal · 2019 [cited by examiner]
US 20190354291A1 · Stabrawa et al. · 2019 [cited by applicant]
US 20200192820A1 · Nair et al. · 2020 [cited by applicant]
US 20200364088A1 · Ashwathnarayan et al. · 2020 [cited by applicant]
US 20200409732A1 · Kovacevic · 2020 [cited by applicant]
US 20210036877A1 · Klenk et al. · 2021 [cited by applicant]
US 20210037107A1 · Klenk et al. · 2021 [cited by applicant]
US 20210067173A1 · Miele · 2021 [cited by applicant]
US 20210081849A1 · Mezaael · 2021 [cited by applicant]
US 20220342710A1 · Vishnuswaroop Ramesh et al. · 2022 [cited by applicant]
US 20230080480A1 · Kayi · 2023 [cited by examiner]
US 20230107660A1 · Rushing · 2023 [cited by examiner]
US 20240345990A1 · Striramassarma et al. · 2024 [cited by applicant]
“OpenCL(TM) 2.0 Shared Virtual Memory Overview” Sep. 10, 2014. Retrieved from https://www.intel.com/content/www/us/en/developer/articles/technical/opencl-20-shared-virtual-memory-overview.html (Year: 2014). [cited by examiner]
Grappa, “Scaling Data-intensive Applications on Commodity Clusters,” retrieved from http://grappa.io/, 2014, 4 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
MPI Project, “MPI_Win_create_dynamic(3) man page (version 3.1.6),” retrieved from https://www.open-mpi.org/doc/v3.1/man3/MPI_Win_create_dynamic.3.php, Mar. 20, 2020, 2 pages. [cited by applicant]
U.S. Appl. No. 17/712,991, filed Apr. 4, 2022. [cited by applicant]
U.S. Appl. No. 17/712,997, filed Apr. 4, 2022. [cited by applicant]
U.S. Appl. No. 17/713,054, filed Apr. 4, 2022. [cited by applicant]
Howes et al., “The OpenCI Specification,” Kronos OpenCL Working Group, retrieved from <https://registry.khronos.org/OpenCL/specs/opencl-2.0.pdf,> 2015, 288 pages. [cited by applicant]