Application programming interface to bind memory to shared virtual memory
Apparatuses, systems, and techniques to facilitate memory management. In at least one embodiment, an application programming interface is performed to enable access to shared virtual memory by a plurality of processors.
1 . One or more processors, comprising:
circuitry to:
in response to an application programming interface (API) call, cause each of a plurality of processors to bind a portion of physical memory to an indicator of a shared memory location provided to the API; and
enable access to the shared memory location when at least a number of processors allocate physical memory for use by the plurality of processors, the number indicated by an input parameter to the API.
2 . The one or more processors of claim 1 , wherein the shared memory location is multicast memory.
3 . The one or more processors of claim 1 , wherein:
a first processor of the plurality of processors is to be on a first device; and
a second processor of the plurality of processors is to be on a second device.
4 . The one or more processors of claim 1 , wherein the location of the shared memory location is to be shared between the plurality of processors.
5 . The one or more processors of claim 1 , wherein the circuitry is to perform a memory manager to coordinate the shared memory location between the plurality of processors.
6 . The one or more processors of claim 1 , wherein at least one processor of the plurality of processors is to designate a physical memory location that corresponds to the shared memory location.
7 . The one or more processors of claim 1 , wherein the circuitry comprises a switch to route a location of the shared memory location to different physical memory locations located on different processors.
8 . The one or more processors of claim 1 , wherein at least one processor of the plurality of processors is a graphics processing unit (GPU).
9 . The one or more processors of claim 1 , wherein:
a first processor of the plurality of processors is of a first node of a compute cluster;
a second processor of the plurality of processors is of a second node of the compute cluster; and
the first processor is to access memory of the second processor using the shared memory location.
10 . The one or more processors of claim 1 , wherein the circuitry is to further allow at least one of the processors of the plurality of processors to use a virtual memory address of the shared memory location, to which the one or more processors can write data to cause the at least one of the processors of the plurality of processors to store the data in physical memory.
11 . The one or more processors of claim 1 , wherein one or more parameters to the API indicate an allocation size for physical memory corresponding to the shared memory location.
12 . A computer-implemented method, comprising:
in response to an application programming interface (API) call comprising a handle indicative of a shared memory location:
causing a plurality of processors to each bind a portion of physical memory to the handle; and
enabling access to the shared memory location when at least a number of processors allocate physical memory for use by the plurality of processors, the number indicated by an input parameter to the API.
13 . The computer-implemented method of claim 12 , wherein the shared memory location is multicast memory.
14 . The computer-implemented method of claim 12 , wherein:
a first processor of the plurality of processors is of a first node of a compute cluster; and
a second processor of the plurality of processors is of a second node of the compute cluster.
15 . The computer-implemented method of claim 12 , further comprising:
sharing a location of the shared memory location between the plurality of processors using the handle.
16 . The computer-implemented method of claim 12 , further comprising: performing a memory manager to coordinate the shared memory location between the plurality of processors.
17 . The computer-implemented method of claim 12 , further comprising, in response to the API call:
selecting a processor of the plurality of processors;
allocating a physical memory location in memory of the processor; and
mapping the physical memory location to a location of the shared memory location.
18 . The computer-implemented method of claim 12 , wherein at least one processor of the plurality of processors is a graphics processing unit (GPU).
19 . The computer-implemented method of claim 12 , further comprising:
routing a location of the shared memory location to different physical memory locations located on different devices.
20 . The computer-implemented method of claim 12 , further comprising: performing one or more compute operations using the shared memory location.
21 . The computer-implemented method of claim 12 , wherein a set processors of the plurality of processors are to share a first memory handle corresponding to the shared memory location.
22 . The computer-implemented method of claim 12 , wherein:
a set of processors, of the plurality of processors, are to share a first memory handle corresponding to the shared memory location; and
a subset of the set of processors are to share a second memory handle corresponding to a second memory location.
23 . A computer system comprising:
one or more processors and memory storing executable instructions that, as a result of being executed by the one or more processors, cause the one or more processors to:
in response to an application programming interface (API) call, cause each of a plurality of processors to bind a portion of physical memory to an indicator of a shared memory location provided to the API; and
enable access to the shared memory location when at least a number of processors allocate physical memory for use by the plurality of processors, the number indicated by an input parameter to the API.
24 . The computer system of claim 23 , wherein the shared memory location is multicast memory.
25 . The computer system of claim 23 , wherein:
a first processor of the plurality of processors is to be on a first device; and
a second processor of the plurality of processors is to be on a second device.
26 . The computer system of claim 23 , wherein a location of the shared memory location is to be shared between the plurality of processors.
27 . The computer system of claim 23 , wherein a location of the shared memory location is to be routed to different physical memory locations located on different devices.
28 . The computer system of claim 23 , wherein the one or more processors are to perform one or more compute operations using the shared memory location.
29 . The computer system of claim 23 , wherein at least one processor of the plurality of processors is a graphics processing unit (GPU).
30 . The computer system of claim 23 , wherein a set processors of the plurality of processors are to share a first memory handle corresponding to the shared memory location.
31 . The computer system of claim 23 , wherein:
a set of processors of the plurality of processors are to share a first memory handle corresponding to the shared memory location; and
a subset of the set of processors are to share a second memory handle corresponding to a second shared memory location.
32 . The computer system of claim 23 , wherein one or more parameters to the API indicate an allocation size for physical memory corresponding to the shared memory location.
33 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to:
in response to an application programming interface (API) call, cause each of a plurality of processors to bind a portion of physical memory to an indicator of a shared memory location provided to the API; and
enable access to the shared memory location when at least a number of processors allocate physical memory for use by the plurality of processors, the number indicated by an input parameter to the API.
34 . The non-transitory machine-readable medium of claim 33 , wherein the shared memory location is multicast memory.
35 . The non-transitory machine-readable medium of claim 33 , wherein the shared memory location is unicast memory.
36 . The non-transitory machine-readable medium of claim 33 , wherein:
a first processor of the plurality of processors is of a first node of a compute cluster; and
a second processor of the plurality of processors is of a second node of the compute cluster.
37 . The non-transitory machine-readable medium of claim 33 , wherein a location of the shared memory location is shared between the plurality of processors.
38 . The non-transitory machine-readable medium of claim 33 , wherein the API receives a memory handle that at least indicates a location of the shared memory location.
39 . The non-transitory machine-readable medium of claim 38 , wherein the API receives a set of properties of the memory handle.
40 . The non-transitory machine-readable medium of claim 33 , wherein the API receives an allocation size that specifies a size of the physical memory that corresponds to the shared memory location.
41 . The non-transitory machine-readable medium of claim 33 , wherein the API receives a sharing descriptor that further enables one or more processors of the plurality of processors to access the shared memory location.
42 . The non-transitory machine-readable medium of claim 33 , wherein the API returns a success indicator to a calling process executing on a processor of the one or more processors.