Application programming interface to translate a tensor
Apparatuses, systems, and techniques to cause a first tensor to be translated into a second tensor according to a tensor map. In at least one embodiment, one or more circuits are to perform an application programming interface (API) to cause a first tensor to be translated into a second tensor according to a tensor map.
1 . One or more processors, comprising:
circuitry to, in response to receiving an application programming interface (API) call that indicates a first tensor and tensor map, cause a thread to initiate a transformation of the first tensor to be translated into a second tensor according to the tensor map, the first tensor and the second tensor having different tensor layouts, wherein the circuitry is to cause the second tensor to be asynchronously stored such that the thread is permitted to perform other operations before storage of the second tensor is complete.
2 . The one or more processors of claim 1 , wherein the circuitry is to, in response to the API call, asynchronously store the second tensor.
3 . The one or more processors of claim 1 , wherein the circuitry is to, in response to the API call, asynchronously store the second tensor in a memory of a graphics processing unit (GPU).
4 . The one or more processors of claim 1 , wherein the first tensor is to be stored in a first memory of a graphics processing unit (GPU), and the circuitry is to, in response to the API call, asynchronously transform the first tensor into the second tensor and store the second tensor in a second memory of the GPU.
5 . The one or more processors of claim 1 , wherein in response to the API call, the circuitry uses automatic transaction accounting.
6 . The one or more processors of claim 1 , wherein the API call receives as an input parameter an indication of a location in which the tensor map is stored, and wherein the tensor map is stored in a data structure that indicates information about the first tensor and the second tensor.
7 . The one or more processors of claim 1 , wherein in the API call receives as an input parameter a portion of the first tensor to be transformed into the second tensor.
8 . A system, comprising:
one or more processors to, in response to receiving an application programming interface (API) call that indicates a first tensor and a tensor map, cause a thread to initiate a transformation of the first tensor into a second tensor according to the tensor map, the first tensor and the second tensor having different tensor layouts, wherein the one or more processors are to cause the second tensor to be asynchronously stored such that the thread is permitted to perform other operations before storage of the second tensor is complete.
9 . The system of claim 8 , wherein in response to the API call, the one or more processors cause the first tensor to be transformed into the second tensor asynchronously.
10 . The system of claim 8 , wherein in response to the API call, the one or more processors asynchronously store the second tensor in a memory of a graphics processing unit (GPU).
11 . The system of claim 8 , wherein in response to the API call, the one or more processors use automatic transaction accounting.
12 . The system of claim 8 , wherein in response to the API call, the one or more processors asynchronously copy data from the first tensor according to the tensor map.
13 . The system of claim 8 , wherein in response to the API call, the one or more processors obtain a portion of the first tensor to be transformed into the second tensor based on an input parameter of the API call.
14 . A method, comprising:
receiving an application programming interface (API) call that indicates a first tensor and a tensor map, wherein the tensor map indicates a layout of a second tensor in memory; and
in response to receipt of the API call, causing a thread to initiate a transformation of the first tensor into the second tensor according to the tensor map, the first tensor and the second tensor having different layouts, wherein the second tensor is to be asynchronously stored such that the thread is permitted to perform other operations before storage of the second tensor is complete.
15 . The method of claim 14 , further comprising, in response to receipt of the API call, initiating one or more memory copy operations to be performed asynchronously.
16 . The method of claim 14 , wherein the API is to use transaction accounting different from manual transaction accounting.
17 . The method of claim 14 , wherein the tensor map is stored in a data structure that indicates information about the first tensor and the second tensor.
18 . The method of claim 14 , further comprising, in response to receipt of the API call, using an input parameter of the API call to obtain the tensor map from memory.
19 . The method of claim 14 , further comprising, in response to receipt of the API call, using an input parameter of the API call to obtain a portion of the first tensor to be transformed into the second tensor.
20 . A non-transitory computer-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least perform the method of claim 14 .