Interconnect mode for computational arrays
A processing engine array is provided with an interconnect mode of operation to use the array as an interconnect to move data elements to different locations in memory such as to perform a matrix transpose operation. In this interconnect mode of operation, although computations are still being performed in the array, the computations are not carried out to modify or change the values of the data elements, but are instead carried out to rearrange the data elements in memory. As such, the computations carried out in the interconnect mode of operation can deviate from the expected behavior of floating-point calculations. A mode selection signal can be used to provide the proper outputs of the processing elements of the array depending on the mode of operation.
1 . A neural network processor comprising:
a state buffer memory;
a results buffer memory; and
a processing engine array having a plurality of processing elements,
wherein the neural network processor is operable to execute a set of instructions to perform a transpose operation to transpose a tensor, wherein the transpose operation includes:
loading data elements of the tensor from the state buffer memory into the processing engine array;
performing a matrix multiplication on the data elements loaded in the processing engine array with an identity matrix to generate a set of outputs; and
storing the set of outputs in the results buffer memory, and
wherein the processing engine array is configured to set a multiplication result to a zero value for the transpose operation in response to multiplying a data element representing infinity stored in the processing engine array with an element of zero from the identity matrix, and
wherein each processing element of the processing engine array includes a multiplication result selection circuit operable to select the multiplication result from a plurality of values including the zero value, a NaN value, and an infinity value.
2 . The neural network processor of claim 1 , wherein the data elements of the tensor are loaded from row partitions of the state buffer memory into corresponding rows of processing elements of the processing engine array, and the set of outputs are stored in column partitions of the results buffer memory, wherein the column partitions in the results buffer memory are mapped to the row partitions in the state buffer memory.
3 . The neural network processor of claim 1 , wherein the multiplication result selection circuit is operable to select the NaN value as a result of multiplying infinity with zero when the processing engine array is not performing a transpose operation.
4 . The neural network processor of claim 1 , wherein the multiplication result selection circuit is operable to select an infinity value as a result of multiplying infinity with a non-zero number.
5 . An integrated circuit device comprising:
a processing element including a multiplication circuit and an adder circuit coupled to the multiplication circuit,
wherein the processing element is operable to generate a multiplication result based on multiplying two multiplication operands using the multiplication circuit, and to generate a partial sum output based on a partial sum input using the adder circuit,
wherein in an interconnect mode of operation, the processing element is operable to generate the partial sum output having a same value as the partial sum input in response to the two multiplication operands being zero and infinity,
wherein the processing element further comprises a multiplication result selection circuit that is operable, in the interconnect mode of operation, to set the multiplication result to a zero value in response to the two multiplication operands being zero and infinity, and
wherein the multiplication result selection circuit is operable in a compute mode of operation, to set the multiplication result to a not-a-number (NaN) value in response to the two multiplication operands being zero and infinity.
6 . The integrated circuit device of claim 5 , wherein the multiplication result selection circuit is operable, in the interconnect mode of operation, to set the multiplication result to the zero value in response to the two multiplication operands being zero and NaN.
7 . The integrated circuit device of claim 5 , wherein the processing element further comprises an output selection circuit that is operable, in the interconnect mode of operation, to select the partial sum input as the partial sum output in response to the two multiplication operands being zero and infinity.
8 . The integrated circuit device of claim 7 , wherein the output selection circuit is operable, in a compute mode of operation, to select a not-a-number (NaN) value as the partial sum output in response to the two multiplication operands being zero and infinity.
9 . The integrated circuit device of claim 5 , wherein the processing element is part of an array of processing elements, and wherein the interconnect mode of operation is used to perform a matrix transpose operation using the array of processing elements.
10 . The integrated circuit device of claim 9 , wherein the matrix transpose operation includes multiplying an input matrix with an identity matrix.
11 . The integrated circuit device of claim 10 , wherein the matrix transpose operation includes preloading the input matrix into weight cache registers of the array, and subsequently shifting the identity matrix into feature map registers of the array.
12 . The integrated circuit device of claim 5 , wherein the processing element is operable to dynamically switch between the interconnect mode of operation and a compute mode of operation based on an instruction being executed by the integrated circuit device.
13 . A method comprising:
receiving a request to perform a transpose operation on a tensor;
in response to receiving the request to perform the transpose operation, asserting a mode signal to cause a computational array to use a zero value as a result of multiplying zero with infinity;
loading data elements of the tensor from a first memory into the computational array;
performing an identity multiplication on the data elements loaded into the computational array, wherein performing the identity multiplication includes shifting an identity matrix into the computational array on feature map input buses;
storing a result of the identity multiplication in a second memory; and
loading the result from the second memory into the first memory.
14 . The method of claim 13 , further comprising:
receiving a request to perform a matrix multiplication operation; and
deasserting the mode signal to cause the computational array to use a not-a-number value as a result of multiplying zero with infinity.
15 . The method of claim 14 , wherein the mode signal controls multiplication result selection circuits in the computational array.
16 . The method of claim 14 , wherein the mode signal controls partial sum output selection circuits in the computational array.
17 . The method of claim 14 , wherein the data elements of the tensor loaded into the computational array are stored in weight cache registers of the computational array.
18 . A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to execute a compiler, the compiler performing operations including:
receiving a set of operations to be performed by an execution engine having a computational array;
determining that the set of operations includes a matrix transpose operation to be performed on a tensor; and
generating a set of compiled instructions operable to transpose the tensor by:
asserting a mode signal to cause the computational array to use a zero value as a result of multiplying zero with infinity;
loading data elements of the tensor from a first memory into the computational array;
performing an identity multiplication on the data elements loaded into the computational array, wherein performing the identity multiplication includes shifting an identity matrix into the computational array on feature map input buses;
storing a result of the identity multiplication in a second memory; and
loading the result from the second memory into the first memory.