DATAFLOW ARCHITECTURE PROCESSOR STATICALLY RECONFIGURABLE TO PERFORM N-DIMENSIONAL AFFINE TRANSFORMATION IN PARALLEL MANNER BY REPLICATING COPIES OF INPUT IMAGE ACROSS SCRATCHPAD MEMORY BANKS
A statically reconfigurable dataflow architecture processor performs an N-dimensional affine transform specified by a matrix on an input image to produce an output image includes at least N+1 statically reconfigurable pattern compute units (PCUs) and pattern memory units (PMUs) each comprising a memory arranged as a vector of L banks. A first PMU writes a copy of the input image into each of the L banks. Each of N of the PCUs associated with the N dimensions is statically reconfigurable to apply a respective row of the transform matrix to N L-vectors of output pixel coordinates to generate a respective L-vector of input pixel coordinates. At least one of the PCUs flattens the N L-vectors of input pixel coordinates to calculate an L-vector of addresses. The first PMU uses the L addresses of the L-vector of addresses to read an L-vector of input pixels from the L banks in parallel.
1 . A statically reconfigurable dataflow architecture processor (SRDAP) to perform an N-dimensional affine transform specified by a matrix on an N-dimensional input image to produce an N-dimensional output image comprising output pixels, each output pixel having a coordinate in each of the N dimensions, wherein N is at least two, comprising:
at least N+1 statically reconfigurable pattern compute units (PCUs); and
one or more statically reconfigurable pattern memory units (PMUs) each comprising a memory arranged as a vector of L banks;
wherein a first PMU of the PMUs is statically reconfigurable to receive the input image and to write a copy of the input image into each bank of the L banks;
wherein each of N of the PCUs associated with the N dimensions is statically reconfigurable to apply a respective row of the transform matrix to N L-vectors of output pixel coordinates to generate a respective L-vector of input pixel coordinates;
wherein at least one of the PCUs is statically reconfigurable to calculate an L-vector of addresses by flattening the N L-vectors of input pixel coordinates; and
wherein the first PMU is statically reconfigurable to use the L addresses of the L-vector of addresses to read an L-vector of input pixels from the L banks in parallel.
2 . The SRDAP of claim 1 , further comprising:
configuration stores loadable with configuration data to statically reconfigure the SRDAP.
3 . The SRDAP of claim 2 ,
wherein to statically reconfigure the SRDAP comprises loading the configuration stores with the configuration data prior to initiation of production of the output image without re-loading the configuration stores with the configuration data until completion of production of the output image.
4 . The SRDAP of claim 1 ,
wherein a second PMU of the PMUs is statically reconfigurable to receive the L-vector of input pixels and to write the L-vector of input pixels to the memory of the second PMU.
5 . The SRDAP of claim 4 ,
wherein each of N of the PCUs is further statically reconfigurable to apply the respective row of the transform matrix to a series of N L-vectors of output pixel coordinates to generate a respective series of L-vectors of input pixel coordinates;
wherein the at least one of the PCUs is further statically reconfigurable to calculate a series of L-vectors of addresses by flattening the series of N L-vectors of input pixel coordinates;
wherein the first PMU is further statically reconfigurable to read a series of L-vectors of input pixels from the L banks in parallel; and
wherein the second PMU is further statically reconfigurable to receive the series of L-vectors of input pixels and to write the series of L-vectors of input pixels to the memory of the second PMU to form the output image.
6 . The SRDAP of claim 5 ,
wherein the first PMU is configured to receive the input image as a series of input pixels;
wherein the first PMU comprises a counter that is statically reconfigurable with a terminal value that is a size of the input image; and
wherein the first PMU comprises address generation logic that is statically reconfigurable to use a series of indexes received from the counter to write a copy of each input pixel of the series of input pixels to each bank of the L banks at the series of indexes.
7 . The SRDAP of claim 5 ,
wherein the first PMU comprises address generation logic that is statically reconfigurable to provide the L addresses of the series of L-vectors of addresses to the L banks in parallel to read the series of L-vectors of input pixels from the L banks in parallel.
8 . The SRDAP of claim 7 ,
wherein the series of L-vectors of input pixels read from the L banks in parallel is a number equal to a quotient of a size of the output image divided by L; and
wherein the first PMU further comprises a counter that is statically reconfigurable to count the number of times to control the first PMU to read the series of L-vectors of input pixels read from the L banks in parallel.
9 . The SRDAP of claim 5 ,
wherein the second PMU comprises a counter that is statically reconfigurable with an initial value of zero, a stride value of one, and a terminal value that is a quotient of a size of the output image divided by L to generate a series of bank indexes; and
wherein the second PMU is statically reconfigurable to use the series of bank indexes received from the counter to write the series of vectors of input pixels to the L banks of the second PMU memory to form the output image.
10 . The SRDAP of claim 5 ,
wherein second PMU is statically reconfigurable to read the output image for writing to a memory external to the SRDAP.
11 . The SRDAP of claim 4 ,
wherein the SRDAP is statically reconfigurable to sustain writing a series of L-vectors of input pixels to the memory of the second PMU at a throughput of at least one L-vector of input pixels per N clock cycles.
12 . The SRDAP of claim 1 ,
wherein first PMU is statically reconfigurable to receive the input image from a memory external to the SRDAP.
13 . A computer-implemented method for performing an N-dimensional affine transform specified by a matrix on an N-dimensional input image to produce an N-dimensional output image comprising output pixels, each output pixel having a coordinate in each of the N dimensions, wherein N is at least two, comprising:
statically reconfiguring a statically reconfigurable dataflow architecture processor (SRDAP) that comprises at least N+1 statically reconfigurable pattern compute units (PCUs) and one or more statically reconfigurable pattern memory units (PMUs) each comprising a memory arranged as a vector of L banks;
receiving, by a first PMU of the PMUs, the input image and writing a copy of the input image into each bank of the L banks;
applying, by each of N of the PCUs associated with the N dimensions, a respective row of the transform matrix to N L-vectors of output pixel coordinates to generate a respective L-vector of input pixel coordinates;
calculating, by at least one of the PCUs, an L-vector of addresses by flattening the N L-vectors of input pixel coordinates; and
using, by the first PMU, the L addresses of the L-vector of addresses to read an L-vector of input pixels from the L banks in parallel.
14 . The method of claim 13 , further comprising:
wherein said statically reconfiguring the SRDAP comprises loading configuration stores of the SRDAP with configuration data.
15 . The method of claim 14 ,
wherein said statically reconfiguring the SRDAP comprises loading the configuration stores with the configuration data prior to initiation of production of the output image without re-loading the configuration stores with the configuration data until completion of production of the output image.
16 . The method of claim 13 , further comprising:
receiving, by a second PMU of the PMUs, the L-vector of input pixels and writing the L-vector of input pixels to the memory of the second PMU.
17 . The method of claim 16 , further comprising:
applying, by each of N of the PCUs, the respective row of the transform matrix to a series of N L-vectors of output pixel coordinates to generate a respective series of L-vectors of input pixel coordinates;
calculating, by the at least one of the PCUs, a series of L-vectors of addresses by flattening the series of N L-vectors of input pixel coordinates;
reading, by the first PMU, a series of L-vectors of input pixels from the L banks in parallel; and
receiving, by the second PMU, the series of L-vectors of input pixels and writing the series of L-vectors of input pixels to the memory of the second PMU to form the output image.
18 . The method of claim 17 , further comprising:
receiving, by the first PMU, the input image as a series of input pixels;
statically reconfiguring a counter of the first PMU with a terminal value that is a size of the input image; and
using, by address generation logic of the first PMU, a series of indexes received from the counter to write a copy of each input pixel of the series of input pixels to each bank of the L banks at the series of indexes.
19 . The method of claim 17 , further comprising:
providing, by address generation logic of the first PMU, the L addresses of the series of L-vectors of addresses to the L banks in parallel to read the series of L-vectors of input pixels from the L banks in parallel.
20 . A non-transitory computer-readable storage medium having computer program instructions stored thereon that are capable of causing or configuring a statically reconfigurable dataflow architecture processor (SRDAP) to perform an N-dimensional affine transform specified by a matrix on an N-dimensional input image to produce an N-dimensional output image comprising output pixels, each output pixel having a coordinate in each of the N dimensions, wherein N is at least two, comprising:
at least N+1 statically reconfigurable pattern compute units (PCUs); and
one or more statically reconfigurable pattern memory units (PMUs) each comprising a memory arranged as a vector of L banks;
wherein a first PMU of the PMUs is statically reconfigurable to receive the input image and to write a copy of the input image into each bank of the L banks;
wherein each of N of the PCUs associated with the N dimensions is statically reconfigurable to apply a respective row of the transform matrix to N L-vectors of output pixel coordinates to generate a respective L-vector of input pixel coordinates;
wherein at least one of the PCUs is statically reconfigurable to calculate an L-vector of addresses by flattening the N L-vectors of input pixel coordinates; and
wherein the first PMU is statically reconfigurable to use the L addresses of the L-vector of addresses to read an L-vector of input pixels from the L banks in parallel.