System and architecture neural network accelerator including filter circuit
View Patent ↗A system and an accelerator circuit includes an internal memory to store data received a memory associated with a processor and a filter circuit block comprising a plurality of circuit stripes, each circuit stripe including a filter processor, a plurality of filter circuits, and a slice of the internal memory assigned to the plurality of filter circuits, where the filter processor is to execute a filter instruction to read data values from the internal memory based on a first memory address, for each of the plurality of circuit stripes: load the data values in weight registers and input registers associated with the plurality of filter circuits of the circuit stripe to generate a plurality of filter results, and write a result generated using the plurality of filter circuits in the internal memory at a second memory address.
1 . An accelerator circuit comprising:
an internal memory to store data received at a memory associated with a processor;
a filter circuit block comprising a plurality of circuit stripes, each circuit stripe comprising:
a filter processor;
a plurality of filter circuits; and
a slice of the internal memory assigned to the plurality of filter circuits,
wherein the filter processor is to execute a filter instruction, to:
read data values from the internal memory based on a first memory address;
for each of the plurality of circuit stripes:
load the data values in weight registers and input registers associated with the plurality of filter circuits of the circuit stripe to generate a plurality of filter results; and
write a result generated using the plurality of filter circuits in the internal memory at a second memory address;
an input circuit block comprises an input processor, a first instruction memory, a first data memory, and a first program counter (PC), wherein the first instruction memory is to store first instructions of a first task specified according to a first instruction set architecture (ISA) of the input processor, wherein the first PC is to store an address of a current first instruction to be executed by a first processor, and wherein the input processor is to execute the first instructions using first data stored in the first data memory to perform the first task;
the filter circuit block comprising a second instruction memory, a second data memory, and a second PC, wherein the second instruction memory is to store second instructions of a second task specified according to a second ISA of the filter processor, wherein the second PC is to store an address of a current second instruction to be executed by a second processor, and
wherein the filter processor is to execute the second instructions using second data stored in the second data memory to perform the second task;
a post-processing circuit block comprising a post-processing processor, a third instruction memory, a third data memory, a third PC, wherein the third instruction memory is to store third instructions of a third task specified according to a third ISA of the post-processing processor, wherein the third PC is to store an address of a current third instruction to be executed by a third processor, and wherein the post-processing processor is to execute the third instructions using third data stored in the third data memory to perform the third task; and
an output circuit block comprising an output processor, a fourth instruction memory, a fourth data memory, a fourth PC, wherein the fourth instruction memory is to store fourth instructions of a fourth task specified according to a fourth ISA of the output processor, wherein the fourth PC is to store an address of a current fourth instruction to be executed by a fourth processor, and wherein the output processor is to execute the fourth instructions using fourth data stored in the fourth data memory to perform the fourth task.
2 . The accelerator circuit of claim 1 , further comprising:
a plurality of general-purpose registers to store a first flag bit indicating a start of execution of at least one of the tasks;
a plurality of shadow registers to store content to be copied to the plurality of general-purpose registers responsive to a copy event; and
an interface circuit comprising:
a control register to store a plurality of interrupts to the processor,
an error register to store a plurality of error flags indicating occurrences of different kinds of errors,
a next register to store a mask for selecting the content of the plurality of shadow registers to be copied to the plurality of general-purpose registers, and
a quality-of-service register to store controls to the memory associated with the processor.
3 . The accelerator circuit of claim 1 , wherein the filter circuit block comprises:
eight identical circuit stripes, wherein a circuit stripe of the eight identical circuit stripes comprises:
four three-by-three convolutional neural network (CNN) filter circuits; and
a slice of the internal memory assigned to the four three-by-three CNN filter circuits.
4 . The accelerator circuit of claim 3 , wherein the circuit stripe of the eight identical circuit stripes comprises the four three-by-three CNN filter circuits, and wherein the four three-by-three CNN filter circuits share a first plurality of registers to store common weight parameters of the four three-by-three CNN filter circuits.
5 . The accelerator circuit of claim 3 , wherein one of the four three-by-three CNN filter circuits comprises:
a first plurality of registers to store weight parameters;
a second plurality of registers to store input values;
a reduce tree comprising a plurality of multiplication circuits to calculate a production between a respective one of the weight parameters and a respective one of the input values; and
a sum circuit to add results from the plurality of multiplication circuits with a carry-over result from another one of the three-by-three CNN filter circuit.
6 . The accelerator circuit of claim 5 , wherein the eight identical circuit stripes are each to output four results generated by the four three-by-three CNN filter circuits therein, and wherein the filter circuit block comprises four addition circuits, each of the four addition circuit is to sum a respective one of the four results from the eight filter circuits.
7 . The accelerator circuit of claim 5 , wherein the input circuit block, the filter circuit block, the post-processing circuit block, and the output circuit block form an execution pipeline, and
wherein the input processor of the input circuit block is to read from the memory associated with the processor and to write a first result generated by the input circuit block to the internal memory, the filter processor is to read the first result from the internal memory and to write a second result generated by the filter circuit block to the internal memory, the post-processing processor is to read the second result from the internal memory and to write a third result generated by the post-processing circuit block to the internal memory, and the output processor is to read the third result from the internal memory and to write a fourth result generated by the output circuit block to the memory associated with the processor.
8 . The accelerator circuit of claim 5 , wherein at least two of the input circuit block, the filter circuit block, the post-processing circuit block, and the output circuit block are to operate concurrently.
9 . The accelerator circuit of claim 5 , wherein the post-processing circuit block is to perform at least one of a compaction function, an activation function, or a top-N function.
10 . The accelerator circuit of claim 5 , wherein the first instructions comprise an input instruction comprising a type value, a mode value, and a size value, and wherein to execute the input instruction, the input processor is to:
concatenate a first local register and a second local register of the input processor to form a first memory address;
read, based on the first memory address, the memory associated with the processor to retrieve data values in a format determined by the type value, wherein the size value determines a number of bytes associated with the data values retrieved from the memory; and
write the data values to the local memory in a mode determined by the mode value.
11 . The accelerator circuit of claim 5 , wherein the second instructions comprise a filter instruction comprising a sum value and a mode value, and wherein to execute the filter instruction, the filter processor is to:
read data values from the internal memory based on a first memory address stored in a first local register of the filter processor;
for each circuit stripe: load the data values in weight registers and input registers associated with the plurality of filter circuits of the circuit stripe to generate a plurality of filter results;
calculate, based on the sum value, sum values of corresponding filter results from each circuit stripe;
determine, based on the mode value, whether the filter circuit block is in a slice mode or in a global mode;
responsive to determining that the filter circuit block is in the slice mode,
write the sum values to each slice with an offset stored in a second local register; and
responsive to determining that the filter circuit block is in a global mode,
write the sum values to the internal memory at a memory address stored in the second local register.
12 . The accelerator circuit of claim 5 , wherein the third instructions comprise a post-processing instruction comprising an identifier of the control register, a column value, a row value, and a kind value, wherein to execute the post-processing instruction, the post-processing processor is to:
responsive to determining that the kind value indicates a compaction mode,
read, based on a first memory address stored in a first local register, a number of data values from the internal memory, wherein the column value specifies the number;
group, based on the column value, the data values into a plurality of groups;
compact, based on the kind value, each of the plurality of groups into one of a maximum value of the group or a first element of the group;
determine a maximum value among the compacted values and a value stored in a stage register;
store the maximum value in the state register;
apply an activation function to the maximum value, wherein the activation function is one of an identity function, a step function, a sigmoid function, or a hyperbolic tangent function; and
write, based on a second memory address stored in a second local register, a result of the activation function to the internal memory; and
responsive to determining that the kind value indicates a top-N mode,
read, based on a third memory address stored in a third local register, an array of data values;
determine a top-N values and their corresponding positions in the array, wherein N is an integer greater than 1;
responsive to determining that the control value indicates a memory write, write the top-N values and their corresponding positions in the internal memory at a fourth memory address stored in a fourth local register; and
responsive to determining that the control value indicates a register write, write the top-N values and their corresponding positions in state registers associated with the post-processing processor.
13 . The accelerator circuit of claim 5 , wherein the fourth instructions comprise an output instruction comprising a type value, a mode value, and a size value, and wherein to execute the output instruction, the output processor is to:
read a plurality of data values from the internal memory based on a first memory address stored in a first local register of the output processor, wherein the size value specifies a number of the plurality of data values;
concatenate a second local register and a third local register of the output processor to form a second memory address; and
write the plurality of data values to the memory associated with the processor based on the second memory address in a format determined by the type value.
14 . A system, comprising:
a memory to store data;
a processor, communicatively coupled to the memory, to execute a neural network application comprising filter operations using the data; and
an accelerator circuit comprising:
an internal memory;
a filter circuit block comprising a plurality of circuit stripes, each circuit stripe comprising:
a filter processor;
a plurality of filter circuits; and
a slice of the internal memory assigned to the plurality of filter circuits,
wherein the filter processor is to execute a filter instruction, to:
read data values from the internal memory based on a first memory address;
for each circuit stripe:
load the data values in weight registers and input registers associated with the plurality of filter circuits of the circuit stripe to generate a plurality of filter results; and
write a result generated using the plurality of filter circuits in the internal memory at a second memory address;
an input circuit block comprises an input processor, a first instruction memory, a first data memory, and a first program counter (PC), wherein the first instruction memory is to store first instructions of a first task specified according to a first instruction set architecture (ISA) of the input processor, wherein the first PC is to store an address of a current first instruction to be executed by a first processor, and wherein the input processor is to execute the first instructions using first data stored in the first data memory to perform the first task;
the filter circuit block comprising a second instruction memory, a second data memory, and a second PC, wherein the second instruction memory is to store second instructions of a second task specified according to a second ISA of the filter processor, wherein the second PC is to store an address of a current second instruction to be executed by a second processor, and wherein the filter processor is to execute the second instructions using second data stored in the second data memory to perform the second task;
a post-processing circuit block comprising a post-processing processor, a third instruction memory, a third data memory, and a third PC, wherein the third instruction memory is to store third instructions of a third task specified according to a third ISA of the post-processing processor, wherein the third PC is to store an address of a current third instruction to be executed by a third processor, and wherein the post-processing processor is to execute the third instructions using third data stored in the third data memory to perform the third task; and
an output circuit block comprising an output processor, a fourth instruction memory, a fourth data memory, and a fourth PC, wherein the fourth instruction memory is to store fourth instructions of a fourth task specified according to a fourth ISA of the output processor, wherein the fourth PC is to store an address of a current fourth instruction to be executed by a fourth processor, and wherein the output processor is to execute the fourth instructions using fourth data stored in the fourth data memory to perform the fourth task.
15 . The system of claim 14 , wherein the filter circuit block comprises:
eight identical circuit stripes, wherein a circuit stripe of the eight identical circuit stripes comprises:
four three-by-three convolutional neural network (CNN) filter circuits; and
a slice of the internal memory assigned to the four three-by-three CNN filter circuits.
16 . The system of claim 15 , wherein the circuit stripe of the eight identical circuit stripes comprises the four three-by-three CNN filter circuits, and wherein the four three-by-three CNN filter circuits share a first plurality of registers to store common weight parameters of the four three-by-three CNN filter circuits.
17 . The system of claim 15 , wherein one of the four three-by-three CNN filter circuits comprises:
a first plurality of registers to store weight parameters;
a second plurality of registers to store input values;
a reduce tree comprising a plurality of multiplication circuits to calculate a production between a respective one of the weight parameters and a respective one of the input values; and
a sum circuit to add results from the plurality of multiplication circuits with a carry-over result from another one of the three-by-three CNN filter circuit.
18 . A method comprising:
receiving, by an accelerator circuit, a task comprising a filter instruction from a processor, wherein the accelerator circuit comprises a filter circuit block comprising a plurality of circuit stripes, each circuit stripe comprising a filter engine, a plurality of filter circuits, an input circuit block comprising an input processor, a first instruction memory, a first data memory, and a first program counter (PC), wherein the first instruction memory is to store first instructions of a first task specified according to a first instruction set architecture (ISA) of the input processor, wherein the first PC is to store an address of a current first instruction to be executed by a first processor, and wherein the input processor is to execute the first instructions using first data stored in the first data memory to perform the first task,
the filter circuit block comprising a second instruction memory, a second data memory, and a second PC, wherein the second instruction memory is to store second instructions of a second task specified according to a second ISA of the filter processor, wherein the second PC is to store an address of a current second instruction to be executed by a second processor, and wherein the filter processor is to execute the second instructions using second data stored in the second data memory to perform the second task,
a post-processing circuit block comprising a post-processing processor, a third instruction memory, a third data memory, a third PC, wherein the third instruction memory is to store third instructions of a third task specified according to a third ISA of the post-processing processor, wherein the third PC is to store an address of a current third instruction to be executed by a third processor, and wherein the post-processing processor is to execute the third instructions using third data stored in the third data memory to perform the third task;
an output circuit block comprising an output processor, a fourth instruction memory, a fourth data memory, a fourth PC, wherein the fourth instruction memory is to store fourth instructions of a fourth task specified according to a fourth ISA of the output processor, wherein the fourth PC is to store an address of a current fourth instruction to be executed by a fourth processor, and wherein the output processor is to execute the fourth instructions using fourth data stored in the fourth data memory to perform the fourth task,
and a slice of an internal memory assigned to the plurality of filter circuits;
reading data values from the internal memory of the accelerator circuit starting from a first memory address;
for each of the plurality of circuit stripes, loading the data values in weight registers and input registers associated with a plurality of filter circuits to generate a plurality of filter results; and
writing a result generated using the plurality of filter circuits in the internal memory at a second memory address.