Convolutional Logic Gate Networks
Systems and methods are described for training and using logic gate tree networks and convolutional logic gate tree networks. A computing system may apply a convolution operation to an input tensor using one or more kernels, each kernel comprising a tree of logic gate nodes. For each kernel placement, leaves of the tree select input activations from a receptive field, and internal nodes generate a kernel output by applying logic operations. During training, each node may be parameterized with differentiable parameters that define a probability distribution over candidate logic operators. The parameters may be shared across kernel placements to provide equivariance. Gradient-based optimization may be used to update the differentiable parameters during training. A fixed tree kernel may be defined by selecting a logic operator for each node, which can be executed by a processor or synthesized for use in programmable logic and/or in an application-specific integrated circuit.
1 . A method for training a convolutional node network, comprising:
receiving, at a computing system, a training data set comprising inputs,
instantiating, in a memory of the computing system, an untrained convolutional node network comprising at least two layers, each having at least two convolutional nodes,
for each convolutional node, directly optimizing the choice of logic gate operation across two or more options using a gradient-based optimization algorithm, the two or more options of logic gate operations including two or more of:
an AND operator, an OR operator, a NAND operator, a NOR operator, an XOR operator, a constant TRUE operator, a constant FALSE operator, an inverter operator, a pass-through operator that outputs one of the node inputs, and entries of a lookup table with two or more inputs; and
generating, after the optimization, a fixed logic gate network based on a selection of a single logic gate operator for at least some of the convolutional nodes.
2 . The method of claim 1 , wherein generating the fixed logic gate network further comprises: generating a gate-level netlist configured for synthesis for a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC).
3 . A method for training and generating a fixed convolutional logic gate network, comprising:
receiving, at a computing system, a training data set comprising inputs, wherein the inputs comprise an input tensor comprising activations arranged across a plurality of input channels and a plurality of positions in a domain;
instantiating, in a memory of the computing system, an untrained convolutional logic gate network including a plurality of convolutional nodes with at least one layer having at least one convolutional node, and wherein each convolutional node is parameterized by a set of differentiable parameters corresponding to a predefined finite set of potential logic gate operators;
iteratively training the untrained convolutional logic gate network via a plurality of training iterations, each training iteration including:
forward-propagating a batch of the inputs through the convolutional logic gate network by, for each convolutional node, computing a differentiable output that is a function of the inputs of the respective node and current differentiable parameters thereof, wherein forward-propagating includes generating an output tensor by, for each of a plurality of output channels, convolving the input tensor with a respective convolutional node across a plurality of kernel placements in the domain;
computing a loss value;
determining, via a training optimization algorithm, updated differentiable parameters for at least one convolutional node; and
applying the updated differentiable parameters to at least one convolutional node; and
selecting, after completion of the plurality of training iterations, for each of at least some of the plurality of convolutional nodes, a single logic gate operator from the predefined finite set of potential logic gate operators based on the differentiable parameters of the respective node; and
generating, based on the selecting, a fixed logic gate network based on the selected single logic gate operators for at least some of the plurality of convolutional nodes.
4 . The method of claim 3 , wherein generating a fixed logic gate network includes unrolling multiple kernel placements of a convolutional node's selected logic gate operator in a circuit.
5 . The method of claim 4 , further comprising, after the unrolling, implementing a logic synthesis process that removes or modifies some of the kernel placements.
6 . The method of claim 3 , wherein generating a fixed logic gate network includes time-multiplexing by, for each convolutional node of a subset of at least some of the convolutional nodes, implementing a plurality of kernel placements of the convolutional node with only a single logic gate for the plurality of kernel placements of the convolutional node.
7 . The method of claim 6 , wherein time-multiplexing is facilitated through flip-flops placed in a circuit.
8 . The method of claim 3 , wherein generating a fixed logic gate network includes, for each convolutional node of a subset of at least some of the convolutional nodes:
(i) unrolling some kernel placements of the convolutional node's selected logic gate operator in a circuit; and
(ii) time-multiplexing by implementing a plurality of kernel placements with only a single logic gate operator.
9 . The method of claim 8 , wherein the domain has at least two spatial dimensions, and one dimension is unrolled, and one dimension is time-multiplexed.
10 . The method of claim 8 , wherein a set of at least two neighboring kernel placements is unrolled, and their collective use is time-multiplexed.
11 . The method of claim 3 , wherein the domain has exactly one spatial dimension.
12 . The method of claim 3 , wherein the domain has at least two spatial dimensions.
13 . The method of claim 3 , wherein, for each convolutional node of a subset of at least some of the convolutional nodes, the plurality of kernel placements of the convolutional node shares the same differentiable parameters.
14 . The method of claim 3 , wherein generating the fixed logic gate network comprises a logic synthesis process that reduces the number of logic gate operators by more than 50%.
15 . The method of claim 14 , wherein generating the fixed logic gate network comprises a logic synthesis process that reduces the number of logic gate operators by more than 75%.
16 . The method of claim 15 , wherein the logic synthesis process comprises one or more of (i) constant propagation, (ii) wire removal or collapse, and (iii) removal of inverters by collapse into one or more connected nodes.
17 . The method of claim 3 , further comprising at least two layers, each having at least one convolutional node, including at least a first convolutional layer and a second convolutional layer, and wherein forward-propagating a batch of the inputs through the convolutional logic gate network comprises:
for each node of the first convolutional layer, computing a real-valued non-binary differentiable output that is a real-valued non-binarizing non-linear differentiable function of:
(i) input activations to the node, and
(ii) current differentiable parameters of the respective node; and
for each node of the second convolutional layer, computing a real-valued non-binary differentiable output that is a real-valued non-binarizing non-linear differentiable function of the:
(i) input activations to the node, at least some of which are the real-valued non-binary differentiable outputs of the first convolutional layer, and
(ii) current differentiable parameters of the respective node.
18 . The method of claim 3 , wherein the predefined set of potential logic gate operators includes entries of a lookup table.
19 . The method of claim 3 , wherein the predefined set of potential logic gate operators includes at least two elements, including one or more of: an AND operator, an OR operator, a NAND operator, a NOR operator, an XOR operator, a constant TRUE operator, a constant FALSE operator, an inverter operator, and a pass-through operator that outputs one of the node inputs.
20 . The method of claim 3 , further comprising:
implementing a logical expression of the fixed logic gate network in an application-specific integrated circuit (ASIC).
21 . An application-specific integrated circuit (ASIC) with logic circuitry that is an implementation of a fixed logic gate network, manufactured using a process comprising:
receiving, at a computing system, a training data set comprising inputs, wherein the inputs comprise an input tensor comprising activations arranged across a plurality of input channels and a plurality of positions in a domain;
instantiating, in a memory of the computing system, an untrained convolutional logic gate network including a plurality of convolutional nodes with at least one layer having at least one convolutional node, and wherein each convolutional node is parameterized by a set of differentiable parameters corresponding to a predefined finite set of potential logic gate operators;
iteratively training the untrained convolutional logic gate network via a plurality of training iterations, each training iteration including:
forward-propagating a batch of the inputs through the convolutional logic gate network by, for each convolutional node, computing a differentiable output that is a function of the inputs of the respective node and current differentiable parameters thereof, wherein forward-propagating includes generating an output tensor by, for each of a plurality of output channels, convolving the input tensor with a respective convolutional node across a plurality of kernel placements in the domain;
computing a loss value;
determining, via a training optimization algorithm, updated differentiable parameters for at least one convolutional node; and
applying the updated differentiable parameters to at least one convolutional node; and
selecting, after completion of the plurality of training iterations, for each of at least some of the plurality of convolutional nodes, a single logic gate operator from the predefined finite set of potential logic gate operators based on the differentiable parameters of the respective node; and
generating a fixed logic gate network based on the selected single logic gate operators for at least some of the plurality of convolutional nodes.
22 . The ASIC of claim 21 , wherein the process further comprises one or more of:
synthesizing the fixed logic gate network;
technology-mapping the fixed logic gate network; and
place-and-routing operations for circuit components.