Heterogeneous relaxations in differentiable logic gate networks
Systems, methods, and computer-readable media are described for training and implementing logic gate networks for inference tasks. An untrained node network includes a first set of nodes parameterized by differentiable parameters, each associated with a predefined set of potential logic gate operators, and a second set of nodes corresponding to predefined logic gate operators. During training, node outputs of the first set are computed using a first relaxation, and node outputs of the second set are computed using a second relaxation different from the first relaxation. Updated differentiable parameters are applied over multiple training iterations, and a fixed logic-gate network is generated by selecting logic gate operators for at least some nodes in the first set while retaining the predefined logic gate operators for the second set. In some embodiments, the second relaxation imposes structural inductive bias. ASICs and FPGAs implementing fixed logic gate networks are also described.
1 . A computing system, comprising:
one or more processors; and
a memory storing instructions that, when executed by the one or more processors, cause the computing system to:
instantiate an untrained node network comprising a plurality of nodes, wherein:
(i) each node of a first set of nodes of the plurality of nodes is parameterized by a set of differentiable parameters associated with a respective predefined set of potential logic gate operators, and
(ii) each node of a second set of nodes of the plurality of nodes corresponds to a respective predefined logic gate operator;
iteratively train the node network over a plurality of training iterations, each training iteration including:
computing, for each node in the first set of nodes, a respective node output based on the differentiable parameters thereof using a first relaxation for selecting a single logic gate operator from the set of potential logic gate operators,
computing, for each node in the second set of nodes, a respective node output based on the respective predefined logic gate operator using a second relaxation that is different from the first relaxation beyond the fact that the first relaxation uses differentiable parameters for selecting a single logic gate operator from the set of potential logic gate operators, and
applying updated differentiable parameters to at least one node in the first set of nodes; and
generate, after completion of the plurality of training iterations, a fixed logic gate network based on (i) the predefined logic gate operators of the second set of nodes, and (ii) a selection of a single logic gate operator from the set of potential logic gate operators based on the differentiable parameters of at least some nodes in the first set of nodes.
2 . The computing system of claim 1 , wherein the first relaxation is a probabilistic relaxation.
3 . The computing system of claim 1 , wherein the second relaxation uses at least one of:
(i) for at least one node of the second set of nodes that corresponds to a logical AND operator, a minimum operation as a relaxation of the logical AND operator; and
(ii) for at least one node of the second set of nodes that corresponds to a logical OR operator, a maximum operation as a relaxation of the logical OR operator.
4 . The computing system of claim 3 , wherein, for backpropagation through at least one node of the second set of nodes that uses the second relaxation, the computing system stores only path-selection information for routing a gradient, without storing an activation value of the at least one node.
5 . The computing system of claim 3 , wherein use of the second relaxation for the second set of nodes reduces memory consumption during training relative to using the first relaxation for all nodes.
6 . The computing system of claim 3 , wherein, at least two nodes of the first set are connected to one node of the second set, the node of the second set receiving the outputs of the at least two nodes of the first set, and wherein for backpropagation, a gradient is only routed through one path of the node of the second set such that the gradient is not backpropagated through all of the at least two nodes of the first set.
7 . The computing system of claim 3 , wherein the first relaxation is a probabilistic relaxation.
8 . The computing system of claim 2 , wherein the probabilistic relaxation comprises probabilities corresponding to each of the respective logic gate operators of the predefined set of logic gate operators based on the differentiable parameters, and wherein inputs to a respective node of the first set are input activation probabilities between 0 and 1, and wherein the training iterations further comprise computing an expectation value of an output under assumption of independent node input probabilities.
9 . The computing system of claim 1 , wherein the first relaxation is based on a Yager t-norm and Yager t-conorm, and wherein the second relaxation is based on the Hamacher t-norm and Hamacher t-conorm.
10 . The computing system of claim 1 , wherein the first relaxation comprises a continuous non-linear function of inputs to the respective node and differentiable parameters thereof.
11 . The computing system of claim 1 , wherein the first relaxation is a stochastic sampling-based relaxation, and wherein the second relaxation is a continuous relaxation.
12 . The computing system of claim 1 , wherein the second relaxation is used to impose a structural inductive bias in the node network.
13 . The computing system of claim 12 , wherein the structural inductive bias comprises using the second relaxation for spatial pooling in a convolutional logic gate network.
14 . The computing system of claim 12 , wherein the structural inductive bias comprises using the second relaxation for at least one temporal connection in a sequential logic gate network.
15 . The computing system of claim 14 , wherein the at least one temporal connection is implemented using at least one of a flip-flop, a latch, or another circuit element configured to convey information from a previous cycle.
16 . The computing system of claim 12 , wherein the structural inductive bias comprises using the second relaxation to combine an activation from one layer of the node network with a residual activation from a different layer of the node network.
17 . The computing system of claim 1 , wherein the predefined set of potential logic gate operators comprises entries of a lookup table.
18 . The computing system of claim 1 , wherein the predefined set of potential logic gate operators includes at least two elements, including one or more of: an AND operator, an OR operator, a NAND operator, a NOR operator, an XOR operator, a constant TRUE operator, a constant FALSE operator, an inverter operator, and a pass-through operator that outputs one of the node inputs.
19 . An application-specific integrated circuit (ASIC), comprising:
logic circuitry implementing a fixed logic gate network for performing an inference task, the fixed logic gate network comprising a plurality of nodes, including:
a first set of nodes, each node of the first set implementing a respective selected logic gate operator selected from a respective predefined set of potential logic gate operators; and
a second set of nodes, each node of the second set implementing a respective predefined logic gate operator, wherein the second set of nodes comprises a plurality of repeating structural inductive bias gates arranged at multiple locations in the fixed logic gate network to impose structural inductive bias on the fixed logic gate network;
wherein the fixed logic gate network implemented by the logic circuitry was generated by a training process comprising:
instantiating an untrained node network comprising a plurality of nodes, wherein:
(i) each node of a first set of nodes of the plurality of nodes is parameterized by a set of differentiable parameters associated with a respective predefined set of potential logic gate operators, and
(ii) each node of a second set of nodes of the plurality of nodes corresponds to a respective predefined logic gate operator;
iteratively training the node network over a plurality of training iterations, each training iteration including:
computing, for each node in the first set of nodes, a respective node output based on the differentiable parameters thereof using a first relaxation for selecting among the set of potential logic gate operators,
computing, for each node in the second set of nodes, a respective node output based on the respective predefined logic gate operator using a second relaxation different from the first relaxation, and
applying updated differentiable parameters to at least one node in the first set of nodes; and
generating, after completion of the plurality of training iterations, the fixed logic gate network based on (i) the predefined logic gate operators of the second set of nodes and (ii) selection of a single logic gate operator from the set of potential logic gate operators based on the differentiable parameters for at least some nodes of the first set of nodes.
20 . The ASIC of claim 19 , wherein at least a subset of the plurality of repeating structural inductive bias gates comprises one or more of:
(i) spatial pooling gates in a convolutional logic gate network,
(ii) temporal connection gates in a sequential logic gate network, and
(iii) gates configured to combine an activation from one layer with a residual activation from a different layer.