Systems and Methods for Efficient Differentiable Logic Gate Networks
Systems and methods are disclosed for implementing and training differentiable logic-gate neural networks. A network includes hyper-logic gate nodes configured as differentiable lookup tables that receive selector signals and LUT entry signals and generate outputs via differentiable selection, enabling LUT-style multiplexing during discretization while remaining trainable by gradient-based optimization. LUT entry signals may be learned parameters, outputs of other nodes, external inputs, or internal memory-state values, enabling aggregation, switching, skip connections, and memory read/write operations. Training may employ constrained or reduced parameterizations that transform unconstrained trainable parameters via elementwise bounding functions and coefficient-mapping matrices to produce coefficient vectors for k-input gate functions, reducing storage and computation relative to enumerating discrete gates. Mixed-precision training may compute forward activations at lower precision while performing backpropagation at higher precision with selective storage or recomputation.
1 . A computing system, comprising:
at least one processor; and
a memory storing instructions that, when executed by the at least one processor, cause the computing system to perform training of a logic gate neural network, the logic gate neural network comprising a plurality of interconnected nodes, wherein at least one node comprises a continuous lookup table (LUT) configured to operate as a hyper-logic gate to:
receive a plurality of real-valued input signals, including data signals and configuration signals, and
produce a real-valued hyper-logic gate output by performing a continuous selection among the configuration signals as a function of the data signals.
2 . A computer-implemented method for training a neural network, comprising:
providing a neural network comprising a plurality of computational nodes, wherein at least one computational node comprises a hyper-logic gate, and wherein the hyper-logic gate, during training, is realized by a differentiable k-input lookup table, where k is an integer greater than or equal to 1 , the differentiable k-input lookup table having:
(i) k data inputs, each receiving a real value; and
(ii) N configuration inputs, where N is an integer greater than k, each configuration input receiving a real value;
computing an output of the at least one computational node as a differentiable function of the k data inputs and the N configuration inputs, wherein, when the k data inputs and the N configuration inputs are restricted to binary values, the output equals the configuration input indexed by the binary values of the k data inputs, wherein at least one of the N configuration inputs is provided by an output of one of the computational nodes in the neural network;
computing a loss from outputs of the neural network;
updating parameters of the neural network based on gradients or approximate gradients of the loss propagated through the differentiable k-input lookup table; and
discretizing the neural network, after training, such that the inputs to the hyper-logic gate are binary values, and the output of the hyper-logic gate equals the configuration input indexed by the binary values of the k data inputs.
3 . The method of claim 2 , wherein the number of configuration inputs, N, is 2{circumflex over ( )}k.
4 . The method of claim 3 , further comprising, after training, discretizing the neural network produces a fixed logic gate network, wherein the differentiable k-input lookup table is reduced to a Boolean k-input lookup table by virtue of all inputs thereto assuming binary values.
5 . The method of claim 4 , further comprising:
implementing a logical expression of the fixed logic gate network in a field-programmable gate array (FPGA) for subsequent use with other inputs.
6 . The method of claim 4 , further comprising:
implementing a logical expression of the fixed logic gate network as an application-specific integrated circuit (ASIC) for subsequent use with other inputs.
7 . The method of claim 4 , wherein the fixed logic gate network is represented in a hardware description language.
8 . The method of claim 3 , further comprising, after training, discretizing the neural network to produce a fixed logic gate network, wherein the differentiable k-input lookup tables is mapped to a multiplexer gate by virtue of all inputs thereto assuming binary values.
9 . The method of claim 3 , wherein, after discretizing, at least one hyper-logic gate operates as a multiplexer circuit wherein the k data inputs select which of the 2{circumflex over ( )}k configuration inputs is passed to the output.
10 . The method of claim 3 , wherein computing the output comprises: for each of the k data inputs, forming a complementary pair comprising the data input value and one minus the data input value; forming 2{circumflex over ( )}k terms, each term computed by applying a differentiable approximation of logical conjunction to a respective configuration input and one element selected from each of the k complementary pairs; and aggregating the 2{circumflex over ( )}k terms using a differentiable approximation of logical disjunction.
11 . The method of claim 2 , wherein the hyper-logic gate comprises an aggregation comprising at least one of: summation, a differentiable disjunction operator, and a t-conorm.
12 . The method of claim 2 , wherein, when the k data inputs are restricted to binary values, the output equals the configuration input indexed by the binary values of the k data inputs.
13 . The method of claim 3 , wherein at least one of the 2{circumflex over ( )}k configuration inputs or k data inputs is provided by an internal memory state stored in a register or flip-flop.
14 . The method of claim 13 , wherein all 2{circumflex over ( )}k configuration inputs to at least one hyper-logic gate are provided by internal memory values stored in registers or flip-flops.
15 . The method of claim 13 , wherein all k data inputs to at least one hyper-logic gate are provided by internal memory values stored in registers or flip-flops.
16 . The method of claim 3 , wherein k is greater than or equal to 2.
17 . The method of claim 3 , wherein at least ten computational nodes each comprise a hyper-logic gate, the method further comprising:
generating, after training, a fixed logic gate network from the neural network, comprising:
performing logic synthesis, wherein, during the logic synthesis, at least five computational nodes that comprised a hyper-logic gate are omitted or simplified.
18 . The method of claim 3 , further comprising, after training:
discretizing the neural network produces a fixed logic gate network, wherein the differentiable k-input lookup table is reduced to a Boolean k-input lookup table by virtue of all inputs thereto assuming binary values,
wherein, when the k data inputs are restricted to binary values, the output equals the configuration input indexed by the binary values of the k data inputs,
wherein k is greater than or equal to 2,
wherein at least ten computational nodes each comprise a hyper-logic gate, the method further comprising:
generating, after training, a fixed logic gate network from the neural network, including performing logic synthesis, during which at least five computational nodes that comprised a hyper-logic gate are omitted or simplified.
19 . The method of claim 3 , wherein k data inputs of a first hyper-logic gate are provided by outputs of other computational nodes in the neural network, and 2{circumflex over ( )}k configuration inputs of the first hyper-logic gate are provided by outputs of other computational nodes in the neural network, such that all k+2{circumflex over ( )}k inputs to the first hyper-logic gate are dynamic signals.
20 . The method of claim 3 , wherein the neural network comprises at least ten successive layers of computational nodes.
21 . The method of claim 3 , wherein the neural network further comprises an internal memory state comprising a plurality of memory locations, wherein the hyper-logic gate is configured to access a selected memory location as a function of the k data inputs and to produce the hyper-logic gate output based at least in part on content stored in the selected memory location.
22 . The method of claim 3 , wherein the hyper-logic gate is configured to implement a skip connection by receiving, as at least one of the configuration inputs, an output of a node that is not in an immediately preceding layer of the neural network, and the hyper-logic gate selecting or combining the configuration inputs and data inputs such that the hyper-logic gate output provides information from the node that is not in the immediately preceding layer.
23 . The method of claim 3 , wherein the hyper-logic gate is configured to aggregate a larger number of outputs of one or more previous layers into fewer elements by receiving, as at least some of each the data inputs and configuration inputs, respective computational node outputs generated in the one or more previous layers and producing the hyper-logic gate output for use by a subsequent layer.
24 . The method of claim 3 , wherein a plurality of the computational nodes comprises hyper-logic gates arranged in a tree structure, wherein a hyper-logic gate output of at least one hyper-logic gate in the tree is provided as at least one of the configuration inputs to a distinct hyper-logic gate in the tree.
25 . The method of claim 3 , wherein the neural network is a differentiable logic gate network.
26 . The method of claim 25 , the neural network further comprising at least one computational node comprising a trainable logic node parameterized by a set of differentiable parameters corresponding to a predefined set of potential logic gate operators,
wherein forward-propagating through the trainable logic node comprises computing a differentiable output that is a function of the outputs of the potential logic gate operators of the respective node according to current differentiable parameters thereof, and
wherein the discretizing further comprises, for each of at least a subset of the trainable nodes, selecting a single logic gate from the set of potential logic gate operators for a plurality of the nodes based on the differentiable parameters thereof.
27 . The method of claim 26 , wherein the predefined set of potential logic gate operators includes at least two elements, including one or more of: an AND operator, an OR operator, a NAND operator, a NOR operator, an XOR operator, a constant TRUE operator, a constant FALSE operator, an inverter operator, and a pass-through operator that outputs one of the node inputs.
28 . The method of claim 3 , wherein initializing the neural network comprises residual initialization, wherein the parameters of the neural network are initialized such that, upon discretization of an initialized model, a majority of computational nodes passes through exactly one input signal unchanged or inverted.
29 . The method of claim 3 , wherein the neural network comprises at least one million computational nodes.
30 . An application-specific integrated circuit (ASIC), comprising:
logic circuitry implementing a fixed logic gate network,
wherein the fixed logic gate network is generated by a process comprising:
providing a neural network comprising a plurality of computational nodes, wherein at least one computational node comprises a hyper-logic gate, and wherein the hyper-logic gate, during training, is realized by a differentiable k-input lookup table, k≥1, the differentiable k-input lookup table having:
(i) k data inputs, each receiving a real value; and
(ii) N configuration inputs, where N is an integer greater than k, each configuration input receiving a real value;
computing an output of the at least one computational node as a differentiable function of the k data inputs and the N configuration inputs, wherein, when the k data inputs and the N configuration inputs are restricted to binary values, the output equals the configuration input indexed by the binary values of the k data inputs, and wherein at least one of the N configuration inputs is provided by an output of one of the plurality of computational nodes in the neural network;
computing a loss from outputs of the neural network;
updating parameters of the neural network based on gradients or approximate gradients of the loss propagated through the differentiable k-input lookup table; and
discretizing the neural network, after training, to produce the fixed logic gate network such that inputs to the hyper-logic gate are binary values and the output of the hyper-logic gate equals the configuration input indexed by the binary values of the k data inputs.