IP Library › Patent Application 19562760
Patent Application
App. No. 19/562,760

Differentiable Logic Gate Networks with Logic Gate Tree Networks

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/562,760
Abstract

Systems and methods are described for training and using logic gate tree networks and convolutional logic gate tree networks. A computing system may apply a convolution operation to an input tensor using one or more kernels, each kernel comprising a tree of logic gate nodes. For each kernel placement, leaves of the tree select input activations from a receptive field, and internal nodes generate a kernel output by applying logic operations. During training, each node may be parameterized with differentiable parameters that define a probability distribution over candidate logic operators. The parameters may be shared across kernel placements to provide equivariance. Gradient-based optimization may be used to update the differentiable parameters during training. A fixed tree kernel may be defined by selecting a logic operator for each node, which can be executed by a processor or synthesized for use in programmable logic and/or in an application-specific integrated circuit.

Claims (62)

1 . A method for training a logic gate tree network, comprising:

receiving, at a computing system, a training data set comprising inputs;

instantiating, in a memory of the computing system, an untrained logic gate tree network comprising at least one layer having at least one logic gate tree, wherein each logic gate tree comprises a plurality of nodes arranged in a tree topology in which outputs of lower nodes provide inputs to higher nodes, and wherein each node is parameterized by a set of differentiable parameters corresponding to a predefined set of potential logic gate operators;

iteratively training the untrained logic gate tree network via a plurality of training iterations, each training iteration including:

forward-propagating a batch of the inputs through the untrained logic gate tree network, wherein forward-propagating includes, for at least one logic gate tree, selecting a plurality of leaf input activations from the batch of inputs and/or from intermediate activations generated by the logic gate tree network, and forward-propagating the plurality of leaf input activations through the logic gate tree to produce an output activation;

computing a loss value;

determining, via a training optimization algorithm, updated differentiable parameters for at least one node; and

applying the updated differentiable parameters to the at least one node; and

generating, after completion of the plurality of training iterations, a fixed logic gate tree network by selecting, for each of at least some of the plurality of nodes, a single logic gate operator from the predefined set of potential logic gate operators for the respective node based on the differentiable parameters thereof.

2 . The method of claim 1 , wherein the leaf input activations are selected from input activations comprising an input activation tensor comprising activations arranged across a plurality of input channels and a plurality of positions in a domain, and wherein one or more output activations of a layer comprise an output activation tensor.

3 . The method of claim 2 , wherein forward-propagating includes generating the output activation tensor by, for each of a plurality of output channels, convolving a respective input activation tensor with a respective logic gate tree across a plurality of kernel placements in the domain.

4 . The method of claim 3 , wherein, for each given logic gate tree, the differentiable parameters of nodes of the given logic gate tree are shared across the plurality of kernel placements for the given logic gate tree.

5 . The method of claim 3 , wherein, for a given kernel placement, convolving comprises selecting, from a receptive field of the input activation tensor corresponding to the given kernel placement, the plurality of leaf input activations for the respective logic gate tree.

6 . The method of claim 5 , wherein selecting the plurality of leaf input activations comprises selecting, for each leaf input activation, an input-channel identifier and a position offset within the receptive field.

7 . The method of claim 3 , wherein, for a given logic gate tree, the plurality of leaf input activations are selected from no more than two distinct input channels.

8 . The method of claim 1 , wherein at least one logic gate tree comprises a binary tree having a depth, d, greater than or equal to two, and wherein the plurality of leaf input activations comprises 2{circumflex over ( )}d leaf input activations for use as inputs to the binary tree.

9 . The method of claim 8 , wherein the depth, d, is three.

10 . The method of claim 1 , wherein at least one logic gate tree comprises a k-ary tree of k-input gates, the k-ary tree having a depth, d, greater than or equal to two, and wherein the plurality of leaf input activations comprises k{circumflex over ( )}d leaf input activations for use as inputs to the k-ary tree, wherein k is an integer greater than or equal to 2.

11 . The method of claim 10 , wherein the depth, d, is three.

12 . The method of claim 6 , wherein the input-channel identifiers and the position offsets are defined by one or more connection-index arrays stored in the memory, and wherein the one or more connection-index arrays are reused across the plurality of kernel placements.

13 . The method of claim 12 , wherein the one or more connection-index arrays are generated pseudo-randomly when instantiating a layer comprising the respective logic gate tree, and thereafter remain fixed during generating the output activation tensor.

14 . The method of claim 1 , wherein each node of at least one logic gate tree is configured to receive exactly two input activations and to output a single output activation.

15 . The method of claim 1 , wherein the input activations are partitioned into a plurality of channel groups, and wherein the plurality of leaf input activations for a given logic gate tree are selected from within a respective one of the channel groups.

16 . The method of claim 1 , wherein the predefined set of potential logic gate operators includes entries of a lookup table.

17 . The method of claim 1 , wherein the predefined set of potential logic gate operators includes at least two elements, including one or more of: an AND operator, an OR operator, a NAND operator, a NOR operator, an XOR operator, a constant TRUE operator, a constant FALSE operator, an inverter operator, and a pass-through operator that outputs one of the node inputs.

18 . The method of claim 1 , wherein, for each node, the computation w 1 +w 2 ·a+w 3 ·b+w 4 ·a·b is performed during a forward propagation, wherein a set of coefficients w 1 , w 2 , w 3 , w 4 that are functions of the differentiable parameters and two inputs a, b to the node.

19 . The method of claim 1 , wherein, for each node, the differentiable parameters define a categorical probability distribution over the predefined set of potential logic gate operators, and wherein forward-propagating the plurality of leaf input activations through the logic gate tree comprises, for each node, computing a node output as a weighted combination of respective outputs of the potential logic gate operators according to the categorical probability distribution, the categorical probability distribution being derived from the differentiable parameters.

20 . The method of claim 1 , wherein generating the fixed logic gate tree network comprises a logic synthesis process that reduces the number of logic gate operators by more than 65%.

21 . The method of claim 20 , wherein the logic synthesis process comprises one or more of (i) constant propagation, (ii) wire removal or collapse, and (iii) removal of inverters by collapse into one or more connected nodes.

22 . The method of claim 1 , wherein at least one logic gate tree comprises a plurality of nodes arranged in a plurality of tree layers, including at least a first tree layer and a second tree layer, and wherein forward-propagating a batch of the inputs through the logic gate tree network comprises:

for each node of the first tree layer, computing a real-valued non-binary relaxed differentiable output that is a real-valued non-binarizing relaxed differentiable function of the real-valued non-binary relaxed differentiable outputs of the potential logic gate operators of the respective node according to a non-linear real-valued non-binarizing relaxed differentiable function of current differentiable parameters thereof; and

for each node of the second tree layer, computing a real-valued non-binary relaxed differentiable output that is a real-valued non-binarizing relaxed differentiable function of the real-valued non-binary relaxed differentiable outputs of the potential logic gate operators of the respective node, at least some of which are real-valued non-binarizing relaxed differentiable functions of one or more inputs, where each input is a real-valued non-binary relaxed differentiable output of the first layer, according to a non-linear real-valued non-binarizing relaxed differentiable function of current differentiable parameters thereof.

23 . The method of claim 1 , wherein at least one logic gate tree comprises a plurality of nodes arranged in a plurality of tree layers, including at least a first tree layer and a second tree layer, and wherein forward-propagating a batch of the inputs through the logic gate tree network comprises:

for each node of the first tree layer, computing a real-valued non-binary differentiable output that is a real-valued non-binarizing non-linear differentiable function of:

(i) input activations to the node, and

(ii) current differentiable parameters of the respective node; and

for each node of the second tree layer, computing a real-valued non-binary differentiable output that is a real-valued non-binarizing non-linear differentiable function of the:

(i) input activations to the node, at least some of which are the real-valued non-binary differentiable outputs of the first tree layer, and

(ii) current differentiable parameters of the respective node.

24 . A system, comprising:

at least one processor; and

a memory coupled to the at least one processor and storing instructions that, when executed by the at least one processor, cause the system to:

receive a training data set comprising inputs;

instantiate, in the memory, an untrained logic gate tree network comprising at least one layer having at least one logic gate tree, wherein each logic gate tree comprises a plurality of nodes arranged in a tree topology in which outputs of lower nodes provide inputs to higher nodes, and wherein each node is parameterized by a set of differentiable parameters corresponding to a predefined set of potential logic gate operators;

iteratively train the untrained logic gate tree network via a plurality of training iterations, each training iteration including:

forward-propagating a batch of the inputs through the untrained logic gate tree network, wherein forward-propagating includes, for at least one logic gate tree, selecting a plurality of leaf input activations from at least one of (i) the batch of inputs and (ii) intermediate activations generated by the logic gate tree network; and

forward-propagating the plurality of leaf input activations through the logic gate tree to produce an output activation;

computing a loss value;

determining, via a training optimization algorithm, updated differentiable parameters for at least one node; and

applying the updated differentiable parameters to the at least one node; and

generate, after completion of the plurality of training iterations, a fixed logic gate tree network by selecting, for each of at least some of the plurality of nodes, a single logic gate operator from the predefined set of potential logic gate operators for the respective node based on the differentiable parameters thereof.

25 . A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations comprising:

receiving a training data set comprising inputs;

instantiating, in a memory of the computing system, an untrained logic gate tree network comprising at least one layer having at least one logic gate tree, wherein each logic gate tree comprises a plurality of nodes arranged in a tree topology in which outputs of lower nodes provide inputs to higher nodes, and wherein each node is parameterized by a set of differentiable parameters corresponding to a predefined set of potential logic gate operators;

iteratively training the untrained logic gate tree network via a plurality of training iterations, each training iteration including:

forward-propagating a batch of the inputs through the untrained logic gate tree network, wherein forward-propagating includes, for at least one logic gate tree, selecting a plurality of leaf input activations from at least one of (i) the batch of inputs and (ii) intermediate activations generated by the logic gate tree network; and

forward-propagating the plurality of leaf input activations through the logic gate tree to produce an output activation;

computing a loss value;

determining, via a training optimization algorithm, updated differentiable parameters for at least one node; and

applying the updated differentiable parameters to the at least one node; and

generating, after completion of the plurality of training iterations, a fixed logic gate tree network by selecting, for each of at least some of the plurality of nodes, a single logic gate operator from the predefined set of potential logic gate operators for the respective node based on the differentiable parameters thereof.

26 . The non-transitory computer-readable medium of claim 25 , wherein forward-propagating includes generating an output tensor by, for each of a plurality of output channels, convolving an input tensor comprising the leaf input activations with a respective logic gate tree across a plurality of kernel placements in a domain, and wherein differentiable parameters of nodes of a given logic gate tree kernel are shared across the plurality of kernel placements.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2026
From: PETERSEN, FELIX
To: DIFFLOGIC, INC.
Reel/Frame 075107/0563 →