IP Library › Patent Application 19562764
Patent Application
App. No. 19/562,764

Student-Teacher Distillation for Training and Deploying Differentiable Logic Gate Networks

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/562,764
Abstract

Systems and methods are described for training a logic gate network using student-teacher learning. In various embodiments, a system obtains a teacher model that produces a target output distribution for an input and records a training dataset of inputs and teacher outputs. The system instantiates a differentiable logic-gate network with nodes parameterized over candidate logic gate operators. The system trains the differentiable logic gate network by minimizing the divergence between its outputs and the teacher outputs. After training, a logic gate operator may be selected for each node. The logic gate network may be implemented as logic circuitry in a programmable device or as a fixed-in-silicon application-specific integrated circuit.

Claims (77)

1 . A computing system, comprising:

at least one processor; and

a memory coupled to the at least one processor and storing instructions that, when executed by the at least one processor, cause the computing system to:

receive a training data set comprising inputs;

obtain a teacher model configured to generate, for each input, corresponding teacher output values comprising teacher predictions;

instantiate, in the memory, an untrained student logic gate network with a plurality of nodes, wherein each node is parameterized by a set of differentiable parameters associated with a predefined set of potential logic gate operators;

iteratively train the student logic gate network via a plurality of training iterations, each training iteration including:

determining, using the teacher model, teacher output values for a batch of the inputs;

forward-propagating the batch of the inputs through the student logic gate network to generate a training network output by, for each node, computing an output activation that is a function of input activations of the respective node according to current differentiable parameters thereof;

computing a loss value that quantifies a difference between the training network output and the teacher output values for the batch;

determining, via a gradient-based training optimization algorithm, updated differentiable parameters for at least one node; and

applying the updated differentiable parameters to the at least one node; and

generating, after completion of the plurality of training iterations, a fixed student logic gate network by selecting a single logic gate operator from the set of potential logic gate operators for a plurality of the nodes based on the differentiable parameters thereof.

2 . The system of claim 1 , wherein the student logic gate network comprises a plurality of nodes arranged in a plurality of layers, including at least a first layer and a second layer, and wherein forward-propagating a batch of the inputs through the student logic gate network comprises:

for each node of the first layer, computing a real-valued non-binary relaxed differentiable output that is a real-valued non-binarizing relaxed differentiable function of the real-valued non-binary relaxed differentiable outputs of the potential logic gate operators of the respective node according to a non-linear real-valued non-binarizing relaxed differentiable function of current differentiable parameters thereof; and

for each node of the second layer, computing a real-valued non-binary relaxed differentiable output that is a real-valued non-binarizing relaxed differentiable function of the real-valued non-binary relaxed differentiable outputs of the potential logic gate operators of the respective node, at least some of which are real-valued non-binarizing relaxed differentiable functions of one or more inputs, where each input is a real-valued non-binary relaxed differentiable output of the first layer, according to a non-linear real-valued non-binarizing relaxed differentiable function of current differentiable parameters thereof.

3 . The system of claim 1 , wherein the student logic gate network comprises a plurality of nodes arranged in a plurality of layers, including at least a first layer and a second layer, and wherein forward-propagating a batch of the inputs through the student logic gate network, comprises:

for each node of the first layer, computing a real-valued non-binary differentiable output that is a real-valued non-binarizing non-linear differentiable function of:

(i) the input activations to the node, and

(ii) current differentiable parameters of the respective node; and

for each node of the second layer, computing a real-valued non-binary differentiable output that is a real-valued non-binarizing non-linear differentiable function of the:

(i) the input activations to the node, at least some of which are the real-valued non-binary differentiable outputs of the first layer, and

(ii) current differentiable parameters of the respective node.

4 . The system of claim 1 , wherein generating the fixed student logic gate network comprises, for each respective node of the plurality of nodes, selecting a respective logic gate operator having a highest probability in the operator-selection probability distribution for the respective node.

5 . The system of claim 1 , wherein the teacher model comprises a conventionally trained artificial neural network, and wherein the student logic gate network comprises a differentiable logic gate network.

6 . The system of claim Error! Reference source not found., wherein at least one processor is configured to, after training the teacher model, execute the teacher model in an environment to generate (i) a set of input observations comprising at least a portion of the inputs and (ii) corresponding teacher output values, and utilize the set of input observations and the corresponding teacher output values as a distillation data set.

7 . The system of claim 6 , wherein at least one processor is configured to generate the distillation data set using at least one of: (i) teacher output values recorded from a beginning of training of the teacher model, (ii) teacher output values recorded starting at a later time during training of the teacher model, and (iii) a mixture of teacher output values recorded at different times during training of the teacher model.

8 . The system of claim 1 , wherein computing the loss value comprises computing the means squared error or an L p norm between the teacher output values and the training network output.

9 . The system of claim 1 , wherein the teacher output values define a teacher probability distribution and the training network output defines a student probability distribution, and wherein computing the loss value comprises computing a Kullback-Leibler divergence between the teacher probability distribution and the student probability distribution.

10 . The system of claim 9 , wherein the teacher probability distribution and the student probability distribution each comprise an action-probability distribution over a discrete action space for reinforcement learning.

11 . The system of claim 9 , wherein the teacher output values comprise logits, and wherein computing the loss value comprises scaling the logits by a post-hoc factor prior to computing the Kullback-Leibler divergence, the post-hoc factor being selected to make teacher output values less certain or more certain.

12 . The system of claim 9 , wherein the student probability distribution comprises a class-probability distribution for a classification task, wherein the student logic gate network generates class scores by aggregating outputs of a plurality of output nodes corresponding to each class, and wherein the class-probability distribution is computed from the class scores.

13 . The system of claim 12 , wherein computing the class-probability distribution comprises applying temperature scaling to logits derived from the class scores.

14 . The system of claim 1 , further comprising a secondary teacher model, wherein the one or more processors are configured to train the secondary teacher model based on residuals between the teacher output values produced by the teacher model and corresponding outputs of the student logic gate network, and to train the student logic gate network based at least in part on outputs of the secondary teacher model.

15 . The system of claim 1 , wherein the one or more processors are configured to, during at least a portion of the plurality of training iterations, utilize a current version of the student logic gate network as an agent in a reinforcement learning setting to obtain at least some of the inputs for which the teacher model generates the teacher output values.

16 . The system of claim 1 , further comprising an encoder configured to encode real-valued input features into a binary representation for input to the student logic gate network using thermometer encoding with thresholds based on quantiles of a data distribution.

17 . The system of claim 1 , wherein generating a fixed logic gate network comprises a logic synthesis process that reduces the number of logic gate operators by more than 65%.

18 . The method of claim 17 , wherein the logic synthesis process comprises one or more of (i) constant propagation, (ii) wire removal or collapse, and (iii) removal of inverters by collapse into one or more connected nodes.

19 . The system of claim 1 , wherein the predefined set of potential logic gate operators includes entries of a lookup table.

20 . The system of claim 1 , wherein the predefined set of potential logic gate operators includes at least two elements, including one or more of: an AND operator, an OR operator, a NAND operator, a NOR operator, an XOR operator, a constant TRUE operator, a constant FALSE operator, an inverter operator, and a pass-through operator that outputs one of the node inputs.

21 . A method for training a differentiable logic gate network, comprising:

receiving, at a computing system, a training data set comprising input vectors for an inference task;

obtaining a teacher model configured to generate, for each input vector, corresponding teacher output values comprising teacher predictions for the inference task;

instantiating, in a memory of the computing system, an untrained student logic gate network with a plurality of nodes, wherein each node is parameterized by a set of differentiable parameters corresponding to a predefined set of potential logic gate operators;

iteratively training the student logic gate network via a plurality of training iterations, each training iteration including:

determining teacher output values for a batch of the input vectors by evaluating the teacher model using the batch of the input vectors;

forward-propagating the batch of the input vectors through the student logic gate network to generate a training network output by, for each node, computing a differentiable output that is a function of the outputs of the potential logic gate operators of the respective node according to current differentiable parameters thereof;

computing a loss value that quantifies a difference between the training network output and the corresponding teacher output values for the batch;

determining, via a training optimization algorithm, updated differentiable parameters for at least one node; and

applying the updated differentiable parameters to the at least one node; and

generating, after completion of the plurality of training iterations, a fixed student logic gate network by identifying a single logic gate from the set of potential logic gate operators for a plurality of the nodes based on the differentiable parameters thereof.

22 . A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations comprising:

receiving a training data set comprising input vectors for an inference task;

obtaining a teacher model configured to generate, for each input vector, corresponding teacher output values comprising teacher predictions for the inference task;

instantiating an untrained student logic gate network with a plurality of nodes, wherein each node is parameterized by a set of differentiable parameters corresponding to a predefined set of potential logic gate operators;

iteratively training the student logic gate network via a plurality of training iterations, each training iteration including:

determining teacher output values for a batch of the input vectors by evaluating the teacher model using the batch of the input vectors;

forward-propagating the batch of the input vectors through the student logic gate network to generate a training network output by, for each node, computing a differentiable output that is a function of outputs of the potential logic gate operators of the respective node according to current differentiable parameters thereof;

computing a loss value that quantifies a difference between the training network output and the corresponding teacher output values for the batch;

determining, via a training optimization algorithm, updated differentiable parameters for at least one node; and

applying the updated differentiable parameters to the at least one node; and

generating, after completion of the plurality of training iterations, a fixed student logic gate network by identifying a single logic gate from the set of potential logic gate operators for a plurality of the nodes based on the differentiable parameters thereof.

23 . An application-specific integrated circuit (ASIC) with logic circuitry that is an implementation of a fixed logic gate network, manufactured using a process comprising:

receiving, at a computing system, a training data set comprising input vectors for an inference task;

obtaining a teacher model configured to generate, for each input vector, corresponding teacher output values comprising teacher predictions for the inference task;

instantiating, in a memory of the computing system, an untrained student logic gate network with a plurality of nodes, wherein each node is parameterized by a set of differentiable parameters corresponding to a predefined set of potential logic gate operators;

iteratively training the student logic gate network via a plurality of training iterations, each training iteration including:

determining teacher output values for a batch of the input vectors by evaluating the teacher model using the batch of the input vectors;

forward-propagating the batch of the input vectors through the student logic gate network to generate a training network output by, for each node, computing a differentiable output that is a function of the outputs of the potential logic gate operators of the respective node according to current differentiable parameters thereof;

computing a loss value that quantifies a difference between the training network output and the corresponding teacher output values for the batch;

determining, via a training optimization algorithm, updated differentiable parameters for at least one node; and

applying the updated differentiable parameters to the at least one node; and

generating, after completion of the plurality of training iterations, a fixed logic gate network by identifying a single logic gate from the set of potential logic gate operators for a plurality of the nodes based on the differentiable parameters thereof.

24 . The ASIC of claim 23 , wherein the process further comprises one or more of:

synthesizing the fixed logic gate network;

technology-mapping the fixed logic gate network; and

place-and-routing operations for circuit components.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2026
From: PETERSEN, FELIX
To: DIFFLOGIC, INC.
Reel/Frame 075107/0563 →