Methods, systems, articles of manufacture and apparatus to train a neural network
Methods, systems, apparatus, and articles of manufacture are disclosed to train a neural network. An example apparatus includes an architecture evaluator to determine an architecture type of a neural network, a knowledge branch implementor to select a quantity of knowledge branches based on the architecture type, and a knowledge branch inserter to improve a training metric by appending the quantity of knowledge branches to respective layers of the neural network.
1 . An apparatus to train a neural network, the apparatus comprising:
interface circuitry;
machine-readable instructions; and
at least one processor circuit to be programmed by the machine-readable instructions to:
determine an architecture type of a neural network, the neural network including intermediate layers between a first layer and a last layer;
select a quantity of network classifiers based on the architecture type; and
improve a training metric of the neural network by:
determining a middle one of the intermediate layers; and
attaching a first one of the network classifiers to the middle one of the intermediate layers of the neural network; and
remove the first one of the network classifiers from a backbone of the neural network prior to deployment of the neural network
wherein a knowledge interaction between the first one of the network classifiers and a second one of the network classifiers includes a knowledge interaction matrix generated based on a binary indicator function, the binary indicator function representative of a set of layer indices of the neural network identifying one or more activation locations of the knowledge interaction.
2 . The apparatus as defined in claim 1 , wherein one or more of the at least one processor circuit is to calculate the quantity of network classifiers based on a quantity of layers associated with the neural network.
3 . The apparatus as defined in claim 2 , wherein one or more of the at least one processor circuit is to calculate the quantity of network classifiers based on dividing the quantity of layers associated with the neural network by a branch factor.
4 . The apparatus of claim 2 , wherein a current class probability output from the first one of the network classifiers represents a soft label used to align a probabilistic prediction output from a second one of the network classifiers with the current class probability output.
5 . The apparatus as defined in claim 1 , wherein one or more of the at least one processor circuit is to identify candidate insertion locations of the neural network.
6 . The apparatus as defined in claim 5 , wherein one or more of the at least one processor circuit is to insert one of the quantity of network classifiers at one of the candidate insertion locations.
7 . The apparatus as defined in claim 1 , wherein one or more of the at least one processor circuit is to attach a second one of the network classifiers to a second layer adjacent to the middle one of the intermediate layers.
8 . The apparatus of claim 1 , wherein a soft cross-entropy loss function defines a knowledge interaction between the first one of the network classifiers and a second one of the network classifiers.
9 . At least one non-transitory computer readable medium comprising instructions to cause at least one processor circuit to:
determine an architecture type of a neural network, the neural network including intermediate layers between a first layer and a last layer;
select a quantity of network classifiers based on the architecture type; and
improve a training metric of the neural network by:
determining a middle one of the intermediate layers; and
attaching a first one of the network classifiers to the middle one of the intermediate layers of the neural network; and
remove the first one of the network classifiers from a backbone of the neural network prior to deployment of the neural network
wherein a knowledge interaction between the first one of the network classifiers and a second one of the network classifiers includes a knowledge interaction matrix generated based on a binary indicator function, the binary indicator function representative of a set of layer indices of the neural network identifying one or more activation locations of the knowledge interaction.
10 . The at least one non-transitory computer readable medium as defined in claim 9 , wherein the computer readable instructions are to cause one or more of the at least one processor circuit to calculate the quantity of network classifiers based on a quantity of layers associated with the neural network.
11 . The at least one non-transitory computer readable medium as defined in claim 10 , wherein the computer readable instructions are to cause one or more of the at least one processor circuit to calculate the quantity of network classifiers based on dividing the quantity of layers associated with the neural network by a branch factor.
12 . The at least one non-transitory computer readable medium as defined in claim 9 , wherein the computer readable instructions are to cause one or more of the at least one processor circuit to identify candidate insertion locations of the neural network.
13 . The at least one non-transitory computer readable medium as defined in claim 12 , wherein the computer readable instructions are to cause the one or more of the at least one processor circuit to insert one of the quantity of network classifiers at one of the candidate insertion locations.
14 . The at least one non-transitory computer readable medium as defined in claim 9 , wherein the computer readable instructions are to cause one or more of the at least one processor circuit to attach a second one of the network classifiers to a second layer adjacent to the middle one of the intermediate layers.
15 . The at least one non-transitory computer readable medium as defined in claim 9 , wherein the computer readable instructions are to cause the one or more of the at least one processor circuit to select a knowledge interaction framework for the quantity of network classifiers.
16 . The at least one non-transitory computer readable medium as defined in claim 15 , wherein the computer readable instructions are to cause the one or more of the at least one processor circuit to implement the knowledge interaction framework as at least one of a top-down knowledge interaction framework, a bottom-up knowledge interaction framework, or a bi-directional knowledge interaction framework.
17 . The at least one non-transitory computer readable medium as defined in claim 15 , wherein the computer readable instructions are to cause the one or more of the at least one processor circuit to define an optimization goal, the optimization goal to include the selected knowledge interaction framework.
18 . A computer implemented method to train a neural network, the method comprising:
determining, by executing an instruction with at least one processor, an architecture type of a neural network, the neural network including intermediate layers between a first layer and a last layer;
selecting, by executing an instruction with one or more of the at least one processor, a quantity of network classifiers based on the architecture type; and
improving, by executing an instruction with one or more of the at least one processor, a training metric of the neural network by:
determining a middle one of the intermediate layers; and
attaching a first one of the network classifiers to the middle one of the intermediate layers of the neural network; and
removing, by executing an instruction with one or more of the at least one processor, the first one of the network classifiers from a backbone of the neural network prior to deployment of the neural network,
wherein a knowledge interaction between the first one of the network classifiers and a second one of the network classifiers includes a knowledge interaction matrix generated based on a binary indicator function, the binary indicator function representative of a set of layer indices of the neural network identifying one or more activation locations of the knowledge interaction.
19 . The method as defined in claim 18 , further including selecting a knowledge interaction framework for the quantity of network classifiers.
20 . The method as defined in claim 19 , further including applying at least one of a top-down knowledge interaction framework, a bottom-up knowledge interaction framework, or a bi-directional knowledge interaction framework.
21 . The method as defined in claim 19 , further including defining an optimization goal, the optimization goal to include the selected knowledge interaction framework.
22 . A system to train a neural network, the system comprising:
means for determining an architecture type of a neural network, the neural network including intermediate layers between a first layer and a last layer;
means for selecting a quantity of network classifiers based on the architecture type; and
means for improving a training metric of the neural network by:
determining a middle one of the intermediate layers; and
attaching a first one of the network classifiers to the middle one of the intermediate layers of the neural network; and
means for implementing a knowledge branch to remove the first one of the network classifiers from a backbone of the neural network prior to deployment of the neural network
wherein a knowledge interaction between the first one of the network classifiers and a second one of the network classifiers includes a knowledge interaction matrix generated based on a binary indicator function, the binary indicator function representative of a set of layer indices of the neural network identifying one or more activation locations of the knowledge interaction.
23 . The system as defined in claim 22 , further including means for calculating the quantity of network classifiers based on a quantity of layers associated with the neural network.
24 . The system as defined in claim 23 , wherein the means for calculating the quantity of network classifiers is based on dividing the quantity of layers associated with the neural network by a branch factor.