Neural network training technique
Apparatuses, systems, and techniques to train one or more neural networks. In at least one embodiment, one or more neural networks are trained based, at least in part, on inferencing output from one or more second neural networks.
1 . One or more processors, comprising:
circuitry to use one or more first neural networks to identify one or more objects in one or more synthetic images based, at least in part, on:
identifying a plurality of layers of the one or more first neural networks based on one or more inputs to the plurality of layers;
a scaling factor, computed based on calculating a loss corresponding to an inferencing response of one or more second neural networks in response to an input of one or more second synthetic images, wherein the one or more second neural networks are trained on non-synthetic images; and
one or more hyperparameters generated using the one or more second neural networks and adjusted for the one or more first neural networks using the scaling factor in a policy corresponding to the plurality of layers of the one or more first neural networks, wherein the one or more first neural networks are trained using the one or more hyperparameters while one or more parameters of the one or more second neural networks are frozen.
2 . The one or more processors of claim 1 , wherein during training, a learning rate of the one or more first neural networks is adjusted according to the scaling factor in the policy associated with the plurality of layers of the one or more first neural networks, based, at least in part, on an inferencing output of the one or more second neural networks.
3 . The one or more processors of claim 2 , wherein the learning rate is adjusted according to the scaling factor in the policy for a region of the one or more first neural networks, the region comprising the plurality of layers of the one or more first neural networks grouped based, at least in part, on input resolution.
4 . The one or more processors of claim 1 , wherein the circuitry adjusts training of the one or more first neural networks based, at least in part, on an inferencing output of the one or more second neural networks in response to the input of the one or more second synthetic images.
5 . The one or more processors of claim 4 , wherein the circuitry is to compute the scaling factor and adjust a learning rate of the one or more first neural networks based, at least in part, on the scaling factor associated with the plurality of layers of the one or more first neural networks.
6 . The one or more processors of claim 1 , wherein the one or more second neural networks are trained to perform an image processing task equivalent to an image processing task performed by the one or more first neural networks.
7 . The one or more processors of claim 1 , wherein the one or more hyperparameters of the one or more second neural networks are not adjusted based on the input of the one or more second synthetic images.
8 . The one or more processors of claim 1 , wherein the one or more first neural networks and the one or more second neural networks have equivalent structures.
9 . A system, comprising:
one or more processors to use one or more first neural networks to identify one or more objects in one or more synthetic images based, at least in part, on:
identifying a plurality of layers of the one or more first neural networks based on one or more inputs to the plurality of layers;
a scaling factor, computed based on calculating a loss corresponding to an inferencing response of one or more second neural networks in response to an input of one or more second synthetic images, wherein the one or more second neural networks are trained on non-synthetic images; and
one or more hyperparameters, generated using the one or more second neural networks and adjusted for the one or more first neural networks using the scaling factor in a policy corresponding to the plurality of layers of the one or more first neural networks, wherein the one or more first neural networks are trained using the one or more hyperparameters while one or more parameters of the one or more second neural networks are frozen.
10 . The system of claim 9 , wherein during training, a learning rate of the one or more first neural networks is adjusted based, at least in part, on an inferencing output of the one or more second neural networks, and wherein the learning rate is further adjusted according to the scaling factor in the policy associated with the plurality of layers of the one or more first neural networks.
11 . The system of claim 10 , wherein the one or more processors are to adjust the learning rate according to the scaling factor in the policy for a region of the one or more first neural networks.
12 . The system of claim 11 , wherein the region comprises the plurality of layers of the one or more first neural networks grouped according to input resolution.
13 . The system of claim 9 , wherein the one or more processors are to adjust training of the one or more first neural networks based, at least in part, on an inferencing output of the one or more second neural networks in response to the input of the one or more second synthetic images.
14 . The system of claim 9 , wherein the one or more processors are to compute a learning rate scaling factor based, at least in part, on a divergence factor, the divergence factor computed, at least in part, based on an inferencing output of the one or more second neural networks.
15 . The system of claim 9 , wherein the one or more second neural networks are frozen during training of the one or more first neural networks.
16 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:
use one or more first neural networks to identify one or more objects in one or more synthetic images based, at least in part, on:
identifying a plurality of layers of the one or more first neural networks based on one or more inputs to the plurality of layers;
a scaling factor, computed based on calculating a loss corresponding to an inferencing response of one or more second neural networks in response to an input of one or more second synthetic images, wherein the one or more second neural networks are trained on non-synthetic images; and
one or more hyperparameters generated using the one or more second neural networks and adjusted for the one or more first neural networks using the scaling factor in a policy corresponding to the plurality of layers of the one or more first neural networks, wherein the one or more first neural networks are trained using the one or more hyperparameters while one or more parameters of the one or more second neural networks are frozen.
17 . The non-transitory machine-readable medium of claim 16 , having stored thereon the set of instructions which if performed by one or more processors, cause the one or more processors to at least:
compute an adjustment to a learning rate of the one or more first neural networks, based, at least in part, on an inferencing output of the one or more second neural networks, wherein the learning rate is adjusted according to the scaling factor in the policy associated with the plurality of layers of the one or more first neural networks.
18 . The non-transitory machine-readable medium of claim 17 , wherein the adjustment to the learning rate is calculated for a region of the one or more first neural networks, the region comprising the plurality of layers of the one or more first neural networks grouped based, at least in part, on input resolution.
19 . The non-transitory machine-readable medium of claim 18 , wherein the adjustment to the learning rate is calculated based, at least in part, by a long short-term memory (“LSTM”) module.
20 . The non-transitory machine-readable medium of claim 16 , wherein an inferencing output of the one or more second neural networks is based, at least in part, on an input comprising the one or more second synthetic images.
21 . The non-transitory machine-readable medium of claim 16 , wherein the one or more second neural networks are trained to perform an image processing task equivalent to an image processing task performed by the one or more first neural networks.
22 . The non-transitory machine-readable medium of claim 16 , wherein the one or more hyperparameters of the one or more second neural networks are not adjusted based on the input derived from the one or more second synthetic images.
23 . The non-transitory machine-readable medium of claim 16 , wherein the one or more first neural networks, once trained, are usable to perform an inferencing task independently of the one or more second neural networks.
24 . A computing device, comprising:
one or more processors to perform an image processing task based, at least in part, on one or more first neural networks to identify one or more objects in one or more synthetic images based, at least in part, on:
identifying a plurality of layers of the one or more first neural networks based on one or more inputs to the plurality of layers;
a scaling factor, computed based on calculating a loss corresponding to an inferencing response of one or more second neural networks in response to an input of one or more second synthetic images, wherein the one or more second neural networks are trained on non-synthetic images; and
one or more hyperparameters generated using the one or more second neural networks and adjusted for the one or more first neural networks using the scaling factor in a policy corresponding to the plurality of layers of the one or more first neural networks, wherein the one or more first neural networks are trained using the one or more hyperparameters while one or more parameters of the one or more second neural networks are frozen.
25 . The computing device of claim 24 , wherein the image processing task comprises at least one of recognition or classification.
26 . The computing device of claim 24 , the one or more processors to compute an adjustment to a learning rate of the one or more first neural networks, based, at least in part, on an inferencing output of the one or more second neural networks, wherein the learning rate is adjusted according to the scaling factor in the policy associated with the plurality of layers of the one or more first neural networks.
27 . The computing device of claim 26 , wherein the adjustment to the learning rate is calculated for a region of the one or more first neural networks, the region comprising the plurality of layers of the one or more first neural networks grouped based, at least in part, on input resolution; and wherein the adjustment to the learning rate is according to the scaling factor in the policy associated with the plurality of layers of the one or more first neural networks.
28 . The computing device of claim 24 , wherein an inferencing output of the one or more second neural networks is based, at least in part, on an input comprising the more second synthetic images.
29 . The computing device of claim 24 , wherein the one or more first neural networks and the one or more second neural networks are each trained to perform the image processing task.
30 . The computing device of claim 29 , wherein the one or more second neural networks are trained only on non-synthetic images.