Partial neural network activation
Apparatuses, systems, and techniques to perform inferencing with neural networks. In at least one embodiment, portions of one or more neural networks are selected for use in inferencing based, at least in part, on information to be inferenced by the one or more neural networks.
1 . A processor, comprising:
one or more circuits to:
cause one or more layers of one or more neural networks to generate output values from performing a portion of inferencing based, at least in part, on input data;
cause one or more sub-networks to be selected from among a plurality of sub-networks of the one or more neural networks to complete the inferencing based, at least in part, on the input data or the output values from the one or more layers; and
provide the output values from the one or more layers to the selected one or more sub-networks, wherein the output values are duplicated to each selected sub-network such that the portion of inferencing is common to the plurality of sub-networks to complete the inferencing.
2 . The processor of claim 1 , wherein the one or more neural networks comprise two or more sub-models to be selected based, at least in part, on one or more router layers of the one or more neural networks.
3 . The processor of claim 2 , wherein the one or more router layers select one of the two or more sub-models based, at least in part, on one or more features detected in one or more inputs to the one or more neural networks.
4 . The processor of claim 2 , wherein a sub-model selected from the two or more sub-models is to be loaded into a memory by the processor.
5 . The processor of claim 2 , wherein the one or more neural networks comprise one or more stem layers to perform inferencing common to two or more of the two or more sub-models.
6 . The processor of claim 5 , wherein the one or more router layers select a sub-model of the two or more sub-models based, at least in part, on one or more outputs of the one or more stem layers.
7 . The processor of claim 2 , wherein the two or more sub-models are respectively trained to recognize text in different environments.
8 . A method, comprising:
performing a portion of inferencing by one or more layers of one or more neural networks to generate output values based, at least in part, on input data;
selecting one or more sub-networks from among a plurality of sub-networks of the one or more neural networks to complete the inferencing based, at least in part, on the input data or the output values from the one or more layers; and
providing the output values from the one or more layers to the selected one or more sub-networks, wherein the output values are duplicated to each selected sub-network such that the portion of inferencing is common to the plurality of sub-networks to complete the inferencing.
9 . The method of claim 8 , wherein the one or more neural networks comprise two or more sub-models to be selected based, at least in part, on one or more router layers of the one or more neural networks.
10 . The method of claim 9 , wherein the one or more router layers select one of the two or more sub-models based, at least in part, on one or more features detected in one or more inputs to the one or more neural networks.
11 . The method of claim 9 , further comprising using one or more processors to load a selected sub-model selected from the two or more sub-models to a memory.
12 . The method of claim 9 , wherein the one or more neural networks comprise one or more stem layers to perform inferencing common to two or more of the two or more sub-models.
13 . The method of claim 12 , wherein the one or more router layers select a sub-model of the two or more sub-models based, at least in part, on one or more outputs of the one or more stem layers.
14 . The method of claim 9 , wherein the two or more sub-models are respectively trained to recognize text in different environments.
15 . A system, comprising:
one or more processors to:
cause one or more layers of one or more neural networks to generate output values from performing a portion of inferencing based, at least in part, on input data;
cause one or more sub-networks to be selected from among a plurality of sub-networks of the one or more neural networks to complete the inferencing based, at least in part, on the input data or the output values from the one or more layers; and
provide the output values from the one or more layers to the selected one or more sub-networks, wherein the output values are duplicated to each selected sub-network such that the portion of inferencing is common to the plurality of sub-networks to complete the inferencing.
16 . The system of claim 15 , wherein the one or more neural networks comprise two or more sub-models to be selected based, at least in part, on one or more router layers of the one or more neural networks.
17 . The system of claim 16 , wherein the one or more router layers select one of the two or more sub-models based, at least in part, on one or more features detected in one or more inputs to the one or more neural networks.
18 . The system of claim 16 , wherein the one or more neural networks comprise one or more stem layers to perform inferencing common to two or more of the two or more sub-models.
19 . The system of claim 18 , wherein the one or more router layers select a sub-model of the two or more sub-models based, at least in part, on one or more outputs of the one or more stem layers.
20 . The system of claim 16 , wherein the two or more sub-models are respectively trained to recognize text in different environments.