Neural modulation codes for multilingual and style dependent speech and language processing
Computer-implemented methods and apparatus that use neural modulation codes as an alternative to training many individual recognition models or to loosing performance by training mixed models. Large neural models are modulated by codes that model the different conditions. The codes directly alter (modulate) the behavior of connections in a multiconditional perceptual classifier, so as to permit the most appropriate neuronal units and their features to be applied to each condition. The approach may be applied to multilingual ASR, where the resulting multilingual network arrangement is able to achieve performance that is competitive or better than individually trained mono-lingual network. Moreover, the approach requires no adaptation data or extensive adaptation/training time to operate in a manner tuned to each condition. Beyond multilingual speech processing systems the approach can be applied to many other perceptual processing problems (e.g. speech recognition, speech synthesis, language translation, image processing) to factor the processing task from the conditioning variables that drive the actual realization. Instead of adapting or retraining neural systems to individual conditions, it modulates a large invariant network to operate in different modes based on conditioning codes that are provided by auxiliary networks that model these conditions.
1 . A machine learning system comprising:
a non-transitory computer-readable medium; and
one or more processors, the one or more processors configured to execute a machine-learning, main task-performance network stored in the non-transitory computer-readable medium, the machine-learning, main task-performance network trained to perform a main machine-learning task, wherein the main task-performance network comprises a deep neural network, and wherein:
the deep neural network comprises:
an input portion comprising an input layer and N hidden layers, where N≥0;
an output portion comprising an output layer and M hidden layers, where M≥0; and
a modulation layer between the input portion and the output portion,
wherein the modulation layer comprises:
a plurality of nodes; and
connects the input portion to the output portion; and
a machine-learning (“ML”) auxiliary network trained to perform an auxiliary machine learning task, wherein features extracted by the auxiliary network are used by the modulation layer of the main task-performance network during training of the main task-performance network to modulate outputs from the input portion to the output portion;
wherein:
the ML auxiliary network is configured to provide modulation codes to the modulation layer, the modulation codes comprise coefficients based on features extracted by the ML auxiliary network;
each of the nodes of the modulation layer multiplies an output from a node of the input portion of the main task-performance network by one of the coefficients of the modulation codes; and
the product of the multiplication being input to a node of the output portion of the main task-performance network.
2 . The machine learning system of claim 1 , where N≥1 and M≥1, and wherein the modulation layer connects a hidden layer of the input portion to a hidden layer of the output portion.
3 . The machine learning system of claim 1 , wherein the auxiliary network comprises a bottleneck layer that extracts the features.
4 . The machine learning system of claim 1 , wherein:
the modulation layer comprises P nodes, where P>1;
wherein the modulation codes comprise P coefficients that are based on the features extracted by the auxiliary network; and
each of the P nodes of the modulation layer, during training of the main classification network, multiplies an output from a node of the input portion of the main task-performance network by one of the P coefficients of the modulation codes from the auxiliary network, with a product of the multiplication being input to a node of the output portion of the main task-performance network.
5 . The machine learning system of claim 4 , wherein the auxiliary network comprises a bottleneck layer that extracts the features.
6 . The machine learning system of claim 1 , wherein the auxiliary network provides a modulation code to the modulation layer, and further comprising a machine-learning converter network connected between the auxiliary network and the modulation layer of the main task-performance network, wherein the converter network receives as input the features extracted by the auxiliary network and outputs the modulation code for the modulation layer.
7 . The machine learning system of claim 6 , wherein the converter network is trained to optimize a loss function of a joint network comprising the main task-performance network and the converter network.
8 . A method for training a machine-learning, main task-performance network to perform a main machine-learning task, wherein the main task-performance network comprises a deep neural network that comprises:
an input portion comprising an input layer and N hidden layers, where N≥0;
an output portion comprising an output layer and M hidden layers, where M≥0; and
a modulation layer between the input portion and the output portion, wherein the modulation layer comprises:
a plurality of nodes; and
connects the input portion to the output portion,
the method comprising:
training, through machine-learning, an auxiliary network to perform an auxiliary machine learning task;
after training the auxiliary network, extracting features by the auxiliary network from input features; and
after extracting the features with the auxiliary network, training the main task-performance network, wherein training the main task-performance network comprises modulating nodes of the modulation layer of the main task-performance network with a modulation code that is based on the extracted features from the auxiliary network;
wherein:
the auxiliary network is configured to provide modulation codes to the modulation laver, the modulation codes comprise coefficients based on features extracted by the auxiliary network;
each of the nodes of the modulation layer multiplies an output from a node of the input portion of the main task-performance network by one of the coefficients of the modulation codes; and
the product of the multiplication being input to a node of the output portion of the main task-performance network.
9 . The method of claim 8 , wherein:
the modulation layer comprises P nodes, where P>1;
wherein the modulation codes comprise P coefficients that are based on the features extracted by the auxiliary network; and
each of the P nodes of the modulation layer, during training of the main classification network, multiplies an output from a node of the input portion of the main task-performance network by one of the P coefficients of the modulation codes from the auxiliary network, with a product of the multiplication being input to a node of the output portion of the main task-performance network.
10 . The method of claim 9 , further comprising, generating, by a machine-learning converter network connected between the auxiliary network and the modulation layer of the main task-performance network, the modulation code from the features extracted by the auxiliary network.
11 . The method of claim 10 , wherein:
the main machine-learning task comprises language-independent speech recognition; and
the auxiliary machine learning task of the auxiliary network comprises language identification.
12 . The method of claim 11 , further comprising a plurality of unique, machine-learning, mono-lingual sub-networks, wherein outputs of the mono-lingual sub-networks are connected to the input portion of the main task performance network.
13 . The method of claim 12 , wherein each of the plurality of unique mono-lingual sub-networks receive as input the features extracted by the auxiliary network.
14 . The method of claim 8 , wherein the machine-learning auxiliary network comprises a plurality of machine-learning auxiliary networks, wherein each of the plurality of machine-learning auxiliary networks performs a separate auxiliary machine-learning task, and wherein features extracted by the plurality of machine-learning auxiliary networks modulate the outputs of the input portion at the modulation layer of the main task-performance network.
15 . A non-transitory computer-readable medium comprising processor executable instructions configured to cause one or more processors to train a machine learning system comprising a machine-learning, main task-performance network to perform a main machine-learning task, wherein the main task-performance network comprises a deep neural network that comprises:
an input portion comprising an input layer and N hidden layers, where N≥0;
an output portion comprising an output layer and M hidden layers, where M≥0; and
a modulation layer between the input portion and the output portion, wherein the modulation layer comprises:
a plurality of nodes; and
connects the input portion to the output portion,
wherein the processor executable instructions are configured to cause one or more processors to:
train, through machine-learning, an auxiliary network to perform an auxiliary machine learning task;
after training the auxiliary network, extract features by the auxiliary network from input features; and
after extracting the features with the auxiliary network, train the main task-performance network, and modulate nodes of the modulation layer of the main task-performance network with a modulation code that is based on the extracted features from the auxiliary network;
wherein:
the ML auxiliary network is configured to provide modulation codes to the modulation layer, the modulation codes comprise coefficients based on features extracted by the ML auxiliary network;
each of the nodes of the modulation layer multiplies an output from a node of the input portion of the main task-performance network by one of the coefficients of the modulation codes; and
the product of the multiplication being input to a node of the output portion of the main task-performance network.
16 . The non-transitory computer-readable medium of claim 15 , wherein:
the modulation layer comprises P nodes, where P>1;
wherein the modulation codes comprise P coefficients that are based on the features extracted by the auxiliary network; and
each of the P nodes of the modulation layer, during training of the main classification network, multiplies an output from a node of the input portion of the main task-performance network by one of the P coefficients of the modulation codes from the auxiliary network, with a product of the multiplication being input to a node of the output portion of the main task-performance network.
17 . The non-transitory computer-readable medium of claim 16 , wherein the processor executable instructions are further configured to cause one or more processors to generate, by a machine-learning converter network connected between the auxiliary network and the modulation layer of the main task-performance network, the modulation code from the features extracted by the auxiliary network.
18 . The non-transitory computer-readable medium of claim 17 , wherein:
the main machine-learning task comprises language-independent speech recognition; and
the auxiliary machine-learning task of the auxiliary network comprises language identification.
19 . The non-transitory computer-readable medium of claim 18 , wherein the machine learning system further comprises a plurality of unique, machine-learning, mono-lingual sub-networks, wherein outputs of the mono-lingual sub-networks are connected to the input portion of the main task performance network.
20 . The non-transitory computer-readable medium of claim 19 , wherein each of the plurality of unique mono-lingual sub-networks receive as input the features extracted by the auxiliary network.