Multi-modal artifical neural network and a self-supervised learning method for training same
A multi-modal artificial neural network and a self-supervised learning method for training that network. The learning method involves processing, using a first modality simple Siamese network, a pair of first modality augmented views of an input; processing, using a second modality simple Siamese network, a pair of second modality augmented views of the input; determining at least one cross-modal loss between the first and second modality simple Siamese networks; determining a total loss from: (i) first and second modality losses respectively determined during the processing using the first and second modality simple Siamese networks; and (ii) the at least one cross-modal loss; and training the first and second modality simple Siamese networks based on the total loss. The trained network may be used to analyze multi-modal content such as video content that has an audio track. A Multi-Modal Multi-Head Network (M3HN) may also be trained to process modality-specific and modality-agnostic representations.
1 . A self-supervised learning method comprising:
(a) processing, using a first modality simple Siamese network, a pair of first modality augmented views of an input;
(b) processing, using a second modality simple Siamese network, a pair of second modality augmented views of the input;
(c) determining at least one cross-modal loss between the first and second modality simple Siamese networks;
(d) determining a total loss from:
(i) first and second modality losses respectively determined during the processing using the first and second modality simple Siamese networks; and
(ii) the at least one cross-modal loss; and
(e) training the first and second modality simple Siamese networks based on the total loss,
wherein determining the at least one cross-modal loss comprises determining a cross-modal loss based on:
an output of a modality-agnostic projector of the first modality simple Siamese network that processes one of the first modality augmented views; and
an output of a modality-agnostic projector of the second modality simple Siamese network that processes one of the second modality augmented views,
wherein the input comprises a video, the first modality comprises images from the video, and the second modality comprises audio from the video, and
wherein a gradient is back propagated to each of the modality-agnostic projectors, and a stop-gradient is applied to respective encoders of the first and second Siamese networks, wherein the outputs of the encoders of the first and second Siamese networks are respectively input to the modality-agnostic projector of the first and second Siamese networks.
2 . The method of claim 1 , wherein predictors of the first and second modality simple Siamese networks share weights.
3 . The method of claim 1 , wherein the encoders of the first and second Siamese networks comprise second and third encoders, and wherein:
the first modality simple Siamese network comprises:
a first encoder and the second encoder that respectively process the pair of first modality augmented views; and
a first modality-specific projector and a second modality-specific projector that respectively receive an output of the first encoder and the output of the second encoder:
the second modality simple Siamese network comprises:
the third encoder and a fourth encoder that respectively process the pair of second modality augmented views; and
a third modality-specific projector and a fourth modality-specific projector that respectively receive the output of the third encoder and an output of the fourth encoder; and
at least one of the i) first and second or ii) third and fourth modality-specific projectors share weights.
4 . The method of claim 3 , wherein the first and second encoders share weights, and the third and fourth encoders share weights.
5 . The method of claim 3 , wherein the first, second, third, and fourth modality-specific projectors share weights.
6 . The method of claim 3 , wherein the first modality-agnostic projector comprises a first network of convolutional layers whose output is input to a projector identical to the first or second modality-specific projector, and wherein the second modality-agnostic projector comprises a second network of convolutional layers whose output is input to a projector identical to the third or fourth modality-specific projector.
7 . The method of claim 1 , wherein the augmented views of the video are generated by performing at least one of random cropping, random rotation, random colorization, temporal shifting, random masking, or random shuffling.
8 . The method of claim 1 , wherein the augmented views of the audio are generated by performing at least one of adding additive white Gaussian noise, temporal shifting, random masking, random shuffling, or adding random silences.
9 . An artificial neural network trained in accordance with a self-supervised learning method, the artificial neural network comprising first and second modality simple Siamese networks, the method comprising:
(a) processing, using the first modality simple Siamese network, a pair of first modality augmented views of an input;
(b) processing, using the second modality simple Siamese network, a pair of second modality augmented views of the input;
(c) determining at least one cross-modal loss between the first and second modality simple Siamese networks;
(d) determining a total loss from:
(i) first and second modality losses respectively determined during the processing using the first and second modality simple Siamese networks; and
(ii) the at least one cross-modal loss; and
(e) training the first and second modality simple Siamese networks based on the total loss,
wherein determining the at least one cross-modal loss comprises determining a cross-modal loss based on:
an output of a modality-agnostic projector of the first modality simple Siamese network that processes one of the first modality augmented views; and
an output of a modality-agnostic projector of the second modality simple Siamese network that processes one of the second modality augmented views,
wherein the input comprises a video, the first modality comprises images from the video, and the second modality comprises audio from the video, and
wherein a gradient is back propagated to each of the modality-agnostic projectors, and a stop-gradient is applied to respective encoders of the first and second Siamese networks, wherein the outputs of the encoders of the first and second Siamese networks are respectively input to the modality-agnostic projector of the first and second Siamese networks.
10 . The artificial neural network of claim 9 , wherein predictors of the first and second modality simple Siamese networks share weights.
11 . The artificial neural network of claim 9 , wherein the encoders of the first and second Siamese networks comprise second and third encoders, and wherein:
the first modality simple Siamese network comprises:
a first encoder and the second encoder that respectively process the pair of first modality augmented views; and
a first modality-specific projector and a second modality-specific projector that respectively receive an output of the first encoder and the output of the second encoder;
the second modality simple Siamese network comprises:
the third encoder and a fourth encoder that respectively process the pair of second modality augmented views; and
a third modality-specific projector and a fourth modality-specific projector that respectively receive the output of the third encoder and an output of the fourth encoder; and
at least one of the i) first and second or ii) third and fourth modality-specific projectors share weights.
12 . The artificial neural network of claim 11 , wherein the first and second encoders share weights, and the third and fourth encoders share weights.
13 . The artificial neural network of claim 11 , wherein the first, second, third, and fourth modality-specific projectors share weights.
14 . The artificial neural network of claim 11 , wherein the first modality-agnostic projector comprises a first network of convolutional layers whose output is input to a projector identical to the first or second modality-specific projector, and wherein the second modality-agnostic projector comprises a second network of convolutional layers whose output is input to a projector identical to the third or fourth modality-specific projector.
15 . The artificial neural network of claim 9 , wherein the augmented views of the video are generated by performing at least one of random cropping, random rotation, random colorization, temporal shifting, random masking, or random shuffling.
16 . The artificial neural network of claim 9 , wherein the augmented views of the audio are generated by performing at least one of adding additive white Gaussian noise, temporal shifting, random masking, random shuffling, or adding random silences.