Systems and methods for vision-language model instruction tuning
Embodiments described herein provide a method of generating a vision-language task output to a text instruction relating to an input image, the method comprising receiving, via a data interface, the input image and the text instruction comprising an instruction relating to the image. The method further includes encoding, via an image encoder, the image into a first image representation. The method further includes generating, by a multimodal encoder, a second image representation based on cross-attending the first image representation to the text instruction. The method further includes generating, by a neural network based language model, a vision-language task output in response to the text instruction based on an input combining the second image representation and the text instruction.
1 . A method of generating a vision-language task output for a text instruction relating to an input image, the method comprising:
receiving, via a data interface, the input image and the text instruction comprising an instruction relating to the input image;
encoding, via an image encoder, the input image into a first image representation;
generating, by a multimodal encoder connected to the image encoder, a second image representation based on cross-attending the first image representation to the text instruction by a cross-attention layer in the multimodal encoder; and
generating, by a neural network based language model connected to the multimodal encoder, the vision-language task output in response to the text instruction based on an input combining the second image representation and the text instruction.
2 . The method of claim 1 , further comprising:
receiving a known-good text output via the data interface;
computing a loss based on the vision-language task output and the known-good text output; and
training the multimodal encoder based on the loss.
3 . The method of claim 2 , further comprising:
keeping the image encoder and the neural network based language model frozen while training the multimodal encoder.
4 . The method of claim 2 , wherein the generating the second image representation is further based on a set of query vectors, further comprising:
updating the set of query vectors based on the vision-language task output and a known-good text output.
5 . The method of claim 1 , further comprising:
pre-training the multimodal encoder based on an output of the multimodal encoder without the neural network based language model.
6 . The method of claim 1 , further comprising:
adapting the second image representation for the neural network based language model via a feed forward neural network.
7 . The method of claim 1 , further comprising:
adapting the text instruction with an instruction template text.
8 . A system for generating a vision-language task output for a text instruction relating to an input image, the system comprising:
a memory that stores a neural network based language model and a plurality of processor executable instructions;
a data interface that receives the input image and the text instruction comprising an instruction relating to the input image; and
one or more hardware processors that read and execute the plurality of processor-executable instructions from the memory to perform operations comprising:
encoding, via an image encoder, the input image into a first image representation;
generating, by a multimodal encoder connected to the image encoder, a second image representation based on cross-attending the first image representation to the text instruction by a cross-attention layer in the multimodal encoder; and
generating, by the neural network based language model connected to the multimodal encoder, the vision-language task output in response to the text instruction based on an input combining the second image representation and the text instruction.
9 . The system of claim 8 , the operations further comprising:
receiving a known-good text output via the data interface;
computing a loss based on the vision-language task output and the known-good text output; and
training the multimodal encoder based on the loss.
10 . The system of claim 9 , the operations further comprising:
keeping the image encoder and the neural network based language model frozen while training the multimodal encoder.
11 . The system of claim 9 , wherein the generating the second image representation is further based on a set of query vectors, the operations further comprising:
updating the set of query vectors based on the vision-language task output and a known-good text output.
12 . The system of claim 8 , the operations further comprising:
pre-training the multimodal encoder based on an output of the multimodal encoder without the neural network based language model.
13 . The system of claim 8 , the operations further comprising:
adapting the second image representation for the neural network based language model via a feed forward neural network.
14 . The system of claim 8 , the operations further comprising:
adapting the text instruction with an instruction template text.
15 . A non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations comprising:
receiving, via a data interface, an input image and a text instruction comprising an instruction relating to the input image;
encoding, via an image encoder, the input image into a first image representation;
generating, by a multimodal encoder connected to the image encoder, a second image representation based on cross-attending the first image representation to the text instruction by a cross-attention layer in the multimodal encoder; and
generating, by a neural network based language model connected to the multimodal encoder, a vision-language task output in response to the text instruction based on an input combining the second image representation and the text instruction.
16 . The non-transitory machine-readable medium of claim 15 , the operations further comprising:
receiving a known-good text output via the data interface;
computing a loss based on the vision-language task output and the known-good text output; and
training the multimodal encoder based on the loss.
17 . The non-transitory machine-readable medium of claim 16 , the operations further comprising:
keeping the image encoder and the neural network based language model frozen while training the multimodal encoder.
18 . The non-transitory machine-readable medium of claim 16 , wherein the generating the second image representation is further based on a set of query vectors, the operations further comprising:
updating the set of query vectors based on the vision-language task output and a known-good text output.
19 . The non-transitory machine-readable medium of claim 15 , the operations further comprising:
pre-training the multimodal encoder based on an output of the multimodal encoder without the neural network based language model.
20 . The non-transitory machine-readable medium of claim 15 , the operations further comprising:
adapting the second image representation for the neural network based language model via a feed forward neural network.