Generating text prompts for digital images utilizing vision-language models and contextual prompt learning
The present invention relates to systems, methods, and non-transitory computer-readable media that implements a vision language machine learning model to generate text representations of digital image from localized context tokens, where the text representations semantically reflect the content of the digital image under consideration. In particular, in some embodiments, the disclosed invention generates image patch feature representations that represent patches from an input image. Further, in some embodiments, the disclosed invention generates localized context tokens from the image patch feature representations and prompt context tokens. Moreover, in some embodiments, by utilizing the localized context tokens, the disclosed invention generates a text representation by utilizing a text encoder of the vision language machine learning model.
1 . A computer-implemented method comprising:
generating, utilizing an image encoder of a vision language machine learning model, image patch feature representations that represent patches from an input image;
identifying learned prompt context tokens, wherein the learned prompt context tokens are learned in training the vision language machine learning model;
utilizing an attention layer to generate localized context tokens by:
generating, utilizing an alignment model of the attention layer, alignment vectors between the learned prompt context tokens and the image patch feature representations by weighing the learned prompt context tokens according to the image patch feature representations;
generating context vectors by combining the alignment vectors and the learned prompt context tokens; and
generating, utilizing an attention layer of the vision language machine learning model, the localized context tokens by combining the context vectors with the learned prompt context tokens; and
generating, utilizing a text encoder of the vision language machine learning model, a text caption describing the input image from the localized context tokens generated based on the alignment vectors and the learned prompt context tokens.
2 . The computer-implemented method of claim 1 , wherein generating the image patch feature representations comprises:
extracting the patches from the input image;
generating, utilizing the image encoder, image patch feature vectors from the patches; and
generating, utilizing a neural network, conditional image patch tokens from the image patch feature vectors.
3 . The computer-implemented method of claim 1 , wherein training the vision language machine learning model further comprises:
initializing prompt context tokens by selecting the prompt context tokens from a distribution; and
utilizing the prompt context tokens as learnable parameters for the vision language machine learning model.
4 . The computer-implemented method of claim 1 , wherein generating the alignment vectors comprises utilizing learned weights to compare the learned prompt context tokens with conditional image patch tokens.
5 . The computer-implemented method of claim 1 , wherein generating, utilizing the attention layer of the vision language machine learning model, the localized context tokens further comprises combining the context vectors for the patches from the input image with the learned prompt context tokens to generate the localized context tokens.
6 . The computer-implemented method of claim 1 , wherein training the vision language machine learning model further comprises:
generating, utilizing the image encoder, an image feature vector of the input image;
determining a measure of loss by comparing the text caption with the image feature vector; and
modifying prompt context tokens initialized from a distribution and weights of the attention layer of the vision language machine learning model based on the determined measure of loss.
7 . The computer-implemented method of claim 6 , further comprising training the vision language machine learning model by generating the text caption, utilizing the text encoder of the vision language machine learning model, from the localized context tokens and a ground truth class corresponding to the input image.
8 . A system comprising:
one or more memory devices comprising an input image, learned prompt context tokens, and a vision language machine learning model comprising an image encoder, an attention layer, and a text encoder; and
one or more processors configured to cause the system to:
generate, utilizing the image encoder, image patch feature representations from a plurality of patches of the input image;
identify learned prompt context tokens, wherein the learned prompt context tokens are learned in training the vision language machine learning model;
utilize the attention layer to generate localized context tokens by:
generating, utilizing an alignment model, alignment vectors between the learned prompt context tokens and the image patch feature representations generated from the plurality of patches of the input image by weighing the learned prompt context tokens according to the image patch feature representations;
generating context vectors for the plurality of patches by combining the alignment vectors generated from weighting the learned prompt context tokens according to the image patch feature representations and the learned prompt context tokens; and
generating the localized context tokens by combining the context vectors generated from the alignment vectors and the learned prompt context tokens with the learned prompt context tokens; and
generate, utilizing the text encoder, a text representation that textually describes the input image from the localized context tokens that are generated from the context vectors and the learned prompt context tokens.
9 . The system of claim 8 , wherein the one or more processors are configured to cause the system to generate the image patch feature representations by:
generating, utilizing the image encoder, image patch feature vectors from the plurality of patches from the input image; and
generating, utilizing a neural network, conditional image patch tokens from the plurality of patches from the input image.
10 . The system of claim 8 , wherein generating the alignment vectors comprises utilizing learned weights of the attention layer to compare the learned prompt context tokens with the image patch feature representations.
11 . The system of claim 8 , wherein the one or more processors are configured to cause the system to train the vision language machine learning model by:
generating the text representation from the localized context tokens and a ground truth class corresponding to the input image; and
determining a measure of loss by comparing the text representation with an image feature vector of the input image.
12 . The system of claim 8 , wherein the one or more processors are configured to cause the system to train the vision language machine learning model by modifying prompt context tokens and learned weights of the attention layer of the vision language machine learning model based on a determined measure of loss.
13 . The system of claim 8 , wherein generating the alignment vectors between the learned prompt context tokens and the image patch feature representations comprises:
generating a first alignment vector for a first learned prompt context token and a first image patch; and
generating a second alignment vector for a second learned prompt context token and a first image patch.
14 . The system of claim 13 , wherein the one or more processors are configured to cause the system to:
combine the first alignment vector and the second alignment vector to generate a first context vector; and
generate a first localized context token by combining the first context vector with the first learned prompt context token.
15 . A non-transitory computer-readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:
generating, utilizing an image encoder, image patch feature representations that represent patches from an input image;
identifying learned prompt context tokens, wherein the learned prompt context tokens are learned in training a vision language machine learning model;
generating, utilizing an alignment model, alignment vectors between the learned prompt context tokens and the image patch feature representations by weighing the learned prompt context tokens according to the image patch feature representations generated from the patches of the input image;
generating localized context tokens from the learned prompt context tokens and from utilizing the alignment vectors generated by weighting the learned prompt context tokens according to the image patch feature representations; and
generating, utilizing a text encoder, a text representation of the input image from the localized context tokens generated from the learned prompt context tokens and the alignment vectors.
16 . The non-transitory computer-readable medium of claim 15 , wherein generating the image patch feature representations further comprises:
generating, utilizing a neural network, conditional image patch tokens; and
generating the alignment vectors by utilizing the conditional image patch tokens and the learned prompt context tokens.
17 . The non-transitory computer-readable medium of claim 15 , wherein determining the alignment vectors further comprises:
generating, utilizing a neural network, conditional image patch tokens from the patches of the input image; and
determining the alignment vectors by applying weights of the alignment model to the conditional image patch tokens and the learned prompt context tokens.
18 . The non-transitory computer-readable medium of claim 15 , wherein generating the localized context tokens further comprises:
generating context vectors by combining the alignment vectors and the learned prompt context tokens; and
generating the localized context tokens by combining the context vectors with the learned prompt context tokens.
19 . The non-transitory computer-readable medium of claim 15 , wherein generating the text representation further comprises processing, utilizing the text encoder, a ground truth class corresponding to the input image.
20 . The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise training the vision language machine learning model by:
generating, utilizing the image encoder, an image feature vector of the input image;
determining a measure of loss by comparing the text representation with the image feature vector; and
modifying prompt context tokens initialized from a distribution and weights of the alignment model based on the determined measure of loss.