IP Library › Granted Patent US 12,737,554
Granted Patent B2
US 12,737,554 · App. 18/342,954 · Granted Sep 15, 2026

Generating text prompts for digital images utilizing vision-language models and contextual prompt learning

Inventors: Koustava Goswami (Bangalore, IN); Srikrishna Karanam (Bangalore, IN); Joseph Koonthanam Jose (Kottayam, IN); Prateksha Udhayanan (Bangalore, IN); Balaji Vasan Srinivasan (Bangalore, IN)
Assignee: Adobe Inc.
G06F40/40G06V10/7715G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,554
App. No.
18/342,954
Granted
Sep 15, 2026
Kind
B2
Abstract

The present invention relates to systems, methods, and non-transitory computer-readable media that implements a vision language machine learning model to generate text representations of digital image from localized context tokens, where the text representations semantically reflect the content of the digital image under consideration. In particular, in some embodiments, the disclosed invention generates image patch feature representations that represent patches from an input image. Further, in some embodiments, the disclosed invention generates localized context tokens from the image patch feature representations and prompt context tokens. Moreover, in some embodiments, by utilizing the localized context tokens, the disclosed invention generates a text representation by utilizing a text encoder of the vision language machine learning model.

Claims (66)

1 . A computer-implemented method comprising:

generating, utilizing an image encoder of a vision language machine learning model, image patch feature representations that represent patches from an input image;

identifying learned prompt context tokens, wherein the learned prompt context tokens are learned in training the vision language machine learning model;

utilizing an attention layer to generate localized context tokens by:

generating, utilizing an alignment model of the attention layer, alignment vectors between the learned prompt context tokens and the image patch feature representations by weighing the learned prompt context tokens according to the image patch feature representations;

generating context vectors by combining the alignment vectors and the learned prompt context tokens; and

generating, utilizing an attention layer of the vision language machine learning model, the localized context tokens by combining the context vectors with the learned prompt context tokens; and

generating, utilizing a text encoder of the vision language machine learning model, a text caption describing the input image from the localized context tokens generated based on the alignment vectors and the learned prompt context tokens.

2 . The computer-implemented method of claim 1 , wherein generating the image patch feature representations comprises:

extracting the patches from the input image;

generating, utilizing the image encoder, image patch feature vectors from the patches; and

generating, utilizing a neural network, conditional image patch tokens from the image patch feature vectors.

3 . The computer-implemented method of claim 1 , wherein training the vision language machine learning model further comprises:

initializing prompt context tokens by selecting the prompt context tokens from a distribution; and

utilizing the prompt context tokens as learnable parameters for the vision language machine learning model.

4 . The computer-implemented method of claim 1 , wherein generating the alignment vectors comprises utilizing learned weights to compare the learned prompt context tokens with conditional image patch tokens.

5 . The computer-implemented method of claim 1 , wherein generating, utilizing the attention layer of the vision language machine learning model, the localized context tokens further comprises combining the context vectors for the patches from the input image with the learned prompt context tokens to generate the localized context tokens.

6 . The computer-implemented method of claim 1 , wherein training the vision language machine learning model further comprises:

generating, utilizing the image encoder, an image feature vector of the input image;

determining a measure of loss by comparing the text caption with the image feature vector; and

modifying prompt context tokens initialized from a distribution and weights of the attention layer of the vision language machine learning model based on the determined measure of loss.

7 . The computer-implemented method of claim 6 , further comprising training the vision language machine learning model by generating the text caption, utilizing the text encoder of the vision language machine learning model, from the localized context tokens and a ground truth class corresponding to the input image.

8 . A system comprising:

one or more memory devices comprising an input image, learned prompt context tokens, and a vision language machine learning model comprising an image encoder, an attention layer, and a text encoder; and

one or more processors configured to cause the system to:

generate, utilizing the image encoder, image patch feature representations from a plurality of patches of the input image;

identify learned prompt context tokens, wherein the learned prompt context tokens are learned in training the vision language machine learning model;

utilize the attention layer to generate localized context tokens by:

generating, utilizing an alignment model, alignment vectors between the learned prompt context tokens and the image patch feature representations generated from the plurality of patches of the input image by weighing the learned prompt context tokens according to the image patch feature representations;

generating context vectors for the plurality of patches by combining the alignment vectors generated from weighting the learned prompt context tokens according to the image patch feature representations and the learned prompt context tokens; and

generating the localized context tokens by combining the context vectors generated from the alignment vectors and the learned prompt context tokens with the learned prompt context tokens; and

generate, utilizing the text encoder, a text representation that textually describes the input image from the localized context tokens that are generated from the context vectors and the learned prompt context tokens.

9 . The system of claim 8 , wherein the one or more processors are configured to cause the system to generate the image patch feature representations by:

generating, utilizing the image encoder, image patch feature vectors from the plurality of patches from the input image; and

generating, utilizing a neural network, conditional image patch tokens from the plurality of patches from the input image.

10 . The system of claim 8 , wherein generating the alignment vectors comprises utilizing learned weights of the attention layer to compare the learned prompt context tokens with the image patch feature representations.

11 . The system of claim 8 , wherein the one or more processors are configured to cause the system to train the vision language machine learning model by:

generating the text representation from the localized context tokens and a ground truth class corresponding to the input image; and

determining a measure of loss by comparing the text representation with an image feature vector of the input image.

12 . The system of claim 8 , wherein the one or more processors are configured to cause the system to train the vision language machine learning model by modifying prompt context tokens and learned weights of the attention layer of the vision language machine learning model based on a determined measure of loss.

13 . The system of claim 8 , wherein generating the alignment vectors between the learned prompt context tokens and the image patch feature representations comprises:

generating a first alignment vector for a first learned prompt context token and a first image patch; and

generating a second alignment vector for a second learned prompt context token and a first image patch.

14 . The system of claim 13 , wherein the one or more processors are configured to cause the system to:

combine the first alignment vector and the second alignment vector to generate a first context vector; and

generate a first localized context token by combining the first context vector with the first learned prompt context token.

15 . A non-transitory computer-readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

generating, utilizing an image encoder, image patch feature representations that represent patches from an input image;

identifying learned prompt context tokens, wherein the learned prompt context tokens are learned in training a vision language machine learning model;

generating, utilizing an alignment model, alignment vectors between the learned prompt context tokens and the image patch feature representations by weighing the learned prompt context tokens according to the image patch feature representations generated from the patches of the input image;

generating localized context tokens from the learned prompt context tokens and from utilizing the alignment vectors generated by weighting the learned prompt context tokens according to the image patch feature representations; and

generating, utilizing a text encoder, a text representation of the input image from the localized context tokens generated from the learned prompt context tokens and the alignment vectors.

16 . The non-transitory computer-readable medium of claim 15 , wherein generating the image patch feature representations further comprises:

generating, utilizing a neural network, conditional image patch tokens; and

generating the alignment vectors by utilizing the conditional image patch tokens and the learned prompt context tokens.

17 . The non-transitory computer-readable medium of claim 15 , wherein determining the alignment vectors further comprises:

generating, utilizing a neural network, conditional image patch tokens from the patches of the input image; and

determining the alignment vectors by applying weights of the alignment model to the conditional image patch tokens and the learned prompt context tokens.

18 . The non-transitory computer-readable medium of claim 15 , wherein generating the localized context tokens further comprises:

generating context vectors by combining the alignment vectors and the learned prompt context tokens; and

generating the localized context tokens by combining the context vectors with the learned prompt context tokens.

19 . The non-transitory computer-readable medium of claim 15 , wherein generating the text representation further comprises processing, utilizing the text encoder, a ground truth class corresponding to the input image.

20 . The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise training the vision language machine learning model by:

generating, utilizing the image encoder, an image feature vector of the input image;

determining a measure of loss by comparing the text representation with the image feature vector; and

modifying prompt context tokens initialized from a distribution and weights of the alignment model based on the determined measure of loss.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2023
From: GOSWAMI, KOUSTAVA; KARANAM, SRIKRISHNA; JOSE, JOSEPH KOONTHANAM; UDHAYANAN, PRATEKSHA; SRINIVASAN, BALAJI VASAN
To: ADOBE INC.
Reel/Frame 064095/0071 →
Continuity (1)
Related Publication 20250005296A1 · Jan 2, 2025
References Cited (69)
US 20220383048A1 · Fei · 2022 [cited by examiner]
US 20220391755A1 · Li · 2022 [cited by examiner]
US 20230281963A1 · Gopalkrishna · 2023 [cited by examiner]
US 20240119257A1 · Guo · 2024 [cited by examiner]
US 20240153258A1 · Mangla · 2024 [cited by examiner]
US 20240220722A1 · Khattak · 2024 [cited by examiner]
US 20250005296A1 · Goswami · 2025 [cited by examiner]
Aishwarya Kamath, et al. MDETR—modulated detection for end-to-end multi-modal understanding. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, Oct. 10-17, 2021, pp. 1760-1770… [cited by applicant]
Alec Radford, et al. Learning transferable visual models from natural language supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, Jul. … [cited by applicant]
Alex Krizhevsky, et al. Imagenet classification with deep convolutional neural networks. In Peter L. Bartlett, Fernando C. N. Pereira, Christopher J. C. Burges, Leon Bottou, and Kilian Q. Weinberger, editors, Advances i… [cited by applicant]
Alexander Kolesnikov, et al. Big transfer (bit): General visual representation learning. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan Michael Frahm, editors, Computer Vision—ECCV 2020—16th European Conference,… [cited by applicant]
Alexey Dosovitskiy, et al. An image is worth 16×16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenR… [cited by applicant]
Alina Kuznetsova, et al. The open images dataset V4: unified image classification, object detection, and visual relationship detection at scale. CoRR, abs/1811.00982, 2018. [cited by applicant]
Andrea Frome, et al. Devise: A deep visual-semantic embedding model. In Christopher J. C. Burges, Lé on Bottou, Zoubin Ghahramani, and Kilian Q. Weinberger, editors, Advances in Neural Information Processing Systems 26:… [cited by applicant]
Andreas Furst, et al. CLOOB: modern hopfield networks with infoloob outperform CLIP. CoRR, abs/2110.11316, 2021. [cited by applicant]
Andy T. Liu, et al. Qaner: Prompting question answering models for few-shot named entity recognition. CoRR, abs/2203.01543, 2022. [cited by applicant]
Ang Li, et al. Learning visual n-grams from web data. In Proceedings of the IEEE International Conference on Computer Vision, pp. 4183-4192, 2017. [cited by applicant]
Armand Joulin, et al. Learning visual features from large weakly supervised data. In Computer Vision—ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, Oct. 11-14, 2016, Proceedings, Part VII 14, pp. 67-84… [cited by applicant]
Ashish Vaswani, et al. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, Advances in Neural Information Proce… [cited by applicant]
Brian Lester, et al. The power of scale for parameter-efficient prompt tuning. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical Metho… [cited by applicant]
Chao Jia, et al. Scaling up visual and vision-language representation learning with noisy text supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, … [cited by applicant]
Chen Sun, et al. Revisiting unreasonable effectiveness of data in deep learning era. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, Oct. 22-29, 2017, pp. 843-852. IEEE Computer Society, 2… [cited by applicant]
Fabio Petroni, et al. Language models as knowledge bases? In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the… [cited by applicant]
Jacob Devlin, et al. BERT: pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North Americ… [cited by applicant]
Jia Deng, et al. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), Jun. 20-25, 2009, Miami, Florida, USA, pp. 248-255. … [cited by applicant]
Jianxiong Xiao, et al. SUN database: Large-scale scene recognition from abbey to zoo. In The Twenty-Third IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2010, San Francisco, CA, USA, Jun. 13-18, 2010, … [cited by applicant]
Jo Plested and Tom Gedeon. Deep transfer learning for image classification: a survey. CoRR, abs/2205.09904, 2022. [cited by applicant]
Jonathan Krause, et al. 3d object representations for fine-grained categorization. In 2013 IEEE International Conference on Computer Vision Workshops, ICCV Workshops 2013, Sydney, Australia, Dec. 1-8, 2013, pp. 554-561.… [cited by applicant]
Kaiming He, et al. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, Jun. 27-30, 2016, pp. 770-778. IEEE Computer Society, 2… [cited by applicant]
Kaiming He, et al. Momentum contrast for unsupervised visual representation learning. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, Jun. 13-19, 2020, pp. 9726-9735.… [cited by applicant]
Kaiyang Zhou, et al. Conditional prompt learning for vision-language models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, Jun. 18-24, 2022, pp. 16795-16804. IEEE, 2… [cited by applicant]
Kaiyang Zhou, et al. Learning to prompt for vision-language models. Int. J. Comput. Vis., 130(9):2337-2348, 2022. [cited by applicant]
Karan Desai and Justin Johnson. Virtex: Learning visual representations from textual annotations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11162-11173, 2021. [cited by applicant]
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San… [cited by applicant]
Khurram Soomro, et al. UCF101: A dataset of 101 human actions classes from videos in the wild. CoRR, abs/1212.0402, 2012. [cited by applicant]
Koustava Goswami, et al. Switchprompt: Learning domain-specific gated soft prompts for classification in low-resource domains. CoRR, abs/2302.06868, 2023. [cited by applicant]
Lei Jimmy Ba, et al. Predicting deep zero-shot convolutional neural networks using textual descriptions. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, Dec. 7-13, 2015, pp. 4247-42… [cited by applicant]
Li Fei-Fei, et al. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In 2004 conference on computer vision and pattern recognition workshop, … [cited by applicant]
Liunian Harold Li, et al. Grounded language-image pre-training. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, Jun. 18-24, 2022, pp. 10955-10965. IEEE, 2022. [cited by applicant]
Liunian Harold Li, et al. Visualbert: A simple and performant baseline for vision and language. CoRR, abs/1908.03557, 2019. [cited by applicant]
Lluis Gomez, et al. Self-supervised learning of visual features through embedding images into text topic spaces. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4230-4239, 2017. [cited by applicant]
Lukas Bossard, et al. Food-101—mining discriminative components with random forests. In David J. Fleet, Tomas Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors, Computer Vision—ECCV 2014—13th European Conference, Zur… [cited by applicant]
Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. In Sixth Indian Conference on Computer Vision, Graphics & Image Processing, ICVGIP 2008, Bhubaneswar, India, Dec… [cited by applicant]
Mihir Parmar, et al. Inboxbart: Get instructions into biomedical multi-task learning. In Marine Carpuat, Marie-Catherine de Marneffe, and Iv á n Vladimir Meza Ruiz, editors, Findings of the Association for Computational… [cited by applicant]
Mingkai Deng, et al. Rlprompt: Optimizing discrete text prompts with reinforcement learning. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natura… [cited by applicant]
Mircea Cimpoi, et al. Describing textures in the wild. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2014, Columbus, OH, USA, Jun. 23-28, 2014, pp. 3606-3613. IEEE Computer Society, 2014. [cited by applicant]
Mitchell Wortsman, et al. Robust fine-tuning of zero-shot models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7959-7971, 2022. [cited by applicant]
Mohamed Elhoseiny, et al. Write a classifier: Zero-shot learning using purely textual descriptions. In IEEE International Conference on Computer Vision, ICCV 2013, Sydney, Australia, Dec. 1-8, 2013, pp. 2584-2591. IEEE … [cited by applicant]
Olivier J. Henaff. Data-efficient image recognition with contrastive predictive coding. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, Jul. 13-18, 2020, Virtual Event, vol. 119 of Pr… [cited by applicant]
Omkar M. Parkhi, et al. Cats and dogs. In 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA, Jun. 16-21, 2012, pp. 3498-3505. IEEE Computer Society, 2012. [cited by applicant]
Patrick Helber, et al. Introducing eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. In 2018 IEEE International Geoscience and Remote Sensing Symposium, IGARSS 2018, Valenc… [cited by applicant]
Peng Gao, et al. Clip-adapter: Better vision-language models with feature adapters. CoRR, abs/2110.04544, 2021. [cited by applicant]
Pengfei Liu, et al. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Comput. Surv., 55(9):195:1195:35, 2023. [cited by applicant]
Richard Socher, et al. Zero-shot learning through cross-modal transfer. In Christopher J. C. Burges, L é on Bottou, Zoubin Ghahramani, and Kilian Q. Weinberger, editors, Advances in Neural Information Processing Systems… [cited by applicant]
Ruosi Wan, et al. Towards making deep transfer learning never hurt. In Jianyong Wang, Kyuseok Shim, and Xindong Wu, editors, 2019 IEEE International Conference on Data Mining, ICDM 2019, Beijing, China, Nov. 8-11, 2019,… [cited by applicant]
Sebastian Ruder. An overview of gradient descent optimization algorithms. CoRR, abs/1609.04747, 2016. [cited by applicant]
Subhransu Maji, et al. Fine-grained visual classification of aircraft. CoRR, abs/1306.5151, 2013. [cited by applicant]
Thang Luong, et al. Effective approaches to attention-based neural machine translation. In Lluis Marquez, Chris Callison-Burch, Jian Su, Daniele Pighin, and Yuval Marton, editors, Proceedings of the 2015 Conference on E… [cited by applicant]
Tianyu Gao, et al. Making pre-trained language models better few-shot learners. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Meeting of the Association for Computat… [cited by applicant]
Timo Schick and Hinrich Schutze. Exploiting cloze questions for few-shot text classification and natural language inference. In Paola Merlo, J'org Tiedemann, and Reut Tsarfaty, editors, Proceedings of the 16th Conferenc… [cited by applicant]
Tom B. Brown, et al. Language models are few-shot learners. In Hugo Larochelle, Marc Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and HsuanTien Lin, editors, Advances in Neural Information Processing Systems 33:… [cited by applicant]
Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Meeting of the Association … [cited by applicant]
Xiao Liu, et al. GPT understands, too. CoRR, abs/2103.10385, 2021. [cited by applicant]
Xiao Liu, et al. P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Asso… [cited by applicant]
Yafu Li, et al. Prompt-driven neural machine translation. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Findings of the Association for Computational Linguistics: ACL 2022, Dublin, Ireland, May 2… [cited by applicant]
Yangguang Li, et al. Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, Apr. 25… [cited by applicant]
Yi Zhu, et al. Prompt-learning for short text classification. CoRR, abs/2202.11345, 2022. [cited by applicant]
Yunhui Guo, et al. Spottune: Transfer learning through adaptive fine-tuning. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, Jun. 16-20, 2019, pp. 4805-4814. Computer Visio… [cited by applicant]
Zhengbao Jiang, et al.. How can we know what language models know. Trans. Assoc. Comput. Linguistics, 8:423-438, 2020. [cited by applicant]