IP Library › Granted Patent US 12,633,087
Granted Patent B2
US 12,633,087 · App. 18/371,688 · Granted May 19, 2026

Scalable prompt learning for large vision-language models

Inventors: Chen Qiu (Pittsburgh, PA); Xingyu Li (New Orleans, LA); Chaithanya Kumar Mummadi (Pittsburgh, PA); Madan Ravi Ganesh (Pittsburgh, PA); Zhenzhen Li (Gibsonia, PA); Wan-Yi Lin (Wexford, PA); Sabrina Schmedding (Tiefenbronn, DE)
Assignee: Robert Bosch GmbH
G06V10/764G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,087
App. No.
18/371,688
Granted
May 19, 2026
Kind
B2
Abstract

A method of generating text-driven prompts and class prediction probabilities using a vision-language model (VLM) includes receiving candidate class names associated with a plurality of candidate classes for images, generating class text tokens based on a text description of the candidate class names, and generating a plurality of context prompt vectors using a prompt generator. The context prompt vectors define context information associated with an image classification task to be performed by the VLM. The method further includes generating prompts for each of the plurality of candidate classes by appending respective class text tokens to the context prompt vectors for each of the plurality of candidate classes, and, using the VLM, generating and outputting a class prediction probability for a sample image based on the plurality of context prompt vectors.

Claims (37)

1 . A method of generating text-driven prompts and class prediction probabilities using a vision-language model (VLM), the method comprising:

receiving candidate class names associated with a plurality of candidate classes for images;

generating class text tokens based on a text description of the candidate class names;

generating, based on the text description and separate from the class text tokens, a plurality of context prompt vectors using a prompt generator, wherein the context prompt vectors define context information associated with an image classification task to be performed by the VLM;

subsequent to generating the plurality of context prompt vectors, generating prompts for each of the plurality of candidate classes by appending respective class text tokens to the context prompt vectors for each of the plurality of candidate classes; and

using the VLM, generating and outputting a class prediction probability for a sample image based on the plurality of context prompt vectors, wherein using the VLM includes providing, to a text encoder of the VLM, the context prompt vectors.

2 . The method of claim 1 , wherein the VLM is a Contrastive Language-Image Pre-training (CLIP) model.

3 . The method of claim 1 , further comprising receiving the text description at an interface.

4 . The method of claim 3 , wherein receiving the text description includes providing a predefined text template and receiving the text description in accordance with the predefined text template.

5 . The method of claim 1 , further comprising providing the text description to a large language model (LLM) and generating the text embeddings using the large language model.

6 . The method of claim 1 , further comprising generating the plurality of context prompt vectors using a prompt generator model f θ .

7 . The method of claim 6 , further comprising, using the prompt generator model f θ , mapping the text embeddings to the plurality of context prompt vectors.

8 . The method of claim 6 , further comprising aggregating respective parameters of a plurality of the prompt generator models f θ and outputting an aggregated prompt generator model based on the aggregated respective parameters.

9 . A computing device configured to generate text-driven prompts and class prediction probabilities using a vision-language model (VLM), the computing device including a processing device configured to execute instructions stored in memory to:

receive candidate class names associated with a plurality of candidate classes for images;

generate class text tokens based on a text description of the candidate class names;

generate, based on the text description and separate from the class text tokens, a plurality of context prompt vectors using a prompt generator, wherein the context prompt vectors define context information associated with an image classification task to be performed by the VLM;

subsequent to generating the plurality of context prompt vectors, generate prompts for each of the plurality of candidate classes by appending respective class text tokens to the context prompt vectors for each of the plurality of candidate classes; and

using the VLM, generate and output a class prediction probability for a sample image based on the plurality of context prompt vectors, wherein using the VLM includes providing, to a text encoder of the VIM, the context prompt vectors.

10 . The computing device of claim 9 , wherein the VLM is a Contrastive Language-Image Pre-training (CLIP) model.

11 . The computing device of claim 9 , further comprising an interface configured to receive the text description.

12 . The computing device of claim 11 , wherein the interface is configured to provide a predefined text template and receive the text description in accordance with the predefined text template.

13 . The computing device of claim 9 , further comprising a large language model (LLM) configured to generate the text embeddings.

14 . The computing device of claim 9 , further comprising a prompt generator model f θ configured to generate the plurality of context prompt vectors.

15 . The computing device of claim 14 , wherein the prompt generator model is configured to map the text embeddings to the plurality of context prompt vectors.

16 . The computing device of claim 9 , wherein the VLM includes the text encoder and an image encoder.

17 . A computer-controlled machine, comprising:

at least one sensor configured to generate an input image;

a control system configured to generate text-driven prompts and class prediction probabilities using a vision-language model (VLM), the control system configured to

receive candidate class names associated with a plurality of candidate classes for the input image,

generate, based on the text description and separate from the class text tokens, a plurality of context prompt vectors using a prompt generator, wherein the context prompt vectors define context information associated with an image classification task to be performed by the VLM,

subsequent to generating the plurality of context prompt vectors, generate prompts for each of the plurality of candidate classes by appending respective class text tokens to the context prompt vectors for each of the plurality of candidate classes, and

using the VLM, generate and output a class prediction probability for the input image based on the plurality of context prompt vectors, wherein using the VLM includes providing, to a text encoder of the VLM, the context prompt vectors; and

an actuator configured to control an operation of the computer-controlled machine based on the class prediction probability.

18 . The computer-controlled machine of claim 17 , wherein the VLM is a Contrastive Language-Image Pre-training (CLIP) model.

19 . The computer-controlled machine of claim 17 , further comprising a prompt generator model f θ configured to generate the plurality of context prompt vectors.

20 . The computer-controlled machine of claim 19 , wherein the prompt generator model is configured to map the text embeddings to the plurality of context prompt vectors.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2023
From: QIU, CHEN; LI, XINGYU; MUMMADI, CHAITHANYA KUMAR; GANESH, MADAN RAVI; LI, ZHENZHEN; LIN, WAN-YI; SCHMEDDING, SABRINA
To: ROBERT BOSCH GMBH
Reel/Frame 065608/0712 →
Continuity (1)
Related Publication 20250104394A1 · Mar 27, 2025
References Cited (33)
US 20240203085A1 · Bangalath · 2024 [cited by examiner]
US 20240386887A1 · Kumar · 2024 [cited by examiner]
Guanghao Li, Wansen Wu, Yan Sun, Li Shen, Baoyuan Wu, Dacheng Tao; “Visual Prompt Based Personalized Federated Learning”; Mar. 15, 2023 https://doi.org/10.48550/arXiv.2303.08678 (Year: 2023). [cited by examiner]
Adrian Bulat and Georgios Tzimiropoulos. Lasp: Text-to-text optimization for language-aware soft prompting of vision language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition … [cited by applicant]
Shengchao Chen, Guodong Long, Tao Shen, Tianyi Zhou, and Jing Jiang. Spatial-temporal prompt learning for federated weather forecasting. arXiv preprint arXiv:2305.14244, 2023. [cited by applicant]
Hongchang Gao, My T Thai, and Jie Wu. When decentralized optimization meets federated learning. IEEE Network, 2023. [cited by applicant]
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. Clip-adapter: Better vision-language models with feature adapters. arXiv preprint arXiv:2110.04544, 2021. [cited by applicant]
Tao Guo, Song Guo, Junxiao Wang, Xueyang Tang, and Wenchao Xu. Promptfl: Let federated participants cooperatively learn prompts instead of models-federated learning in age of foundation model. IEEE Transactions on Mobil… [cited by applicant]
Shaunak Halbe, James Seale Smith, Junjiao Tian, and Zsolt Kira. Hepco: Data-free heterogeneous prompt consolidation for continual federated learning. arXiv preprint arXiv:2306.09970, 2023. [cited by applicant]
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Inte… [cited by applicant]
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. In European Conference on Computer Vision, pp. 709-727. Springer, 2022. [cited by applicant]
Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shah-baz Khan. Maple: Multi-modal prompt learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVP… [cited by applicant]
Guanghao Li, Wansen Wu, Yan Sun, Li Shen, Baoyuan Wu, and Dacheng Tao. Visual prompt based personalized federated learning. arXiv preprint arXiv:2303.08678, 2023. [cited by applicant]
Xingyu Li, Zhe Qu, Shangqing Zhao, Bo Tang, Zhuo Lu, and Yao Liu. Lomar: A local defense against poisoning attack on federated learning. IEEE Transactions on Dependable and Secure Computing, 2021. [cited by applicant]
Wang Lu, Xixu Hu, Jindong Wang, and Xing Xie. Fedclip: Fast generalization and personalization for clip in federated learning. arXiv preprint arXiv:2302.13485, 2023. [cited by applicant]
Yuning Lu, Jianzhuang Liu, Yonggang Zhang, Yajing Liu, and Xinmei Tian. Prompt distribution learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni-tion, pp. 5206-5215, 2022. [cited by applicant]
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communication-efficient learning of deep networks from decentralized data. In Artificial intelli-gence and statistics, pp. 1273-1282.… [cited by applicant]
Zhe Qu, Xingyu Li, Rui Duan, Yao Liu, Bo Tang, and Zhuo Lu. Generalized federated learning via sharpness aware minimization. In International Conference on Machine Learning, pp. 18250-18280. PMLR, 2022a. [cited by applicant]
Zhe Qu, Xingyu Li, Jie Xu, Bo Tang, Zhuo Lu, and Yao Liu. On the convergence of multi-server federated learning with overlapping area. IEEE Transactions on Mobile Computing, 2022b. [cited by applicant]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agar-wal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual mod… [cited by applicant]
Shangchao Su, Mingzhao Yang, Bin Li, and Xiangyang Xue. Cross-domain federated adaptive prompt tuning for clip. arXiv preprint arXiv:2211.07864, 2022. [cited by applicant]
Jiamian Wang, Zongliang Wu, Yulun Zhang, Xin Yuan, Tao Lin, and Zhiqiang Tao. Cooperative hardware-prompt earning for snapshot compressive imaging. arXiv preprint arXiv:2306.01176, 2023. [cited by applicant]
Hantao Yao, Rui Zhang, and Changsheng Xu. Visual-language prompt tuning with knowledge-guided context optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6757-6… [cited by applicant]
Tao Yu, Zhihe Lu, Xin Jin, Zhibo Chen, and Xinchao Wang. Task residual for tuning vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni-tion, pp. 10899-10909, 2023. [cited by applicant]
Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, and Chen Change Loy. Unified vision and language prompt learning. arXiv preprint arXiv:2210.07225, 2022. [cited by applicant]
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Conditional prompt learning for vision-language models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022a. [cited by applicant]
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision (IJCV), 2022b. [cited by applicant]
Beier Zhu, Yulei Niu, Yucheng Han, Yue Wu, and Hanwang Zhang. Prompt-aligned gradient for prompt tuning. arXiv preprint arXiv:2205.14865, 2022. [cited by applicant]
Kirillov, Alexander, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao et al. “Segment anything.” arXiv preprint arXiv:2304.02643 (2023). [cited by applicant]
Oquab, Maxime, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez et al. “Dinov2: Learning robust visual features without supervision.” arXiv preprint arXiv:2304.07193 (2023). [cited by applicant]
Minderer, M., A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, and Z. Shen. “Simple open-vocabulary object detection with vision transformers. arXiv 2022.” arXiv p… [cited by applicant]
Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. “Bert: Pre-training of deep bidirectional transformers for language understanding.” arXiv preprint arXiv:1810.04805 (2018). [cited by applicant]
Brown, Tom, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan et al. “Language models are few-shot learners.” Advances in neural information processing systems 33 (2020):… [cited by applicant]