IP Library Granted Patent US 12,682,595
Granted Patent B2
US 12,682,595 · App. 18/540,267 · Granted Jul 14, 2026

System and method for transfer semantic segmentation via learnable image prompting of foundation models

Inventors: Jonathan Francis (Pittsburgh, PA); Rajshekhar Das (Pittsburgh, PA); Sanket Vaibhav Mehta (Pittsburgh, PA); Tanmay Kulkarni (Pittsburgh, PA)
Assignee: Robert Bosch GmbH
G06V10/26G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,682,595
App. No.
18/540,267
Filed
Dec 14, 2023
Granted
Jul 14, 2026
Kind
B2
Art Unit
2665
USPC
382/159
Abstract

A system includes a controller configured to receive one or more fixed text prompts and one or more images, wherein the fixed text prompts are associated with the one or more images. The controller is further configured to, in response to utilizing the fixed text prompt and the one or more images at a foundation model associated with a machine-learning network, output an intermediate representation from generating a series of objects and a task, decode the intermediate representation utilizing a decoder associated with the foundation model to generate a matrix associated with the task associated with the fixed text prompt and the image, and in response to identifying a highest probability associated with the matrix utilizing label selection, output a final label associated with a visual based prediction task.

Claims (35)

1 . A computer-implemented method, comprising:

receiving one or more fixed text prompts and one or more images, wherein the fixed text prompts are associated with the one or more images;

in response to utilizing the fixed text prompt and the one or more images at a foundation model associated with a machine-learning network, outputting an intermediate representation from generating a series of objects and a task, wherein the foundational model includes an encoder, a decoder, and a prediction head;

decoding, utilizing the decoder of the foundation model, the intermediate representation utilizing the decoder to generate a matrix associated with the task associated with the fixed text prompt and the image; and

in response to identifying a highest probability associated with the matrix utilizing label selection, outputting a final label associated with a visual based prediction task.

2 . The computer-implemented method of claim 1 , wherein the foundational model is a multimodal model.

3 . The computer-implemented method of claim 1 , wherein the method includes utilizing, at an image prompt network, the one or more images to generate a fixed-dimensional continuous latent vector.

4 . The computer-implemented method of claim 1 , wherein the decoder is utilizing a task-specific decoder.

5 . The computer-implemented method of claim 4 , wherein the method includes utilizing a task-specific encoder.

6 . The computer-implemented method of claim 1 , wherein the decoder is a task-specific decoder that includes a learnable functional map that utilizes an input as a vector from an intermediation latent representation space from an encoder.

7 . The computer-implemented method of claim 1 , wherein the visual based prediction task is not including in pretraining of the foundation model.

8 . The computer-implemented method of claim 1 , wherein the visual based prediction task is semantic segmentation and the final label includes a semantic segmentation image.

9 . A method, comprising:

receiving one or more fixed text prompts at a foundational model, wherein the foundational model includes an encoder, a decoder, and a prediction head;

receiving one or more images at a learnable image prompt network, wherein the fixed text prompts are associated with the one or more images;

generating a fixed-dimensional continuous latent vector at the learnable image prompt network utilizing the one or more images;

in response to utilizing the fixed text prompt and the fixed-dimensional continuous latent vector at the foundation model associated with a machine-learning network, outputting an intermediate representation from generating a series of objects and a task;

decoding the intermediate representation utilizing a decoder associated with the foundation model to generate a matrix associated with the task associated with the fixed text prompt and the image; and

in response to identifying a highest probability associated with the matrix utilizing label selection, outputting a final label associated with a visual based prediction task.

10 . The method of claim 9 , wherein the method further includes combining, utilizing a fusion model, representations from the foundation model with a task-specific representations from an encoder.

11 . The method of claim 9 , wherein the visual based prediction task is semantic segmentation and the final label includes a semantic segmentation image.

12 . The method of claim 9 , wherein the visual based prediction task is not including in pretraining of the foundation model.

13 . The method of claim 9 , wherein the task is a single task.

14 . The system of claim 9 , wherein the foundation model is configured to output an intermediate representation from one of a foundation model encoder, a foundation model decoder, a representation from an output of the head prediction head of the found model, or a combination of multiple foundation model prediction heads.

15 . A system, comprising:

a controller configured to:

receive one or more fixed text prompts and one or more images, wherein the fixed text prompts are associated with the one or more images;

in response to utilizing the fixed text prompt and the one or more images at a foundation model associated with a machine-learning network, output an intermediate representation from generating a series of objects and a task, wherein the foundational model includes an encoder a decoder, and a prediction head;

decode the intermediate representation utilizing the decoder of the foundation model to generate a matrix associated with the task associated with the fixed text prompt and the image; and

in response to identifying a highest probability associated with the matrix utilizing label selection, output a final label associated with a visual based prediction task.

16 . The system of claim 15 , wherein the visual based prediction task is semantic segmentation and the final label includes a semantic segmentation image.

17 . The system of claim 15 , wherein the visual based prediction task is not including in pretraining of the foundation model.

18 . The system of claim 15 , wherein the foundation model includes a language interface and visual interface.

19 . The system of claim 15 , wherein the system includes a task encoder configured to send one or more visual representations to the task decoder.

20 . The system of claim 15 , wherein the one or more images include a red green blue (RGB) image, sound image, video image, or radar image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2024
From: FRANCIS, JONATHAN; DAS, RAJSHEKHAR; MEHTA, SANKET VAIBHAV; KULKARNI, TANMAY
To: ROBERT BOSCH GMBH
Reel/Frame 066081/0259 →
Continuity (1)
Related Publication 20250200928A1 · Jun 19, 2025
References Cited (37)
US 10430946B1 · Zhou · 2019 [cited by examiner]
US 20210343014A1 · Haghighi · 2021 [cited by examiner]
US 20220391636A1 · Lian · 2022 [cited by examiner]
US 20230162481A1 · Yuan · 2023 [cited by examiner]
US 20230325725A1 · Lester · 2023 [cited by examiner]
US 20230351102A1 · Tran · 2023 [cited by examiner]
US 20240265718A1 · Zhang · 2024 [cited by examiner]
CN 116311254A · 2023 [cited by examiner]
CN 116311271A · 2023 [cited by examiner]
CN 116452472A · 2023 [cited by examiner]
CN 116933740A · 2023 [cited by examiner]
CN 117034933A · 2023 [cited by examiner]
CN 113657253B · 2023 [cited by examiner]
CN 117636473A · 2024 [cited by examiner]
Alberti, E., Tavera, A., Masone, C., and Caputo, B. IDDA: A large-scale multi-domain dataset for autonomous driving. IEEE Robotics and Automation Letters, 5(4):5526-5533, Oct. 2020. doi: 10.1109/Ira.2020.3009075. URL ht… [cited by applicant]
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. End-to-end object detection with transformers, 2020. URL https://arxiv.org/abs/2005.12872. [cited by applicant]
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. The cityscapes dataset for semantic urban scene understanding, 2016. URL https://arxiv.org/abs/1604.01685. [cited by applicant]
Devlin, J., Chang, M .- W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding, 2018. URL https://arxiv.org/abs/1810.04805. [cited by applicant]
Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V. Carla: An open urban driving simulator, 2017. URL https://arxiv.org/abs/1711.03938. [cited by applicant]
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for … [cited by applicant]
Ge, C., Huang, R., Xie, M., Lai, Z., Song, S., Li, S., and Huang, G. Domain adaptation via prompt learning, 2022. URL https://arxiv.org/abs/2202.06687. [cited by applicant]
Goodfellow, I. J., Mirza, M., Xiao, D., Courville, A., and Bengio, Y. An empirical investigation of catastrophic forgetting in gradient-based neural networks, 2013. URL https://arxiv.org/abs/1312.6211. [cited by applicant]
Herman, J., Francis, J., Ganju, S., Chen, B., Koul, A., Gupta, A., Skabelkin, A., Zhukov, I., Kumskoy, M., and Nyberg, E. Learn-to-race: A multimodal control environment for autonomous racing. In Proceedings of the IEEE… [cited by applicant]
Huo, X., Xie, L., Hu, H., Zhou, W., Li, H., and Tian, Q. Domain-agnostic prior for transfer semantic segmentation, 2022. URL https://arxiv.org/abs/2204.02684. [cited by applicant]
Langley, P. Crafting papers on machine learning. In Langley, P. (ed.), Proceedings of the 17th International Conference on Machine Learning (ICML 2000), pp. 1207-1216, Stanford, CA, 2000. Morgan Kaufmann. [cited by applicant]
Li, B., Weinberger, K. Q., Belongie, S., Koltun, V., and Ranftl, R. Language-driven semantic segmentation, 2022. URL https://arxiv.org/abs/2201.03546. [cited by applicant]
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586, 2021a. [cited by applicant]
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows, 2021b. URL https://arxiv.org/abs/2103.14030. [cited by applicant]
Lu, J., Clark, C., Zellers, R., Mottaghi, R., and Kembhavi, A. Unified-io: A unified model for vision, language, and multi-modal tasks. arXiv preprint arXiv:2206.08916, 2022. [cited by applicant]
Luddecke, T. and Ecker, A. S. Image segmentation using text and image prompts, 2021. URL https://arxiv.org/abs/2112.10003. [cited by applicant]
Mokady, R., Hertz, A., and Bermano, A. H. Clipcap: Clip prefix for image captioning, 2021. URL https://arxiv.org/abs/2111.09734. [cited by applicant]
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervisio… [cited by applicant]
Richter, S. R., Vineet, V., Roth, S., and Koltun, V. Playing for data: Ground truth from computer games. In Leibe, B., Matas, J., Sebe, N., and Welling, M. (eds.), European Conference on Computer Vision (ECCV), vol. 990… [cited by applicant]
Ros, G., Sellart, L., Materzynska, J., Vazquez, D., and Lopez, A. M. The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes. In 2016 IEEE Conference on Computer Vision and … [cited by applicant]
Sun, C., Shrivastava, A., Singh, S., and Gupta, A. Revisiting unreasonable effectiveness of data in deep learning era. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Oct. 2017. [cited by applicant]
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need, 2017. URL https://arxiv.org/abs/1706.03762. [cited by applicant]
Xu, J., De Mello, S., Liu, S., Byeon, W., Breuel, T., Kautz, J., & Wang, X. (2022). GroupViT: Semantic Segmentation Emerges from Text Supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern… [cited by applicant]