IP Library › Granted Patent US 12,657,400
Granted Patent B2
US 12,657,400 · App. 18/505,982 · Granted Jun 16, 2026

Systems and methods for vision-language model instruction tuning

Inventors: Wenliang Dai (Singapore, SG); Junnan Li (Singapore, SG); Chu Hong Hoi (Singapore, SG); Dongxu Li (Singapore, SG)
Assignee: Salesforce, Inc.
G06F40/40G06V10/774G06V10/82G06V20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,400
App. No.
18/505,982
Granted
Jun 16, 2026
Kind
B2
Abstract

Embodiments described herein provide a method of generating a vision-language task output to a text instruction relating to an input image, the method comprising receiving, via a data interface, the input image and the text instruction comprising an instruction relating to the image. The method further includes encoding, via an image encoder, the image into a first image representation. The method further includes generating, by a multimodal encoder, a second image representation based on cross-attending the first image representation to the text instruction. The method further includes generating, by a neural network based language model, a vision-language task output in response to the text instruction based on an input combining the second image representation and the text instruction.

Claims (57)

1 . A method of generating a vision-language task output for a text instruction relating to an input image, the method comprising:

receiving, via a data interface, the input image and the text instruction comprising an instruction relating to the input image;

encoding, via an image encoder, the input image into a first image representation;

generating, by a multimodal encoder connected to the image encoder, a second image representation based on cross-attending the first image representation to the text instruction by a cross-attention layer in the multimodal encoder; and

generating, by a neural network based language model connected to the multimodal encoder, the vision-language task output in response to the text instruction based on an input combining the second image representation and the text instruction.

2 . The method of claim 1 , further comprising:

receiving a known-good text output via the data interface;

computing a loss based on the vision-language task output and the known-good text output; and

training the multimodal encoder based on the loss.

3 . The method of claim 2 , further comprising:

keeping the image encoder and the neural network based language model frozen while training the multimodal encoder.

4 . The method of claim 2 , wherein the generating the second image representation is further based on a set of query vectors, further comprising:

updating the set of query vectors based on the vision-language task output and a known-good text output.

5 . The method of claim 1 , further comprising:

pre-training the multimodal encoder based on an output of the multimodal encoder without the neural network based language model.

6 . The method of claim 1 , further comprising:

adapting the second image representation for the neural network based language model via a feed forward neural network.

7 . The method of claim 1 , further comprising:

adapting the text instruction with an instruction template text.

8 . A system for generating a vision-language task output for a text instruction relating to an input image, the system comprising:

a memory that stores a neural network based language model and a plurality of processor executable instructions;

a data interface that receives the input image and the text instruction comprising an instruction relating to the input image; and

one or more hardware processors that read and execute the plurality of processor-executable instructions from the memory to perform operations comprising:

encoding, via an image encoder, the input image into a first image representation;

generating, by a multimodal encoder connected to the image encoder, a second image representation based on cross-attending the first image representation to the text instruction by a cross-attention layer in the multimodal encoder; and

generating, by the neural network based language model connected to the multimodal encoder, the vision-language task output in response to the text instruction based on an input combining the second image representation and the text instruction.

9 . The system of claim 8 , the operations further comprising:

receiving a known-good text output via the data interface;

computing a loss based on the vision-language task output and the known-good text output; and

training the multimodal encoder based on the loss.

10 . The system of claim 9 , the operations further comprising:

keeping the image encoder and the neural network based language model frozen while training the multimodal encoder.

11 . The system of claim 9 , wherein the generating the second image representation is further based on a set of query vectors, the operations further comprising:

updating the set of query vectors based on the vision-language task output and a known-good text output.

12 . The system of claim 8 , the operations further comprising:

pre-training the multimodal encoder based on an output of the multimodal encoder without the neural network based language model.

13 . The system of claim 8 , the operations further comprising:

adapting the second image representation for the neural network based language model via a feed forward neural network.

14 . The system of claim 8 , the operations further comprising:

adapting the text instruction with an instruction template text.

15 . A non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations comprising:

receiving, via a data interface, an input image and a text instruction comprising an instruction relating to the input image;

encoding, via an image encoder, the input image into a first image representation;

generating, by a multimodal encoder connected to the image encoder, a second image representation based on cross-attending the first image representation to the text instruction by a cross-attention layer in the multimodal encoder; and

generating, by a neural network based language model connected to the multimodal encoder, a vision-language task output in response to the text instruction based on an input combining the second image representation and the text instruction.

16 . The non-transitory machine-readable medium of claim 15 , the operations further comprising:

receiving a known-good text output via the data interface;

computing a loss based on the vision-language task output and the known-good text output; and

training the multimodal encoder based on the loss.

17 . The non-transitory machine-readable medium of claim 16 , the operations further comprising:

keeping the image encoder and the neural network based language model frozen while training the multimodal encoder.

18 . The non-transitory machine-readable medium of claim 16 , wherein the generating the second image representation is further based on a set of query vectors, the operations further comprising:

updating the set of query vectors based on the vision-language task output and a known-good text output.

19 . The non-transitory machine-readable medium of claim 15 , the operations further comprising:

pre-training the multimodal encoder based on an output of the multimodal encoder without the neural network based language model.

20 . The non-transitory machine-readable medium of claim 15 , the operations further comprising:

adapting the second image representation for the neural network based language model via a feed forward neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2026
From: DAI, WENLIANG; LI, JUNNAN; HOI, CHU HONG; LI, DONGXU
To: SALESFORCE, INC.
Reel/Frame 074067/0431 →
Continuity (4)
Continuation In Part 18160664 · Jan 27, 2023
Provisional Application 63500551 · May 5, 2023
Provisional Application 63424413 · Nov 10, 2022
Related Publication 20240160858A1 · May 16, 2024
References Cited (16)
US 12198048B2 · Singh et al. · 2025 [cited by applicant]
US 20230281400A1 · Zirui · 2023 [cited by applicant]
US 20230368510A1 · Chen et al. · 2023 [cited by applicant]
US 20240087265A1 · Park et al. · 2024 [cited by applicant]
US 20240282094A1 · Tsimpoukelli · 2024 [cited by examiner]
Alayrac et al., “Flamingo: a Visual Language Model for Few-Shot Learning”, DeepMind, Apr. 28, 2022., arXiv: 2204.14198v1, pp. 1-66. [cited by applicant]
Deyao Zhu et al: “MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 20, 2023 (Apr.… [cited by applicant]
Jean-Baptiste Alayrac et al: “Flamingo: a Visual Language Model for Few-Shot Learning”, A arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 29, 2022 (Apr. 29, 2022), XP091… [cited by applicant]
Junnan Li et al: “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, May… [cited by applicant]
International Search Report and Written Opinion mailed Jul. 19, 2024, International Patent Application No. PCT/US2024/027695, 110 pages. [cited by applicant]
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Jungi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. “Instructblip: Towards general-purpose vision-language models with instruction tuning,” 20… [cited by applicant]
Chen et al,“Subject-driven Text-to-Image Generation via Apprenticeship Learning”, arxiv (Cornell University), Apr. 14, 2023, pp. 1-18, XP093186354, DOI: 10.48550/arxiv.2304.00186 Retrieved from the Internet: URL:https:/… [cited by applicant]
Jia et al., “Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 5, 2023, XP091… [cited by applicant]
Li et al.,“BLIP-Diffusion: Pre-trained Subject Representation for Controllable Text-to-Image Generation and Editing”, arxiv.org, May 24, 2023, XP093180794, DOI: 10.48550/arxiv.2305.14720 Retrieved from the Internet: URL… [cited by applicant]
Ruiz et al.,“DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation”, arxiv (Cornell University), Mar. 15, 2023, XP093179964, DOI: 10.48550/arxiv.2208.12242 Retrieved from the Internet: URL… [cited by applicant]
International Search Report and Written Opinion for PCT/US2024/027830, dated Jul. 17, 2024, 14 pages. [cited by applicant]