IP Library Granted Patent US 12,288,380
Granted Patent B2
US 12,288,380 · App. 17/745,634 · Granted Apr 29, 2025

Systems and methods for unified vision-language understanding and generation

Inventors: Junnan Li (Singapore, SG); Chu Hong Hoi (Singapore, SG)
Assignee: Salesforce, Inc.
G06V10/774G06F40/126G06F40/284G06T9/00G06V10/764G06V10/803
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,288,380
App. No.
17/745,634
Granted
Apr 29, 2025
Kind
B2
Abstract

Embodiments described herein provide systems, methods, and devices for generating enhanced vison-language training data. A method may include: receiving, from a communication interface, a first training dataset of image-text pairs and a second training dataset of annotated image-text pairs; fine-tuning an image-grounded text decoder and an image-grounded text encoder using the second training dataset of annotated image-text pairs; generating, by the fine-tuned image-grounded text decoder, a predicted text based on a training image from the first training dataset; generating, by the fine-tuned image-grounded text encoder, a filtering decision based on the training image and the predicted text; adding the training image and the predicted text to form a third training dataset of image-text pairs depending on the filter decision; and training a vision-language model using the third training dataset of image-text pairs.

Claims (58)

1. A method of generating enhanced vison-language training data, the method comprising:

receiving, from a communication interface, a first training dataset of image-text pairs and a second training dataset of annotated image-text pairs;

fine-tuning an image-grounded text decoder and an image-grounded text encoder using the second training dataset of annotated image-text pairs;

generating, by the fine-tuned image-grounded text decoder, a predicted text based on a training image from the first training dataset;

generating, by the fine-tuned image-grounded text encoder, a filtering decision based on the training image and the predicted text;

adding the training image and the predicted text to form a third training dataset of image-text pairs depending on the filter decision; and

training a vision-language model using the third training dataset of image-text pairs.

2. The method of claim 1 , wherein the image-grounded text decoder is finetuned by:

generating a predicted text in response to an image in the second training dataset; and

computing a language modeling loss comparing the predicted text with an annotated text paired with the image.

3. The method of claim 1 , wherein the image-grounded text encoder is finetuned by:

generating a text encoding of a text from the second training dataset;

generating an image encoding of an image pairing the text from the second training dataset; and

computing an image-text contrastive loss based on a positive pair of the text encoding and the image encoding, and negative pairs of the image encoding paired with other text encodings.

4. The method of claim 1 , wherein the image-grounded text encoder is finetuned by:

generating a binary classification indicating whether a text and an image from the second training dataset are a match; and

computing an image-text matching loss comparing the binary classification and a ground truth.

5. The method of claim 1 , wherein the filter decision is generated by generating a binary classification indicating whether an input image and an input text matches.

6. The method of claim 5 , wherein the third training dataset is formed by adding the input image and the input text as a training pair when the binary classification indicates the input image and the input text match.

7. The method of claim 1 , further comprising:

adding the second training dataset to the third training dataset.

8. The method of claim 1 , wherein the vision-language model includes any combination of an image encoder, a text encoder, an image-grounded text encoder or an image grounded text decoder.

9. The method of claim 1 , wherein the image-grounded text decoder and the image-grounded text encoder at least partially share parameters.

10. A system of generating enhanced vison-language training data, the system comprising:

a communication interface that receives a first training dataset of image-text pairs and a second training dataset of annotated image-text pairs;

a memory storing a plurality of processor-executable instructions; and

a processor executing the plurality of processor-executable instructions to perform operations comprising:

fine-tuning an image-grounded text decoder and an image-grounded text encoder using the second training dataset of annotated image-text pairs;

generating, by the fine-tuned image-grounded text decoder, a predicted text based on a training image from the first training dataset;

generating, by the fine-tuned image-grounded text encoder, a filtering decision based on the training image and the predicted text;

adding the training image and the predicted text to form a third training dataset of image-text pairs depending on the filter decision; and

training a vision-language model using the third training dataset of image-text pairs.

11. The system of claim 10 , wherein the image-grounded text decoder is finetuned by:

generating a predicted text in response to an image in the second training dataset; and

computing a language modeling loss comparing the predicted text with an annotated text paired with the image.

12. The system of claim 10 , wherein the image-grounded text encoder is finetuned by:

generating a text encoding of a text from the second training dataset;

generating an image encoding of an image pairing the text from the second training dataset; and

computing an image-text contrastive loss based on a positive pair of the text encoding and the image encoding, and negative pairs of the image encoding paired with other text encodings.

13. The system of claim 10 , wherein the image-grounded text encoder is finetuned by:

generating a binary classification indicating whether a text and an image from the second training dataset are a match; and

computing an image-text matching loss comparing the binary classification and a ground truth.

14. The system of claim 10 , wherein the filter decision is generated by generating a binary classification indicating whether an input image and an input text matches.

15. The system of claim 14 , wherein the third training dataset is formed by adding the input image and the input text as a training pair when the binary classification indicates the input image and the input text match.

16. The system of claim 10 , wherein the operations further comprise:

adding the second training dataset to the third training dataset.

17. The system of claim 10 , wherein the vision-language model includes any combination of an image encoder, a text encoder, an image-grounded text encoder or an image grounded text decoder.

18. A non-transitory processor-readable storage medium storing a plurality of processor-executable instructions for generating enhanced vison-language training data, the instructions being executed by a processor to perform operations comprising:

receiving, from a communication interface, a first training dataset of image-text pairs and a second training dataset of annotated image-text pairs;

fine-tuning an image-grounded text decoder and an image-grounded text encoder using the second training dataset of annotated image-text pairs;

generating, by the fine-tuned image-grounded text decoder, a predicted text based on a training image from the first training dataset;

generating, by the fine-tuned image-grounded text encoder, a filtering decision based on the training image and the predicted text;

adding the training image and the predicted text to form a third training dataset of image-text pairs depending on the filter decision; and

training a vision-language model using the third training dataset of image-text pairs.

19. The non-transitory processor-readable storage medium of claim 18 , wherein the image-grounded text encoder is finetuned by:

generating a binary classification indicating whether a text and an image from the second training dataset are a match; and

computing an image-text matching loss comparing the binary classification and a ground truth.

20. The non-transitory processor-readable storage medium of claim 18 , wherein the vision-language model includes any combination of an image encoder, a text encoder, an image-grounded text encoder or an image grounded text decoder.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2022
From: LI, JUNNAN; HOI, CHU HONG
To: SALESFORCE, INC
Reel/Frame 060837/0210 →
Continuity (2)
Provisional Application 63301978 · Jan 21, 2022
Related Publication 20230237773A1 · Jul 27, 2023
References Cited (17)
US 20210082398A1 · Hori et al. · 2021 [cited by applicant]
US 20220147838A1 · Gu · 2022 [cited by examiner]
US 20220156527A1 · Ramasamy et al. · 2022 [cited by applicant]
US 20220207001A1 · Ong et al. · 2022 [cited by applicant]
US 20230081171A1 · Zhang · 2023 [cited by examiner]
US 20230124389A1 · Wang et al. · 2023 [cited by applicant]
US 20230177810A1 · Xu et al. · 2023 [cited by applicant]
US 20230222285A1 · Zhang · 2023 [cited by examiner]
US 20230281400A1 · Wang · 2023 [cited by examiner]
US 20230281963A1 · Gopalkrishna · 2023 [cited by examiner]
US 20230351149A1 · Yu · 2023 [cited by examiner]
US 20240185602A1 · Liu · 2024 [cited by examiner]
CN 113792113A · 2021 [cited by applicant]
Li, Junnan, et al. “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.” International conference on machine learning. PMLR, Feb. 15, 2022 (Year: 2022). [cited by examiner]
Gan, Zhe, et al. “Vision-language pre-training: Basics, recent advances, and future trends.” Foundations and Trends® in Computer Graphics and Vision 14.3-4 (2022): 163-352. (Year: 2022). [cited by examiner]
Cho, Jaemin, et al. “Unifying vision-and-language tasks via text generation.” International Conference on Machine Learning. PMLR, 2021. (Year: 2021). [cited by examiner]
C. Zhang, Z. Yang, X. He and L. Deng, “Multimodal Intelligence: Representation Learning, Information Fusion, and Applications,” in IEEE Journal of Selected Topics in Signal Processing, vol. 14, No. 3, pp. 478-493, Mar. … [cited by examiner]
Cited By (1)
US 12,482,231