IP Library › Patent Application 19183526
Patent Application
App. No. 19/183,526

SYSTEMS AND METHODS FOR UNIFIED VISION-LANGUAGE UNDERSTANDING AND GENERATION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/183,526
Abstract

Embodiments described herein provide bootstrapping language-images pre-training for unified vision-language understanding and generation (BLIP), a unified VLP framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP enables a wider range of downstream tasks, improving on both shortcomings of existing models.

Claims (68)

1 - 20 . (canceled)

21 . A method of generating a text description for an input image, the method comprising:

training a multi-modal model implemented one or more hardware processors using a training dataset of images and noisy text captions;

loading, from the trained multi-modal model, a first set of parameters and a second set of parameters into an image-grounded text decoder and an image-grounded text encoder, respectively;

generating, by the image-grounded text decoder, a predicted text description based on a training image from the training dataset;

generating, by the image-grounded text encoder, a filtering decision based on the training image, a corresponding noisy text caption and the predicted text description;

updating the training dataset including removing the corresponding noisy text caption depending on the filter decision;

re-training the multi-modal model using the updated training dataset; and

generating, by the re-trained multi-modal model, the text description in response to the input image.

22 . The method of claim 21 , further comprising:

fine-tuning the image-grounded text decoder and/or the image-grounded text encoder using annotated image-text pairs.

23 . The method of claim 21 , wherein the training dataset of images and noisy text captions are obtained from web images and corresponding texts.

24 . The method of claim 22 , wherein the image-grounded text decoder is finetuned by:

generating a predicted text in response to an image in an annotated image-text pair; and

computing a language modeling loss comparing the predicted text with an annotated text paired with the image.

25 . The method of claim 22 , wherein the image-grounded text encoder is finetuned by:

generating a text encoding of a text from an annotated image-text pair;

generating an image encoding of an image pairing the text;

computing an image-text contrastive loss based on a positive pair of the text encoding and the image encoding, and negative pairs of the image encoding paired with other text encodings.

26 . The method of claim 22 , wherein the image-grounded text encoder is finetuned by:

generating a binary classification indicating whether a text and an image from an annotated image-text pair are a match; and

computing an image-text matching loss comparing the binary classification and a ground truth.

27 . The method of claim 21 , wherein the filter decision is generated by a binary classification indicating whether the training image and the predicted text description matches.

28 . The method of claim 27 , wherein updating the training dataset further includes:

adding the training image and the predicted text description as a training pair when the binary classification indicates a match.

29 . The method of claim 21 , further comprising:

removing the corresponding noisy text caption from the training dataset when a binary classification indicates the corresponding noisy text caption and the training image do not match.

30 . The method of claim 21 , further comprising performing, using the re-trained multi-modal model, vision-language tasks including one or more of:

image-to-text retrieval;

text-to-image retrieval;

image captioning; and

visual question answering.

31 . A system of generating a text description for an input image, the system comprising:

a memory storing a multi-modal model, an image-grounded text decoder, an image-grounded text encoder, and a plurality of processor-executed instructions; and

one or more hardware processors reading and executing plurality of processor-executed instructions to perform operations including:

training the multi-modal model implemented one or more hardware processors using a training dataset of images and noisy text captions;

loading, from the trained multi-modal model, a first set of parameters and a second set of parameters into the image-grounded text decoder and the image-grounded text encoder, respectively;

generating, by the image-grounded text decoder, a predicted text description based on a training image from the training dataset;

generating, by the image-grounded text encoder, a filtering decision based on the training image, a corresponding noisy text caption and the predicted text description;

updating the training dataset including removing the corresponding noisy text caption depending on the filter decision;

re-training the multi-modal model using the updated training dataset; and

generating, by the re-trained multi-modal model, the text description in response to the input image.

32 . The system of claim 31 , wherein the operations further comprise:

fine-tuning the image-grounded text decoder and/or the image-grounded text encoder using annotated image-text pairs.

33 . The system of claim 31 , wherein the training dataset of images and noisy text captions are obtained from web images and corresponding texts.

34 . The system of claim 32 , wherein the image-grounded text decoder is finetuned by:

generating a predicted text in response to an image in an annotated image-text pair; and

computing a language modeling loss comparing the predicted text with an annotated text paired with the image.

35 . The system of claim 32 , wherein the image-grounded text encoder is finetuned by:

generating a text encoding of a text from an annotated image-text pair;

generating an image encoding of an image pairing the text;

computing an image-text contrastive loss based on a positive pair of the text encoding and the image encoding, and negative pairs of the image encoding paired with other text encodings.

36 . The system of claim 32 , wherein the image-grounded text encoder is finetuned by:

generating a binary classification indicating whether a text and an image from an annotated image-text pair are a match; and

computing an image-text matching loss comparing the binary classification and a ground truth.

37 . The system of claim 31 , wherein the filter decision is generated by a binary classification indicating whether the training image and the predicted text description matches.

38 . The system of claim 37 , wherein the operation of updating the training dataset further includes:

adding the training image and the predicted text description as a training pair when the binary classification indicates a match.

39 . The system of claim 31 , wherein the operations further comprise:

removing the corresponding noisy text caption from the training dataset when a binary classification indicates the corresponding noisy text caption and the training image do not match.

40 . A non-transitory processor-readable storage medium storing a plurality of processor-executable instructions for generating a text description for an input image, the instructions being executed by a processor to perform operations comprising:

training a multi-modal model implemented one or more hardware processors using a training dataset of images and noisy text captions;

loading, from the trained multi-modal model, a first set of parameters and a second set of parameters into an image-grounded text decoder and an image-grounded text encoder, respectively;

generating, by the image-grounded text decoder, a predicted text description based on a training image from the training dataset;

generating, by the image-grounded text encoder, a filtering decision based on the training image, a corresponding noisy text caption and the predicted text description;

updating the training dataset including removing the corresponding noisy text caption depending on the filter decision;

re-training the multi-modal model using the updated training dataset; and

generating, by the re-trained multi-modal model, the text description in response to the input image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2025
From: LI, JUNNAN; HOI, CHU HONG
To: SALESFORCE, INC.
Reel/Frame 070897/0989 →