IP Library Granted Patent US 12,374,099
Granted Patent B2
US 12,374,099 · App. 17/934,671 · Granted Jul 29, 2025

Systems and methods for visual question answering

Inventors: Anthony Meng Huat Tiong (Singapore, SG); Junnan Li (Singapore, SG); Chu Hong Hoi (Singapore, SG)
Assignee: Salesforce, Inc.
G06V10/86G06N3/045G06V10/26G06V10/774G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,374,099
App. No.
17/934,671
Granted
Jul 29, 2025
Kind
B2
Abstract

Embodiments described herein provide a zero-shot visual question answering (VQA) framework, which conjoins foundation network models with zero additional training. A first image and a question relating to the first image are received. The first image is divided into a plurality of image patches. A plurality of relevant image patches that are relevant to the question are determined, using a first neural network model, from the plurality of image patches. A plurality of image captions are generated, using a second neural network model, based on the plurality of relevant image patches. An answer to the question is generated based on the plurality of image captions.

Claims (52)

1. A method of zero-shot visual question answering, the method comprising:

receiving, via a data interface, a first image and a question relating to the first image;

dividing the first image into a plurality of image patches;

determining, using a first neural network model, relevance of each image patch of the plurality of image patches to the question;

generating, using a second neural network model, a plurality of image captions based on the relevance of each image patch to the question, wherein the generating the plurality of image captions includes:

sampling the plurality of image patches to generate a plurality of sampled image patches based on the relevance of each image patch; and

generating an image caption of the plurality of image captions based on the plurality of sampled image patches; and

generating, using a third neural network model, an answer in response to an input of the question and the plurality of image captions.

2. The method of claim 1 , wherein prior to receiving the first image and the question, each of the first neural network model, the second neural network model, and the third neural network model is pretrained using a separate training dataset.

3. The method of claim 1 , wherein the relevance of each image patch is determined based on a cross-attention score matrix of the plurality of image patches and the question.

4. The method of claim 1 , wherein the second neural network model uses stochastic decoding for generating the plurality of captions.

5. The method of claim 1 , wherein the third neural network model includes a question-answering encoder-decoder model.

6. The method of claim 1 , wherein the generating the answer further comprises:

for each image caption, concatenating the image caption with the question to generate a question and caption combination of a plurality of question and caption combinations;

encoding each of the plurality of question and caption combinations to generate a plurality of corresponding encoded representation;

combining the plurality of encoded representations to provide a concatenated encoded representation; and

generating the answer by decoding the concatenated encoded representation.

7. A non-transitory machine-readable medium comprising a plurality of machine-readable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform a method comprising:

receiving, via a data interface, a first image and a question relating to the first image;

dividing the first image into a plurality of image patches;

determining, using a first neural network model, relevance of each image patch of the plurality of image patches to the question;

generating, using a second neural network model, a plurality of image captions based on the relevance of each image patch to the question; and

generating, using a third neural network model, an answer in response to an input of the question and the plurality of image captions, wherein the generating the answer further comprises:

for each image caption, concatenating the image caption with the question to generate a question and caption combination of a plurality of question and caption combinations;

encoding each of the plurality of question and caption combinations to generate a plurality of corresponding encoded representation;

combining the plurality of encoded representations to provide a concatenated encoded representation; and

generating the answer by decoding the concatenated encoded representation.

8. The non-transitory machine-readable medium of claim 7 , wherein prior to receiving the first image and the question, each of the first neural network model, the second neural network model, and the third neural network model is pretrained using a separate training dataset.

9. The non-transitory machine-readable medium of claim 7 , wherein the relevance of each image patch is determined based on a cross-attention score matrix of the plurality of image patches and the question.

10. The non-transitory machine-readable medium of claim 7 , wherein the generating the plurality of image captions includes:

sampling the plurality of image patches to generate a plurality of sampled image patches based on the relevance of each image patch; and

generating an image caption of the plurality of image captions based on the plurality of sampled image patches.

11. The non-transitory machine-readable medium of claim 7 , wherein the second neural network model uses stochastic decoding for generating the plurality of captions.

12. The non-transitory machine-readable medium of claim 7 , wherein the third neural network model includes a question-answering encoder-decoder model.

13. A system, comprising:

a non-transitory memory; and

one or more hardware processors coupled to the non-transitory memory and configured to read instructions from the non-transitory memory to cause the system to perform a method comprising:

receiving, via a data interface, a first image and a question relating to the first image;

dividing the first image into a plurality of image patches;

determining, using a first neural network model, relevance of each image patch of the plurality of image patches to the question;

generating, using a second neural network model, a plurality of image captions based on the relevance of each image patch to the question, wherein the generating the plurality of image captions includes:

sampling the plurality of image patches to generate a plurality of sampled image patches based on the relevance of each image patch; and

generating an image caption of the plurality of image captions based on the plurality of sampled image patches; and

generating, using a third neural network model, an answer in response to an input of the question and the plurality of image captions.

14. The system of claim 13 , wherein prior to receiving the first image and the question, each of the first neural network model, the second neural network model, and the third neural network model is pretrained using a separate training dataset.

15. The system of claim 13 , wherein the relevance of each image patch is determined based on a cross-attention score matrix of the plurality of image patches and the question.

16. The system of claim 13 , wherein the second neural network model uses stochastic decoding for generating the plurality of captions.

17. The system of claim 13 , wherein the generating the answer further comprises:

for each image caption, concatenating the image caption with the question to generate a question and caption combination of a plurality of question and caption combinations;

encoding each of the plurality of question and caption combinations to generate a plurality of corresponding encoded representation;

combining the plurality of encoded representations to provide a concatenated encoded representation; and

generating the answer by decoding the concatenated encoded representation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2022
From: MENG HUAT TIONG, ANTHONY; LI, JUNNAN; HOI, CHU HONG
To: SALESFORCE, INC.
Reel/Frame 061192/0879 →
Continuity (2)
Provisional Application 63355298 · Jun 24, 2022
Related Publication 20230419652A1 · Dec 28, 2023
References Cited (24)
US 20180025271A1 · Sawada · 2018 [cited by examiner]
US 20180204111A1 · Zadeh · 2018 [cited by examiner]
US 20190042867A1 · Chen · 2019 [cited by examiner]
US 20190073353A1 · Yu · 2019 [cited by examiner]
US 20190130110A1 · Lee · 2019 [cited by examiner]
US 20190133510A1 · el Kaliouby · 2019 [cited by examiner]
US 20190171929A1 · Abadi · 2019 [cited by examiner]
US 20190180144A1 · Tsishkou · 2019 [cited by examiner]
US 20200020117A1 · Daehler · 2020 [cited by examiner]
US 20200104670A1 · Seo · 2020 [cited by examiner]
US 20210081715A1 · Rosman · 2021 [cited by examiner]
US 20210150118A1 · Le · 2021 [cited by examiner]
US 20210150350A1 · Gao · 2021 [cited by examiner]
US 20210216862A1 · Liu · 2021 [cited by examiner]
US 20210390700A1 · Lee · 2021 [cited by examiner]
US 20220121702A1 · Kale · 2022 [cited by examiner]
US 20220147838A1 · Gu · 2022 [cited by examiner]
US 20230154188A1 · Li · 2023 [cited by examiner]
US 20230237773A1 · Li · 2023 [cited by examiner]
US 20230281400A1 · Wang · 2023 [cited by examiner]
US 20230368529A1 · Wu · 2023 [cited by examiner]
Yunan Ye et al., Video question answering via grounded cross-attention network learning,Apr. 16, 2020, Information Processing and Management 57 (2020) 102265,pp. 1-10. [cited by examiner]
Vishvak Murahari et al., “Large-Scale Pretraining for Visual Dialog: A Simple State-of-the-Art Baseline,” Dec. 4, 2020, Computer Vision—ECCV 2020,pp. 336-349. [cited by examiner]
Jiasen Lu et al., “ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks,” Aug. 6, 2019, Computer Vision and Pattern Recognition,arXiv:1908.02265,pp. 1-8. [cited by examiner]