IP Library › Granted Patent US 12,198,048
Granted Patent B2
US 12,198,048 · App. 17/153,130 · Granted Jan 14, 2025

Modality adaptive information retrieval

Inventors: Hrituraj Singh (Uttar Pradesh, IN); Jatin Lamba (Haryana, IN); Denil Pareshbhai Mehta (Gujarat, IN); Balaji Vasan Srinivasan (Karnataka, IN); Anshul Nasery (Maharashtra, IN); Aishwarya Agarwal (Uttar Pradesh, IN)
Assignee: Adobe Inc.
G06N3/08G06F16/243G06F18/214G06F18/22G06F40/20G06N3/045G06V30/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,048
App. No.
17/153,130
Granted
Jan 14, 2025
Kind
B2
Abstract

In some embodiments, a multimodal computing system receives a query and identifies, from source documents, text passages and images that are relevant to the query. The multimodal computing system accesses a multimodal question-answering model that includes a textual stream of language models and a visual stream of language models. Each of the textual stream and the visual stream contains a set of transformer-based models and each transformer-based model includes a cross-attention layer using data generated by both the textual stream and visual stream of language models as an input. The multimodal computing system identifies text relevant to the query by applying the textual stream to the text passages and computes, using the visual stream, relevance scores of the images to the query, respectively. The multimodal computing system further generates a response to the query by including the text and/or an image according to the relevance scores.

Claims (67)

1. A computer-implemented method, comprising:

receiving and storing in a memory device, by a multimodal query subsystem, a text-based query;

converting a set of images into a first embedding representation and one or more source documents into a second embedding representation;

determining, based on a comparison of the first embedding representation and the second embedding representation, by the multimodal query subsystem in one or more source documents, a text passage and one or more images from the set of images that are relevant to the text-based query;

accessing, by the multimodal query subsystem, a multimodal question-answering model comprising (a) a first stream of language models comprising a first set of transformer-based models concatenated with each other and (b) a second stream of language models comprising a second set of transformer-based models concatenated with each other, wherein each transformer- based model comprises a respective cross-attention layer using data generated by both the first stream of language models and the second stream of language models as input;

generating, by the multimodal query subsystem, an indication of a portion of the text passage that is relevant to the text-based query by, at least, applying the first stream of language models to the text passage;

computing and storing in the memory device, by the multimodal query subsystem and with the second stream of language models, relevance scores of the text-based query for the one or more images, respectively, wherein the relevance scores are computed based on data received from the first stream of language models via cross-attention layers of the second stream of language models;

generating, by the multimodal query subsystem, a response to the text-based query comprising at least one of (a) the portion of the text passage or (b) an image among the one or more images from the set of images according to the respective relevance scores; and

transmitting the response for display on a display device.

2. The computer-implemented method of claim 1 , wherein determining the text passage and the one or more images that are relevant to the text- based query comprises:

determining the text passage from a text portion of the one or more source documents;

extracting multiple images from the one or more source documents;

determining a similarity between each of the multiple images and the text passage; and

identifying the one or more of images that are relevant to the text-based query as images among the multiple images having a similarity higher than a pre- determined threshold.

3. The computer-implemented method of claim 1 , wherein each of the first set of transformer-based models in the first stream of language models further comprises a self-attention layer using data generated by the first stream of language models as an input.

4. The computer-implemented method of claim 3 , wherein each of the first set of transformer-based models in the first stream of language models and the second set of transformer-based models in the second stream of language models further comprises a feedforward neural network.

5. A computer-implemented method for generating a multimodal question-answering model for providing a multimodal answer to a query, the method comprising:

a training data generation module configured for generating and storing in a memory device training data for the multimodal question-answering model, wherein the multimodal question-answering model comprises a first stream of language models configured for text content and a second stream of language models configured for image content, generating the training data comprising:

accessing a dataset comprising queries and text-based answers for the respective queries;

identifying, from the queries in the dataset, a query whose text-based answer is contained in a document including both textual content and image content;

converting the image content from the document into an embedding representation;

extracting the image content from the embedding representation;

determining a relevance score indicating a relevance of the image content to the query using one or more embedding representations of the image content, a caption of the image content, the text-based answer of the query, or a source passage containing the text-based answer in the document; and

generating an entry of the training data that includes the query, the source passage, the image content, the text-based answer, and the relevance score of the image content; and

a model training module configured for training the multimodal question-answering model using the training data.

6. The computer-implemented method of claim 5 , wherein determining the relevance score comprises:

calculating a proximity distance between the image content and the source passage using a number of tokens as a distance unit;

calculating a first term frequency-inverse document frequency (TF-IDF) score of the caption of the image content with the query;

calculating a second TF-IDF score of the caption of the image content with the text-based answer;

calculating a third TF-IDF score of the caption of the image content with the source passage; and

computing the relevance score of the image content by combining the proximity distance, the first TF-IDF score of the caption, the second TF-IDF score of the caption, and

the third TF-IDF score of the caption.

7. The computer-implemented method of claim 6 , wherein a loss function used in training the multimodal question-answering model comprises a loss term for the image content, the loss term for the image content comprising a weighted binary cross-entropy with a weight as the relevance score.

8. The computer-implemented method of claim 5 , wherein:

the first stream of language models is configured to accept a query and a text passage as input and output an indication of a portion of the text passage that is relevant to the query; and

the second stream of language models is configured to accept multiple images as input and output a relevance score to the query for each of the multiple images.

9. The computer-implemented method of claim 5 , wherein the second stream of language models comprising a plurality of transformer-based models concatenated with each other, each of the plurality of transformer-based models comprising a cross-attention layer that uses data generated by both the first stream of language models and the second stream of language models as an input.

10. The computer-implemented method of claim 5 , further comprising a pre- training module configured for:

pre-training the first stream of language models using a first training dataset comprising queries and corresponding text-based answers; and

pre-training the second stream of language models using a second training dataset, wherein training the multimodal question-answering model by the model training module comprises training the pre-trained first stream of language models and the pre-trained second stream of language models.

11. The computer-implemented method of claim 10 , wherein the pre-training module is further configured for generating the second training dataset by:

accessing an image and a caption of the image;

associating an irrelevant image with the caption of the image; and

generating an entry of the second training dataset, the entry comprising the caption of the image, the image, an indication that the image is relevant to the caption, the irrelevant image, and an additional indication indicating the caption is irrelevant to the irrelevant image.

12. The computer-implemented method of claim 5 , wherein the first stream of language models comprises a plurality of transformer-based models concatenated with each other, each of the plurality of transformer-based models comprising (i) a self-attention layer using data generated by the first stream of language models as input and (ii) a cross-attention layer using data generated by both the first stream of language models and the second stream of language models as inputs.

13. The computer-implemented method of claim 12 , wherein each of the plurality of transformer-based models in the first stream of language models and a second plurality of transformer-based models in the second stream of language models comprises a feedforward neural network.

14. The computer-implemented method of claim 5 , wherein training the multimodal question-answering model comprises adjusting parameters of the multimodal question-answering model to minimize a loss function, wherein the loss function comprises a first term calculated for the first stream of language models and a second loss term calculated for the second stream of language models.

15. A system, comprising:

one or more processing devices; and

a non-transitory computer-readable medium having program code that is stored thereon, the program code executable by one or more processing devices for performing operations comprising:

identifying, from one or more source documents, a text passage and a plurality of images that are relevant to a query;

converting the text passage into a first embedding representation and the plurality of images into a second embedding representation;

applying, to the first embedding representation, a first stream of language models from a multimodal question-answering model to obtain an indication of a portion of the text passage that is relevant to the query;

applying, to the second embedding representation, a second stream of language models from the multimodal question-answering model to obtain relevance scores of the query for the plurality of images, respectively, wherein each of the first stream of language models and the second stream of language models comprises a set of transformer-based models concatenated with each other, each of the set of transformer-based models comprising a cross-attention layer using data generated by both the first stream of language models and the second stream of language models as input; and

generating a response to the query comprising at least one of (a) the portion of the text passage or (b) an image among the plurality of images according to the respective relevance scores; and

transmitting the response for display on a display device.

16. The system of claim 15 , wherein determining the text passage and the plurality of images that are relevant to the query comprises:

determining the text passage from a text portion of the one or more source documents;

extracting multiple images from the one or more source documents;

determining a similarity between each of the multiple images and the text passage; and

identifying the plurality of images that are relevant to the query as images among the multiple images having a similarity higher than a pre-determined threshold.

17. The system of claim 16 , wherein determining a similarity between each of the multiple images and the text passage comprises:

generating an embedding for each of the multiple images;

generating an embedding for the text passage; and

calculating the similarity as a similarity between the embedding of each of the multiple images and the embedding for the text passage.

18. The system of claim 15 , wherein each of the transformer-based models in the first stream of language models further comprises a self-attention layer using data generated by the first stream of language models as an input.

19. The system of claim 15 , wherein each of the transformer-based models in the first stream of language models and the second stream of language models further comprises a feedforward neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2021
From: SINGH, HRITURAJ; LAMBA, JATIN; MEHTA, DENIL PARESHBHAI; SRINIVASAN, BALAJI VASAN; NASERY, ANSHUL; AGARWAL, AISHWARYA
To: ADOBE INC.
Reel/Frame 054966/0687 →
Continuity (1)
Related Publication 20220230061A1 · Jul 21, 2022
References Cited (55)
US 20210064879A1 · Gupta · 2021 [cited by examiner]
US 20220230061A1 · Singh · 2022 [cited by examiner]
US 20230116969A1 · Zhao · 2023 [cited by examiner]
US 20230280985A1 · Hayashi · 2023 [cited by examiner]
Kelvin Guu; Real: Retrieval-Augmented Language Model Pre-Training; PMLR; 2020. pp. 1-10 (Year: 2020). [cited by examiner]
Guido Zuccon; Integrating and Evaluating Neural Word Embeddings in Information Retrieval; ACM; pp. 1-8 (Year: 2015). [cited by examiner]
Nishida et al., Multi-Style Generative Reading Comprehension, In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy: Association for Computational Linguistics, 2019,… [cited by applicant]
Anderson et al., Bottom-up and Top-down Attention for Image Captioning and Visual Question Answering, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-23, 2018, pp. 6077-6086. [cited by applicant]
Antol et al., VQA: Visual Question Answering, Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 2425-2433. [cited by applicant]
Chaudhry et al., Leaf-qa: Locate, Encode & Attend for Figure Question Answering, 2020 IEEE Winter Conference on Applications of Computer Vision (WACV), Mar. 2020, pp. 1-10. [cited by applicant]
Chen et al., Uniter: Learning Universal Image-text Representations, arXiv:1909.11740 [cs. CV], Available Online at: https://arxiv.org/abs/1909.11740, 2019, 26 pages. [cited by applicant]
Cho et al., Adversarial Tableqa: Attention Supervision for Question Answering on Tables, Proceedings of Machine Learning Research, vol. 95, 2018, pp. 391-406. [cited by applicant]
Choi et al., QuAC: Question Answering in Context, Empirical Methods in Natural Language Processing., 2018, pp. 2174-2184. [cited by applicant]
Dale, Audiovisual Methods in Teaching, New York: Dryden, 1969, pp. 1-3. [cited by applicant]
Das et al., Visual Dialog, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 326-335. [cited by applicant]
Devlin et al., Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding, arXiv:1810.04805, Available Online at: https://arxiv.org/abs/1810.04805, 2018, 16 pages. [cited by applicant]
Fayek et al., Temporal Reasoning Via Audio Question Answering, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2020, 12 pages. [cited by applicant]
Feng et al., Visual Information in Semantic Representation, The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, Jun. 2010, pp. 91-99. [cited by applicant]
Goyal et al., Making the V in vqa Matter: Elevating the Role of Image Understanding in Visual Question Answering, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 6904-6913. [cited by applicant]
Hannan et al., Modality Disambiguation and qa Over Diverse Inputs, AAAI, Available Online at: https://arxiv.org/pdf/2001.08034.pdf, 2020, pp. 7879-7886. [cited by applicant]
Hermann et al., Teaching Machines to Read and Comprehend, NIPS'15: Proceedings of the 28th International Conference on Neural Information Processing Systems, vol. 1, Dec. 2015, pp. 1-14. [cited by applicant]
Hirschman et al., Deep Read: A Reading Comprehension System, Proceedings of the 37th annual meeting of the Association for Computational Linguistics, Jun. 1999, pp. 325-332. [cited by applicant]
Horev, Bert Explained: State of the Art Language Model for NLP, Towards Data Science, Available Online at: https://towardsdatascience.com/bert-explained-state-of-the-art-language-model-for-nlp-f8b21a9b6270, Nov. 10, 201… [cited by applicant]
Jauhar, Tables as Semi-structured Knowledge for Question Answering, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, vol. 1, 2016, pp. 474-483. [cited by applicant]
Kafle et al., Answering Questions About Data Visualizations Using Efficient Bimodal Fusion, The IEEE Winter Conference on Applications of Computer Vision, 2020, pp. 1498-1507. [cited by applicant]
Kafle et al., Dvqa: Understanding Data Visualizations via Question Answering, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 5648-5656. [cited by applicant]
Kahou et al., Figureqa: An Annotated Figure Dataset for Visual Reasoning, International Conference on Learning Representations, 2018, 20 pages. [cited by applicant]
Kocisky et al., The NarrativeQA Reading Comprehension Challenge, Transactions of the Association for Computational Linguistics, vol. 6, May 2018, pp. 317-328. [cited by applicant]
Kwiatkowski et al., Natural Questions: A Benchmark for Question Answering Research, Transactions of the Association of Computational Linguistics, 2019, 14 pages. [cited by applicant]
Lei et al., Tvqa: Localized, Compositional Video Question Answering, arXiv:1809.01696, Available Online at:, 2018, 13 pages. [cited by applicant]
Lu et al., Hierarchical Question-image Co-attention for Visual Question Answering, Advances in neural information processing systems, 2016, pp. 289-297. [cited by applicant]
Lu et al., Vilbert; Pretraining Task-agnostic Visi-olinguistic Representations for Vision-and-language Tasks, Advances in Neural Information Processing Systems, 2019, pp. 13-23. [cited by applicant]
Mihaylov et al., Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering, arXiv:1809.02789, Available Online at: https://arxiv.org/pdf/1809.02789.pdf, 2018, 14 pages. [cited by applicant]
Moreno et al., Interactive Multimodal Learning Environments, Educational psychology review, vol. 19, No. 3, 2007, pp. 309-326. [cited by applicant]
Li, et al., Multi-Modal Sentence Summarization With Modality Attention And Image Filtering, Proceedings of the Twenty-Seventh International Joint Conference on Artificial, 2018, pp. 4152-4158. [cited by applicant]
Nguyen et al., Ms Marco: A Human-generated Machine Reading Comprehension Dataset, arXiv:1611.09268, Available Online at: http://ceurws.org/Vol1773/CoCoNIPS_2016_paper9.pdf, 2016, 10 pages. [cited by applicant]
Nogueira et al., Passage Re-ranking with Bert, arXiv:1901.04085, 2019, 5 pages. [cited by applicant]
Pasupat et al., Compositional Semantic Parsing on Semi-structured Tables, arXiv:1508.0030, Available Online at: https://arxiv.org/pdf/1508.00305.pdf, 2015, 11 pages. [cited by applicant]
Rajpurkar et al., Know What You Don't Know: Unanswerable Questions for SQuAD, arXiv:1806.03822, Available Online at: https://arxiv.org/pdf/1806.03822.pdf, Jun. 11, 2018, 9 pages. [cited by applicant]
Rajpurkar et al., SQuAD: 100,000+ Questions for Machine Comprehension of Text, Proceedings of the Conference on Empirical Methods in Natural Language Processing, Available online at: https://doi.org/10.18653/v1/D16-1264… [cited by applicant]
Reddy et al., Coqa: A Conversational Question Answering Challenge, Transactions of the Association for Computational Linguistics, vol. 7, Mar. 2019, pp. 249-266. [cited by applicant]
Richardson et al., Mctest: A Challenge Dataset for the Open-domain Machine Comprehension of Text, Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Oct. 2013, pp. 193-203. [cited by applicant]
Sankey et al., Engaging Students Through Multimodal Learning Environments: The Journey Continues, Proceedings ASCILITE 2010: 27th annual conference of the Australasian Society for Computers in Learning in Tertiary Educa… [cited by applicant]
Sharma et al., Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset for Automatic Image Captioning, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, vol. 1, 2018, pp… [cited by applicant]
Silberer et al., Learning Grounded Meaning Representations with Autoencoders, Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics, vol. 1, Jun. 2014, pp. 721-732. [cited by applicant]
Simonyan et al., Very Deep Convolutional Networks for Large-Scale Image Recognition, ICLR, Available online at: https://arxiv.org/abs/1409.1556, Apr. 10, 2015, 14 pages. [cited by applicant]
Srivastava, Multimodal Learning with Deep Boltzmann Machines, Advances in neural information processing systems, 2012, pp. 2222-2230. [cited by applicant]
Tan et al., Lxmert: Learning Cross-modality Encoder Representations from Transformers, arXiv:1908.07490, Available Online at: https://arxiv.org/pdf/1908.07490.pdf, 2019, 14 pages. [cited by applicant]
Vaswani et al., Attention is all You Need, 31st Conference on Neural Information Processing Systems, Available Online at: arXiv.1706.03762v5, Dec. 6, 2017, 15 pages. [cited by applicant]
Welbl et al., Constructing Datasets for Multi-hop Reading Comprehension Across Documents, Transactions of the Association for Computational Linguistics, vol. 6, 2018, pp. 287-302. [cited by applicant]
Yagcioglu et al., Recipeqa: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Oct.-Nov. 2018, 11 pages. [cited by applicant]
Yang et al., Hotpotqa: A Dataset for Diverse, Explainable Multi-hop Question Answering, arXiv:1809.09600, Available Online at: https://arxiv.org/abs/1809.09600, 2018, 12 pages. [cited by applicant]
Zhu et al., MSMO: Multimodal Summarization with Multimodal Output, Proceedings of the 2018 conference on empirical methods in natural language processing, 2018, pp. 4154-4164. [cited by applicant]
Ren et al., Faster R-Cnn: Towards Real-Time Object Detection With Region Proposal Networks, In Advances in neural information processing systems, 2015 , pp. 1-9. [cited by applicant]
Bajaj et al., Ms Marco: A Human Generated Machine Reading Comprehension Dataset, 30th Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain, 2016, arXiv preprint arXiv:1611.09268, 2018, pp. 1… [cited by applicant]
Cited By (3)
US 12,536,725 US 12,657,400 US 12,705,910