IP Library › Granted Patent US 12,475,692
Granted Patent B2
US 12,475,692 · App. 18/306,071 · Granted Nov 18, 2025

Training multi-modal models on documents using multiple instance learning

Inventors: Amit Alfassy (Haifa, IL); Assaf Arbelle (Lehvot Haviva, IL); Leonid Karlinsky (Acton, MA)
Assignee: International Business Machines Corporation
G06V10/803G06V10/774G06V10/82G06V20/62G06V30/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,692
App. No.
18/306,071
Granted
Nov 18, 2025
Kind
B2
Abstract

An example system includes a processor to automatically extract text and images from a document. The processor can automatically generate text bags including a number of nearest texts for each of the extracted images. The processor can then train a multi-modal model based on the automatically generated text bags using a CLIP-MIL loss that computes, for each of the extracted images, a correlation between each of the different texts in the texts bags using a CLIP feature space at each gradient step of the gradient descent-based multiple instance learning (MIL) algorithm.

Claims (33)

1 . A computer system, comprising:

a processor set;

one or more computer-readable storage media; and

program instructions stored on the one or more computer-readable storage media to cause the processor set to perform operations comprising:

automatically extracting text and images from a document;

automatically generating positive text bags comprising a plurality of collocated texts for each of the extracted images in the document;

automatically generating negative text bags comprising a plurality of collocated texts for extracted images in other documents; and

training a multi-modal model based on the automatically generated positive text bags using a Contrastive Language-Image Pre-Training (CLIP) multiple instance learning (MIL) loss that computes, for each of the extracted images, a correlation between each of the texts in the positive text bags using a CLIP feature space at each gradient step of a gradient descent-based MIL algorithm, wherein the CLIP-MIL loss uses negative text bags and a MIL-Softmax loss function that comprises an average weighted by similarity over each positive text bag.

2 . The computer system of claim 1 , wherein the CLIP-MIL loss comprises a MIL-MAX loss function that selects a text with maximal similarity from each positive text bag associated with an image.

3 . The computer system of claim 1 , wherein the average is normalized using both positive text matches and negative text matches.

4 . The computer system of claim 1 , wherein the CLIP-MIL loss comprises MIL-NCE loss with a function is adapted to a CLIP contrastive loss using the image text matches and text images positive match bags and negative match bags.

5 . The computer system of claim 1 , wherein the operations further comprise executing a low-rank adaptation of large models (LoRA) to fine-tune the multi-modal model by reducing a number of trainable parameters during training.

6 . The computer system of claim 5 , wherein the LoRA is adapted to work on a residual neural network (ResNet) of the multi-modal model, wherein each of a linear layer, a convolutional layer, and an embedding layer of the ResNet is adapted into a LoRA layer by a addition of a set of trainable parameters.

7 . The computer system of claim 5 , wherein the LoRA is applied to all layers of the multi-modal model.

8 . The computer system of claim 1 , wherein the processor is to train the multi-modal model using image-texts bags comprising an image and a plurality of associated proximate texts.

9 . The computer system of claim 1 , wherein the operations further comprise training the multi-modal model using text-images bags comprising a text and a plurality of associated images.

10 . A computer-implemented method, comprising: automatically generating a plurality of positive image-texts bags including number of collocated extracted texts for each extracted image from a document;

automatically generating negative text bags comprising a plurality of collocated texts for extracted images in other documents; and

fine-tuning an unlocked subset of parameters of a pretrained multi-modal model based on generated positive text bags using a Contrastive Language-Image Pre-Training (CLIP) multiple instance learning (MIL) loss that computes, for each extracted image, a correlation between different texts in positive text bags using a CLIP feature space at each gradient step of gradient descent-based MIL algorithm, wherein pretrained parameters of the pretrained multi-modal model are locked during the fine-tuning, wherein the CLIP-MIL loss uses negative text bags and a MIL-Softmax loss function that comprises an average weighted by similarity over each positive text bag.

11 . The computer-implemented method of claim 10 , further comprising pretraining the multi-modal model on a large set of image-text pairs.

12 . The computer-implemented method of claim 10 , wherein fine-tuning the unlocked subset of parameters comprises using a CLIP-MIL Max loss that selects a text with a highest correlation for each extracted image from the document.

13 . The computer-implemented method of claim 10 , wherein fine-tuning the unlocked subset of parameters comprises using a CLIP-MIL NCE loss that calculates a weighted average between texts in each of a plurality of positive text bags.

14 . The computer-implemented method of claim 13 , comprising normalizing the weighted average using both positive text matches and negative text matches.

15 . The computer-implemented method of claim 10 , wherein fine-tuning the unlocked subset of parameters comprises executing a low-rank adaptation of large models (LoRA) to reduce a number of trainable parameters during fine-tuning.

16 . The computer-implemented method of claim 10 , comprising receiving a text as input at the multi-modal model and retrieving an associated image via the multi-modal model.

17 . The computer-implemented method of claim 10 , comprising receiving an image as input at the multi-modal model and retrieving an associated text via the multi-modal model.

18 . A computer program product comprising:

one or more computer-readable storage media; and

program instructions stored on the one or more computer-readable storage media to perform operations comprising:

automatically generate a plurality of positive image-texts bags including number of collocated extracted texts for each extracted image from a document;

automatically generating negative text bags comprising a plurality of collocated texts for extracted images in other documents; and

fine-tuning an unlocked subset of parameters of a pretrained multi-modal model based on generated positive text bags using a Contrastive Language-Image Pre-Training (CLIP) multiple instance learning (MIL) loss that computes, for each extracted image, a correlation between different texts in positive text bags using a CLIP feature space at each gradient step of gradient descent-based MIL algorithm, wherein pretrained parameters of the pretrained multi-modal model are locked during the fine-tuning, wherein the CLIP-MIL loss uses negative text bags and a MIL-Softmax loss function that comprises an average weighted by similarity over each positive text bag.

19 . The computer program product of claim 18 , wherein the operations further comprise pretraining the multi-modal model on a large set of image-text pairs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2023
From: ALFASSY, AMIT; ARBELLE, ASSAF; KARLINSKY, LEONID
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 063421/0675 →
Continuity (1)
Related Publication 20240355104A1 · Oct 24, 2024
References Cited (80)
US 9807473B2 · Mei et al. · 2017 [cited by applicant]
US 10360631B1 · Jezewski · 2019 [cited by examiner]
US 20140079297A1 · Tadayon · 2014 [cited by examiner]
US 20170193545A1 · Zhou · 2017 [cited by examiner]
US 20190147320A1 · Mattyus · 2019 [cited by examiner]
US 20230153522A1 · Cho · 2023 [cited by examiner]
US 20230161808A1 · Bursztyn · 2023 [cited by examiner]
US 20230196716A1 · Feng · 2023 [cited by examiner]
US 20230259787A1 · David · 2023 [cited by examiner]
US 20230298224A1 · Aggarwal · 2023 [cited by examiner]
US 20240160917A1 · Xue · 2024 [cited by examiner]
X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval, Ma et al. (Year: 2022). [cited by examiner]
LORA: Low-RANK Adaptation of Large Language Models, Hu et al. (Year: 2021). [cited by examiner]
Chao Jia, et al., “Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision”, Published In: arXiv, Jun. 11, 2021, 14 pages. [cited by applicant]
Edward Hu, et al., “LoRA: Low-Rank Adaptation of Large Language Models”, Published In ArXiv, Oct. 16, 2021, 26 pages. [cited by applicant]
Hanoona Rasheed, et al., “Fine-tuned CLIP Models are Efficient Video Learners”, Published In ArXiv, Dec. 6, 2022, 13 pages. [cited by applicant]
Hongwei Xue, et al., “CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment”, Published In ArXiv, Sep. 23, 2022, 11 pages. [cited by applicant]
Huansha Wang, et al., Research Progress on Vision-Language Multimodal Pretraining Technology, Published in MDPI, Oct. 31, 2022, 25 pages. [cited by applicant]
Thien-Tri Cao, et al., “HCMUS at MediaEval 2021: Fine-tuning CLIP for Automatic News-Images Re-Matching”, Published CEUR, Dec. 13-15, 2021, 2 pages. [cited by applicant]
Weijie Su, et al., “VL-BERT: Pre-Training of Generic Visual-Linguistic Representations”, Published In: arXiv, Feb. 18, 2020, 16 pages. [cited by applicant]
Miech et al. “End-to-End Learning of Visual Representations from Uncurated Instructional Videos”, arXiv:1912.06430, Dec. 13, 2019, 14 pages, https://doi.org/10.48550/arXiv.1912.06430. [cited by applicant]
Radford et al. “Learning Transferable Visual Models From Natural Language Supervision”, arXiv:2103.00020, Feb. 26, 2021, 47 pages, https://doi.org/10.48550/arXiv.2103.00020. [cited by applicant]
Wang et al., “Learning Deep Structure-Preserving Image-Text Embeddings”, arXiv:1511.06078v2 [cs.CV], Apr. 14, 2016, 9 pages. [cited by applicant]
Yao et al., “FILIP: Fine-grained Interactive Language-image Pre-training”, arXiv:2111.07783v1 [cs.CV], Nov. 9, 2021, 18 pages. [cited by applicant]
Ye et al., “Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection”, arXiv:1907.10164v3 [cs.CV], Aug. 16, 2019, 17 pages. [cited by applicant]
Yu et al., “Coca: Contrastive captioners are image-text foundation models”, arXiv:2205.01917v2 [cs.CV], Jun. 14, 2022, 19 pages. [cited by applicant]
Yuan et al., “Florence: A new foundation model for computer vision”, arXiv:2111.11432v1 [cs.CV], Nov. 22, 2021, 17 pages. [cited by applicant]
Zhai et al., “LiT : Zero-Shot Transfer with Locked-image text Tuning”, arXiv:2111.07991v3 [cs.CV], Jun. 22, 2022, 26 pages. [cited by applicant]
Zhang et al., “Contrastive Learning of Medical Visual Representations from Paired Images and Text”, arXiv:2010.00747v2 [cs.CV], Sep. 19, 2022, 24 pages. [cited by applicant]
Zhou et al., “Learning Deep Features for Discriminative Localization”, arXiv:1512.04150v1 [cs.CV], Dec. 14, 2015, 10 pages. [cited by applicant]
Alayrac et al., “Flamingo: a Visual Language Model for Few-Shot Learning”, arXiv:2204.14198v2 [cs.CV], Nov. 15, 2022, 54 pages. [cited by applicant]
Alfassy et al., “Feta: Towards specializing foundation models for expert task applications”, arXiv:2209.03648v2 [cs.CV], Dec. 19, 2022, 29 pages. [cited by applicant]
Andrews et al., “Support Vector Machines for Multiple-Instance Learning”, Advances in Neural Information Processing Systems 15 (NIPS 2002), Jan. 1, 2002, 8 pages. [cited by applicant]
Auer et al., “Delivering document conversion as a cloud service with high throughput and responsiveness”, arXiv:2206.00785v1 [cs.DL], Jun. 1, 2022, 11 pages. [cited by applicant]
Bach et al., “DIFFRAC : a discriminative and flexible framework for clustering”, NIPS'07: Proceedings of the 21st International Conference on Neural Information Processing System, Dec. 3, 2007, pp. 49-56. [cited by applicant]
Bojanowski et al., “Finding actors and actions in movies”, 2013 IEEE International Conference on Computer Vision, Dec. 2013, pp. 2280-2287. [cited by applicant]
Bommasani et al., “On the opportunities and risks of foundation models”, arXiv:2108.07258v3 [cs.LG], Jul. 12, 2022, 214 pages. [cited by applicant]
Brown et al., “Language models are few-shot learners”, arXiv:2005.14165v4 [cs.CL], Jul. 22, 2020, 75 pages. [cited by applicant]
Chéron et al., “A flexible model for training action localization with varying levels of supervision”, arXiv:1806.11328v2 [cs.CV], Nov. 27, 2018, 17 pages. [cited by applicant]
Desai et al., “VirTex: Learning Visual Representations from Textual Annotations”, arXiv:2006.06666v3 [cs.CV], Sep. 25, 2021, 17 pages. [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, arXiv:1810.04805v2 [cs.CL], May 24, 2019, 16 pages. [cited by applicant]
Dietterich et al., “Solving the multiple instance problem with axis-parallel rectangles”, Artificial Intelligence, Jan. 1997, pp. 31-71. [cited by applicant]
Geng et al., “Multimodal masked autoencoders learn transferable representations”, arXiv:2205.14204v3 [cs.CV], Oct. 21, 2022, 15 pages. [cited by applicant]
Github, “mifoundations / open_clip”, available online at <https://web.archive.org/web/20210729164745/https://github.com/mlfoundations/open_clip>, Jul. 29, 2021, 8 pages. [cited by applicant]
Hinton et al., “Distilling the Knowledge in a Neural Network”, arXiv:1503.02531 [stat.ML], Mar. 9, 2015, 9 pages. [cited by applicant]
Ilse et al., “Attention-based Deep Multiple Instance Learning”, arXiv:1802.04712v4 [cs.LG], Jun. 28, 2018, 16 pages. [cited by applicant]
Jerbi et al., “Learning Object Detection from Captions via Textual Scene Attributes”, arXiv:2009.14558v1 [cs.CV], Sep. 30, 2020, 10 pages. [cited by applicant]
Keeler et al., “Integrated segmentation and recognition of hand-printed numerals”, NIPS-3: Proceedings of the 1990 conference on Advances in neural information processing systems 3, Oct. 1, 1990, pp. 557-563. [cited by applicant]
Kim et al., “Vilt: Vision-and-language transformer without convolution or region supervision”, arXiv:2102.03334v2 [stat.ML], Jun. 10, 2021, 12 pages. [cited by applicant]
Kiros et al., “Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models”, arXiv:1411.2539v1 [cs.LG], Nov. 10, 2014, 13 pages. [cited by applicant]
Leung et al., “Handling Label Noise in Video Classification via Multiple Instance Learning”, ICCV '11: Proceedings of the 2011 International Conference on Computer Vision, Nov. 6, 2011, pp. 2056-2063. [cited by applicant]
Li et al., Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, arXiv:2107.07651v2 [cs.CV], Oct. 7, 2021, 16 pages. [cited by applicant]
Li et al., “BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation”, arXiv:2201.12086v2 [cs.CV], Feb. 15, 2022. [cited by applicant]
Li et al., “Supervision Exists Everywhere: A Data Efficient Contrastive Language-image Pre-training Paradigm”, arXiv:2110.05208v2 [cs.CV], Mar. 14, 2022. [cited by applicant]
Livathinos et al., “Robust pdf document conversion using recurrent neural networks”, arXiv:2102.09395v1 [cs.LG], Feb. 18, 2021, 9 pages. [cited by applicant]
Lu et al., “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks”, arXiv:1908.02265v1 [cs.CV], Aug. 6, 2019, 11 pages. [cited by applicant]
Maron et al., “A framework for multiple-instance learning”, NIPS '97: Proceedings of the 1997 conference on Advances in neural information processing systems 10, Jul. 31, 1998, pp. 570-576. [cited by applicant]
Miech et al., “HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips”, arXiv:1906.03327v2 [cs.CV], Jul. 31, 2019, 14 pages. [cited by applicant]
Miech et al., “Learning from Video and Text via Large-Scale Discriminative Clustering”, arXiv:1707.09074v1 [cs.CV], Jul. 27, 2017, 12 pages. [cited by applicant]
Nassar et al., “TableFormer: Table Structure Understanding with Transformers”, arXiv:2203.01017v2 [cs.CV], Mar. 11, 2022, 16 pages. [cited by applicant]
Nichol et al., “GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models”, arXiv:2112.10741v3 [cs.CV], Mar. 8, 2022, 20 pages. [cited by applicant]
No Author, “Deep Search Toolkit”, https://web.archive.org/web/20220714170549/https://github.com/DS4SD/deepsearch-toolkit, Jul. 14, 2022, 3 pages. [cited by applicant]
No Author, “East Automotive Archives”, https://web.archive.org/web/20230307201157/https://www.workshopservicemanual.com/, Mar. 7, 2023, 1 page. [cited by applicant]
No. Author, “OpenCLIP”, https://web.archive.org/web/20210729164745/https://github.com/mlfoundations/open_clip, dated Jun. 13, 2016, 8 pages. [cited by applicant]
Oquab et al., “Is object localization for free?—Weakly-supervised learning with convolutional neural networks”, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2015, 10 pages. [cited by applicant]
Pfitzmann et al., “DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis”, arXiv:2206.01062v1 [cs.CV], Jun. 2, 2022, 9 pages. [cited by applicant]
Quellec et al., “Multiple-Instance Learning for Medical Image and Video Analysis”, IEEE Reviews in Biomedical Engineering, Jan. 10, 2017, pp. 213-234. [cited by applicant]
Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”, arXiv:1910.10683v4 [cs.LG], Sep. 19, 2023, 67 pages. [cited by applicant]
Ramesh et al., “Zero-shot text-to-image generation”, arXiv:2102.12092v2 [cs.CV], Feb. 26, 2021, 20 pages. [cited by applicant]
Reed et al., “A generalist agent”, arXiv:2205.06175v3 [cs.AI], Nov. 11, 2022, 42 pages. [cited by applicant]
Saharia et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, arXiv:2205.11487v1 [cs.CV], May 23, 2022, 46 pages. [cited by applicant]
Schuhmann et al., “LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs”, arXiv:2111.02114v1 [cs.CV], Nov. 3, 2021, 5 pages. [cited by applicant]
Shapovalova et al., “Similarity Constrained Latent Support Vector Machine: An Application to Weakly Supervised Action Classification”, ECCV'12: Proceedings of the 12th European conference on Computer Vision—vol. Part VI… [cited by applicant]
Singh et al., “Flava: A foundational language and vision alignment model”, arXiv:2112.04482v3 [cs.CV], Mar. 29, 2022, 17 pages. [cited by applicant]
Sirinukunwattana et al., “Locality Sensitive Deep Learning for Detection and Classification of Nuclei in Routine Colon Cancer Histology Images”, IEEE Transactions on Medical Imaging, Feb. 4, 2016, pp. 1196-1206. [cited by applicant]
Staar et al., “Corpus Conversion Service: A Machine Learning Platform to Ingest Documents at Scale”, KDD '18: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Jul. 19, 20… [cited by applicant]
Sun et al., “Videobert: A joint model for video and language representation learning”, arXiv:1904.01766v2 [cs.CV], Sep. 11, 2019, 13 pages. [cited by applicant]
Thrush et al., “Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality”, arXiv:2204.03162v2 [cs.CV], Apr. 22, 2022, 18 pages. [cited by applicant]
Tishby et al., “Deep Learning and the Information Bottleneck Principle”, arXiv:1503.02406v1 [cs.LG], Mar. 9, 2015, 5 pages. [cited by applicant]
Vendrov et al., “Order-embeddings of Images and Language”, arXiv:1511.06361v6 [cs.LG], Mar. 1, 2016, 12 pages. [cited by applicant]