IP Library › Granted Patent US 12,511,876
Granted Patent B2
US 12,511,876 · App. 18/175,906 · Granted Dec 30, 2025

Single stream multi-level alignment for vision-language pretraining

Inventors: Vijay Kumar Baikampady Gopalkrishna (Santa Clara, CA); Xiang Yu (Mountain View, CA); Samuel Schulter (Long Island City, NY)
Assignee: NEC Corporation
G06V10/7715G06F40/205G06F40/284G06F40/30G06V10/774G06V10/776G06V10/806G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,876
App. No.
18/175,906
Granted
Dec 30, 2025
Kind
B2
Abstract

A method is provided for pretraining vision and language models that includes receiving image-text pairs, each including an image and a text describing the image. The method encodes an image into a set of feature vectors corresponding to input image patches and a CLS token which represents a global image feature. The method parses, by a text tokenizer, the text into a set of feature vectors as tokens for each word in the text. The method encodes the CLS token from the NN based visual encoder and the tokens from the text tokenizer into a set of features by a NN based text and multimodal encoder that shares weights for encoding both the CLS token and the tokens. The method accumulates the weights from multiple iterations as an exponential moving average of the weights during the pretraining until a predetermined error threshold is reduced to be under a threshold amount.

Claims (44)

1 . A computer-implemented method for pretraining vision and language models, comprising:

receiving image-text pairs, each including an image and a text describing the image;

encoding, by a neural network (NN) based visual encoder, an image into a set of feature vectors corresponding to input image patches and a Classification (CLS) token which represents a global feature of the image;

parsing, by a text tokenizer, the text into a set of feature vectors as tokens for each word in the text;

encoding the CLS token from the NN based visual encoder and the tokens from the text tokenizer into a set of features by a NN based text and multimodal encoder that shares weights for encoding both the CLS token and the tokens;

accumulating, by a NN based momentum encoder, the weights from multiple iterations as an exponential moving average of the weights during the pretraining until a predetermined error threshold is reduced to be under a threshold amount including minimizing a divergence between predicted pseudolabels and momentum pseudolabels that are concepts obtained by projecting the CLS token to a V-dimensional space with a fully-connected layer of the NN based momentum encoder to obtain a pre-trained NN-based visual and language model; and

controlling an autonomous vehicle to avoid collisions based on a trajectory generated with the pre-trained NN-based visual and language model that utilizes input images from a traffic scene.

2 . The computer-implemented method of claim 1 , wherein accumulating the weights further comprises performing for multiple iterations using an image-text contrastive loss to align image and text features, such that the image and the text features from a same pair are encouraged to be closer in a feature space and the image and the text features from a different pair are encouraged to be farther in the feature space.

3 . The computer-implemented method of claim 1 , wherein accumulating the weights further comprises performing for multiple iterations using an image-text matching loss implemented as a binary classification task that encourages a higher score for matching image-text pairs from the NN based text and multimodal encoder than for non-matching image-text pairs.

4 . The computer-implemented method of claim 1 , wherein wherein accumulating the weights further comprises performing for multiple iterations using a masked image and masked text modeling loss that comprises masking some parts of at least one of the image and the text to obtain masked parts and providing the masked parts to corresponding ones of the NN based visual encoder and the NN based text and multimodal encoder for reconstruction and minimization of a difference between the some parts and reconstruction versions of the some parts.

5 . The computer-implemented method of claim 1 , wherein accumulating the weights further comprises performing for multiple iterations using a concept alignment loss that encourages the NN based visual encoder to predict semantic concepts present in at least one of the image and the text by minimizing the divergence between the predicted pseudolabels and the momentum pseudolabels.

6 . The computer-implemented method of claim 1 , wherein the image-text pairs are noisy by at least one of missing one or more concepts, being abstract and being irrelevant.

7 . The computer-implemented method of claim 1 , wherein accumulating the weights further comprises masking the tokens from one modality selected from the image and the text and using cross-modal information to reconstruct masked tokens using an unmasked modality as context to obtain a fine-grained alignment between the image and the text.

8 . The computer-implemented method of claim 1 , further comprising:

computing a global alignment by capturing modality-invariant information; and

computing a loss value based on the modality-invariant information.

9 . A computer program product for pretraining vision and language models, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

receiving image-text pairs by a hardware processor of the computer, each including an image and a text describing the image;

encoding, by a neural network (NN) based visual encoder implemented by the hardware processor, an image into a set of feature vectors corresponding to input image patches and a Classification (CLS) token which represents a global feature of the image;

parsing, by a text tokenizer implemented by the hardware processor, the text into a set of feature vectors as tokens for each word in the text;

encoding the CLS token from the NN based visual encoder and the tokens from the text tokenizer into a set of features by a NN based text and multimodal encoder implemented by the hardware processor that shares weights for encoding both the CLS token and the tokens;

accumulating, by a NN based momentum encoder implemented by the hardware processor, the weights from multiple iterations as an exponential moving average of the weights during the pretraining until a predetermined error threshold is reduced to be under a threshold amount including minimizing a divergence between predicted pseudolabels and momentum pseudolabels that are concepts obtained by projecting the CLS token to a V-dimensional space with a fully-connected layer of the NN based momentum encoder to obtain a pre-trained NN-based visual and language model; and

controlling an autonomous vehicle to avoid collisions based on a trajectory generated with the pre-trained NN-based visual and language model that utilizes input images from a traffic scene.

10 . The computer program product of claim 9 , wherein accumulating the weights further comprises performing for multiple iterations using an image-text contrastive loss to align features of the image and the text, such that the image and the text features from a same pair are encouraged to be closer in a feature space and the image and the text features from a different pair are encouraged to be farther in the feature space.

11 . The computer program product of claim 9 , wherein accumulating the weights further comprises performing for multiple iterations using an image-text matching loss implemented as a binary classification task that encourages a higher score for matching image-text pairs from the NN based text and multimodal encoder than for non-matching image-text pairs.

12 . The computer program product of claim 9 , wherein accumulating the weights further comprises performing for multiple iterations using a masked image and masked text modeling loss that comprises masking some parts of at least one of the image and the text to obtain masked parts and providing the masked parts to corresponding ones of the NN based visual encoder and the NN based text and multimodal encoder for reconstruction and minimization of a difference between the some parts and reconstruction versions of the some parts.

13 . The computer program product of claim 9 , wherein accumulating the weights further comprises performing for multiple iterations using a concept alignment loss that encourages the NN based visual encoder to predict semantic concepts present in at least one of the image and the text.

14 . The computer program product of claim 9 , wherein the image-text pairs are noisy by at least one of missing one or more concepts, being abstract and being irrelevant.

15 . The computer program product of claim 9 , wherein accumulating the weights further comprises performing further comprises masking the tokens from one modality selected from the image and the text and using cross-modal information to reconstruct masked tokens to obtain a fine-grained alignment between the image and the text.

16 . The computer program product of claim 9 , further comprising:

computing a global alignment by capturing modality-invariant information; and

computing a loss value based on the modality-invariant information.

17 . A computer processing system, comprising:

a memory device for storing program code; and

a hardware processor operatively coupled to the memory device for running the program code for pretraining vision and language models including:

receive image-text pairs, each including an image and a text describing the image;

encode, by a neural network (NN) based visual encoder implemented by the hardware processor, an image into a set of feature vectors corresponding to input image patches and a Classification (CLS) token which represents a global feature of the image;

parse, by a text tokenizer implemented by the hardware processor, the text into a set of feature vectors as tokens for each word in the text;

encode the CLS token from the NN based visual encoder and the tokens from the text tokenizer into a set of features by a NN based text and multimodal encoder implemented by the hardware processor that shares weights for encoding both the CLS token and the tokens;

accumulate, by a NN based momentum encoder implemented by the hardware processor, the weights from multiple iterations as an exponential moving average of the weights during the pretraining until a predetermined error threshold is reduced to be under a threshold amount including minimizing a divergence between predicted pseudolabels and momentum pseudolabels that are concepts obtained by projecting the CLS token to a V-dimensional space with a fully-connected layer of the NN based momentum encoder to obtain a pre-trained NN-based visual and language model; and

control an autonomous vehicle to avoid collisions based on a trajectory generated with the pre-trained NN-based visual and language model that utilizes input images from a traffic scene.

18 . The computer processing system of claim 17 , wherein accumulating the weights further comprises performing for multiple iterations using an image-text contrastive loss to align features of the image and the text, such that the image and the text features from a same pair are encouraged to be closer in a feature space and the image and the text features from a different pair are encouraged to be farther in the feature space.

19 . The computer processing system of claim 17 , wherein accumulating the weights further comprises performing for multiple iterations using an image-text matching loss implemented as a binary classification task that encourages a higher score for matching image-text pairs from the NN based text and multimodal encoder than for non-matching image-text pairs.

20 . The computer processing system of claim 17 , wherein accumulating the weights further comprises performing for multiple iterations using a masked image and masked text modeling loss that comprises masking some parts of at least one of the image and the text to obtain masked parts and providing the masked parts to corresponding ones of the NN based visual encoder and the NN based text and multimodal encoder for reconstruction and minimization of a difference between the some parts and reconstruction versions of the some parts.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 072938/0913 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2023
From: GOPALKRISHNA, VIJAY KUMAR BAIKAMPADY; YU, XIANG; SCHULTER, SAMUEL
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 062828/0205 →
Continuity (2)
Provisional Application 63317499 · Mar 7, 2022
Related Publication 20230281963A1 · Sep 7, 2023
References Cited (48)
US 20210271707A1 · Lin · 2021 [cited by examiner]
US 20210294945A1 · Müller · 2021 [cited by examiner]
US 20220172080A1 · Chaudhury · 2022 [cited by examiner]
US 20230135659A1 · Wu · 2023 [cited by examiner]
US 20230154188A1 · Li · 2023 [cited by examiner]
US 20230162490A1 · Zhang · 2023 [cited by examiner]
US 20230237772A1 · Li · 2023 [cited by examiner]
US 20230237773A1 · Li · 2023 [cited by examiner]
US 20240371082A1 · Goyal · 2024 [cited by examiner]
Xiaosong Wang et al., ‘Self-supervised Image-text Pre-training With Mixed Data In Chest X-rays’, arXiv:2103.16022 <https://arxiv.org/pdf/2103.16022v1.pdf> (pp. 2-5; figures 2-3). [cited by applicant]
Youwei Liang et al., ‘Not All Patches Are What You Need: Expediting Vision Transformers Via Token Reorganizations’, arXiv:2202.07800v1, Feb. 2022 [retrieved on May 26, 2023]. Retrieved from <https://arxiv.org/pdf/2202.0… [cited by applicant]
Jing Yu et al., ‘Learning cross-modal correlations by exploring inter-word semantics and stacked coattention’, Pattern Recognition Letters, vol. 130, Feb. 2020 (pp. 193, 197; and figure 1). [cited by applicant]
Zaid Khan et al., ‘Single-Stream Multi-Level Alignment for Vision-Language Pretraining’, arXiv: 2203.14395v3, Jul. 2022 [retrieved on May 26, 2023]. Retrieved from <https://arxiv.org/pdf/2203.14395v3.pdf> (whole documen… [cited by applicant]
Bao, H., Dong, L., & Wei, F. (Jun. 15, 2021). Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254. [cited by applicant]
Chen, Y. C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., . . . & Liu, J. (Aug. 23, 2020). Uniter: Universal image-text representation learning. In European conference on computer vision (pp. 104-120). Springer, Ch… [cited by applicant]
Cho, J., Lei, J., Tan, H., & Bansal, M. (Jul. 1, 2021). Unifying vision-and-language tasks via text generation. In International Conference on Machine Learning (pp. 1931-1942). PMLR. [cited by applicant]
Cubuk, E. D., Zoph, B., Shlens, J., & Le, Q. V. (Jun. 13, 2020). Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF conference on computer vision and pattern r… [cited by applicant]
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (Oct. 11, 2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. [cited by applicant]
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., . . . & Houlsby, N. (Oct. 22, 2020). An image is worth 16×16 words: Transformers for image recognition at scale. arXiv preprint arX… [cited by applicant]
Gan, Z., Chen, Y. C., Li, L., Zhu, C., Cheng, Y., & Liu, J. (Dec. 6, 2020). Large-scale adversarial training for vision-and-language representation learning. Advances in Neural Information Processing Systems, 33, 6616-6… [cited by applicant]
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., & Parikh, D. (Jul. 22, 2017). Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference o… [cited by applicant]
He, K., Fan, H., Wu, Y., Xie, S., & Girshick, R. (Jun. 14, 2020). Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (… [cited by applicant]
Hinton, G., Vinyals, O., & Dean, J. (Mar. 9, 2015). Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7). [cited by applicant]
Jia, C., Yang, Y., Xia, Y., Chen, Y. T., Parekh, Z., Pham, H., . . . & Duerig, T. (Jul. 1, 2021). Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on… [cited by applicant]
Karpathy, A., & Fei-Fei, L. (Jun. 7, 2015). Deep visual-semantic alignments for generating image descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3128-3137). [cited by applicant]
Kim, J. H., Jun, J., & Zhang, B. T. (Dec. 3, 2018). Bilinear attention networks. Advances in neural information processing systems, 31. [cited by applicant]
Kim, W., Son, B., & Kim, I. (Jul. 1, 2021). Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning (pp. 5583-5594). PMLR. [cited by applicant]
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., . . . & Fei-Fei, L. (May 1, 2017). Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of … [cited by applicant]
Li, J., Selvaraju, R., Gotmare, A., Joty, S., Xiong, C., & Hoi, S. C. H. (Dec. 6, 2021). Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processi… [cited by applicant]
Li, L. H., Yatskar, M., Yin, D., Hsieh, C. J., & Chang, K. W. (Aug. 9, 2019). Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557. [cited by applicant]
Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., . . . & Gao, J. (Aug. 23, 2020). Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision (pp. 121-137). Sp… [cited by applicant]
Lin, T. Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., . . . & Zitnick, C. L. (Sep. 6, 2014). Microsoft coco: Common objects in context. In European conference on computer vision (pp. 740-755). Springer… [cited by applicant]
Liu, X., Li, L., Wang, S., Zha, Z. J., Meng, D., & Huang, Q. (Jun. 15, 2019). Adaptive reconstruction network for weakly supervised referring expression grounding. In Proceedings of the IEEE/CVF International Conference… [cited by applicant]
Loshchilov, I., & Hutter, F. (Nov. 14, 2017). Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101. [cited by applicant]
Lu, J., Batra, D., Parikh, D., & Lee, S. (Dec. 8, 2019). Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32. [cited by applicant]
Lu, J., Goswami, V., Rohrbach, M., Parikh, D., & Lee, S. (Jun. 13, 2020). 12-in-1: Multi-task vision and language representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni… [cited by applicant]
Van den Oord, A., Li, Y., & Vinyals, O. (Jan. 22, 2018). Representation learning with contrastive predictive coding. CoRR. arXiv preprint arXiv:1807.03748. [cited by applicant]
Van Den Oord, A., & Vinyals, O. (2017, Dec. 4, 2017). Neural discrete representation learning. Advances in neural Information processing systems, 30. [cited by applicant]
Ordonez, V., Kulkarni, G., & Berg, T. (Dec. 12, 2011). Im2text: Describing images using 1 million captioned photographs. Advances in neural information processing systems, 24. [cited by applicant]
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., & Lazebnik, S. (Dec. 11, 2015). Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In Procee… [cited by applicant]
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., . . . & Sutskever, I. (Jul. 1, 2021). Learning transferable visual models from natural language supervision. In International Conference on Machine… [cited by applicant]
Sharma, P., Ding, N., Goodman, S., & Soricut, R. (Jul. 15, 2018). Conceptual captions: A cleaned, hypernymed, Image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Assoc… [cited by applicant]
Su, W., Zhu, X., Cao, Y., Li, B., Lu, L., Wei, F., & Dai, J. (Aug. 22, 2019). VI-bert: Pre-training of generic visual-inguistic representations. arXiv preprint arXiv:1908.08530. [cited by applicant]
Tan, H., & Bansal, M. (Aug. 20, 2019). Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490. [cited by applicant]
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., . . . & Polosukhin, I. (Dec. 4, 2017). Attention is all you need. Advances in neural information processing systems, 30. [cited by applicant]
Xie, N., Lai, F., Doran, D., & Kadav, A. (Jan. 20, 2019). Visual entailment: A novel task for fine-grained image understanding. arXiv preprint arXiv:1901.06706. [cited by applicant]
Yu, L., Poirson, P., Yang, S., Berg, A. C., & Berg, T. L. (Oct. 8, 2016). Modeling context in referring expressions. In European Conference on Computer Vision (pp. 69-85). Springer, Cham. [cited by applicant]
Zhang, Z., Zhao, Z., Lin, Z., & He, X. (Dec. 6, 2020). Counterfactual contrastive learning for weakly-supervised vision-language grounding. Advances in Neural Information Processing Systems, 33, 18123-18134. [cited by applicant]