IP Library Granted Patent US 12,223,439
Granted Patent B2
US 12,223,439 · App. 17/190,668 · Granted Feb 11, 2025

Visual-semantic representation learning via multi-modal contrastive training

Inventors: Xin Yuan (Chicago, IL); Zhe Lin (Bellevue, WA); Jason Wen Yong Kuen (Santa Clara, CA); Jianming Zhang (Campbell, CA); Yilin Wang (Sunnyvale, CA); Ajinkya Kale (San Jose, CA); Baldo Faieta (San Francisco, CA)
Assignee: ADOBE INC.
G06N5/04G06N3/08G06N20/00G06T2207/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,439
App. No.
17/190,668
Granted
Feb 11, 2025
Kind
B2
Abstract

Systems and methods for multi-modal representation learning are described. One or more embodiments provide a visual representation learning system trained using machine learning techniques. For example, some embodiments of the visual representation learning system are trained using cross-modal training tasks including a combination of intra-modal and inter-modal similarity preservation objectives. In some examples, the training tasks are based on contrastive learning techniques.

Claims (47)

1. A method of training a machine learning model, the method comprising:

identifying a training set comprising a plurality of images and a plurality of captions corresponding to the images;

encoding the images using an image encoder to produce encoded images;

encoding the captions using a text encoder to produce encoded text;

computing a multi-modal loss function based on the encoded images and the encoded text, the multi-modal loss function comprising at least one image loss term, at least one text loss term, and at least one cross-modal term; and

training the image encoder and the text encoder based on the multi-modal loss function.

2. The method of claim 1 , further comprising:

computing an image self-supervised contrastive loss, wherein the at least one image loss term includes the image self-supervised contrastive loss.

3. The method of claim 1 , further comprising:

computing a tag-supervised contrastive loss, wherein the at least one image loss term includes the image tag-supervised contrastive loss.

4. The method of claim 1 , further comprising:

computing a caption self-supervised contrastive loss, wherein the at least one text loss term includes the caption self-supervised contrastive loss.

5. The method of claim 1 , further comprising:

computing a caption-image contrastive loss, wherein the at least one cross-modal term includes the caption-image contrastive loss.

6. The method of claim 1 , further comprising:

computing an image-caption contrastive loss, wherein the at least one cross-modal term includes the image-caption contrastive loss.

7. The method of claim 1 , further comprising:

encoding the images using a momentum image encoder to produce momentum encoded images, wherein the at least one image loss term is based on the encoded images and the momentum encoded images.

8. The method of claim 7 , wherein:

the at least one cross-modal term is based on the momentum encoded images and the encoded text.

9. The method of claim 1 , further comprising:

encoding the captions using a momentum text encoder to produce momentum encoded text, wherein the at least one text loss term is based on the encoded text and the momentum encoded text.

10. The method of claim 9 , wherein:

the at least one cross-modal term is based on the encoded images and the momentum encoded text.

11. The method of claim 1 , wherein:

the multi-modal loss function is based on a contrastive learning framework.

12. The method of claim 1 , wherein:

the encoded images and the encoded text are represented in a same embedding space.

13. The method of claim 1 , further comprising:

adjusting one or more of the images to produce an augmented training set, wherein the training is based on the augmented training set.

14. An apparatus comprising:

an image encoder configured to encode images to produce encoded images;

a text encoder configured to encode captions corresponding to the images to produce encoded text; and

a training component configured to compute a multi-modal loss function based on the encoded images and the encoded text and to train the image encoder and the text encoder based on the multi-modal loss function, wherein the multi-modal loss function comprises at least one image loss term, at least one text loss term, and at least one cross-modal term.

15. The apparatus of claim 14 , further comprising:

a momentum image encoder configured to encode the images to produce momentum encoded images, wherein the at least one image loss term is based on the encoded images and the momentum encoded images.

16. The apparatus of claim 14 , further comprising:

a momentum text encoder configured to encode the captions to produce momentum encoded text, wherein the at least one text loss term is based on the encode text and the momentum encoded text.

17. The apparatus of claim 14 , wherein:

the image encoder comprises a first image output head and a second image output head, wherein the at least one image loss term is based on the first image output head and the at least one cross-modal term is based on the second image output head.

18. The apparatus of claim 14 , wherein:

the text encoder comprises a first text output head and a second text output head, wherein the at least one text loss term is based on the first text output head and the at least one cross-modal term is based on the second text output head.

19. A method of image search comprising:

encoding an image using an image encoder to produce an encoded image;

encoding text using a text encoder to produce encoded text, wherein the image encoder and the text encoder are jointly trained based on a multi-modal loss function comprising at least one image loss term, at least one text loss term, and at least one cross-modal term; and

performing an image search based on the encoded image and the encoded text.

20. The method of claim 19 , wherein: the image search comprises retrieving search text corresponding to a query image, retrieving a search image corresponding to a query text, or retrieving the search image corresponding to the query image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2021
From: YUAN, XIN; LIN, ZHE; KUEN, JASON WEN YONG; ZHANG, JIANMING; WANG, YILIN; KALE, AJINKYA; FAIETA, BALDO
To: ADOBE INC.
Reel/Frame 055476/0304 →
Continuity (1)
Related Publication 20220284321A1 · Sep 8, 2022
References Cited (68)
US 11562039B2 · Renders · 2023 [cited by examiner]
US 20130080426A1 · Chen · 2013 [cited by examiner]
US 20160321522A1 · Yuan · 2016 [cited by examiner]
US 20170330068A1 · Yu · 2017 [cited by examiner]
US 20180108124A1 · Guo · 2018 [cited by examiner]
US 20200302340A1 · Durand · 2020 [cited by examiner]
US 20220245179A1 · Dernoncourt · 2022 [cited by examiner]
US 20230260164A1 · Yuan · 2023 [cited by examiner]
US 20240070440A1 · Hoehne · 2024 [cited by examiner]
CN 106485270A · 2017 [cited by examiner]
CN 106777402A · 2017 [cited by examiner]
Translation CN-106777402-A (Year: 2024). [cited by examiner]
Translation CN-106485270-A (Year: 2024). [cited by examiner]
Udandarao et al., COBRA: Contrastive Bi-Modal Representation Algorithm, 2020, arXiv:2005.03687 [cs.LG] https://doi.org/10.48550/arXiv.2005.03687, Entire document (Year: 2020). [cited by examiner]
1Alayrac, et al., “Self-Supervised MultiModal Versatile Networks”, CoRR, abs/2006.16228, 2020, 19 pages. [cited by applicant]
2Bachman, et al., “Learning Representations by Maximizing Mutual Information Across Views”, In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d'Alche-Buc, Emily B. Fox, and Roman Garnett, editors, NeurIP… [cited by applicant]
3Caron, et al., “Unsupervised Pre-Training of Image Features on Non-Curated Data”, In ICCV, pp. 2959-2968. IEEE, 2019. [cited by applicant]
4Chen, et al., “A Simple Framework for Contrastive Learning of Visual Representations”, CoRR, abs/2002.05709, 2020, 11 pages. [cited by applicant]
5Chen, et al., “Improved Baselines with Momentum Contrastive Learning”, CoRR, abs/2003.04297, 2020, arXiv:2003.04297v1 [cs.CV] Mar. 9, 2020, 3 pages. [cited by applicant]
6Church, et al., Word Association Norms, Mutual Information, and Lexicography, Comput. Linguistics, 16(1):22-29, 1990. [cited by applicant]
7Deng, et al., “ImageNet: A Large-Scale Hierarchical Image Database”, In CVPR, pp. 248-255. IEEE Computer Society, 2009. [cited by applicant]
8Desai, et al., “VirTex: Learning Visual Representations from Textual Annotations”, CoRR, abs/2006.06666, arXiv:2006.06666v1 [cs.CV] Jun. 11, 2020, 27 pages. [cited by applicant]
9Devlin, et al., “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding”, In Jill Burstein, Christy Doran, and Thamar Solorio, editors, NAACL-HLT, pp. 4171-4186, Association for Computational … [cited by applicant]
10Doersch, et al., “Unsupervised Visual Representation Learning by Context Prediction”, In ICCV, pp. 1422-1430, IEEE Computer Society, 2015. [cited by applicant]
11Edunov, et al., “Understanding Back-Translation at Scale”, In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun'ichi Tsujii, editors, EMNLP, pp. 489-500, Association for Computational Linguistics, arXiv:1808.0938… [cited by applicant]
12Everingham, et al., “The Pascal Visual Object Classes Challenge: A Retrospective”, Int. J. Comput. Vis., 111(1):98-136, 2015. [cited by applicant]
13Faghri, et al., “VSE++: Improving Visual-Semantic Embeddings with Hard Negatives”, In BMVC, p. 12. BMVA Press, arXiv:1707.05612v4 [sc.LG] Jul. 29, 2018. [cited by applicant]
14Gattupalli, et al., “Weakly Supervised Deep Image Hashing through Tag Embeddings”, In CVPR, pp. 10375-10384, Computer Vision Foundation /IEEE, 2019. [cited by applicant]
15Girshick, et al., “Rich feature hierarchies for accurate object detection and semantic segmentation”, In CVPR, pp. 580-587. IEEE Computer Society, 2014. [cited by applicant]
16Gomez, et al., “Self-supervised learning of visual features through embedding images into text topic spaces”, In CVPR, pp. 2017-2026, IEEE Computer Society, 2017. [cited by applicant]
17Gordo, et al., “Beyond instance-level image retrieval: Leveraging captions to learn a global visual representation for semantic retrieval”, In CVPR, pp. 6589-6598, IEEE Computer Society, 2017. [cited by applicant]
18Gu, et al., “Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models”, In CVPR, pp. 7181-7189, IEEE Computer Society, 2018. [cited by applicant]
19Gu, et al., “An Empirical Study of Language CNN for Image Captioning”, In ICCV, pp. 1222-1231, IEEE Computer Society, 2017. [cited by applicant]
20Guo, et al., “Visual Attention Consistency under Image Transforms for Multi-Label Image Classification”, In CVPR, pp. 729-739, Computer Vision Foundation / IEEE, 2019. [cited by applicant]
21Gutmann, et al., “Noise-contrastive estimation: A new estimation principle for unnormalized statistical models”, In Yee Whye Teh and D. Mike Titterington, editors, AIStats, vol. 9 of JMLR Proceedings, pp. 297-304, JML… [cited by applicant]
22He, et al., “Momentum Contrast for Unsupervised Visual Representation Learning”, In CVPR, pp. 9729-9738, IEEE, 2020. [cited by applicant]
23He, et al., “Mask R-CNN”, In ICCV, pp. 2961-2969, IEEE Computer Society, 2017. [cited by applicant]
24He, et al., “Deep Residual Learning for Image Recognition”, In CVPR, pp. 770-778, IEEE Computer Society, 2016. [cited by applicant]
25Henaff, et al., “Data-Efficient Image Recognition with Contrastive Predictive Coding”, CoRR, abs/1905.09272, 11 pages, 2019. [cited by applicant]
26Hjelm, et al., “Learning Deep Representations by Mutual Information Estimation and Maximization”, In ICLR, OpenReview.net, arXiv:1808.06670v5 [stat.ML] Feb. 22, 2019, 24 pages. [cited by applicant]
27Huang, et al., “Densely Connected Convolutional Networks”, In CVPR, pp. 4700-4708, IEEE Computer Society, 2017. [cited by applicant]
[28] Andrej Karpathy and Fei-Fei Li. Deep visual-semantic alignments for generating image descriptions. In CVPR, pp. 3128-3137, IEEE Computer Society, 2015. [cited by applicant]
29Khosla, et al., “Supervised Contrastive Learning”, CoRR, abs/2004.11362, arXiv:2004.11362v4 [cs.LG] Dec. 10, 2020, 23 pages. [cited by applicant]
30Kingma, et al., ADAM: A Method for Stochastic Optimization, In Yoshua Bengio and Yann LeCun, editors, ICLR, arXiv:1412.6980v9 [cs.LG] Jan. 30, 2017, 15 pages, 2015. [cited by applicant]
31Kiros, et al., “Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models”, CoRR, abs/1411.2539, arXiv:1411.2539v1 [cs.LG] Nov. 10, 2014, 13 pages. [cited by applicant]
32Li, et al., “Learning Visual N-Grams from Web Data”, In ICCV, pp. 4183-4192, IEEE Computer Society, 2017. [cited by applicant]
33Li, et al., “Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks”, In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, ECCV, vol. 12375 of Lecture Notes in Computer Scienc… [cited by applicant]
34Lin, et al., “Feature Pyramid Networks for Object Detection”, In CVPR, pp. 2117-2125, IEEE Computer Society, 2017. [cited by applicant]
35Lin, et al., “Microsoft COCO: Common Objects in Context”, In David J. Fleet, Tomas Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors, ECCV, vol. 8693 of Lecture Notes in Computer Science, pp. 740-755. Springer, 201… [cited by applicant]
36Liu, et al., “SSD: Single Shot MultiBox Detector”, In Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling, editors, ECCV, vol. 9905 of Lecture Notes in Computer Science, pp. 21-37, Springer, arXiv:1512.02325v5 [cs.C… [cited by applicant]
37Long, et al., “Fully Convolutional Networks for Semantic Segmentation”, In CVPR, pp. 3431-3440, IEEE Computer Society, 2015. [cited by applicant]
38Miech, et al., “End-to-End Learning of Visual Representations from Uncurated Instructional Videos”, In CVPR, pp. 9879-9889, IEEE, 2020. [cited by applicant]
39Ott, et al., Scaling Neural Machine Translation, In Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno-Yepes, Philipp Koehn, Christof Monz, Mat… [cited by applicant]
40Peng, et al., “MegDet: A Large Mini-Batch Object Detector”, In CVPR, pp. 6181-6189, IEEE Computer Society, 2018. [cited by applicant]
41Quattoni, et al., “Learning Visual Representations using Images with Captions”, In CVPR, IEEE Computer Society, 2007, 8 pages. [cited by applicant]
42Redmon, et al., “You Only Look Once: Unified, Real-Time Object Detection”, In CVPR, pp. 779-788, IEEE Computer Society, 2016. [cited by applicant]
43Ren, et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”, In Corinna Cortes, Neil D. Lawrence, Daniel D. Lee, Masashi Sugiyama, and Roman Garnett, editors, NIPS, pp. 91-99, 2015, a… [cited by applicant]
44Sariyildiz, et al., “Learning Visual Representations with Caption Annotations”, In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, ECCV, vol. 12353 of Lecture Notes in Computer Science, pp.… [cited by applicant]
45Simonyan et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, In Yoshua Bengio and Yann LeCun, editors, ICLR, 2015, 14 pages. [cited by applicant]
46Su, et al., “VL-BERT: Pre-Training of Generic VisualLinguistic Representations”, In ICLR, OpenReview.net, 2020, 16 pages. [cited by applicant]
47Tian, et al., “Contrastive Multiview Coding”, CoRR, abs/1906.05849, 2019, 16 pages, arXiv:1906.05849v5 [cs.CV] Dec. 18, 2020. [cited by applicant]
48Oord, et al., “Representation Learning with Contrastive Predictive Coding”, CoRR, abs/1807.03748, 2018, 13 pages, arXiv:1807.03748v2 [cs.LG] Jan. 22, 2019. [cited by applicant]
49Vaswani, et al., “Attention Is All You Need”, In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, NIPS, pp. 5998-6008, arXiv:1706.03762v… [cited by applicant]
50Wang, et al., “Learning Deep Structure-Preserving Image-Text Embeddings”, In CVPR, pp. 5005-5013, IEEE Computer Society, 2016. [cited by applicant]
51Wu, et al., “Unsupervised Feature Learning via Non-Parametric Instance Discrimination”, In CVPR, pp. 3733-3742, IEEE Computer Society, 2018. [cited by applicant]
52Zhang, et al., “Salient Object Subitizing”, Int. J. Comput. Vis., 124(2):169-186, 2017. [cited by applicant]
53Zhang, et al., “Colorful Image Colorization”, In Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling, editors, ECCV, vol. 9907 of Lecture Notes in Computer Science, pp. 649-666, Springer, arXiv:1603.08511v5 [cs.CV] … [cited by applicant]
54Zhuang, et al., “Local Aggregation for Unsupervised Learning of Visual Embeddings”, In ICCV, pp. 6002-6012, IEEE, 2019. [cited by applicant]
Cited By (2)
US 12,347,158 US 12,633,092