IP Library › Granted Patent US 12,462,524
Granted Patent B2
US 12,462,524 · App. 17/920,623 · Granted Nov 4, 2025

Supervised contrastive learning with multiple positive examples

Inventors: Dilip Krishnan (Arlington, MA); Prannay Khosla (Cambridge, MA); Piotr Teterwak (Boston, MA); Aaron Yehuda Sarna (Cambridge, MA); Aaron Joseph Maschinot (Somerville, MA); Ce Liu (Cambridge, MA); Philip John Isola (Cambridge, MA); Yonglong Tian (Cambridge, MA); Chen Wang (Jersey City, NJ)
Assignee: GOOGLE LLC
G06V10/454G06F18/214G06F18/2178G06F18/22G06F18/2431G06N3/08G06N3/09G06V10/761G06V10/764G06V10/774G06V10/776G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,524
App. No.
17/920,623
Granted
Nov 4, 2025
Kind
B2
Abstract

The present disclosure provides an improved training methodology that enables supervised contrastive learning to be simultaneously performed across multiple positive and negative training examples. In particular, example aspects of the present disclosure are directed to an improved, supervised version of the batch contrastive loss, which has been shown to be very effective at learning powerful representations in the self-supervised setting. Thus, the proposed techniques adapt contrastive learning to the fully supervised setting and also enable learning to occur simultaneously across multiple positive examples.

Claims (56)

1 . A computing system to perform supervised contrastive learning of visual representations, the computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store:

a base encoder neural network configured to process an input image to generate an embedding representation of the input image;

a projection head neural network configured to process the embedding representation of the input image to generate a projected representation of the input image; and

instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

obtaining an anchor image associated with a first class of a plurality of classes, a plurality of positive images associated with the first class, and one or more negative images associated with one or more other classes of the plurality of classes, the one or more other classes being different from the first class, wherein:

the anchor image corresponds to a first image from a training dataset;

the plurality of positive images respectively correspond to a plurality of second images from the training dataset; and

the one or more negative images respectively correspond to one or more third images from the training dataset;

processing, with the base encoder neural network, the anchor image to obtain an anchor embedding representation for the anchor image, the plurality of positive images to respectively obtain a plurality of positive embedding representations, and the one or more negative images to respectively obtain one or more negative embedding representations;

processing, with the projection head neural network, the anchor embedding representation to obtain an anchor projected representation for the anchor image, the plurality of positive embedding representations to respectively obtain a plurality of positive projected representations, and the one or more negative embedding representations to respectively obtain one or more negative projected representations;

evaluating a loss function that evaluates a similarity metric between the anchor projected representation and each of the plurality of positive projected representations and each of the one or more negative projected representations; and

modifying one or more values of one or more parameters of at least the base encoder neural network based at least in part on the loss function.

2 . The computing system of claim 1 , wherein the anchor image and at least one of the plurality of positive images depicts different subjects belonging to the same first class of the plurality of classes.

3 . The computing system of claim 1 , wherein the plurality of positive images comprise all images contained within a training batch that are associated with the first class, and wherein the one or more negative images comprise all images contained within the training batch that are not associated with any of the plurality of classes other than the first class.

4 . The computing system of claim 1 , wherein the operations further comprise respectively augmenting each of the anchor image, the plurality of positive images, and the one or more negative images prior to processing each of the anchor image, the plurality of positive images, and the one or more negative images with the base encoder neural network.

5 . The computing system of claim 1 , wherein the projection head neural network comprises a normalization layer that normalizes the projected representation for the input image.

6 . The computing system of claim 1 , wherein the similarity metric comprises an inner product.

7 . The computing system of claim 1 , wherein the loss function comprises a normalization term times a sum, across all images in a training batch, of a contrastive loss term, wherein the normalization term normalizes for a number of images included in the first class of the anchor image.

8 . The computing system of claim 7 , wherein the normalization term comprises negative one divided by two times the number of images included in the first class of the anchor image minus one.

9 . The computing system of claim 7 , wherein the contrastive loss term comprises, when the image under evaluation by the sum is included in the first class, a log of a first term divided by a second term, wherein:

the first term comprises an exponential of the similarity metric between the anchor image and the image under evaluation; and

the second term comprises a sum, across each respective image in the training batch not included in the first class, of an exponential of the similarity between the anchor and the respective image.

10 . The computing system of claim 1 , wherein the operations further comprise, after modifying one or more values of one or more parameters of at least the base encoder neural network based at least in part on the loss function:

adding a classification head to the base encoder neural network; and

finetuning the classification head based on a set of supervised training data.

11 . The computing system of claim 1 , wherein the operations further comprise, after modifying one or more values of one or more parameters of at least the base encoder neural network based at least in part on the loss function:

providing an additional input to the base encoder neural network;

receiving an additional embedding representation for the additional input as an output of the base encoder neural network; and

generating a prediction for the additional input based at least in part on the additional embedding representation.

12 . The computing system of claim 11 , wherein the prediction comprises a classification prediction, a detection prediction, a recognition prediction, a regression prediction, a segmentation prediction, or a similarity search prediction.

13 . The computing system of claim 1 , wherein the anchor image comprises an x-ray image.

14 . The computing system of claim 1 , wherein the anchor image comprises a set of LiDAR data.

15 . The computing system of claim 1 , wherein the anchor image comprises a video.

16 . A computer-implemented method, the method comprising:

obtaining, by a computing system comprising one or more computing devices, an anchor image associated with a first class of a plurality of classes, a plurality of positive images associated with the first class, and one or more negative images associated with one or more other classes of the plurality of classes, the one or more other classes being different from the first class, wherein:

the anchor image corresponds to a first image from a training dataset;

the plurality of positive images respectively correspond to a plurality of second images from the training dataset; and

the one or more negative images respectively correspond to one or more third images from the training dataset;

processing, by the computing system, with a base encoder neural network, the anchor image to obtain an anchor embedding representation for the anchor image, the plurality of positive images to respectively obtain a plurality of positive embedding representations, and the one or more negative images to respectively obtain one or more negative embedding representations;

processing, by the computing system, with a projection head neural network, the anchor embedding representation to obtain an anchor projected representation for the anchor image, the plurality of positive embedding representations to respectively obtain a plurality of positive projected representations, and the one or more negative embedding representations to respectively obtain one or more negative projected representations;

evaluating, by the computing system, a loss function that evaluates a similarity metric between the anchor projected representation and each of the plurality of positive projected representations and each of the one or more negative projected representations; and

modifying, by the computing system, one or more values of one or more parameters of at least the base encoder neural network based at least in part on the loss function.

17 . The computer-implemented method of claim 16 , wherein the plurality of positive images comprise all images contained within a training batch that are associated with the first class, and wherein the one or more negative images comprise all images contained within the training batch that are not associated with any of the plurality of classes other than the first class.

18 . The computer-implemented method of claim 16 , wherein the anchor image and at least one of the plurality of positive images depicts different subjects belonging to the same first class of the plurality of classes.

19 . One or more non-transitory computer-readable media that collectively store at least a base encoder neural network that has been trained by:

obtaining, by a computing system comprising one or more computing devices, an anchor image associated with a first class of a plurality of classes, a plurality of positive images associated with the first class, and one or more negative images associated with one or more other classes of the plurality of classes, the one or more other classes being different from the first class, wherein:

the anchor image corresponds to a first image from a training dataset;

the plurality of positive images respectively correspond to a plurality of second images from the training dataset; and

the one or more negative images respectively correspond to one or more third images from the training dataset;

processing, by the computing system, with a base encoder neural network, the anchor image to obtain an anchor embedding representation for the anchor image, the plurality of positive images to respectively obtain a plurality of positive embedding representations, and the one or more negative images to respectively obtain one or more negative embedding representations;

processing, by the computing system, with a projection head neural network, the anchor embedding representation to obtain an anchor projected representation for the anchor image, the plurality of positive embedding representations to respectively obtain a plurality of positive projected representations, and the one or more negative embedding representations to respectively obtain one or more negative projected representations;

evaluating, by the computing system, a loss function that evaluates a similarity metric between the anchor projected representation and each of the plurality of positive projected representations and each of the one or more negative projected representations; and

modifying, by the computing system, one or more values of one or more parameters of at least the base encoder neural network based at least in part on the loss function.

20 . The one or more non-transitory computer-readable media of claim 19 , wherein the anchor image and at least one of the plurality of positive images depicts different subjects belonging to the same first class of the plurality of classes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2022
From: KRISHNAN, DILIP; KHOSLA, PRANNAY; TETERWAK, PIOTR; SARNA, AARON YEHUDA; MASCHINOT, AARON JOSEPH; LIU, CE; ISOLA, PHILLIP JOHN; TIAN, YONGLONG; WANG, CHEN
To: GOOGLE LLC
Reel/Frame 061658/0555 →
Continuity (2)
Provisional Application 63013153 · Apr 21, 2020
Related Publication 20230153629A1 · May 18, 2023
References Cited (77)
US 10565496B2 · Sohn · 2020 [cited by applicant]
US 10592732B1 · Sather et al. · 2020 [cited by applicant]
US 10832062B1 · Evans et al. · 2020 [cited by applicant]
US 10922574B1 · Tariq · 2021 [cited by applicant]
US 20030206228A1 · Trevers · 2003 [cited by examiner]
US 20170124711A1 · Chandraker et al. · 2017 [cited by applicant]
US 20180136314A1 · Taylor · 2018 [cited by examiner]
US 20180137642A1 · Malisiewicz et al. · 2018 [cited by applicant]
US 20200090039A1 · Song et al. · 2020 [cited by applicant]
US 20200097742A1 · Kumar et al. · 2020 [cited by applicant]
US 20210022698A1 · Vignon · 2021 [cited by examiner]
CN 110837836 · 2020 [cited by applicant]
Baum et al., “Supervised Learning of Probability Distributions by Neural Networks”, Advances in Neural Information Processing Systems, 1988, Denver, Colorado, United States, pp. 52-61. [cited by applicant]
Cao et al., “Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss”, arXiv:1906.07413v2, Oct. 27, 2019, 18 pages. [cited by applicant]
Chopra et al., “Learning a Similarity Metric Discriminatively, with Application to Face Verification”, 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 20-26, 2006, San Diego… [cited by applicant]
Cubuk et al., “AutoAugment: Learning Augmentation Strategies from Data”, Conference on Computer Vision and Pattern Recognition, Jun. 16-20, 2019, Long Beach, California, 11 pages. [cited by applicant]
Cubuk et al., “RandAugment: Practical Automated Data Augmentation with a Reduced Search Space”, arXiv:1909.13719v2, Nov. 14, 2019, 13 pages. [cited by applicant]
Deng et al., “ImageNet: A Large-Scale Hierarchical Image Database”, 2009 Conference on Computer Vision and Pattern Recognition, Jun. 20-25, 2009, Miami, Florida, United States, 8 pages. [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, arXiv:1810.04805v1, Oct. 11, 2018, 14 pages. [cited by applicant]
Doersch et al., “Unsupervised Visual Representation Learning by Context Prediction”, International Conference on Computer Vision, Dec. 11-18, 2015, Santiago, Chile, pp. 1422-1430. [cited by applicant]
El Sayed et al., “Large Margin Deep Networks for Classification”, Thirty-second Conference on Neural Information Processing Systems, Dec. 3-8, 2018, Montreal, Canada, 11 pages. [cited by applicant]
Frosst et al., “Analyzing and Improving Representations with the Soft Nearest Neighbor Loss”, 36th International Conference on Machine Learning, Jun. 9-15, 2019, Long Beach, California, United States, 9 pages. [cited by applicant]
Gutmann et al., “Noise-Contrastive Estimation: A New Estimation for Unnormalized Statistical Models”, Thirteenth International Conference on Artificial Intelligence and Statistics, May 13-15, 2010, Sardinia, Italy, 8 pa… [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, Conference on Computer Vision and Pattern Recognition, Jun. 26-Jul. 1, 2016, Las Vegas, Nevada, United States, pp. 770-778. [cited by applicant]
He et al., “Momentum Contrast for Unsupervised Visual Representation Learning”, arXiv:1911.05722v2. Nov. 14, 2019, 11 pages. [cited by applicant]
He et al., “Rethinking ImageNet Pre-training”, International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, Seoul, Korea, pp. 4918-4927. [cited by applicant]
Henaff et al., “Data-Efficient Image Recognition with Contrastive Predictive Coding”, arXiv:1905.09272v2, Dec. 6, 2019, 15 pages. [cited by applicant]
Hendrycks et al., “Benchmarking Neural Network Robustness to Common Corruptions and Perturbations”, arXiv:1903.12261v1, Mar. 28, 2019, 16 pages. [cited by applicant]
Hinton et al., “Distilling the Knowledge in a Neural Network”, arXiv:1503.02531v1, Mar. 9, 2015, 9 pages. [cited by applicant]
Hinton et al., “Neural Networks for Machine Learning Lecture 6a Overview of Mini-Batch Gradient Descent”, University of Toronto Department of Computer Science, Lecture Slides, 2012, 31 pages. [cited by applicant]
Hjelm et al., “Learning Deep Representations by Mutual Information Estimation and Maximization”. Seventh International Conference on Learning Representations, May 6-9, 2019, New Orleans, Louisiana, United States, 24 pag… [cited by applicant]
Kamnitsas et al., “Semi-Supervised Learning via Compact Latent Space Clustering”, arXiv:1806.02679v2, Jul. 29, 2018, 10 pages. [cited by applicant]
Kolesnikov et al., “Large Scale Learning of General Visual Representations for Transfer”, arXiv:1912.11370v1. Dec. 24, 2019, 23 pages. [cited by applicant]
Kornblith et al., “Do Better ImageNet Models Transfer Better?”, Conference on Computer Vision and Pattern Recognition, Jun. 16-20, 2019, Long Beach, California, United States, pp. 2661-2671. [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks”, Twenty-sixth Conference on Neural Information Processing Systems, Dec. 3-8, 2012, Lake Tahoe, Nevada, United States, 9 pages. [cited by applicant]
Krizhevsky et al., “Learning Multiple Layers of Features from Tiny Images”, University of Toronto, Technical Report, 2009, 60 pages. [cited by applicant]
Lim et al., “Fast AutoAugment”, arXiv:1905.00397v2, May 25, 2019, 10 pages. [cited by applicant]
Liu et al., “Lage-Margin Softmax Loss for Convolutional Neural Networks”, 33rd International Conference on Machine Learning, Jun. 19-24, 2016, New York City, New York, United States, 10 pages. [cited by applicant]
Mikolov et al., “Distributed Representations of Words and Phrases and their Compositionality”, Twenty-seventh Conference on Neural Information Processing Systems, Dec. 5-10, 2013, Lake Tahoe, Nevada, United States, 9 pa… [cited by applicant]
Mnih et al., “Learning Word Embeddings Efficiently with Noise-Contrastive Estimation”, Twenty-seventh Conference on Neural Information Processing Systems, Dec. 5-10, 2013, Lake Tahoe, Nevada, United States, 9 pages. [cited by applicant]
Muller et al., “When Does Label Smoothing Help?”, Thirty-third Conference on Neural Information Processing Systems, Dec. 8-14, 2019, Vancouver, Canada, 10 pages. [cited by applicant]
Nar et al., “Cross-Entropy Loss and Low-Rank Features Have Responsibility for Adversarial Examples”, arXiv:1901.08360v1, Jan. 24, 2019, 10 pages. [cited by applicant]
Noroozi et al., “Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles”, arXiv:1603.09246v2, Jun. 26, 2016, 17 pages. [cited by applicant]
Oord et al., “Representation Learning with Contrastive Predictive Coding”, arXiv:1807.03748v1, Jul. 10, 2018, 13 pages. [cited by applicant]
Ruder, “An Overview of Gradient Descent Optimization Algorithms”, arXiv:1609.04747v1, Sep. 15, 2016, 12 pages. [cited by applicant]
Rumelhart et al., “Learning Representations by Back-Propagating Errors”, Nature, vol. 323, Oct. 9, 1986, pp. 533-536. [cited by applicant]
Salakhutdinov et al., “Learning a Nonlinear Embedding by Preserving Class Neighbourhood Structure”, Proceedings of Machine Learning Research (PMLR), vol. 2, 2017, 8 pages. [cited by applicant]
Schroff et al., “FaceNet: A Unified Embedding for Face Recognition and Clustering”, Conference on Computer Vision and Pattern Recognition, Jun. 7-12, 2015, Boston, Massachusetts, United States, 9 pages. [cited by applicant]
Sermanet et al., “Time-Contrastive Networks: Self-Supervised Learning from Video”, arXiv:1704.06888v3, Mar. 20, 2018, 15 pages. [cited by applicant]
Simonyan et al., “Very Deep Convolutional Networks For Large-Scale Image Recognition”, arXiv:1409.1556v6, Apr. 10, 2015, 14 pages. [cited by applicant]
Sohn, “Improved Deep Metric Learning with Multi-class N-pair Loss Objective”, Thirtieth Conference on Neural Information Processing Systems, Dec. 5-10, 2016, Barcelona, Spain, 9 pages. [cited by applicant]
Solla et al., “Accelerated Learning in Layered Neural Networks”, Complex Systems, vol. 2, 1988, pp. 625-640. [cited by applicant]
Sukhbaatar et al., “Training Convolutional Networks with Noisy Labels”, arXiv:1406.2080v2, Dec. 20, 2014, 11 pages. [cited by applicant]
Szegedy et al., “Rethinking the Inception Architecture for Computer Vision”, Conference on Computer Vision and Pattern Recognition, Jun. 26-Jul. 1, 2016, Las Vegas, Nevada, United States, pp. 2818-2826. [cited by applicant]
Tian et al., “Contrastive Multiview Coding”, arXiv:1906.05849v3, Oct. 21, 2019, 21 pages. [cited by applicant]
Tian et al., “What Makes for Good Views for Contrastive Learning”, arXiv:2005.10243v3, Dec. 18. 2020. 24 pages. [cited by applicant]
Tschannen et al., “On Mutual Information Maximization for Representation Learning”, arXiv:1907.13625v2, Jan. 23, 2020, 16 pages. [cited by applicant]
Wang et al., “Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere”, arXiv:2005.10242v9, Nov. 10, 2020, 41 pages. [cited by applicant]
Weinberger et al., “Distance Metric Learning for Large Margin Nearest Neighbor Classification”, Journal of Machine Learning Research, vol. 10, Feb. 2009, pp. 207-244. [cited by applicant]
Wu et al., “Improving Generalization via Scalable Neighborhood Component Analysis”, European Conference on Computer Vision, Sep. 8-14, 2018, Munich, Germany, 17 pages. [cited by applicant]
Wu et al., “Unsupervised Feature Learning via Non-Parametric Instance Discrimination”. Conference on Computer Vision and Pattern Recognition, Jun. 18-22, 2018, Salt Lake City, Utah, United States, 10 pages. [cited by applicant]
Xie et al., “Self-training with Noisy Student Improves ImageNet Classification”, arXiv:1911.04252v1. Nov. 11, 2019, 13 pages. [cited by applicant]
Yang et al., “Deep Representation Learning with Target Coding”, Conference on Artificial Intelligence, Jan. 25-30, 2015, Austin, Texas, United States, 7 pages. [cited by applicant]
You et al., “Large Batch Training of Convolutional Networks”, arXiv:1708.03888v3, Sep. 13, 2017, 8 pages. [cited by applicant]
Yun et al., “CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features”, International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, Seoul, Korea, pp. 6023-6032. [cited by applicant]
Zhang et al., “Colorful Image Colorization”, European Conference on Computer Vision, Oct. 8-16, 2016, Amsterdam, The Netherlands, pp. 649-666. [cited by applicant]
Zhang et al., “Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels”, Thirty-second Conference on Neural Information Processing Systems, Dec. 2-8, 2018, Montreal, Canada, 11 pages. [cited by applicant]
Zhang et al., “mixup: Beyond Empirical Risk Minimization”, arXiv:1710.09412v1, Oct. 25, 2017, 11 pages. [cited by applicant]
Zhang et al., “Split-Brain Autoencoders: Unsupervised Learning by Cross-Channel Prediction”, Conference on Computer Vision and Pattern Recognition, Jul. 21-26, 2017, Honolulu, Hawaii, United States, pp. 1058-1067. [cited by applicant]
Zhilin et al., “XLNet: Generalized Autoregressive Pretraining for Language Understanding”, Thirty-third Conference on Neural Information Processing Systems, Dec. 8-14, 2019, Vancouver, Canada, 11 pages. [cited by applicant]
Zhu et al., “A New Loss Function for CNN Classifier Based on Predefined Evenly-Distributed Class Centroids”, IEEE Access, vol. 8, 2019, pp. 10888-10895. [cited by applicant]
International Preliminary Report on Patentability for Application No. PCT/US2021/026836, mailed Nov. 3, 2022, 13 pages. [cited by applicant]
Chen, et al., “A Simple Framework for Contrastive Learning of Visual Representations” arxiv.org, Feb. 13, 2020, XP081632474, 18 pages. [cited by applicant]
International Search Report for Application No. PCT/US2021/026836, mailed on Aug. 16, 2021, 3 pages. [cited by applicant]
Khosla, et al., “Supervised Contrastive Learning”, arxiv.org, Mar. 10, 2021, XP081895539, 24 pages. [cited by applicant]
Qiuyu et al., “A New Loss Function for CNN Classifier Based on Predefined Evenly-Distributed Class Centroids”, IEEE Access, vol. 8, Dec. 14, 2019, pp. 10888-10895. [cited by applicant]
Chinese Search Report Corresponding to Application No. 2021800071804 on Dec. 11, 2024. [cited by applicant]