IP Library › Granted Patent US 12,620,212
Granted Patent B2
US 12,620,212 · App. 18/051,106 · Granted May 5, 2026

Locked-model multimodal contrastive tuning

Inventors: Daniel Keysers (Stallikon, CH); Xiaohua Zhai (Rüschlikon, CH); Xiao Wang (Kilchberg, CH); Lucas Beyer (Zürich, CH); Basil Mustafa (Zürich, CH); Andreas Steiner (Zürich, CH); Alexander Kolesnikov (Zürich, CH)
Assignee: Google LLC
G06V10/778
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,212
App. No.
18/051,106
Granted
May 5, 2026
Kind
B2
Abstract

A method may include obtaining a pretrained image encoder and a training sample comprising a training image and a training text string corresponding to the training image. The method may also include initializing a text encoder in an untrained state, determining, using the pretrained image encoder and based on the training image, a first latent representation of the training image, and determining, using the text encoder and based on the training text string, a second latent representation of the training text string. The method may further include determining a loss value based on the first latent representation and the second latent representation, updating, based on the loss value, one or more parameters of the text encoder while holding fixed parameters of the pretrained image encoder, and outputting the text encoder in a trained state.

Claims (84)

1 . A computer-implemented method comprising:

obtaining (i) a pretrained image encoder and (ii) a training sample comprising a training image and a training text string corresponding to the training image;

initializing a text encoder in an untrained state;

determining, by the pretrained image encoder and based on the training image, a first latent representation of the training image;

determining, by the text encoder and based on the training text string, a second latent representation of the training text string;

determining a loss value based on the first latent representation as determined by the pretrained image encoder and the second latent representation as determined by the text encoder;

updating, based on the loss value, one or more parameters of the text encoder while holding fixed all parameters of the pretrained image encoder throughout training of the text encoder; and

outputting the text encoder in a trained state.

2 . The computer-implemented method of claim 1 , wherein the pretrained image encoder and the text encoder form a multimodal contrastive learning pair, and wherein updating the one or more parameters of the text encoder is configured to train the text encoder to determine latent representations that, for a given training sample, converge to latent representations determined by the pretrained image encoder.

3 . The computer-implemented method of claim 1 , wherein initializing the text encoder comprises:

initializing parameters of the text encoder using substantially randomly selected values.

4 . The computer-implemented method of claim 1 , wherein a size of the first latent representation is equal to a size of the second latent representation.

5 . The computer-implemented method of claim 4 , wherein an output layer of the pretrained image encoder has a first size, and wherein the text encoder comprises a final projection layer configured to project an output of a penultimate layer of the text encoder to the first size.

6 . The computer-implemented method of claim 1 , wherein determining the loss value comprises:

determining the loss value using a contrastive loss function configured to determine a similarity between the first latent representation as determined by the pretrained image encoder and the second latent representation as determined by the text encoder.

7 . The computer-implemented method of claim 6 , wherein updating the one or more parameters of the text encoder based on the loss value determined by the contrastive loss function is configured to train the text encoder to determine latent representations that (i), for training samples comprising matched image-text pairs, converge to latent representations determined by the pretrained image encoder and (ii), for training samples comprising unmatched image-text pairs, diverge from latent representations determined by the pretrained image encoder.

8 . The computer-implemented method of claim 1 , further comprising:

obtaining a second training sample comprising the training image and a second training text string that does not correspond to the training image;

determining, using the text encoder and based on the second training text string, a third latent representation of the second training text string;

determining a second loss value based on the first latent representation and the third latent representation; and

updating, based on the second loss value, one or more additional parameters of the text encoder while holding fixed all the parameters of the pretrained image encoder.

9 . The computer-implemented method of claim 1 , wherein obtaining the pretrained image encoder comprises:

initializing an image encoder in a second untrained state; and

training the image encoder using a training image data set and independently of the text encoder.

10 . The computer-implemented method of claim 1 , further comprising:

obtaining a text query after updating the one or more parameters of the text encoder;

generating, using the text encoder and based on the text query, a third latent representation of the text query; and

retrieving one or more images, wherein each respective image of the one or more images is associated with a corresponding latent representation that (i) has been generated by the pretrained image encoder and (ii) has at least a threshold extent of similarity to the third latent representation.

11 . The computer-implemented method of claim 1 , further comprising:

obtaining an image query;

generating, using the pretrained image encoder and based on the image query, a third latent representation of the image query; and

retrieving one or more text strings, wherein each respective text string of the one or more text strings is associated with a corresponding latent representation that (i) has been generated by the text encoder after updating the one or more parameters of the text encoder and (ii) has at least a threshold extent of similarity to the third latent representation.

12 . The computer-implemented method of claim 1 , further comprising:

obtaining an image, a first text string, and a second text string;

generating, using the pretrained image encoder and based on the image, a third latent representation of the image;

generating, using the text encoder, (i) a fourth latent representation of the first text string based on the first text string and (ii) a fifth latent representation of the second text string based on the second text string;

determining (i) a first similarity between the third latent representation and the fourth latent representation and (ii) a second similarity between the third latent representation and the fifth latent representation;

determining that the first similarity exceeds the second similarity; and

based on determining that the first similarity exceeds the second similarity, determining that the image belongs to a class corresponding to the first text string.

13 . The computer-implemented method of claim 1 , wherein determining the first latent representation of the training image comprises:

precomputing the first latent representation by the pretrained image encoder prior to training of the text encoder.

14 . The computer-implemented method of claim 13 , wherein determining the loss value comprises:

reusing, throughout the training of the text encoder, the first latent representation as precomputed by the pretrained image encoder.

15 . The computer-implemented method of claim 14 , wherein:

obtaining the training sample comprises obtaining a plurality of training samples, wherein each respective training sample of the plurality of training samples comprises a respective training image and a respective training text string corresponding to the respective training image;

determining the first latent representation of the training image comprises precomputing, for each respective training sample, by the pretrained image encoder, and based on the respective training image, a corresponding latent representation of the respective training image prior to training of the text encoder; and

determining the loss value comprises reusing, throughout the training of the text encoder, the corresponding latent representation as precomputed by the pretrained image encoder for the respective training sample.

16 . A system comprising:

a processor; and

a non-transitory computer-readable medium having stored thereon instructions that, when executed by the processor, cause the processor to perform operations comprising:

obtaining (i) a pretrained image encoder and (ii) a training sample comprising a training image and a training text string corresponding to the training image;

initializing a text encoder in an untrained state;

determining, by the pretrained image encoder and based on the training image, a first latent representation of the training image;

determining, by the text encoder and based on the training text string, a second latent representation of the training text string;

determining a loss value based on the first latent representation as determined by the pretrained image encoder and the second latent representation as determined by the text encoder;

updating, based on the loss value, one or more parameters of the text encoder while holding fixed all parameters of the pretrained image encoder throughout training of the text encoder; and

outputting the text encoder in a trained state.

17 . The system of claim 16 , wherein the pretrained image encoder and the text encoder form a multimodal contrastive learning pair, and wherein updating the one or more parameters of the text encoder is configured to train the text encoder to determine latent representations that, for a given training sample, converge to latent representations determined by the pretrained image encoder.

18 . The system of claim 16 , wherein determining the first latent representation of the training image comprises:

precomputing the first latent representation by the pretrained image encoder prior to training of the text encoder.

19 . A computer-implemented method comprising:

obtaining an image, a text string, a pretrained image encoder, and a text encoder, wherein the text encoder has been trained by a training process comprising:

obtaining (i) the pretrained image encoder and (ii) a training sample comprising a training image and a training text string corresponding to the training image;

initializing the text encoder in an untrained state;

determining, by the pretrained image encoder and based on the training image, a first latent representation of the training image;

determining, by the text encoder and based on the training text string, a second latent representation of the training text string;

determining a loss value based on the first latent representation as determined by the pretrained image encoder and the second latent representation as determined by the text encoder; and

updating, based on the loss value, one or more parameters of the text encoder while holding fixed all parameters of the pretrained image encoder throughout training of the text encoder;

determining, by the pretrained image encoder and based on the image, a third latent representation of the image;

determining, by the text encoder and based on the text string, a fourth latent representation of the text string;

determining a similarity between the third latent representation and the fourth latent representation; and

generating an output based on the similarity.

20 . A non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor, cause the processor to perform operations comprising:

obtaining an image, a text string, a pretrained image encoder, and a text encoder, wherein the text encoder has been trained by a training process comprising:

obtaining (i) the pretrained image encoder and (ii) a training sample comprising a training image and a training text string corresponding to the training image;

initializing the text encoder in an untrained state;

determining, by the pretrained image encoder and based on the training image, a first latent representation of the training image;

determining, by the text encoder and based on the training text string, a second latent representation of the training text string;

determining a loss value based on the first latent representation as determined by the pretrained image encoder and the second latent representation as determined by the text encoder; and

updating, based on the loss value, one or more parameters of the text encoder while holding fixed all parameters of the pretrained image encoder throughout training of the text encoder;

determining, by the pretrained image encoder and based on the image, a third latent representation of the image;

determining, by the text encoder and based on the text string, a fourth latent representation of the text string;

determining a similarity between the third latent representation and the fourth latent representation; and

generating an output based on the similarity.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2026
From: ZHAI, XIAOHUA
To: GOOGLE LLC
Reel/Frame 074300/0141 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2022
From: KEYSERS, DANIEL; ZHAI, XIAOHAU; WANG, XIAO; BEYER, LUCAS; MUSTAFA, BASIL; STEINER, ANDREAS; KOLESNIKOV, ALEXANDER
To: GOOGLE LLC
Reel/Frame 061648/0141 →
Continuity (1)
Related Publication 20240153256A1 · May 9, 2024
References Cited (13)
US 20180373979A1 · Wang · 2018 [cited by examiner]
US 20220147838A1 · Gu · 2022 [cited by examiner]
US 20220172080A1 · Chaudhury · 2022 [cited by examiner]
US 20220353522A1 · Ding · 2022 [cited by examiner]
US 20230019211A1 · Wang · 2023 [cited by examiner]
US 20230177810A1 · Xu · 2023 [cited by examiner]
US 20230325685A1 · Caba Heilbron · 2023 [cited by examiner]
Zhang, Yuhao, et al. “Contrastive learning of medical visual representations from paired images and text.” Machine learning for healthcare conference. PMLR, 2022. (Year: 2022). [cited by examiner]
Radford, Alec, et al. “Learning transferable visual models from natural language supervision.” International conference on machine learning. PmLR, 2021. (Year: 2021). [cited by examiner]
Lee, Janghyeon, et al. “Uniclip: Unified framework for contrastive language-image pre-training.” Advances in Neural Information Processing Systems 35 (2022): 1008-1019. (Year: 2022). [cited by examiner]
Chen, Tianlang, and Jiebo Luo. “Expressing objects just like words: Recurrent visual embedding for image-text matching.” Proceedings of the AAAI conference on artificial intelligence. vol. 34. No. 07. 2020. (Year: 2020). [cited by examiner]
Steiner et al., “Locked-Image Tuning: Adding Language Understanding to Image Models,” Google Research, Apr. 14, 2022, 7 pages. [cited by applicant]
Zhai et al., “LiT: Zero-Shot Transfer with Locked-image text Tuning,” arXiv:2111.07991v3, Jun. 22, 2022, first submitted on Nov. 15, 2021, 28 pages. [cited by applicant]