IP Library Granted Patent US 12,724,820
Granted Patent B2
US 12,724,820 · App. 18/791,745 · Granted Sep 1, 2026

Text-based image retrieval

Inventors: Hyunjae Kim (Seoul, KR); Seunghyun Yoon (San Jose, CA); Trung Huu Bui (San Jose, CA); Handong Zhao (Cupertino, CA); Quan Tran (San Jose, CA); Franck Dernoncourt (Spokane, WA)
Assignee: ADOBE INC.
G06F16/535G06F40/40G06T11/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,724,820
App. No.
18/791,745
Granted
Sep 1, 2026
Kind
B2
Abstract

A method, apparatus, non-transitory computer readable medium, and system for media processing include obtaining a text prompt describing content, generating, using a multi-modal encoder, a text embedding based on the text prompt, and obtaining an image depicting the content based on the text embedding. The multi-modal encoder is trained to encode image descriptions based on a similarity between a caption of a training image and a paraphrase of the caption.

Claims (46)

1 . A method for media processing, comprising:

obtaining a text prompt describing content;

generating, using a text encoder of a multi-modal encoder, a text embedding in a multi-modal embedding space based on the text prompt, wherein the text encoder is trained to encode image descriptions in the multi-modal embedding space by computing a paraphrase-caption loss based on a similarity between a caption of a training image and a paraphrase of the caption and updating parameters of the text encoder based on the paraphrase-caption loss, wherein the caption comprises a first text string describing the training image and the paraphrase comprises a second text string describing the training image, and wherein the first text string is different from the second text string; and

obtaining an image depicting the content based on the text embedding.

2 . The method of claim 1 , wherein obtaining the image comprises:

identifying an image embedding of the image; and

retrieving the image from a database based on a comparison of the text embedding and the image embedding.

3 . The method of claim 2 , wherein:

the text embedding and the image embedding comprise vectors in the multi-modal embedding space.

4 . The method of claim 1 , further comprising:

retrieving a plurality of images from a database based on the text embedding.

5 . The method of claim 1 , wherein obtaining the image comprises:

generating the image using an image generation model conditioned on the text embedding.

6 . The method of claim 1 , further comprising:

tokenizing the text prompt to obtain a sequence of tokens representing the content, wherein the text embedding is generated based on the sequence of tokens.

7 . A method for training a machine learning model, comprising:

obtaining a training set comprising a training image, a caption of the training image, and a paraphrase of the caption, wherein the caption comprises a first text string describing the training image and the paraphrase comprises a second text string describing the training image, and wherein the first text string is different from the second text string;

encoding, using an image encoder of a multi-modal encoder, the training image to obtain an image embedding in a multi-modal embedding space;

encoding, using a text encoder of the multi-modal encoder, the caption and the paraphrase to obtain a caption embedding and a paraphrase embedding, respectively, in the multi-modal embedding space; and

training the text encoder of the multi-modal encoder by computing an image-caption loss based on a first similarity between the image embedding and the caption embedding, computing a paraphrase-caption loss based on a second similarity between the caption embedding and the paraphrase embedding, and updating parameters of the text encoder based on the image-caption loss and the paraphrase-caption loss.

8 . The method of claim 7 , wherein training the text encoder comprises:

computing a paraphrase-paraphrase loss based on a third similarity between the paraphrase embedding and an additional paraphrase embedding of an additional paraphrase of the paraphrase and updating the parameters of the text encoder based on the paraphrase-paraphrase loss.

9 . The method of claim 7 , wherein obtaining the training set comprises:

generating the caption based on the training image.

10 . The method of claim 7 , wherein obtaining the training set comprises:

generating the paraphrase based on the caption.

11 . The method of claim 10 , wherein generating the paraphrase comprises:

generating a prompt requesting a variant of the caption using different language; and

providing the prompt to a large language model.

12 . The method of claim 7 , wherein obtaining the training set comprises:

generating an additional paraphrase based on the paraphrase.

13 . The method of claim 7 , wherein training the text encoder comprises:

fine-tuning a pre-trained multi-modal encoder.

14 . The method of claim 7 , wherein training the text encoder comprises:

freezing the image encoder.

15 . A system for media processing, comprising:

at least one processor;

at least one memory storing instructions executable by the at least one processor; and

a multi-modal encoder comprising a text encoder comprising encoding parameters stored in the at least one memory, the text encoder configured to generate a text embedding in a multi-modal embedding space based on a text prompt, wherein the text encoder is trained to encode image descriptions in the multi-modal embedding space by computing a paraphrase-caption loss based on a similarity between a caption of a training image and a paraphrase of the caption and updating parameters of the text encoder based on the paraphrase-caption loss, wherein the caption comprises a first text string describing the training image and the paraphrase comprises a second text string describing the training image, and wherein the first text string is different from the second text string.

16 . The system of claim 15 , the system further comprising:

a language generation model comprising text generation parameters stored in the at least one memory, the language generation model trained to generate the paraphrase.

17 . The system of claim 15 , the system further comprising:

a database storing an image embedding; and

a retrieval component configured to retrieve an image from the database based on the text embedding and the image embedding.

18 . The system of claim 15 , the system further comprising:

an image generation model comprising image generation parameters stored in the at least one memory, the image generation model trained to generate an image based on the text embedding.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 1, 2024
From: KIM, HYUNJAE; YOON, SEUNGHYUN; BUI, TRUNG HUU; ZHAO, HANDONG; TRAN, QUAN; DERNONCOURT, FRANCK
To: ADOBE INC.
Reel/Frame 068152/0691 →
Continuity (1)
Related Publication 20260037572A1 · Feb 5, 2026
References Cited (54)
US 9691095B2 · McCluskey · 2017 [cited by examiner]
US 10467703B2 · Olson · 2019 [cited by examiner]
US 10699323B1 · Tang · 2020 [cited by examiner]
US 20020107747A1 · Gerogianni · 2002 [cited by examiner]
US 20040267610A1 · Gossett · 2004 [cited by examiner]
US 20100274775A1 · Fontes · 2010 [cited by examiner]
US 20120173692A1 · Lakes · 2012 [cited by examiner]
US 20130238612A1 · Tsongas · 2013 [cited by examiner]
US 20140143095A1 · McCluskey · 2014 [cited by examiner]
US 20140279275A1 · Burgiss · 2014 [cited by examiner]
US 20160307174A1 · Marcelle · 2016 [cited by examiner]
US 20170255983A1 · McCluskey · 2017 [cited by examiner]
US 20180219849A1 · Jones · 2018 [cited by examiner]
US 20180349988A1 · Shebesta · 2018 [cited by examiner]
US 20210120297A1 · Deshpande · 2021 [cited by examiner]
US 20220121702A1 · Kale · 2022 [cited by examiner]
US 20220309546A1 · Zheng · 2022 [cited by examiner]
US 20240168992A1 · Shu · 2024 [cited by examiner]
US 20240290065A1 · Park · 2024 [cited by examiner]
US 20240370718A1 · Panagopoulou · 2024 [cited by examiner]
US 20240378369A1 · Chugh · 2024 [cited by examiner]
US 20240386049A1 · Agrawal · 2024 [cited by examiner]
US 20250173911A1 · Amadori · 2025 [cited by examiner]
Image Captioning using Deep Learning: Text Augmentation by Paraphrasing via Backtranslation; Ingrid Ravn Turkerud, Ole Jakob Mengshoel, Department of Computer Science NTNU, Trondheim, Norway, (Year: 2011). [cited by examiner]
Generating Diverse and Descriptive Image Captions Using Visual Paraphrases; Lixin Liu, Jiajun Tang, Xiaojun Wan, Zongming Guo; Institute of Computer Science and Technology, Peking University Center for Data Science, Pek… [cited by examiner]
Agirre, et al., “SemEval-2015 Task 2: Semantic Textual Similarity, English, Spanish, and Pilot on Interpretability”, In Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval 2015), pp. 252-263, J… [cited by applicant]
Agirre, et al., “SemEval-2014 Task 10: Multilingual Semantic Textual Similarity”, In Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014), pp. 81-91, Aug. 23, 2014, 11 pages. [cited by applicant]
Agirre, et al., “SemEval-2016 Task 1: Semantic Textual Similarity, Monolingual and Cross-lingual Evaluation”, In Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016), pp. 497-511, Jun. 16… [cited by applicant]
Agirre, et al., “SemEval-2012 Task 6: A Pilot on Semantic Textual Similarity”, In Proceedings of the First Joint Conference on Lexical and Computational Semantics (*SEM), pp. 385-393, Jun. 7, 2012, 9 pages. [cited by applicant]
Agirre, et al., “*SEM 2013 shared task: Semantic textual similarity”, In Second Joint Conference on Lexical and Computational Semantics (*SEM), vol. 1: Proceedings of the Main Conference and the Shared Task: Semantic Te… [cited by applicant]
Brown, et al., “Language Models are Few-Shot learners”, 34th Conference on Neural Information Processing Systems (NeurIPS 2020), arXiv preprint arXiv:2005.14165v4 [cs.CL] Jul. 22, 2020, 25 pages. [cited by applicant]
Cer, et al., “SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Cross-lingual Focused Evaluation”, In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pp. 1-14, Aug. … [cited by applicant]
Cherti, et al., “Reproducible scaling laws for contrastive language-image learning”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2818-2829, 2023, 12 pages. [cited by applicant]
Deng, et al., “ImageNet: A Large-Scale Hierarchical Image Database”, In 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), Jun. 20, 2009, 8 pages. [cited by applicant]
Dosovitskiy, et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale”, arXiv preprint arXiv:2010.11929v1 [cs.CV] Oct. 22, 2020, 21 pages. [cited by applicant]
Fagin, et al., “Comparing Top k Lists”, In SIAM Journal on Discrete Mathematics, vol. 17, No. 1, pp. 134-160, 2003, 27 pages. [cited by applicant]
Fan, et al., “Improving CLIP Training with Language Rewrites”, arXiv preprint arXiv:2305.20088v2 [cs.CV] Oct. 28, 2023, 32 pages. [cited by applicant]
Feng, et al., “Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis”, arXiv preprint arXiv:2212.05032v3 [cs.CV] Feb. 28, 2023, 21 pages. [cited by applicant]
Gao, et al., “SimCSE: Simple Contrastive Learning of Sentence Embeddings”, In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 6894-6910, Nov. 7, 2021, 17 pages. [cited by applicant]
Jaccard, “The Distribution of the Flora in the Alpine Zone”, In New Phytologist, vol. 11, No. 2, pp. 37-50, Feb. 29, 1912, 14 pages. [cited by applicant]
Lin, et al., “Microsoft COCO: Common Objects in Context”, In Computer Vision—ECCV 2014: 13th European Conference, Zurich, Switzerland, Proceedings, Part V, pp. 740-755, Sep. 6, 2014, 16 pages. [cited by applicant]
Liu, et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach”, arXiv preprint arXiv:1907.11692v1 [cs.CL] Jul. 26, 2019, 13 pages. [cited by applicant]
Loshchilov and Hutler., “Decoupled Weight Decay Regularization”, arXiv preprint arXiv:1711.05101v3 [cs.LG] Jan. 4, 2019, 19 pages. [cited by applicant]
Marelli, et al., “A SICK cure for the evaluation of compositional distributional semantic models”, In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14), pp. 216-223, 2014, … [cited by applicant]
Van den Oord, et al., “Representation Learning with Contrastive Predictive Coding”, arXiv preprint arXiv:1807.03748v2 [cs.LG] Jan. 22, 2019, 13 pages. [cited by applicant]
OpenAI, “Introducing ChatGPT”, Nov. 30, 2022, available at https://openai.com/index/chatgpt/. [cited by applicant]
Plummer, et al., “Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models”, In Proceedings of the IEEE International Conference on Computer Vision, pp. 2641-2649, 2015, 9 page… [cited by applicant]
Radford, et al., “Learning Transferable Visual Models from Natural Language Supervision”, arXiv preprint arXiv:2103.00020v1 [cs.CV] Feb. 26, 2021, 48 pages. [cited by applicant]
Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684-10695, 2022, 12 pages. [cited by applicant]
Saharia, et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, arXiv preprint arXiv:2205.11487v1 [cs.CV] May 23, 2022, 46 pages. [cited by applicant]
Schuhmann, et al., “Laion-5B: An open large-scale dataset for training next generation image-text models”, arXiv preprint arXiv:2210.08402v1 [cs.CV] Oct. 16, 2022, 50 pages. [cited by applicant]
Schuhmann, et al., “Laion-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs”, arXiv preprint arXiv:2111.02114v1 [cs.CV] Nov. 3, 2021, 5 pages. [cited by applicant]
Touvron, et al., “LLaMA: Open and Efficient Foundation Language Models”, arXiv preprint arXiv:2302.13971v1 [cs.CL], Feb. 27, 2023, 27 pages. [cited by applicant]
Yuksekgonul, et al., “When and Why Vision-Language Models Behave Like Bags-Of-Words, and What to Do About It?”, In The Eleventh International Conference on Learning Representations (ICLR 2023), pp. 1-20, May 1, 2023, 20… [cited by applicant]