IP Library › Granted Patent US 12,488,042
Granted Patent B2
US 12,488,042 · App. 17/847,848 · Granted Dec 2, 2025

Data compatibility for text-enhanced visual retrieval

Inventors: Rafael Sampaio De Rezende (Grenoble, FR); Diane Larlus (La Tronche, FR); Ginger Delmas (Grenoble, FR); Gabriela Csurka Khedari (Crolles, FR)
Assignee: NAVER CORPORATION
G06F16/5866G06F16/532G06F16/56G06F40/30G06V10/774G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,042
App. No.
17/847,848
Granted
Dec 2, 2025
Kind
B2
Abstract

An interaction module includes: a first text-image interaction module configured to generate a vector representation of a first text-image pair based on an encoded representation of a reference image and an encoded representation of a text modifier, the reference image and the text modifier received from a computing device. A second text-image interaction module is configured to generate a vector representation of a second text-image pair based on the encoded representation of the text modifier and an encoded representation of a candidate target image. A compatibility module is configured to compute, based on the vector representation of the first text-image pair and the vector representation of the second text-image pair, a compatibility score for a triplet including the reference image, the text modifier, and the candidate target image. A ranking module is configured to rank a set of candidate target images including the candidate target image by compatibility scores.

Claims (115)

1 . A system for image retrieval with neural networks, the system comprising:

an interaction module comprising:

a first text-image interaction neural network trained to generate a vector representation of a first text-image pair based on (a) an encoder representation of a reference image and (b) an encoder representation of a text modifier, the reference image and the text modifier received from a computing device,

wherein (i) the text modifier describes at least one visual difference from the reference image, and (ii) the first text-image interaction neural network is trained to use an attention mechanism to weigh dimensions of the encoder representation of the reference image; and

a second text-image interaction neural network trained to generate a vector representation of a second text-image pair based on (a) the encoder representation of the text modifier and (b) an encoder representation of a candidate target image,

wherein the second text-image interaction neural network is trained to use the attention mechanism to weigh dimensions of the encoder representation of the candidate target image;

a compatibility module configured to compute, based on (a) the vector representation of the first text-image pair and (b) the vector representation of the second text-image pair, a compatibility score for a triplet including the reference image, the text modifier, and the candidate target image; and

a ranking module configured to;

rank a set of candidate target images including the candidate target image by compatibility scores of the candidate target images, respectively, of the set; and

select at least one of the candidate target images from the set based on the ranks of the candidate target images, respectively; and

transmit the selected at least one of the candidate target images to the computing device;

wherein the attention mechanism is the text modifier.

2 . The system of claim 1 wherein the compatibility module is configured to:

compute a cosine similarity between the vector representation of the first text-image pair and the vector representation of the second text-image pair; and

determine the compatibility score for the triplet based on the cosine similarity.

3 . The system of claim 2 wherein the compatibility module is configured to:

compute a second cosine similarity between the vector representation of the second text-image pair and the encoder representation of the text modifier; and

determine the compatibility score for the triplet further based on the second cosine similarity.

4 . The system of claim 3 wherein the compatibility module is configured to set the compatibility score for the triplet based on a sum of the cosine similarity and the second cosine similarity.

5 . The system of claim 1 wherein:

the first text-image interaction neural network applies a first attention mechanism on the text modifier and the reference image to generate the vector representation of the first text-image pair; and

the second text-image interaction neural network applies a second attention mechanism on the text modifier and the candidate target image to generate the vector representation of the second text-image pair.

6 . The system of claim 1 , wherein the compatibility module includes a compatibility neural network configured to compute the compatibility score for the triplet.

7 . The system of claim 1 wherein:

the interaction module further comprises an image-image interaction neural network trained to generate a vector representation of an image-image pair based on the encoder representation of the reference image and the encoder representation of the candidate target image; and

the compatibility module is configured to determine the compatibility score for the triplet further based on the vector representation of the image-image pair.

8 . The system of claim 1 , wherein the interaction module is trained based on compatible triplets and incompatible triplets,

wherein each compatible triplet includes a candidate target image that corresponds to a reference image of the respective compatible triplet modified according to a text modifier of the respective compatible triplet, and

wherein an incompatible triplet comprises a candidate target image that does not correspond to a reference image of the respective incompatible triplet modified according to a text modifier of the respective incompatible triplet.

9 . The system of claim 1 wherein the ranking module is configured to rank the rank set of candidate target images including the candidate target image from highest compatibility score to lowest compatibility score.

10 . The system of claim 1 further comprising a pre-ranking module configured to select the set of candidate target images from a second set of candidate target images that includes more candidate target images than the set of candidate target images based on the reference image and the candidate target images of the second set of candidate target images.

11 . The system of claim 10 wherein the pre-ranking module is configured to select the set of candidate target images from the second set using one of a cross-modal retrieval model and a query-composition retrieval model.

12 . A method for image retrieval with neural networks, the method comprising:

generating a vector representation of a first text-image pair based on (a) an encoder representation of a reference image and (b) an encoder representation of a text modifier using a first text-image interaction neural network, the reference image and the text modifier received from a computing device,

wherein the text modifier describes at least one visual difference from the reference image; wherein the first text-image interaction neural network is trained to use an attention mechanism to weigh dimensions of the encoder representation of the reference image;

generating a vector representation of a second text-image pair based on (a) the encoder representation of the text modifier and (b) an encoder representation of a candidate target image using a second text-image interaction neural network; wherein the second text-image interaction neural network is trained to use the attention mechanism to weigh dimensions of the encoder representation of the candidate target image;

computing, based on (a) the vector representation of the first text-image pair and (b) the vector representation of the second text-image pair, a compatibility score for a triplet including the reference image, the text modifier, and the candidate target image; and

ranking a set of candidate target images including the candidate target image by compatibility scores of the candidate target images, respectively, of the set;

selecting at least one of the candidate target images from the set based on the ranks of the candidate target images, respectively; and

transmitting the selected at least one of the candidate target images to the computing device;

wherein the attention mechanism is the text modifier.

13 . The method of claim 12 wherein determining the compatibility score includes:

computing a cosine similarity between the vector representation of the first text-image pair and the vector representation of the second text-image pair; and

determining the compatibility score for the triplet based on the cosine similarity.

14 . The method of claim 13 wherein determining the compatibility score includes:

computing a second cosine similarity between the vector representation of the second text-image pair and the encoder representation of the text modifier; and

determining the compatibility score for the triplet further based on the second cosine similarity.

15 . The method of claim 14 wherein determining the compatibility score includes setting the compatibility score for the triplet based on a sum of the cosine similarity and the second cosine similarity.

16 . The method of claim 12 wherein:

generating the vector representation of the first text-image pair includes generating the vector representation of the first text-image pair by applying a first attention mechanism on the text modifier and the reference image; and

generating the vector representation of the second text-image pair includes generating the vector representation of the second text-image interaction neural network by applying a second attention mechanism on the text modifier and the candidate target image.

17 . The method of claim 12 , wherein determining the compatibility score includes determining the compatibility score for the triplet by a compatibility neural network.

18 . The method of claim 12 wherein determining the compatibility score includes:

generating a vector representation of an image-image pair based on the encoder representation of the reference image and the encoder representation of the candidate target image; and

determining the compatibility score for the triplet further based on the vector representation of the image-image pair.

19 . The method of claim 12 further comprising training to generate the vector representations based on compatible triplets and incompatible triplets,

wherein each compatible triplet includes a candidate target image that corresponds to a reference image of the respective compatible triplet modified according to a text modifier of the respective compatible triplet, and

wherein an incompatible triplet comprises a candidate target image that does not correspond to a reference image of the respective incompatible triplet modified according to a text modifier of the respective incompatible triplet.

20 . The method of claim 12 wherein the ranking includes ranking the set of candidate target images including the candidate target image from highest compatibility score to lowest compatibility score.

21 . The method of claim 12 further comprising selecting the set of candidate target images from a second set of candidate target images that includes more candidate target images than the set of candidate target images based on the reference image and the candidate target images of the second set of candidate target images.

22 . The method of claim 21 wherein selecting the set of candidate target images includes selecting the set of candidate target images from the second set using one of a cross-modal retrieval model and a query-composition retrieval model.

23 . A method of training a-neural networks to retrieve images, the method:

generating a training dataset, comprising:

obtaining a first set of triplets, each triplet of the first set of triplets comprising a reference image, a text modifier, and a target image, corresponding to the reference image modified according to the text modifier, wherein the text modifier describes at least one visual difference from the reference image; and

generating, from the first set of triplets, a second set of triplets, each triplet of the second set of triplets comprising the reference image, a text modifier, and a target image that does not correspond to the reference image modified according to the text modifier;

processing all the triplets of the training dataset by:

selecting a triplet of the training dataset;

inputting an encoder representation of the reference image of the selected triplet and an encoder representation of the text modifier of the selected triplet into a first text-image interaction neural network to generate a vector representation of a first text-image pair, the first text-image interaction neural network is trained to use an attention mechanism to weigh dimensions of the encoder representation of the reference image;

inputting an encoder representation of the target image of the selected triplet and the encoder representation of the text modifier of the selected triplet into a second text-image interaction neural network to generate a vector representation of a second text-image pair, the second text-image interaction neural network is trained to use the attention mechanism to weigh dimensions of the encoder representation of the candidate target image;

generating, using a compatibility neural network and based on the first vector representation of the first text-image pair and the vector representation of the second text-image pair, a prediction whether the selected triplet is obtained from the first set of triplets or the second set of triplets; and

calculating, based on the prediction, a loss value using a loss function; and

when all of the triplets have been processed, selectively modifying one or more learnable parameters of the neural networks based on the loss value;

wherein the attention mechanism is the text modifier.

24 . The method of claim 23 , wherein generating, from the first set of triplets, a second set of triplets, <r b , t b , m b >, includes at least one of:

pairing the triplets within the first set of triplets and, for each pair of triplets, swapping the respective target images;

pairing the triplets within the first set of triplets and, for each pair of triplets, swapping the respective reference images;

pairing the triplets within the first set of triplets and, for each pair of triplets, swapping the respective text modifiers;

for a selected triplet within the first set of triplets, swapping respective target images with another triplet within the first set of triplets;

for a selected triplet within the first set of triplets, swapping respective reference images with another triplet within the first set of triplets; and

for a selected triplet within the first set of triplets, swapping respective text modifiers with another triplet within the first set of triplets.

25 . The method of claim 23 , wherein generating the prediction comprises computing a cosine similarity between the vector representation of the first text-image pair and the vector representation of the second text-image pair.

26 . The method of claim 25 , wherein generating the prediction further comprises computing a cosine similarity between the vector representation of the second text-image pair and the encoder representation of the text modifier, and

wherein the prediction is based on a sum of the cosine similarity between the vector representation of the first text-image pair and the vector representation of the second text-image pair and the cosine similarity between the vector representation of the second text-image pair and the encoder representation of the text modifier.

27 . The method of claim 23 , wherein the first text-image interaction neural network uses a first attention mechanism on the text modifier and the reference image to generate the vector representation of the first text-image pair, and the second text-image interaction neural network uses a second attention mechanism on the text modifier and applied to the a respective candidate target image to generate the vector representation of the second text-image pair.

28 . The method of claim 23 , wherein generating the prediction comprises, by a compatibility neural network, generating the prediction based on the vector representation of the first text-image pair and the vector representation of the second text-image pair.

29 . A system for image retrieval with neural networks, the system comprising:

an interaction module comprising:

a first text-image interaction module configured to generate a vector representation of a first text-image pair based on (a) an encoder representation of a reference image and (b) an encoder representation of a text modifier, the reference image and the text modifier received from a computing device,

wherein (i) the text modifier describes at least one visual difference from the reference image, and (ii) the first text-image interaction module is configured to generate the vector representation of the first text-image pair using a first neural network with attention that guides according to its input; and

a second text-image interaction module configured to generate a vector representation of a second text-image pair based on (a) the encoder representation of the text modifier and (b) an encoder representation of a candidate target image,

wherein the second text-image interaction module is configured to generate the vector representation of the second text-image pair using a second neural network with attention that guides according to the complement of its input;

a compatibility module configured to compute, based on (a) the vector representation of the first text-image pair and (b) the vector representation of the second text-image pair, a compatibility score for a triplet including the reference image, the text modifier, and the candidate target image; and

a ranking module configured to;

rank a set of candidate target images including the candidate target image by compatibility scores of the candidate target images, respectively, of the set; and

select at least one of the candidate target images from the set based on the ranks of the candidate target images, respectively; and

transmit the selected at least one of the candidate target images to the computing device.

30 . A method for image retrieval with neural networks, the method comprising:

generating a vector representation of a first text-image pair based on (a) an encoder representation of a reference image and (b) an encoder representation of a text modifier using a first neural network with attention that guides according to its input, the reference image and the text modifier received from a computing device,

wherein the text modifier describes at least one visual difference from the reference image;

generating a vector representation of a second text-image pair based on (a) the encoder representation of the text modifier and (b) an encoder representation of a candidate target image using a second neural network with attention that guides according to the complement of its input;

computing, based on (a) the vector representation of the first text-image pair and (b) the vector representation of the second text-image pair, a compatibility score for a triplet including the reference image, the text modifier, and the candidate target image; and

ranking a set of candidate target images including the candidate target image by compatibility scores of the candidate target images, respectively, of the set;

selecting at least one of the candidate target images from the set based on the ranks of the candidate target images, respectively; and

transmitting the selected at least one of the candidate target images to the computing device.

31 . A method of training neural networks to retrieve images, the method:

generating a training dataset, comprising:

obtaining a first set of triplets, each triplet of the first set of triplets comprising a reference image, a text modifier, and a target image, corresponding to the reference image modified according to the text modifier, wherein the text modifier describes at least one visual difference from the reference image; and

generating, from the first set of triplets, a second set of triplets, each triplet of the second set of triplets comprising the reference image, a text modifier, and a target image that does not correspond to the reference image modified according to the text modifier;

processing all the triplets of the training dataset by:

selecting a triplet of the training dataset;

inputting an encoder representation of the reference image of the selected triplet and an encoder representation of the text modifier of the selected triplet into a first text-image interaction neural network to generate a vector representation of a first text- image pair, the first text-image interaction neural network configured to generate the vector representation of the first text-image pair using a first neural network with attention that guides according to its input;

inputting an encoder representation of the target image of the selected triplet and the encoder representation of the text modifier of the selected triplet into a second text-image interaction neural network to generate a vector representation of a second text-image pair, the second text-image interaction neural network configured to generate the vector representation of the second text-image pair using a second neural network with attention that guides according to the complement of its input;

generating, using a compatibility neural network and based on the first vector representation of the first text-image pair and the vector representation of the second text-image pair, a prediction whether the selected triplet is obtained from the first set of triplets or the second set of triplets; and

calculating, based on the prediction, a loss value using a loss function; and

when all of the triplets have been processed, selectively modifying one or more learnable parameters of the neural networks based on the loss value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2022
From: SAMPAIO DE REZENDE, RAFAEL; LARLUS, DIANE; DELMAS, GINGER; CSURKA KHEDARI, GABRIELA
To: CORPORATION, NAVER
Reel/Frame 060294/0170 →
Priority Claims (1)
EP 21306114 · Aug 12, 2021 · regional
Continuity (1)
Related Publication 20230073843A1 · Mar 9, 2023
References Cited (49)
US 10678846B2 · Gordo Soldevila et al. · 2020 [cited by applicant]
US 11138469B2 · Almazan et al. · 2021 [cited by applicant]
US 20180373955A1 · Soldevila et al. · 2018 [cited by applicant]
US 20210073252A1 · Guo et al. · 2021 [cited by applicant]
CN 110222560A · 2019 [cited by examiner]
CN 112818140A · 2021 [cited by examiner]
EP 3754548A1 · 2020 [cited by applicant]
WO WO2019011936A1 · 2019 [cited by applicant]
Liu, Z.—“Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models”—arXiv—Aug. 9, 2021—pp. 1-20 (Year: 2021). [cited by examiner]
Antol, Stanislaw, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. “VQA: Visual Question Answering.” In 2015 [cited by applicant]
Anwaar M. U. et al., “Compositional learning of text-image query for image retrieval”, Proc. WACV, pp. 1140-1149, 2021. [cited by applicant]
Berg. T. L. et al., “Automatic attribute discovery and characterization from noisy web data”, [cited by applicant]
Chen Y. et al., “Learning join visual semantic matching embeddings for language-guided retrieval”, Proc. ECCV, 2020. [cited by applicant]
Chen, Yanbei, Shaogang Gong, and Loris Bazzani. “Image Search With Text Feedback by Visiolinguistic Attention Learning.” In 2020 [cited by applicant]
Cho, Kyunghyun, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. “Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation.” [cited by applicant]
Chun S. et al., “Probabilistic embeddings for cross-modal retrieval, [cited by applicant]
Culpepper, Jack, Eric Dodds, and Simao Herdade. “Transforming Image Representations via Attribute Operators,” Technical report from the Fashion IQ Challenge. 2019. [cited by applicant]
European Search Report for European Application No. 21306114.6 mailed Feb. 4, 2022. [cited by applicant]
Faghri, F. et al., “VSE++: Improving visual-semantic embeddings with hard negatives”, [cited by applicant]
Frome. A. et al., Devise: A deep visual-semantic embedding model, [cited by applicant]
Gao Y. et al., “Fashion IQ Challenge”, Third Workshop on Computer Vision for Fashion, Art and Design, CVPRW, 2020. https://sites.google.com/view/cvcreative2020/fashion-ig. [cited by applicant]
Guo X. et al., “The Fashion IQ Dataset: Retrieving images by combining side information and relative natural language feedback”, [cited by applicant]
Guo Xioaxiao et al., “Dialog-based interactive image retrieval” in S. Bengio et al., editors, [cited by applicant]
Han, Xintong, Zuxuan Wu, Phoenix X Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S Davis. “Automatic Spatially-Aware Fashion Concept Discovery,” ICCV, 2017. [cited by applicant]
He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. “Deep Residual Learning for Image Recognition.” In 2016 [cited by applicant]
Katrien Laenen et al., “Cross-modal search for fashion attributes”, Proceedings of the KDD 2017 Workshop on Machine Learning Meets Fashion, Aug. 14, 2017, XP055502261. [cited by applicant]
Kim, Jongseok, Youngjae Yu, and Seunghwan Lee. “Cycled Compositional Learning between Images and Text,” Technical report from the Fashion IQ Challenge. 2020. [cited by applicant]
Kingma D.P. and Ba J., “Adam: A method for stochastic optimization”, [cited by applicant]
Lee, Kuang-Huei, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. “Stacked Cross Attention for Image-Text Matching.” In [cited by applicant]
Li, Jianri, Jae-whan Lee, Woo-sang Song, Ki-young Shin, and Byung-hyun Go. “Designovel's System Description for Fashion-IQ Challenge 2019.” [cited by applicant]
Liu Y. et al., “Use what you have: Video retrieval using representations from collaborative experts”, Proc. BMVC, 2019. [cited by applicant]
Liu, Xin, Jiancheng Li, Jiaqi Wang, and Ziwei Liu. “MMFashion: An Open-Source Toolbox for Visual Fashion Analysis.” [cited by applicant]
Liu, Zheyuan, Cristian Rodriguez-Opazo, and Stephen Gould. “ResEFNet: Technical Report for the Fashion-IQ Interactive Image Retrieval Challenge,” Technical report from the Fashion IQ Challenge. 2019. [cited by applicant]
Liu, Ziwei, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. “DeepFashion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations.” In 2016 [cited by applicant]
Musgrave K. et al., “A metric learning reality check”, [cited by applicant]
Pennington, Jeffrey, Richard Socher, and Christopher Manning. “Glove: Global Vectors for Word Representation.” In [cited by applicant]
Perez, Ethan, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron Courville. “FiLM: Visual Reasoning with a General Conditioning Layer.” [cited by applicant]
Rostamzadeh, Negar, Seyedarian Hosseini, Thomas Boquet, Wojciech Stokowiec, Ying Zhang, Christian Jauvin, and Chris Pal. “Fashion-Gen: The Generative Fashion Dataset and Challenge.” [cited by applicant]
Russakovsky, Olga, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, et al. “ImageNet Large Scale Visual Recognition Challenge.” [cited by applicant]
Santoro, Adam, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap. “A Simple Neural Network Module for Relational Reasoning,” In Advances in neural information proc… [cited by applicant]
Selvaraju R.R. et al., Grad-cam: Visual explanations from deep networks via gradient-based localization, [cited by applicant]
Shin, Minchul, Yoonjae Cho, and Seongwuk Hong. “Fashion-IQ 2020 Challenge 2nd Place Team's Solution,” Technical report from the Fashion IQ Challenge. 2020. [cited by applicant]
Song, Yale, and Mohammad Soleymani. “Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval.” In 2019 [cited by applicant]
Tan M et al., “Efficientnet: Rethinking model scaling for convolutional neural networks”, [cited by applicant]
Vinyals, Oriol, Alexander Toshev, Samy Bengio, and Dumitru Erhan. “Show and Tell: A Neural Image Caption Generator.” In 2015 [cited by applicant]
Vo Nam et al., “Composing Text and Image for Image Retrieval—an Empirical Odyssey”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 15, 2019, pp. 6432-6441, XP033686476. [cited by applicant]
Wu, Hui, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah, Steven Rennie, Kristen Grauman, and Rogerio Feris. “Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback.” ArXiv:1905.12794 [Cs], Nov. 25, 20… [cited by applicant]
Yu, Youngjae, Seunghwan Lee, and Yuncheol Choi. “CurlingNet: Compositional Learning between Images and Text for Fashion Data,” Technical report from the Fashion IQ Challenge. 2019. [cited by applicant]
Zhao Y. et al., RUC-AIM3: Improved TIRG Model for Fashion-IQ Challenge 2020, Technical report from the Fashion IQ Challenge, 2020. [cited by applicant]