IP Library › Granted Patent US 12,592,059
Granted Patent B2
US 12,592,059 · App. 17/983,327 · Granted Mar 31, 2026

Global embedding learning from different modalities

Inventors: Baohao Liao (Aachen, DE); Sanjika Hewavitharana (San Jose, CA); Michael Damian Kozielski (Aachen, DE); Friedrich Leonard Dahlmann (Aachen, DE); Shahram Khadivi (Herzogenrath, DE)
Assignee: EBAY INC.
G06V10/774G06F40/284G06Q30/0643G06V10/761G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,592,059
App. No.
17/983,327
Granted
Mar 31, 2026
Kind
B2
Abstract

A system may receive, via a first user interface associated with an online marketplace, a multi-modality request to retrieve a listing for an item, the multi-modality request comprising at least a first image and a first natural language text associated with the item. The system may generate an item embedding based on inputting the first image and the first natural language text to a machine learning model and may generate a first vector associated with the first image and the first natural language text included in the multi-modality request. The system may cause presentation, via the first user interface associated with the online marketplace, of one or more listings for the item retrieved based at least in part on a similarity metric between the first vector and a second vector of a plurality of vectors associated with a plurality of listings.

Claims (87)

1 . A computer-implemented method, comprising:

generating training data comprising training titles and training images, one or more training titles having a masked title portion and one or more training images have a masked image portion, the generating of the training data comprising:

masking a portion of a first training image;

generating a predicted portion for the first training image based at least in part on masking the portion of the first training image and a title of the first training image; and

generating a predicted listing category for the first training image based at least in part on the predicted portion for the first training image;

training a machine learning model based on the training data;

receiving, via a first user interface associated with an online marketplace, a multi-modality request to retrieve a listing for an item, the multi-modality request comprising at least a first image and a first natural language text associated with the item;

generating, by the machine learning model, a first image token based at least in part on the first image and a first title token based at least in part on the first natural language text;

generating, by the machine learning model, a first vector based on the first image token and the first title token; and

causing presentation, via the first user interface associated with the online marketplace, of one or more listings for the multi-modality request based at least in part on a similarity metric between the first vector and each vector of a plurality of vectors associated with a plurality of listings.

2 . The computer-implemented method of claim 1 , wherein generating the training data further comprises:

masking a portion of the title of the first training image;

generating a predicted portion for a title of the listing based at least in part on the portion of the title of the first training image and the first training image; and

generating a predicted listing category based at least in part on the predicted portion for the title.

3 . The computer-implemented method of claim 1 , further comprising:

determining, by the machine learning model, a similarity metric between the first vector and a plurality of vectors associated with a plurality of categories; and

associating, by the machine learning model, the first vector with a first category based at least in part on a similarity metric between the first vector and one or more vectors classified by the machine learning model as being associated with the first category.

4 . The computer-implemented method of claim 3 , further comprising:

generating, by the machine learning model, a second vector associated with a second image and a second natural language text included in a received multi-modality query;

associating, by the machine learning model, the second vector with the first category; and

comparing the second vector with a plurality of vectors that include the first vector, wherein the first vector is associated with the first category based at least in part on the first vector and the second vector satisfying a similarity metric.

5 . The computer-implemented method of claim 1 , wherein causing presentation of the one or more listings for the item comprises:

causing presentation, via a second user interface associated with the online marketplace, of one or more listings for the item from a first category, wherein the one or more listings that comprise a second image that is different than the first image, a second natural language text that differs from the first natural language text, or both.

6 . The computer-implemented method of claim 1 , further comprising:

comparing the portion of the training image with the portion of the training title based at least in part on reconstructing the portion of the training image and reconstructing the portion of the training title; and

classifying the listing for the item using the training image and the training title based at least in part on the portion of the training image being associated with the portion of the training title.

7 . The computer-implemented method of claim 1 , further comprising:

associating the first vector with a product category based at least in part on the similarity metric between the first vector and vectors classified as being associated with the product category.

8 . An apparatus, comprising:

a processor;

memory coupled with the processor; and

instructions stored in the memory and executable by the processor to cause the apparatus to:

generating training data comprising training titles and training images, one or more training titles having a masked title portion and one or more training images have a masked image portion, the generating of the training data comprising:

masking a portion of a first training image:

masking a portion of a first training image;

generating a predicted portion for the first training image based at least in part on masking the portion of the first training image and a title of the first training image; and

generating a predicted listing category for the first training image based at least in part on the predicted portion for the first training image;

train a machine learning model based on the training data;

receive, via a first user interface associated with an online marketplace, a multi-modality request to retrieve a listing for an item, the multi-modality request comprising at least a first image and a first natural language text associated with the item;

generate a first image token based at least in part on the first image and a first title token based at least in part on the first natural language text;

generate, by the machine learning model, a first vector based on the first image token and the first title token; and

cause presentation, via the first user interface associated with the online marketplace, of one or more listings for the multi-modality request based at least in part on a similarity metric between the first vector and each vector of a plurality of vectors associated with a plurality of listings.

9 . The apparatus of claim 8 , wherein generating the training data further comprises:

mask a portion of the title of the first training image;

generate a predicted portion for a title of the listing based at least in part on the portion of the title of the first training image and the first training image; and

generate a predicted listing category based at least in part on the predicted portion for the title.

10 . The apparatus of claim 8 , wherein the instructions are further executable by the processor to cause the apparatus to:

determine, by the machine learning model, a similarity metric between the first vector and a plurality of vectors associated with a plurality of categories; and

associate, by the machine learning model, the first vector with a first category based at least in part on a similarity metric between the first vector and one or more vectors classified by the machine learning model as being associated with the first category.

11 . The apparatus of claim 10 , wherein the instructions are further executable by the processor to cause the apparatus to:

generate, by the machine learning model, a second vector associated with a second image and a second natural language text included in a received multi-modality query;

associate, by the machine learning model, the second vector with the first category; and

compare the second vector with a plurality of vectors that include the first vector, wherein the first vector is associated with the first category based at least in part on the first vector and the second vector satisfying a similarity metric.

12 . The apparatus of claim 8 , wherein the instructions to cause presentation of the one or more listings for the item are executable by the processor to cause the apparatus to:

cause presentation, via a second user interface associated with the online marketplace, of one or more listings for the item from a first category, wherein the one or more listings that comprise a second image that is different than the first image, a second natural language text that differs from the first natural language text, or both.

13 . The apparatus of claim 8 , wherein the instructions are further executable by the processor to cause the apparatus to:

compare the portion of the training image with the portion of the training title based at least in part on reconstructing the portion of the training image and reconstructing the portion of the training title; and

classify the listing for the item using the training image and the training title based at least in part on the portion of the training image being associated with the portion of the training title.

14 . The apparatus of claim 8 , wherein the instructions are further executable by the processor to cause the apparatus to:

associate the first vector with a product category based at least in part on the similarity metric between the first vector and vectors classified as being associated with the product category.

15 . A non-transitory computer-readable medium storing code, the code comprising instructions executable by a processor to perform operations comprising:

generating training data comprising training titles and training images, one or more training titles having a masked title portion and one or more training images have a masked image portion, the generating of the training data comprising:

masking a portion of a first training image;

masking a portion of a first training image;

generating a predicted portion for the first training image based at least in part on masking the portion of the first training image and a title of the first training image; and

generating a predicted listing category for the first training image based at least in part on the predicted portion for the first training image;

training a machine learning model based on the training data;

receiving, via a first user interface associated with an online marketplace, a multi-modality request to retrieve a listing for an item, the multi-modality request comprising at least a first image and a first natural language text associated with the item;

generating a first image token based at least in part on the first image and a first title token based at least in part on the first natural language text;

generating, by the machine learning model, a first vector based on the first image token and the first title token; and

causing presentation, via the first user interface associated with the online marketplace, of one or more listings for the multi-modality request based at least in part on a similarity metric between the first vector and each vector of a plurality of vectors associated with a plurality of listings.

16 . The non-transitory computer-readable medium of claim 15 , wherein generating the training data further comprises:

mask a portion of the title of the first training image;

generate a predicted portion for a title of the listing based at least in part on the portion of the title of the first training image and the first training image; and

generate a predicted listing category based at least in part on the predicted portion for the title.

17 . The non-transitory computer-readable medium of claim 15 , wherein the instructions are further executable by the processor to perform operations comprising:

determining, by the machine learning model, a similarity metric between the first vector and a plurality of vectors associated with a plurality of categories; and

associating, by the machine learning model, the first vector with a first category based at least in part on a similarity metric between the first vector and one or more vectors classified by the machine learning model as being associated with the first category.

18 . The non-transitory computer-readable medium of claim 17 , wherein the instructions are further executable by the processor to perform operations comprising:

generating, by the machine learning model, a second vector associated with a second image and a second natural language text included in a received multi-modality query;

associating, by the machine learning model, the second vector with the first category; and

comparing the second vector with a plurality of vectors that include the first vector, wherein the first vector is associated with the first category based at least in part on the first vector and the second vector satisfying a similarity metric.

19 . The non-transitory computer-readable medium of claim 15 , wherein the instructions to cause presentation of the one or more listings for the item are executable by the processor to:

cause presentation, via a second user interface associated with the online marketplace, of one or more listings for the item from a first category, wherein the one or more listings that comprise a second image that is different than the first image, a second natural language text that differs from the first natural language text, or both.

20 . The non-transitory computer-readable medium of claim 15 , wherein the instructions are further executable by the processor to perform operations comprising:

comparing the portion of the training image with the portion of the training title based at least in part on reconstructing the portion of the training image and reconstructing the portion of the training title; and

classifying the listing for the item using the training image and the training title based at least in part on the portion of the training image being associated with the portion of the training title.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2024
From: LIAO, BAOHAO; HEWAVITHARANA, SANJIKA; KOZIELSKI, MICHAEL DAMIAN; DAHLMANN, FRIEDRICH LEONARD; KHADIVI, SHAHRAM
To: EBAY INC.
Reel/Frame 066677/0447 →
Continuity (1)
Related Publication 20240153246A1 · May 9, 2024
References Cited (65)
US 10970768B2 · Zheng et al. · 2021 [cited by applicant]
US 11301732B2 · Hu et al. · 2022 [cited by applicant]
US 11341366B2 · Niu et al. · 2022 [cited by applicant]
US 12141236B1 · Arici · 2024 [cited by examiner]
US 12210516B1 · Zhu · 2025 [cited by examiner]
US 12354011B2 · Biswas · 2025 [cited by examiner]
US 20140330822A1 · Makadia et al. · 2014 [cited by applicant]
US 20200356592A1 · Yada et al. · 2020 [cited by applicant]
US 20210350081A1 · De Peuter · 2021 [cited by examiner]
US 20220130499A1 · Zhou et al. · 2022 [cited by applicant]
US 20230401274A1 · Denninghoff · 2023 [cited by examiner]
CN 113792112A · 2021 [cited by applicant]
CN 113792113A · 2021 [cited by applicant]
CN 114399769A · 2022 [cited by applicant]
CN 114821223A · 2022 [cited by examiner]
CN 118012980 · 2024 [cited by applicant]
Antol et al., “VQA: Visual Question Answering”, In 2015 IEEE International Conference on Computer Vision (ICCV), Computer Vision Foundation (CVF), pp. 2425-2433 (9 pages). [cited by applicant]
Arici et al., “MLIM: Vision-and-Language Model Pre-training with Masked Language and Image Modeling”, arXiv:2109.12178v1 [cs.CV] Sep. 24, 2021, 7 pages. [cited by applicant]
Ba et al., “Layer Normalization”, arXiv:1607.06450v1 [stat.ML] Jul. 21, 2016, 14 pages. [cited by applicant]
Bianchi et al., “Query2Proc2Vec: Grounded Word Embeddings for eCommerce”, In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies… [cited by applicant]
Chen et al., “Microsoft COCO Captions: Data Collection and Evaluation Servicer”, arXiv:1504.00325v1 [cs.CV] Apr. 1, 2015, 7 pages. [cited by applicant]
Chen et al., “UNITER: Learning Universal Image-Text Representations”, arXiv:1909.11740v1 [cs.CV] Sep. 25, 2019, 13 pages. [cited by applicant]
Desai et al., “VirTex: Learning Visual Representations from Textual Annotations”, In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, Jun. 19-25, 2021, Computer Vision Foundation, pp. 11162-11173 (… [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Hu… [cited by applicant]
Dosovitskiy et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale”, In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021, 16 pages. [cited by applicant]
Fan et al., “Automatic Generation of Product-Image Sequence in E-commerce”, arXiv:12994v1 [cs.CV] Jun. 26, 2022, 9 pages. [cited by applicant]
Goyal et al., “Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering”, International Journal of Computer Vision, 127(4):398-414, Springer, 17 pages. [cited by applicant]
Grbovic et al., “E-commerce in Your Inbox: Product Recommendations at Scale”, In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '15, New York, NY, USA, Associatio… [cited by applicant]
Gui et al., “Training Vision-Language Transformers from Captions Alone”, arXiv:2205.09256v1 [cs.CV] May 19, 2022, 15 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, Jun. 27-30, 2016, Computer Vision Foundation, pp. 770-778 (9 … [cited by applicant]
Hendrycks et al., “Bridging Nonlinearities and Stochastic Regularizers with Gaussian Error Linear Units”, arX14:1606.08415v1 [cs.LG] Jun. 27, 2016, 6 pages. [cited by applicant]
Jing et al., “Visual search at Pinterest”, In Proceedings of the 21st ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Sydney, NSW, Australia, Aug. 10-13, 2015, arXiv:1505.07647, 10 pages. [cited by applicant]
Joshi et al., SpanBERT: Improving Pre-training by Representing and Predicting Spans, Transactions of the Association for Computational Linguistics (TACL), 8:64-77, 2020, 14 pages. [cited by applicant]
Kazemzadeh et al., “ReferItGame: Referring to Objects in Photographs of Natural Scenes”, In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP 2014), Doha, Qatar, Association f… [cited by applicant]
Keenan et al., “Global Ecommerce Explained: Stats and Trends to Watch in 2021”, downloaded from The Wayback Machine—https://web.archive/org/web/20220201005436/https://www.shopify.com/enterprise/global-ecommerce-statisti… [cited by applicant]
Kiefer et al., “Stochastic Estimation of the Maximum of a Regression Function”, The Annals of Mathematical Statistics, 22(3):462-466, 1952, 5 pages. [cited by applicant]
Kingma et al., “Adam: A Method for Stochastic Optimization,” In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, arXiv:1412.6980, 13 p… [cited by applicant]
Li et al., “Vision-Language Intelligence: Tasks, Representation Learning, and large Models”, arX14:2203.01922v1 [cs.CV] Mar. 3, 2022, 19 pages. [cited by applicant]
Li et al., “Unsupervised Vision-and-Language Pre-training Without Parallel Images and Captions”, arXiv:2010.12831v2 [cs.CL] Apr. 11, 2021, 12 pages. [cited by applicant]
Li et al., “VisualBERT: A Simple and Performant Baseline for Vision and Language”, arXiv:1908.03557v1 [cs.CV] Aug. 9, 2019, 14 pages. [cited by applicant]
Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach”, arXiv:1907.11692 v1 [cs.CL] Jul. 26, 2019, 13 pages. [cited by applicant]
Loshchilov et al., “Decoupled Weight Decay Regularization”, In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, 18 pages. [cited by applicant]
Lu et al., “VILBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks”, In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing … [cited by applicant]
Ma et al., “Are Multimodal Transformers Robust to Missing Modality?”, 2022, Computer Vision Foundation, pp. 18177-18186 (10 pages). [cited by applicant]
Manning et al., “Chapter 8: Evaluation in Information Retrieval”, An Introduction to Information Retrieval, Draft Apr. 1, 2009, online edition 2009 Cambridge University Press, Cambridge, England, 55 pages. [cited by applicant]
Ott et al., “FAIRSEQ: A Fast, Extensible Toolkit for Sequence Modeling”, In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,… [cited by applicant]
Plummer et al., “Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models”, In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, Dec. 7-13, 201… [cited by applicant]
Qi et al., “ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-text Data”, arXiv:2001.07966v2 [cs.CV] Jan. 23, 2020, 12 pages. [cited by applicant]
Radford et al., “Learning Transferable Visual Models from Natural Language Supervision”, In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, Jul. 18-24, 2021, Virtual Event, vol. 139 of P… [cited by applicant]
Sennrich et al., “Neural Machine Translation of Rare Words with Subword Units”, In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, Aug. 7-12, 2016, Berlin, Germany, vol… [cited by applicant]
Shankar et al., “Deep Learning Based Large Scale Visual Recommendation and Search for e-Commerce”, arXiv:1703.02344v1 [cs.CV] Mar. 7, 2017, 9 pages. [cited by applicant]
Sharma et al., “Conceptual captions: A Cleaned, Hypernymed, Image Alt-text Dataset for Automatic Image Captioning”, In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol. 1: Lon… [cited by applicant]
Song et al., “Deep Metric Learning via Lifted Structured Feature Embedding”, In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, Jun. 27-30, 2016, Computer Vision Foundatio… [cited by applicant]
Su et al., “VL-BERT: Pre-training of Generic Visual-Linguistic Representations”, In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, Apr. 26-30, 2020, 16 pages. [cited by applicant]
Suhr et al., “A Corpus for Reasoning About Natural Language Grounded in Photographs”, In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28-Aug. 2, 20… [cited by applicant]
Tan et al., “LXMERT: Learning Cross-Modality Encoder Representations from Transformers”, In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conferen… [cited by applicant]
Vaswani et al., “Attention is All You Need”, In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, Dec. 4-9, 2017, Long Beach, CA, USA, pp. 5998-6008 (… [cited by applicant]
Yang et al., “Visual search at eBay”, In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Halifax, NS, Canada, Aug. 13-17, 2017, arXiv:1706.03154, 10 pages. [cited by applicant]
Yang et al., “An Evaluation of Statistical Approaches to Text Categorization”, School of Computer Science, Pittsburgh, PA, Apr. 10, 1997, 12 pages. [cited by applicant]
You et al., “Image Captioning with Semantic Attention”, In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, Jun. 27-30, 2016, Computer Vision Foundation, pp. 4651-4659 (9 pages). [cited by applicant]
Yu et al., “CommerceMM: Large-Scale Commerce MultiModal Representation Learning with Omni Retrieval”, arXiv:2202.07247v1 [cs.CV] Feb. 15, 2022, 10 pages. [cited by applicant]
Zellers et al., “From recognition to cognition: Visual Commonsense Reasoning”, In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, Jun. 16-20, 2019, Computer Vision Foundation, pp. … [cited by applicant]
“European Application Serial No. 23206238.0, Extended European Search Report mailed Feb. 7, 2024”, 9 pgs. [cited by applicant]
Chen, Nawei, “A Survey of Indexing and Retrieval of Multimodal Documents: Text and Images”, [Online]. Retrieved from the Internet: URL: https: citeseerx.ist.psu.edu document?repid=replandtype=pdfanddoi=b324f14986cc10bf8… [cited by applicant]
“European Application Serial No. 23206238.0, Communication Pursuant to Article 94(3) EPC mailed Feb. 18, 2025”, 9 pgs. [cited by applicant]