IP Library Granted Patent US 12,700,028
Granted Patent B2
US 12,700,028 · App. 18/167,002 · Granted Aug 4, 2026

Multi-modal product embedding generator

Inventors: Paul Baltescu (San Mateo, CA); Andrew Huan Zhai (San Mateo, CA); Haoyu Chen (Sunnyvale, CA); Jurij Leskovec (Stanford, CA); Nikil Pancha (San Francisco, CA); Charles Joseph Rosenberg (Cupertino, CA)
Assignee: Pinterest, Inc.
G06Q30/0631G06F16/9024G06F16/9038
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,700,028
App. No.
18/167,002
Filed
Feb 9, 2023
Granted
Aug 4, 2026
Kind
B2
Art Unit
3688
USPC
705/26.7
Abstract

Described are systems and methods for providing a multi-tasked trained machine learning model that may be configured to generate product embeddings from multiple types of product information. The exemplary product embeddings may be generated for a corpus of products (e.g., products included in a product catalog, etc.) based on both image information and text information associated with each respective product. Accordingly, the generated product embeddings may be compatible with learned representations of the different types of product information (e.g., image information, text information, etc.) and may be used to create a product index, which can be used to determine and serve product recommendations in connection with multiple different recommendation services that may be configured to receive different types of inputs (e.g., a single image, multiple images, text-based information, etc.).

Claims (48)

1 . A computer-implemented method, comprising:

accessing a plurality of training data including a plurality of query and engaged product pairs, wherein the plurality of query and engaged product pairs includes queries of different modalities including a plurality of image queries, and a plurality of text queries, and a plurality of engaged products associated with a plurality of engagement types;

training a first machine learning model using the plurality of training data to configure the first machine learning model to generate, based at least in part on at least one of image information or text information associated with a product, a product embedding representative of the product according to multiple task objectives, the product embedding being compatible with multiple modalities such that a distance measure between the product embedding and respective embeddings of different modalities represents a relevance of the product embedding to the respective embeddings;

generating an index based at least in part on a plurality of product embeddings generated by the first trained machine learning model, wherein the plurality of product embeddings corresponds to a first plurality of products, and wherein the plurality of product embeddings are compatible with different query modalities;

receiving a query that includes a representation of at least one of an image query, a text query, or a product query; and

determining, based at least in part on an embedding of the representation and the index, a second plurality of products from the first plurality of products as recommended products, the determining comprising calculating a distance measure between the embedding of the representation and one or more product embeddings in the index.

2 . The computer-implemented method of claim 1 , wherein the plurality of product embeddings are compatible with a first learned representation of an image and a second learned representation of a text string.

3 . The computer-implemented method of claim 2 , wherein the first learned representation of the image and the second learned representation of the text string are incompatible.

4 . The computer-implemented method of claim 1 , wherein one or more product embeddings of the plurality of product embeddings are generated based at least in part on a plurality of image information associated with a respective product of the first plurality of products and a plurality of text information associated with the respective product of the first plurality of products.

5 . The computer-implemented method of claim 1 , wherein determining the second plurality of products includes performing a nearest neighbor technique.

6 . A computing system, comprising:

one or more processors; and

a memory storing program instructions that, when executed by the one or more processors, cause the one or more processors to at least:

receive a plurality of product information associated with a plurality of products, wherein one or more products of the plurality of products is associated with product information that includes image product information and text product information;

determine, using a trained product embedding generator and based at least in part on the image product information and text product information associated with one or more products of the plurality of products, a respective product embedding for one or more products of the plurality of products, the respective product embeddings being compatible with multiple modalities such that a distance measure between a product embedding and respective embeddings of different modalities represents a relevance of the product embedding to the respective embeddings;

generate, based at least in part on the respective product embeddings for the plurality of products, a product index;

receive a request for recommended products from one of a plurality of recommendation services, the plurality of recommendation services including two or more recommendation services configured to receive respective inputs of different modalities, wherein the request includes one of a plurality of modalities including: an image query, a text query, or a product query;

determine, based at least in part on an embedding of the image query or an embedding of the text query and the product index, a plurality of recommended products, the determining comprising calculating a distance measure between the embedding of the image query or the text query and one or more product embeddings in the index; and

provide at least a portion of the plurality of recommended products as a response to the request.

7 . The computing system of claim 6 , wherein the image product information includes a plurality of image embeddings corresponding to a plurality of images associated with the product and the text product information includes a plurality of text strings associated with the product.

8 . The computing system of claim 7 , wherein the plurality of text strings are processed using a hash embedding technique.

9 . The computing system of claim 6 , wherein:

the trained product embedding generator is trained using positive engagements associated with a plurality of engagement types; and

the plurality of engagement types includes at least one of:

clicking a first product;

saving a second product;

adding a third product to a cart; or

purchasing a fourth product.

10 . The computing system of claim 6 , wherein:

the product index supports a plurality of recommendation services; and

at least one of the plurality of recommendation services determines the plurality of recommended products.

11 . The computing system of claim 6 , wherein respective product embeddings are compatible with a first learned representation of an image and a second learned representation of a text string, such that a distance between a first product embedding and the first learned representation and the second learned representation represents a relevance of the first learned representation and the second learned representation to the first product embedding.

12 . The computing system of claim 6 , wherein the product index includes a hierarchical navigable small worlds (HNSW) graph.

13 . The computing system of claim 6 , wherein determining the plurality of recommended products includes performing a nearest neighbor search technique.

14 . The computing system of claim 6 , wherein the instructions, when executed by the one or more processors, further cause the one or more processors to at least:

process at least a portion of the respective product embeddings using a ranking service to determine a ranking associated with the plurality of recommended products; and

cause at least a portion of the plurality of recommended products to be presented to a user in accordance with the ranking.

15 . The computing system of claim 6 , wherein the instructions, when executed by the one or more processors, further cause the one or more processors to at least process at least a portion of the respective product embeddings using a trained classifier to determine a further product information associated with at least one of the plurality of products.

16 . One or more non-transitory computer readable storage media having instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

accessing a plurality of training data including a plurality of query and engaged product pairs, wherein the plurality of query and engaged product pairs includes queries of different modalities including a plurality of image queries, and a plurality of text queries, and a plurality of engaged products associated with a plurality of engagement types;

training a first machine learning model using the plurality of training data to configure the first machine learning model to generate, based at least in part on at least one of image information or text information associated with a product, a product embedding representative of the product according to multiple task objectives, the product embedding being compatible with multiple modalities such that a distance measure between the product embedding and respective embeddings of different modalities represents a relevance of the product embedding to the respective embeddings;

generating an index based at least in part on a plurality of product embeddings generated by the first trained machine learning model, wherein the plurality of product embeddings corresponds to a first plurality of products, and wherein the plurality of product embeddings are compatible with different query modalities;

receiving a query that includes a representation of at least one of an image query, a text query, or a product query; and

determining, based at least in part on an embedding of the representation and the index, a second plurality of products from the first plurality of products as recommended products, the determining comprising calculating a distance measure between the embedding of the representation and one or more product embeddings in the index.

17 . The one or more non-transitory computer readable storage media of claim 16 , wherein the plurality of product embeddings are compatible with a first learned representation of an image and a second learned representation of a text string.

18 . The one or more non-transitory computer readable storage media of claim 17 , wherein the first learned representation of the image and the second learned representation of the text string are incompatible.

19 . The one or more non-transitory computer readable storage media of claim 16 , wherein one or more product embeddings of the plurality of product embeddings are generated based at least in part on a plurality of image information associated with a respective product of the first plurality of products and a plurality of text information associated with the respective product of the first plurality of products.

20 . The one or more non-transitory computer readable storage media of claim 16 , wherein determining the second plurality of products includes performing a nearest neighbor technique.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2023
From: BALTESCU, PAUL; ZHAI, ANDREW HUAN; CHEN, HAOYU; LESKOVEC, JURIJ; PANCHA, NIKIL; ROSENBERG, CHARLES JOSEPH
To: PINTEREST, INC.
Reel/Frame 062646/0845 →
Continuity (2)
Provisional Application 63308921 · Feb 10, 2022
Related Publication 20230252550A1 · Aug 10, 2023
References Cited (50)
US 10282462B2 · Magnani et al. · 2019 [cited by applicant]
US 10642887B2 · Chen et al. · 2020 [cited by applicant]
US 10956787B2 · Rothberg et al. · 2021 [cited by applicant]
US 11301774B2 · Duran et al. · 2022 [cited by applicant]
US 11620512B2 · Jain · 2023 [cited by examiner]
US 11769193B2 · Klein · 2023 [cited by examiner]
US 20190065594A1 · Lytkin · 2019 [cited by examiner]
US 20190236394A1 · Price et al. · 2019 [cited by applicant]
US 20200311798A1 · Forsyth · 2020 [cited by examiner]
US 20210241343A1 · Arora · 2021 [cited by examiner]
US 20230074782A1 · Tendulkar · 2023 [cited by examiner]
US 20260017681A1 · Ruan · 2026 [cited by examiner]
Beal, J., et al., “Billion-Scale Pretraining with Vision Transformers for Multi-Task Visual Representations,” CoRR abs/2108.05887 (2021). arXiv:2108.05887 URL: https://arxiv.org/abs/2108.05887, 10 pages. [cited by applicant]
Chen, YC., et al., “UNITER: Learning Universal Image-Text Representations,” (2019). arXiv:1909.11740v1 URL:https://arxiv.org/abs/1909.11740v1, 13 pages. [cited by applicant]
Cormode, Graham and Shan Muthukrishnan, “An Improved Data Stream Summary: The Count-Min Sketch and its Applications,” Journal of Algorithms 55, 1 (2005), 58-75. https://doi.org/10.1016/j.jalgor.2003.12.001. [cited by applicant]
Covington, P. et al., Deep Neural Networks for YouTube Recommendations. In RecSys, pp. 191-198, 2016, https://cseweb.ucsd.edu/classes/fa17/cse291-b/reading/p191-covington.pdf. [cited by applicant]
Devlin, J et al., 2018, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. CoRR abs/1810.04805 (2018), arXiv:1810.04805, Retrieved: https://arxiv.org/pdf/1810.04805v1.pdf, 14 pages. [cited by applicant]
Eksombatchai, C., et al., “Pixie: A System for Recommending 3+ Billion Items to 200+ Million Users in Real-Time,” In Proceedings of the 2018 World Wide Web Conference. 1775-1784. [cited by applicant]
Guo, W., et al., “Deep Multimodal Representation Learning: A Survey,” IEEE Access 7 (2019), 63373-63394. [cited by applicant]
Guo, W., et al., “DeText: A Deep Text Ranking Framework with BERT,” CoRR abs/2008.02460 (2020). arXiv:2008.02460 URL: https://arxiv.org/abs/2008.02460, 8 pages. [cited by applicant]
Hamilton, W.L., et al., “Inductive Representation Learning on Large Graphs,” 2018. arXiv:1706.02216 [cs.SI] URL: https://arxiv.org/abs/1706.02216, 19 pages. [cited by applicant]
Hu, Z., et al., “Heterogeneous Graph Transformer,” WWW '20, Apr. 20-24, 2020, Taipei, Taiwan arXiv:2003.01332 [cs.LG] URL: https://arxiv.org/abs/2003.01332, 11 pages. [cited by applicant]
Huang, Jing, and Brian Kingsbury, “Audio-visual Deep Learning for Noise Robust Speech Recognition,” In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 7596-7599. [cited by applicant]
Huang, JT., et al., “Embedding-based Retrieval in Facebook Search,” In Proceedings of the 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '20), Aug. 23-27, 2020, Virtual Event, CA, USA. ACM, New Y… [cited by applicant]
Huang, PS., et al., “Learning Deep Structured Semantic Models for Web Search Using Clickthrough Data,” In Proceedings of the 22nd ACM International Conference on Information and Knowledge Management (San Francisco, Cali… [cited by applicant]
Jiang, YG., et al., “Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video Classification,” IEEE Transactions on Multimedia 20, 11 (2018), 3137-3147. [cited by applicant]
Li, L.H., et al., “VisualBERT: A Simple and Performant Baseline for Vision and Language,” arXiv preprint arXiv:1908.03557 (2019) URL: https://arxiv.org/abs/1908.03557, 14 pages. [cited by applicant]
Liu, X., et al., “Representation Learning Using Multi-Task Deep Neural Networks for Semantic Classification and Information Retrieval,” Human Language Technologies: The 2015 Annual Conference of the North American Chapt… [cited by applicant]
Liu, Y., et al., “Multimodal Video Classification with Stacked Contractive Autoencoders,” Signal Processing 120 (2016), 761-766. [cited by applicant]
Lu, H., et al., , “Graph-based Multilingual Product Retrieval in E-commerce Search,” CoRR abs/2105.02978 (2021). arXiv:2105.02978 URL: https://arxiv.org/abs/2105.02978, 8 pages. [cited by applicant]
Malkov, Yu. A., and D. A. Yashunin, “Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs,” IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 4 (2018)… [cited by applicant]
Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H. and Ng, A. Y., “Multimodal Deep Learning,” In International Conference on Machine Learning (ICML), pp. 689-696, 2011, 8 pages. [cited by applicant]
Nigam, P., et al., “Semantic Product Search,” KDD '19, Aug. 4-8, 2019, Anchorage, AK, USA. CoRR abs/1907.00937 arXiv:1907.00937 URL: http://arxiv.org/abs/1907.00937, 10 pages. [cited by applicant]
Pal, A., et al., “PinnerSage: Multi-Modal User Embedding Framework for Recommendations at Pinterest,” In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Jul. 2020), (KDD… [cited by applicant]
Ruder, Sebastian, “An Overview of Multi-Task Learning in Deep Neural Networks,” arXiv preprint arXiv:1706.05098 (2017) URL: https://arxiv.org/abs/1706.05098, 14 pages. [cited by applicant]
Sanh, V., et al., “DistilBERT, A Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter,” 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada, arXiv preprint arXiv:1910.01… [cited by applicant]
Saqur, R. and K. Narasimhan, “Multimodal Graph Networks for Compositional Generalization in Visual Question Answering,” NIPS'20, Proceedings of the 34th International Conference on Neural Information Processing Systems,… [cited by applicant]
Svenstrup, D.T., et al., “Hash Embeddings for Efficient Word Representations,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA. arXiv preprint arXiv:1709.03933, URL: https://arx… [cited by applicant]
Tan, Hao and Mohit Bansal, “LXMERT: Learning Cross-Modality Encoder Representations from Transformers,” EMNLP 2019, Conference on Empirical Methods in Natural Language Processing, arXiv preprint arXiv:1908.07490 (2019) … [cited by applicant]
Tang, H., et al., “Progressive Layered Extraction (PLE): A Novel Multi-Task Learning (MTL) Model for Personalized Recommendations,” RecSys '20: Proceedings of the 14th ACM Conference on Recommender Systems, Sep. 2020, 2… [cited by applicant]
Tian, H., et al., “Multimodal Deep Representation Learning for Video Classification,” World Wide Web 22, 3 (2019), 1325-1341. [cited by applicant]
Virani, A., et al., “Lessons Learned Addressing Dataset Bias in Model-Based Candidate Generation at Twitter,” KDD IRS2020, Aug. 23-28, 2020, San Diego, CA, CoRR abs/2105.09293 (2021), arXiv:2105.09293, URL: https://arxi… [cited by applicant]
Yang, J., et al., “Mixed Negative Sampling for Learning Two-tower Neural Networks in Recommendations,” In WWW 20: Companion Proceedings of the Web Conference 2020, Apr. 2020, pp. 441-447, URL: https://doi.org/10.1145/33… [cited by applicant]
Yi, X., et al., “Sampling-Bias-Corrected Neural Modeling for Large Corpus Item Recommendations,” In Thirteenth ACM Conference on Recommender Systems (RecSys '19), Sep. 16-20, 2019, Copenhagen, Denmark. ACM, New York, NY… [cited by applicant]
Ying, R., et al., “Graph Convolutional Neural Networks for Web-Scale Recommender Systems,” In KDD '18: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2018, pp. 974-983,… [cited by applicant]
Zhai, A. et al., 2019, Learning a Unified Embedding for Visual Search at Pinterest. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2019, Anchorage, AK, USA, Aug.… [cited by applicant]
Zhang, H., et al., “Towards Personalized and Semantic Retrieval: An End-to-End Solution for E-commerce Search via Embedding Learning,” CoRR abs/2006.02282 (2020). arXiv:2006.02282 URL: https://arxiv.org/abs/2006.02282, … [cited by applicant]
Zhao, Z., et al., “Recommending What Video to Watch Next: A Multitask Ranking System,” In RecSys '19: Proceedings of the 13th ACM Conference on Recommender Systems, Sep. 2019, pp. 43-51, URL: https://doi.org/10.1145/329… [cited by applicant]
Zhou, L., et al., “Unified Vision-Language Pre-Training for Image Captioning and VQA,” In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34. 13041-13049. [cited by applicant]
Zhuang, J., et al., “PinText 2: Attentive Bag of Annotations Embedding,” In DLP-KDD 2020, Aug. 24, 2020, San Diego, California, USA, 9 pages. [cited by applicant]