IP Library › Granted Patent US 12,235,911
Granted Patent B2
US 12,235,911 · App. 18/239,791 · Granted Feb 25, 2025

Computer-based systems and methods for training and using a machine learning model for improved processing of user queries based on inferred user intent

Inventors: Srivatsa Mallapragada (Atlanta, GA); Ying Xie (Marietta, GA); Varsha Rani Chawan (Atlanta, GA); Zeyad Hailat (Atlanta, GA); Simon Hughes (Atlanta, GA); Yuanbo Wang (Austin, TX)
Assignee: Home Depot Product Authority, LLC
G06F16/9532G06F16/9535
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,235,911
App. No.
18/239,791
Granted
Feb 25, 2025
Kind
B2
Abstract

A method for providing document category recommendations may include training a machine learning algorithm; receiving, from a user, a selection of an anchor document; retrieving a user co-viewing sequence; generating, via the trained machine learning algorithm, a sequence embeddings set based on the anchor document and the user co-viewing sequence; comparing the generated embeddings set to respective embeddings for a plurality of candidate sets; determining the candidate set closest to the generated embeddings; and presenting at least one category from the closest candidate set.

Claims (94)

1. A computer-implemented method for document recommendation comprising:

receiving a query from a user;

generating, via a first trained machine learning model portion, a first embeddings vector representative of the query;

generating, via a second trained machine learning model portion, a second embeddings vector representative of a predicted feature of a potential document from a feature embeddings set that is responsive to the first embeddings vector of the query;

generating, via a third trained machine learning model portion, a third embeddings vector representative of an intent of the query that is responsive to the first embeddings vector of the query;

combining the first, second, and third embeddings vectors to generate a combined embeddings vector;

determining a document from a plurality of documents based on the combined embeddings vector; and

presenting the determined document, wherein:

the first trained machine learning model portion is trained using a text encoder, the second trained machine learning model portion is trained using the text encoder and an image encoder, and

the third trained machine learning model portion is trained based on the first and second trained machine learning model portions.

2. The method of claim 1 , wherein the first query embeddings vector is representative of at least one of a text of the query or of a projected image of the query.

3. The method of claim 1 , further comprising training the first machine learning model portion by:

retrieving a set of query-document interactions, each interaction comprising a past query and a text and an image of a past document responsive to the past query;

generating, for each interaction and by the first trained machine learning model portion, a query embeddings vector;

generating, for each interaction, a first document embeddings vector in a text vector space and a second document embeddings vector in an image vector space; and

training the first trained machine learning model based on a loss function that uses the first document embeddings vector, the second document embeddings vector, and the query embeddings vector.

4. The method of claim 3 , further comprising training the third machine learning model portion by:

quantizing a set of intent training embeddings vectors from a set of mean embeddings vectors for the set of query-document interactions, each mean embeddings vector determined based on the query embeddings vector, the first document embeddings vector, and the second document embeddings vector associated with a respective query-document interaction;

generating, by the third trained machine learning model portion, an intent embeddings vector;

retrieving, from the set of intent embeddings vectors, an intent training embeddings vector closest to the generated intent embeddings vector; and

adjusting the third trained machine learning model portion based on a weighted distance between the retrieved intent training embeddings vector and the generated intent embeddings vector.

5. The method of claim 4 , further comprising selecting a weight to adjust the trained third machine learning model portion closer to the retrieved intent training embeddings vector than to the generated intent embeddings vector and applying the selected weight to generate the weighted distance.

6. The method of claim 3 , further comprising training the second machine learning model portion by:

quantizing a set of first feature training embeddings vectors from a set of first document embeddings vectors for the set of query-document interactions;

quantizing a set of second feature training embeddings vectors from a set of second document embeddings vectors for the set of query-document interactions;

generating, by the second trained machine learning model portion, a first feature embeddings vector and a second feature embeddings vector; and

adjusting the second trained machine learning model portion based on a first distance between the first feature embeddings vector and a closest first feature training embeddings vector and on a second distance between the second feature embeddings vector and a closest second feature training embeddings vector.

7. The method of claim 1 , wherein combining the first, second, and third embeddings vectors comprises concatenating the first, second, and third embeddings vectors, and wherein the combined embeddings vector comprises a concatenated embeddings vector.

8. The method of claim 7 , wherein determining the document from the plurality of documents comprises:

inputting the concatenated embeddings vector into a causal transformer model to generate an output embeddings vector; and

determining the document as a closest document of the plurality of documents to the output embeddings vector.

9. A computer-implemented method for document recommendation comprising:

receiving a text query;

generating text embeddings representative of the text query;

projecting the text embeddings to generate image embeddings representative of a predicted feature from a feature embeddings set based on the text embeddings of the query;

retrieving user intent embeddings from a set of user intent embeddings based on the text embeddings of the query;

retrieving text feature embeddings from a set of text feature embeddings based on the text embeddings;

retrieving image feature embeddings from a set of image feature embeddings based on the image embeddings;

generating predicted document embeddings based on the text embeddings, the intent embeddings, the text feature embeddings, and the image feature embeddings; and

presenting a document from a set of documents based on the predicted document embeddings.

10. The method of claim 9 , wherein retrieving the user intent embeddings comprises:

retrieving a set of query-document interactions;

generating, for each interaction, interaction query embeddings and interaction document embeddings;

training an intent machine learning model based on a difference between an intent embeddings generated by the intent machine learning model and a closest mean embeddings of the interaction query embedding and the respective interaction document embeddings; and

inputting the received text query embeddings into the trained intent machine learning model to generate the intent embeddings.

11. The method of claim 9 , wherein retrieving the text feature embeddings and the image feature embeddings comprises:

retrieving a set of query-document interactions;

generating, for each interaction, an interaction text document embeddings and an interaction image document embeddings;

determining a first difference between a text feature embeddings generated by a text feature machine learning model and each of the interaction text document embeddings;

determining a second difference between an image feature embeddings generated by an image feature machine learning model and each of the interaction image document embeddings;

training both the text feature machine learning model and the image feature machine learning model based on the first and second differences; and

inputting the received text query embeddings into both the trained text feature machine learning model and the trained image feature machine learning model,

wherein a resultant output of the trained text feature machine learning model is the retrieved text feature embeddings, and a resultant output of the trained image feature machine learning model is the retrieved image feature embeddings.

12. The method of claim 9 , wherein generating the predicted document embeddings comprises:

concatenating the text feature embeddings and the image feature embeddings to generate feature embeddings;

concatenating the text embeddings, the intent embeddings, and the feature embeddings to generate concatenated query embeddings; and

inputting the concatenated query embeddings to an attention-based transformer model,

wherein an output of the attention-based transformer model comprises the predicted document embeddings.

13. A system for document recommendation, the system comprising:

a processor; and

a computer-readable media storing instructions that, when executed by the processor, cause the system to:

receive a query from a user;

generate, via a first trained machine learning model portion, a first embeddings vector representative of the query;

generate, via a second trained machine learning model portion, a second embeddings vector representative of a predicted feature of a potential document from a feature embeddings set based on the first embeddings vector that is responsive to the query;

generate, via a third trained machine learning model portion, a third embeddings vector representative of an intent of the query based on the first embeddings vector that is responsive to the query;

combine the first, second, and third embeddings vectors to generate a combined embeddings vector;

determine a document from a plurality of documents based on the combined embeddings vector; and

present the retrieved determined document,

wherein:

the first trained machine learning model portion is trained using a text encoder,

the second trained machine learning model portion is trained using the text encoder and an image encoder, and

the third trained machine learning model portion is trained based on the first and second trained machine learning model portions.

14. The system of claim 13 , wherein the first query embeddings vector is representative of at least one of a text of the query or of a projected image of the query.

15. The system of claim 13 , wherein the first trained machine learning model portion is trained by:

retrieving a set of query-document interactions, each interaction comprising a text of a past query and a text and an image of a past document corresponding to the past query;

generating, for each interaction and by the first trained machine learning model portion, a query embeddings vector;

generating, for each interaction, a first document embeddings vector in a text space and a second document embeddings vector in an image space;

projecting, for each interaction, the query embeddings vector into the text space and into the image space; and

training the first trained machine learning model based on a loss function that uses the first document embeddings vector, the second document embeddings vector, and the projected query embeddings vector.

16. The system of claim 15 , further configured to train the third machine learning model portion by:

quantizing a set of intent training embeddings vectors from a set of mean embeddings vectors for the set of query-document interactions, each mean embeddings vector determined based on the query embeddings vector, the first document embeddings vector, and the second document embeddings vector associated with a respective query-document interaction;

generating, by the third trained machine learning model portion, an intent embeddings vector;

retrieving, from the set of intent embeddings vectors, an intent training embeddings vector closest to the generated intent embeddings vector; and

adjusting the third trained machine learning model portion based on a weighted distance between the retrieved intent training embeddings vector and the generated intent embeddings vector.

17. The system of claim 16 , further configured to select a weight to adjust the trained third machine learning model portion closer to the retrieved intent training embeddings vector than to the generated intent embeddings vector, and to apply the selected weight to generate the weighted distance.

18. The system of claim 16 , further configured to train the second machine learning model portion by:

quantizing a set of first feature training embeddings vectors from a set of first document embeddings vectors for the set of query-document interactions;

quantizing a set of second feature training embeddings vectors from a set of second document embeddings vectors for the set of query-document interactions;

generating, by the second trained machine learning model portion, a first feature embeddings vector and a second feature embeddings vector; and

adjusting the second trained machine learning model portion based on a first distance between the first feature embeddings vector and a closest first feature training embeddings vector and on a second distance between the second feature embeddings vector and a closest second feature training embeddings vector.

19. The system of claim 13 , wherein combining the first, second, and third embeddings vectors comprises concatenating the first, second, and third embeddings, and wherein the combined embeddings vector comprises a concatenated embeddings vector.

20. The system of claim 19 , wherein determining the document from the plurality of documents comprises:

inputting the concatenated embeddings vector into a causal transformer model to generate an output embeddings vector; and

determining the document as a closest document of the plurality of documents to the output embeddings vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2023
From: MALLAPRAGADA, SRIVATSA; XIE, YING; CHAWAN, VARSHA RANI; HAILAT, ZEYAD; HUGHES, SIMON; WANG, YUANBO
To: HOME DEPOT PRODUCT AUTHORITY, LLC
Reel/Frame 065224/0901 →
Continuity (3)
Provisional Application 63424946 · Nov 13, 2022
Provisional Application 63422300 · Nov 3, 2022
Related Publication 20240152561A1 · May 9, 2024
References Cited (61)
US 11720942B1 · Loris · 2023 [cited by applicant]
US 20170031904A1 · Legrand · 2017 [cited by examiner]
US 20170357896A1 · Tsatsin · 2017 [cited by examiner]
US 20180267976A1 · Bordawekar · 2018 [cited by examiner]
US 20180268024A1 · Bandyopadhyay · 2018 [cited by examiner]
US 20200279105A1 · Muffat · 2020 [cited by applicant]
US 20200410157A1 · Van De Kerkhof · 2020 [cited by applicant]
US 20210149980A1 · Pavlini · 2021 [cited by applicant]
US 20220179871A1 · Ahmed · 2022 [cited by applicant]
Gianni Amati and Cornelis Joost Van Rijsbergen. Probabilistic models of information retrieval based on measuring the divergence from randomness. ACM Transactions on Information Systems (TOIS), 20(4):357-389, 2002. [cited by applicant]
Hamed Bonab, Mohammad Aliannejadi, Ali Vardasbi, Evangelos Kanoulas, and James Allan. Cross-market product recommendation. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management. A… [cited by applicant]
Min Cao, Shiping Li, Juntao Li, Liqiang Nie, and Min Zhang. Image-text retrieval: A survey on recent research and development. arXiv preprint arXiv:2203.14713, 2022. [cited by applicant]
Wei-Cheng Chang, Felix X Yu, Yin-Wen Chang, Yiming Yang, and Sanjiv Kumar. Pre-training tasks for embedding-based large-scale retrieval. arXiv preprint arXiv:2002.03932, 2020. [cited by applicant]
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In European conference on computer vision, pp. 104-120. Spri… [cited by applicant]
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Micha… [cited by applicant]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248-255. Ieee, 2009. [cited by applicant]
Xiao Dong, Xunlin Zhan, Yangxin Wu, Yunchao Wei, Michael C Kampffmeyer, Xiaoyong Wei, Minlong Lu, Yaowei Wang, and Xiaodan Liang. M5product: Selfharmonized contrastive learning for e-commercial multimodal pretraining. I… [cited by applicant]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16×16 words: Transfo… [cited by applicant]
Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 12873-12883, 2021. [cited by applicant]
Dehong Gao, Linbo Jin, Ben Chen, Minghui Qiu, Peng Li, Yi Wei, Yi Hu, and Hao Wang. Fashionbert: Text and image matching with adaptive loss for cross-modal retrieval. In Proceedings of the 43rd International ACM SIGIR C… [cited by applicant]
Mariya Hendriksen, Maurits Bleeker, Svitlana Vakulenko, Nanne van Noord, Ernst Kuiper, and Maarten de Rijke. Extending clip for category-to-image retrieval in e-commerce. In European Conference on Information Retrieval,… [cited by applicant]
Jui-Ting Huang, Ashish Sharma, Shuying Sun, Li Xia, David Zhang, Philip Pronin, Janani Padmanabhan, Giuseppe Ottaviano, and Linjun Yang. Embedding-based retrieval in facebook search. In Proceedings of the 26th ACM SIGKD… [cited by applicant]
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. Unsupervised dense information retrieval with contrastive learning. arXiv preprint arXiv:2112.09118,… [cited by applicant]
Young Kyun Jang and Nam Ik Cho. Self-supervised product quantization for deep unsupervised image retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 12085-12094, 2021. [cited by applicant]
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Inte… [cited by applicant]
Shubhra Kanti Karmaker Santu, Parikshit Sondhi, and ChengXiang Zhai. On application of learning to rank for ecommerce search. In Proceedings of the 40th international ACM SIGIR conference on research and development in … [cited by applicant]
Makoto P. Kato, Takehiro Yamamoto, Hiroaki Ohshima, and Katsumi Tanaka. Cognitive search intents hidden behind queries: a user study on query formulations. In Proceedings of the 23rd International Conference on World Wi… [cited by applicant]
Omar Khattab and Matei Zaharia. Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR conference on research and development in… [cited by applicant]
Lakshya Kumar and Sagnik Sarkar. Neural search: Learning query and product representations in fashion e-ommerce. arXiv preprint arXiv:2107.08291, 2021. [cited by applicant]
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural info… [cited by applicant]
Jimmy Lin, Rodrigo Nogueira, and Andrew Yates. Pretrained transformers for text ranking: Bert and beyond. Synthesis Lectures on Human Language Technologies, 14(4):1-325, 2021. [cited by applicant]
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32, 2019. [cited by applicant]
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee. 12-in-1: Multi-task vision and language representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni… [cited by applicant]
Haoyu Ma, Handong Zhao, Zhe Lin, Ajinkya Kale, Zhangyang Wang, Tong Yu, Jiuxiang Gu, Sunav Choudhary, and Xiaohui Xie. Ei-clip: Entity-aware interventional contrastive learning for e-commerce cross-modal retrieval. In P… [cited by applicant]
Priyanka Nigam, Yiwei Song, Vijai Mohan, Vihan Lakshman, Weitian Ding, Ankit Shingavi, Choon Hui Teo, Hao Gu, and Bing Yin. Semantic product search. In Proceedings of the 25th ACM SIGKDD International Conference on Know… [cited by applicant]
Yiming Qiu, Chenyu Zhao, Han Zhang, Jingwei Zhuo, Tianhao Li, Xiaowei Zhang, Songlin Wang, Sulong Xu, Bo Long, and Wen-Yun Yang. Pre-training tasks for user intent detection and embedding retrieval in e-commerce search.… [cited by applicant]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language superv… [cited by applicant]
Ali Razavi, Aaron Van den Oord, and Oriol Vinyals. Generating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems, 32, 2019. [cited by applicant]
Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084, 2019. [cited by applicant]
Stephen Robertson, Hugo Zaragoza, et al. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends® in Information Retrieval, 3(4):333-389, 2009. [cited by applicant]
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019. [cited by applicant]
Keiji Shinzato, Naoki Yoshinaga, Yandi Xia, and Wei-Te Chen. Simple and effective knowledge-driven query expansion for qa-based product attribute extraction. In Proceedings of the 60th Annual Meeting of the Association … [cited by applicant]
Riku Togashi and Tetsuya Sakai. Visual intents vs. clicks, likes, and purchases in e-commerce. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 1869… [cited by applicant]
Aaron van den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017. [cited by applicant]
Xiao Wang, Craig Macdonald, Nicola Tonellotto, and Iadh Ounis. Pseudo-relevance feedback for multiple representation dense retrieval. In Proceedings of the 2021 ACM SIGIR International Conference on Theory of Informatio… [cited by applicant]
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk. Approximate nearest neighbor negative contrastive learning for dense text retrieval. arXiv preprint arXiv:200… [cited by applicant]
Peng Xu, Xiatian Zhu, and David A Clifton. Multimodal learning with transformers: A survey. arXiv preprint arXiv:2206.06488, 2022. [cited by applicant]
HongChien Yu, Chenyan Xiong, and Jamie Callan. Improving query representations for dense retrieval with pseudo relevance feedback. arXiv preprint arXiv:2108.13454, 2021. [cited by applicant]
Licheng Yu, Jun Chen, Animesh Sinha, Mengjiao Wang, Yu Chen, Tamara L Berg, and Ning Zhang. Commercemm: Large-scale commerce multimodal representation learning with omni retrieval. In Proceedings of the 28th ACM SIGKDD … [cited by applicant]
Tan Yu, Junsong Yuan, Chen Fang, and Hailin Jin. Product quantization network for fast image retrieval. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 186-201, 2018. [cited by applicant]
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. Learning discrete representations via constrained clustering for effective and efficient dense retrieval. In Proceedings of the Fifteenth ACM… [cited by applicant]
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Min Zhang, and Shaoping Ma. Learning to retrieve: How to train a dense retrieval model effectively and efficiently. arXiv preprint arXiv:2010.10469, 2020. [cited by applicant]
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Min Zhang, and Shaoping Ma. Repbert: Contextualized text embeddings for first-stage retrieval. arXiv preprint arXiv:2006.15498, 2020. [cited by applicant]
Boxuan Zhang, Chao Wei, Yan Jin, and Weiru Zhang. Acebert: Adversarial cross-modal enhanced bert for e-commerce retrieval. arXiv preprint arXiv:2112.07209, 2021. [cited by applicant]
Mengxiao Zhang, Yongning Wu, Raif Rustamov, Hongyu Zhu, Haoran Shi, Yuqi Wu, Lei Tang, Zuohua Zhang, and Chu Wang. Advancing query rewriting in e-commerce via shopping intent learning. 2022. [cited by applicant]
Shengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang, Zhou Zhao, Jianke Zhu, Jin Yu, Hongxia Yang, and Fei Wu. Devlbert: Learning deconfounded visio-linguistic representations. In Proceedings of the 28th ACM International Conf… [cited by applicant]
Ting Zhang and Jingdong Wang. Collaborative quantization for cross-modal similarity search. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2036-2045, 2016. [cited by applicant]
Mingchen Zhuge, Dehong Gao, Deng-Ping Fan, Linbo Jin, Ben Chen, Haoming Zhou, Minghui Qiu, and Ling Shao. Kaleido-bert: Vision-language pre-training on fashion domain. In Proceedings of the IEEE/CVF Conference on Comput… [cited by applicant]
Hsiang-Fu Yu, Kai Zhong, Jiong Zhang, Wei-Cheng Chang and Inderjit S Dhillon. Pecos: Prediction for enormous and correlated output spaces. Journal of Machine Learning Research, 23(98):1-32, 2022. [cited by applicant]
International Search Report and Written Opinion of international application No. PCT/US2023/078474, dated Feb. 14, 2024, 11 pp. [cited by applicant]
Wei Wei, Chao Huang, Lianghao Xia, Yong Xu, Jiashu Zhao, Dawei Yin, Contrastive Meta Learning with Behavior Multiplicity for Recommendation, arXiv:2202.08523v1 Feb. 17, 2022. [cited by applicant]