IP Library Granted Patent US 12,657,870
Granted Patent B2
US 12,657,870 · App. 17/591,861 · Granted Jun 16, 2026

Product image classification

Inventors: Estelle Afshar (Atlanta, GA); Tianlong Xu (Atlanta, GA); Le Yu (Atlanta, GA); Yuanbo Wang (Austin, TX); James Morgan White (Atlanta, GA)
Assignee: Home Depot Product Authority, LLC
G06V10/764G06Q30/0643G06V10/761G06V10/87
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,870
App. No.
17/591,861
Granted
Jun 16, 2026
Kind
B2
Abstract

A method includes retrieving, by a processor in an image classification system, a plurality of product images from a memory. A first image classifier is applied to a first of the plurality of product images. The first of the plurality of product images is associated with a first category. A second image classifier is applied to a second of the plurality of product images. The second of the plurality of product images is associated with a second category. A first result of the first image classifier for the first of the plurality of product images is stored in the memory. A second result of the second image classifier for the second of the plurality of product images is stored in the memory.

Claims (53)

1 . A method, comprising:

retrieving, by a processor in an image classification system, a plurality of product images from a memory;

determining a category of a plurality of categories for a product in each of the plurality of product images based on a product information associated with respective product images;

applying one or more image classifiers of a plurality of image classifiers to each of the plurality of product images based on the respective determined category to predict labels for each of the plurality of product images, the applying comprising:

providing a first label prediction based on applying a first image classifier to features of a first product image of the plurality of product images, wherein the first image classifier is trained to identify features specific to the determined category of the first product image; and

providing a second label prediction based on applying a second image classifier to features of a second product image of the plurality of product images, wherein the second product image is from a determined category that does not include specifically trained image classifiers of the plurality of image classifiers, and the second image classifier is trained to identify features in product images that are from categories that do not include the specifically trained image classifiers;

storing the labels including the first label prediction of the first image classifier for the first product image and the second label prediction of the second image classifier for the second product image; and

selecting, in response to a captured image from a user computing device, one of the plurality of product images as a recommendation based on the predicted labels for each of the plurality of product images and outputting the selected one of the plurality of product images for display on the user computing device,

wherein the plurality of image classifiers are configured to predict at least one of a content label or a view label for the product in each of the plurality of product images based on the determined category,

wherein the first image classifier is at least one of a plurality of category-specific image classifiers and the second image classifier is one of a plurality of a generic image classifiers for categories not associated with a corresponding one of the category-specific image classifiers.

2 . The method of claim 1 , wherein the first label prediction by the first image classifier is based on the first product image being in a first category of the determined categories and the second label prediction by the second image classifier is based on the second product image being in a second category not associated with one of the determined categories.

3 . The method of claim 2 , further comprising:

providing a third label prediction based on applying the second image classifier to features of a third product image of the plurality of product images, wherein the third product image of the plurality of product images is associated with a third category different from the second category; and

storing the third label prediction of the second image classifier for the third product image of the plurality of product images.

4 . The method of claim 2 , wherein the first image classifier is a category-specific image classifier associated with the first category.

5 . The method of claim 1 , wherein applying the first image classifier to the first product image of the plurality of product images further comprising:

providing a fourth label prediction based on applying a fourth image classifier to the first product image of the plurality of product images, wherein the fourth image classifier is trained to identify specific features based on the determined category of the first product image;

wherein the first label prediction is further based on the fourth label prediction of the fourth image classifier.

6 . A system, comprising:

a processor and a memory, wherein the processor is configured to:

retrieve a plurality of product images from the memory;

determine a category of a product in each of the plurality of product images based on a product information associated with respective product images;

apply a respective image classifier of a plurality of image classifiers to each of the plurality of product images based on the respective determined category to predict labels for each of the plurality of product images, the applying comprising:

provide a first label from applying a first image classifier to one or more features of a first product image of the plurality of product images, wherein the first product image is from a first product category; and

provide a second label from applying a second image classifier to one or more features of a second product image of the plurality of product images, wherein the second product image is from a second product category;

store the labels from the plurality of image classifiers including the first label and the second label in the memory; and

select, in response to a captured image from a user computing device, one of the plurality of product images as a recommendation based on the labels including the first label and the second label and output the selected one of the plurality of product images for display on the user computing device,

wherein the plurality of image classifiers are configured to predict at least one of a content label or a view label for the product in each of the plurality of product images based on the determined category,

wherein the plurality of image classifiers comprises category-specific image classifiers specific to respective ones of the determined categories and generic image classifiers for categories not associated with the category-specific image classifiers,

wherein the first image classifier comprises at least one of the category-specific image classifiers and the second image classifier comprises at least one of the generic image classifiers.

7 . The system of claim 6 , wherein the first product image is in a first category and the second product image is in a second category, and wherein the first category and the second category are different.

8 . The system of claim 6 , wherein the processor is further configured to:

apply the second image classifier to a third product image of the plurality of product images, wherein the third product image of the plurality of product images is associated with a third category different from the second category; and

store a third label of the second image classifier for the third product image of the plurality of product images.

9 . The system of claim 6 , wherein the first image classifier is a category-specific image classifier associated with the first category.

10 . The system of claim 6 , wherein applying the first image classifier to the first of the plurality of product images further comprising:

apply a fourth image classifier to the first product image of the plurality of product images, wherein the fourth image classifier is trained to identify specific features based on the determined category of the first product image;

wherein the first label is further based on a fourth label prediction of the fourth image classifier.

11 . A system, comprising: a server; and

an image classification system, the image classification system comprising: a processor and a memory, wherein the processor is configured to:

retrieve a plurality of product images from the memory;

determine a category of a product in each of the plurality of product images;

predict labels for each of the plurality of product images based on applying a respective image classifier of a plurality of image classifiers to each of the plurality of product images based on the determined category, the applying comprising:

predict a first label for a first product image of the plurality of product images based on applying a first image classifier to features of the first product image, wherein the first image classifier is configured to identify features specific to the determined category of the first product image; and

predict a second label for a second product image of the plurality of product images based on applying a second image classifier to features of the second product image, wherein the second product image is from a determined category that does not include specifically trained image classifiers, and wherein the second image classifier is configured to identify features in product images that are from categories that do not include the specifically trained image classifiers;

store the labels including the first label and the second label in the memory; and

select, in response to receiving a captured image from a user computing device, one of the plurality of product images based on the predicted labels for each of the plurality of product images and output the selected one of the plurality of product images for display on the user computing device,

wherein the plurality of image classifiers are configured to predict at least one of a content label or a view label for the product in each of the plurality of product images based on the determined category,

wherein the first image classifier comprises at least one of a plurality of category-specific image classifiers for the determined categories and the second image classifier comprises at least one of a plurality of generic image classifiers for categories not associated with a corresponding one of the category-specific image classifiers.

12 . The system of claim 11 , wherein the server is configured to receive the captured image from the user computing device over a network.

13 . The system of claim 12 , wherein the server is configured to compute a visual similarity score between the captured image and the plurality of product images based on the predicted labels of the plurality of image classifiers including the first label of the first image classifier and the second label of the second image classifier.

14 . The system of claim 13 , wherein the server is configured to select one of the plurality of product images having a highest visual similarity score with the captured image and output the selected one of the plurality of product images for display on the user computing device.

15 . The system of claim 11 , wherein the first image classifier is a category-specific image classifier.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2022
From: AFSHAR, ESTELLE; XU, TIANLONG; YU, LE; WANG, YUANBO; WHITE, JAMES MORGAN
To: HOME DEPOT PRODUCT AUTHORITY, LLC
Reel/Frame 059294/0907 →
Continuity (2)
Provisional Application 63146339 · Feb 5, 2021
Related Publication 20220254144A1 · Aug 11, 2022
References Cited (75)
US 7349917B2 · Forman · 2008 [cited by examiner]
US 9443320B1 · Gaidon · 2016 [cited by examiner]
US 9792530B1 · Wu · 2017 [cited by examiner]
US 10706326B2 · Tsunoda · 2020 [cited by examiner]
US 10776626B1 · Lin · 2020 [cited by examiner]
US 10832400B1 · Tang · 2020 [cited by examiner]
US 11636639B2 · Adamson, III · 2023 [cited by examiner]
US 11776118B2 · Green · 2023 [cited by examiner]
US 11783602B2 · Futatsugi · 2023 [cited by examiner]
US 20050249401A1 · Bahlmann · 2005 [cited by examiner]
US 20110052063A1 · McAuley · 2011 [cited by examiner]
US 20120016892A1 · Wang et al. · 2012 [cited by applicant]
US 20120314941A1 · Kannan · 2012 [cited by examiner]
US 20140136549A1 · Surya · 2014 [cited by examiner]
US 20160260014A1 · Hagawa · 2016 [cited by examiner]
US 20170076225A1 · Zhang et al. · 2017 [cited by applicant]
US 20170351914A1 · Zavalishin · 2017 [cited by examiner]
US 20170372169A1 · Li · 2017 [cited by examiner]
US 20180012110A1 · Souche · 2018 [cited by examiner]
US 20180239990A1 · Philipose · 2018 [cited by examiner]
US 20180348785A1 · Zheng · 2018 [cited by examiner]
US 20190012568A1 · Kumar · 2019 [cited by examiner]
US 20190130214A1 · N et al. · 2019 [cited by applicant]
US 20200005087A1 · Sewak · 2020 [cited by examiner]
US 20200042796A1 · Kim · 2020 [cited by examiner]
US 20200151499A1 · Kumar · 2020 [cited by examiner]
US 20200279008A1 · Jain · 2020 [cited by examiner]
US 20200286275A1 · Seo · 2020 [cited by examiner]
US 20200294234A1 · Rance · 2020 [cited by examiner]
US 20210011477A1 · Byun · 2021 [cited by examiner]
US 20210072763A1 · Chen · 2021 [cited by examiner]
US 20210326641A1 · Tsai · 2021 [cited by examiner]
US 20220101112A1 · Brown · 2022 [cited by examiner]
US 20220108163A1 · Sokhandan Asl · 2022 [cited by examiner]
US 20230222783A1 · Bowman · 2023 [cited by examiner]
US 20230237802A1 · Falendysz · 2023 [cited by examiner]
US 20230260254A1 · Huang · 2023 [cited by examiner]
US 20230316702A1 · Kar · 2023 [cited by examiner]
Sean Bell, Yiqun Liu, Sami Alsheikh, Yina Tang, Edward Pizzi, M. Henning, Karun Singh, Omkar Parkhi, and Fedor Borisyuk. 2020. GrokNet: Unified Computer Vision Model Trunk and Embeddings for Commerce. 2608-2616. https:/… [cited by applicant]
Alessandro Bergamo and Lorenzo Torresani. 2010. Exploiting weakly-labeled web images to improve object classification: a domain adaptation approach. Advances in neural information processing systems 23 (2010), 181-189. [cited by applicant]
Tianqi Chen and Carlos Guestrin. 2016. XGBoost. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Aug. 2016). https://doi.org/10.1145/2939672.2939785. [cited by applicant]
François Chollet. 2017. Xception: Deep Learning with Depthwise Separable Convolutions. arXiv:1610.02357 [cs.CV]. [cited by applicant]
N. Dalal and B. Triggs. 2005. Histograms of oriented gradients for human detection. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05). [cited by applicant]
Wei Di, Neel Sundaresan, Robinson Piramuthu, and Anurag Bhardwaj. 2014. Is a Picture Really Worth a Thousand Words?—On the Role of Images in e-Commerce. [cited by applicant]
Ehsan Elhamifar, Guillermo Sapiro, Allen Yang, and S Shankar Sasrty. 2013. A convex optimization framework for active learning. In Proceedings of the IEEE International Conference on Computer Vision. 209-216. [cited by applicant]
Yarin Gal, Riashat Islam, and Zoubin Ghahramani. 2017. Deep bayesian active learning with image data. arXiv preprint arXiv:1703.02910 (2017). [cited by applicant]
Rohith Gandhi. 2018. Siamese Network Triplet Loss. https://towardsdatascience.com/siamese-network-triplet-loss-b4ca82c1aec8. [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Deep Residual Learning for Image Recognition. arXiv:1512.03385 [cs.CV]. [cited by applicant]
Houdong Hu, Yan Wang, Linjun Yang, Pavel Komlev, Li Huang, Xi Chen, Jiapei Huang, YeWu, Meenaz Merchant, and Arun Sacheti. 2018. Web-Scale Responsive Visual Search at Bing. arXiv:1802.04914 [cs.CV]. [cited by applicant]
Jie Hu, Li Shen, Samuel Albanie, Gang Sun, and Enhua Wu. 2019. Squeeze-and-Excitation Networks. arXiv:1709.01507 [cs.CV]. [cited by applicant]
Armand Joulin, Laurens van der Maaten, Allan Jabri, and Nicolas Vasilache. 2015. Learning Visual Features from Large Weakly Supervised Data. arXiv:1511.02251 [cs.CV]. [cited by applicant]
Gregory R. Koch. 2015. Siamese Neural Networks for One-Shot Image Recognition. [cited by applicant]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In Advances in Neural Information Processing Systems, F. Pereira, C. J. C. Burges, L. Bottou… [cited by applicant]
Eileen Li, Eric Kim, Andrew Zhai, Josh Beal, and Kunlong Gu. 2020. Bootstrapping Complete The Look at Pinterest. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery Data Mining (Virtual… [cited by applicant]
Fengzi Li, Shashi Kant, Shunichi Araki, Sumer Bangera, and Swapna Samir Shukla. 2020. Neural Networks for Fashion Image Classification and Visual Search. arXiv:2005.08170 [cs.CV]. [cited by applicant]
D.G. Lowe. 2004. Lowe, D.G. Distinctive Image Features from Scale-Invariant Keypoints. International Journal of Computer Vision 60, 91-110 (2004). https://doi.org/10.1023/B:VISI.0000029664.99615.94. [cited by applicant]
Alexander Schindler, Thomas Lidy, Stephan Karner, and Matthias Hecker. 2018. Fashion and Apparel Classification using Convolutional Neural Networks. arXiv:1811.04374 [cs.CV]. [cited by applicant]
Matthew Schultz and Thorsten Joachims. 2004. Learning a Distance Metric from Relative Comparisons. In Advances in Neural Information Processing Systems, S. Thrun, L. Saul, and B. Scholkopf (Eds.), vol. 16. MIT Press, 41… [cited by applicant]
Ozan Sener and Silvio Savarese. 2017. Active learning for convolutional neural networks: A core-set approach. arXiv preprint arXiv:1708.00489 (2017). [cited by applicant]
Devashish Shankar, Sujay Narumanchi, HA Ananya, Pramod Kompalli, and Krishnendu Chaudhury. 2017. Deep learning based large scale visual recommendation and search for e-commerce. arXiv preprint arXiv:1703.02344 (2017). [cited by applicant]
Raymond Shiau, Hao-Yu Wu, Eric Kim, Yue Du, Anqi Guo, Zhiyuan Zhang, Eileen Li, Kunlong Gu, Charles Rosenberg, and Andrew Zhai. 2020. Shop The Look: Building a Large Scale Visual Shopping System at Pinterest. 3203-3212.… [cited by applicant]
Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv:1409.1556 [cs.CV]. [cited by applicant]
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alex Alemi. 2016. Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning. arXiv:1602.07261 [cs.CV]. [cited by applicant]
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2014. Going Deeper with Convolutions. arXiv:1409.4842 [cs.CV]. [cited by applicant]
Yina Tang, Fedor Borisyuk, Siddarth Malreddy, Yixuan Li, Yiqun Liu, and Sergey Kirshner. 2019. MSURU: Large Scale E-commerce Image Classification with Weakly Supervised Search Data. In Proceedings of the 25th ACM SIGKDD… [cited by applicant]
Jiang Wang, Yang Song, Thomas Leung, Chuck Rosenberg, Jinbin Wang, James Philbin, Bo Chen, and Ying Wu. 2014. Learning Fine-grained Image Similarity with Deep Ranking. arXiv:1404.4661 [cs.CV]. [cited by applicant]
Kai Wei, Yuzong Liu, Katrin Kirchhoff, Chris Bartels, and Jeff Bilmes. 2014. Submodular subset selection for large-scale speech training data. In 2014 IEEE International Conference on Acoustics, Speech and Signal Proces… [cited by applicant]
Wikipedia.org. 2020. XGBoost. https://en.wikipedia.org/wiki/XGBoost. [cited by applicant]
Stephen Zakrewsky, Kamelia Aryafar, and Ali Shokoufandeh. 2016. Item Popularity Prediction in E-commerce Using Image Quality Feature Vectors.8 arXiv:1605.03663 [cs.CV]. [cited by applicant]
J. Zhang, S. Lazebnik, and C. Schmid. 2007. Local features and kernels for classification of texture and object categories: a comprehensive study. International Journal of Computer Vision 73 (2007), 2007. [cited by applicant]
Yanhao Zhang, Pan Pan, Yun Zheng, Kang Zhao, Yingya Zhang, Xiaofeng Ren, and Rong Jin. 2018. Visual Search at Alibaba. [cited by applicant]
Fedor Zhdanov. 2019. Diverse mini-batch Active Learning. arXiv preprint arXiv:1901.05954 (2019). [cited by applicant]
Yun Zhu, Sayyed M. Zahiri, Jiaqi Wang, Han-Yu Chen, and Faizan Javed. 2020. Active Learning for Product Type Ontology Enhancement in E-commerce. arXiv:2009.09143 [cs.LG]. [cited by applicant]
Zhen Zuo, L. Wang, Michinari Momma, W. Wang, Yikai Ni, Jianfeng Lin, and Y. Sun. 2020. A flexible large-scale similar product identification system in e-commerce. [cited by applicant]
ISA/US, ISR/WO issued in PCT/US22/15173, dated Apr. 25, 2022, 9 pgs. [cited by applicant]