IP Library › Granted Patent US 12,287,824
Granted Patent B2
US 12,287,824 · App. 18/642,447 · Granted Apr 29, 2025

System and method for determining item labels based on item images

Inventors: Binwei Yang (Milpitas, CA); Cun Mu (Jersey City, NJ)
Assignee: Walmart Apollo, LLC
G06F16/532G06F16/55G06F16/56G06F16/953G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,287,824
App. No.
18/642,447
Granted
Apr 29, 2025
Kind
B2
Abstract

A method including automatically determining, by a machine learning model trained based at least in part on sample items stored in a sample database, a query embedding vector for a query image of a query item. The method further can include determining, based on a respective embedding distance between the query image of the query item and a respective image of each of the sample items, neighboring items from among the sample items. The respective embedding distance can be calculated based on the query embedding vector for the query image and a respective embedding vector for the respective image of each of the sample items. Each of the sample items can include the respective image and at least one respective item label. The method also can include determining a respective normalized weight for each of the neighboring items based on the respective embedding distance between the query image and the respective image of the each of the neighboring items. The method additionally can include determining a query item label of the query item based on a weighted majority vote by the neighboring items via the respective normalized weight for the each of the neighboring items. The method further can include upon determining that the query item label of the query item is different from a first item label of the query item. storing the query item with the query item label in a product database. The method also can include selectively updating the sample items stored in the sample database from items in the product database. In addition, the method can include re-training the machine learning model based at least in part on the sample items in the sample database, as updated. Other embodiments are disclosed.

Claims (133)

1. A system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing computing instructions that, when executed on the one or more processors, cause the one or more processors to perform:

automatically determining, by a machine learning model trained based at least in part on sample items stored in a sample database, a query embedding vector for a query image of a query item;

determining, based on a respective embedding distance between the query image of the query item and a respective image of each of the sample items, neighboring items from among the sample items, wherein:

the respective embedding distance is calculated based on the query embedding vector for the query image and a respective embedding vector for the respective image of each of the sample items; and

each of the sample items comprises the respective image and at least one respective item label;

determining a respective normalized weight for each of the neighboring items based on the respective embedding distance between the query image and the respective image of the each of the neighboring items;

determining a query item label of the query item based on a weighted majority vote by the neighboring items via the respective normalized weight for the each of the neighboring items;

upon determining that the query item label of the query item is different from a first item label of the query item, storing the query item with the query item label in a product database;

selectively updating the sample items stored in the sample database from items in the product database; and

re-training the machine learning model based at least in part on the sample items in the sample database, as updated.

2. The system in claim 1 , wherein:

the computing instructions, when executed on the one or more processors, further cause the one or more processors to perform:

before retrieving the sample items:

selectively determining a new sample image associated with at least one new image label for a new sample item of the sample items, wherein:

the respective image of the new sample item comprises the new sample image; and

the at least one respective item label of the new sample item comprises the at least one new image label;

determining the respective embedding vector for the new sample item; and

storing the new sample item in the sample database.

3. The system in claim 1 , wherein:

determining the neighboring items further comprises using a k-nearest neighbors (K-NN) search engine to search for the neighboring items for the query item.

4. The system in claim 1 , wherein:

the respective embedding distance between the query image and the respective image of the each of the neighboring items is a Euclidean distance between the query embedding vector for the query image and the respective embedding vector of the each of the neighboring items.

5. The system in claim 1 , wherein:

the machine learning model further comprises a neural network model.

6. The system in claim 5 , wherein:

the computing instructions, when executed on the one or more processors, further cause the one or more processors to perform:

before determining automatically the query embedding vector, training the neural network model based at least in part on the respective image and the at least one respective item label of the each of the sample items.

7. The system in claim 1 , wherein:

determining the respective normalized weight for the each of the neighboring items further comprises calculating the respective normalized weight by:

exp

⁡

(

d

i

)

Σ

i

⁢

exp

⁡

(

d

i

)

,

∀

i

∈

[

K

]

wherein:

d i : the respective embedding distance between the query image and its i-th neighbor; and

K: the neighboring items.

8. The system in claim 1 , wherein:

determining the query item label of the query item based on the weighted majority vote further comprises:

tallying at least one respective weighted group vote from each group of one or more groups of the neighboring items by summing up the respective normalized weight for each of one or more respective items of the each group of the one or more groups, wherein:

the neighboring items comprise the one or more respective items; and

the at least one respective item label of the each of the one or more respective items is identical within the each group; and

determining the query item label of the query item to be the at least one respective item label of each of the one or more respective items of a wining group of the one or more groups based on the at least one respective weighted group vote from the wining group.

9. The system in claim 1 , wherein:

the computing instructions, when executed on the one or more processors, further cause the one or more processors to perform:

upon determining that a designation error exists and before storing the query item in the product database, transmitting an alert to a user to be displayed on a user interface executed on a user device for the user based at least in part on the designation error.

10. The system in claim 9 , wherein:

the computing instructions, when executed on the one or more processors, further cause the one or more processors to perform:

before transmitting the alert to the user, determining a confidence level status based on whether a confidence level for the query item label is at least as great as a predetermined threshold; and

transmitting the alert to the user further comprises transmitting the alert to the user further based on the confidence level status.

11. A method being implemented via execution of computing instructions configured to run at-one or more processors and stored at one or more non-transitory computer-readable media, the method comprising:

automatically determining, by a machine learning model trained based at least in part on sample items stored in a sample database, a query embedding vector for a query image of a query item;

determining, based on a respective embedding distance between the query image of the query item and a respective image of each of the sample items, neighboring items from among the sample items, wherein:

the respective embedding distance is calculated based on the query embedding vector for the query image and a respective embedding vector for the respective image of each of the sample items; and

each of the sample items comprises the respective image and at least one respective item label;

determining a respective normalized weight for each of the neighboring items based on the respective embedding distance between the query image and the respective image of the each of the neighboring items;

determining a query item label of the query item based on a weighted majority vote by the neighboring items via the respective normalized weight for the each of the neighboring items;

upon determining that the query item label of the query item is different from a first item label of the query item, storing the query item with the query item label in a product database;

selectively updating the sample items stored in the sample database from items in the product database; and

re-training the machine learning model based at least in part on the sample items in the sample database, as updated.

12. The method in claim 11 , further comprising:

before retrieving the sample items:

selectively determining a new sample image associated with at least one new image label for a new sample item of the sample items, wherein:

the respective image of the new sample item comprises the new sample image; and

the at least one respective item label of the new sample item comprises the at least one new image label;

determining the respective embedding vector for the new sample item; and

storing the new sample item in the sample database.

13. The method in claim 11 , wherein:

determining the neighboring items further comprises using a k-nearest neighbors (K-NN) search engine to search for the neighboring items for the query item.

14. The method in claim 11 , wherein:

the respective embedding distance between the query image and the respective image of the each of the neighboring items is a Euclidean distance between the query embedding vector for the query image and the respective embedding vector of the each of the neighboring items.

15. The method in claim 11 , wherein:

the machine learning model further comprises a neural network model.

16. The method in claim 15 , further comprising, before determining automatically the query embedding vector:

training the neural network model based at least in part on the respective image and the at least one respective item label of the each of the sample items.

17. The method in claim 11 , wherein:

determining the respective normalized weight for the each of the neighboring items further comprises calculating the respective normalized weight by:

exp

⁡

(

d

i

)

Σ

i

⁢

exp

⁡

(

d

i

)

,

∀

i

∈

[

K

]

wherein:

d i : the respective embedding distance between the query image and its i-th neighbor; and

K: the neighboring items.

18. The method in claim 11 , wherein:

determining the query item label of the query item based on the weighted majority vote further comprises:

tallying at least one respective weighted group vote from each group of one or more groups of the neighboring items by summing up the respective normalized weight for each of one or more respective items of the each group of the one or more groups, wherein:

the neighboring items comprise the one or more respective items; and

the at least one respective item label of the each of the one or more respective items is identical within the each group; and

determining the query item label of the query item to be the at least one respective item label of each of the one or more respective items of a wining group of the one or more groups based on the at least one respective weighted group vote from the wining group.

19. The method in claim 11 , further comprising:

upon determining that a designation error exists and before storing the query item in the product database, transmitting an alert to a user to be displayed on a user interface executed on a user device for the user based at least in part on the designation error.

20. The method in claim 19 , further comprising:

before transmitting the alert to the user, determining a confidence level status based on whether a confidence level for the query item label is at least as great as a predetermined threshold,

wherein:

transmitting the alert to the user further comprises transmitting the alert to the user further based on the confidence level status.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2024
From: YANG, BINWEI; MU, CUN
To: WALMART APOLLO, LLC
Reel/Frame 067405/0145 →
Continuity (2)
Continuation 17174662 · Feb 12, 2021
Related Publication 20240273133A1 · Aug 15, 2024
References Cited (19)
US 9104968B2 · Wang · 2015 [cited by applicant]
US 9311644B2 · Liu et al. · 2016 [cited by applicant]
US 10726060B1 · Dutta et al. · 2020 [cited by applicant]
US 10783167B1 · Dutta et al. · 2020 [cited by applicant]
US 20100177956A1 · Cooper et al. · 2010 [cited by applicant]
US 20210019343A1 · Singh et al. · 2021 [cited by applicant]
US 20210034657A1 · Kale et al. · 2021 [cited by applicant]
CN 105701476 · 2016 [cited by applicant]
CN 109145901 · 2019 [cited by applicant]
He, et al., “Deep Residual Learning for Image Recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2016), 12 pages, retrieved from https://arxiv.org/pdf/1512.03385.pdf, 2016. 2016. [cited by applicant]
Huang, et al., “Densely Connected Convolutional Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017), 9 pages, retrieved from https://arxiv.org/pdf/1608.06993.pdf, 2017. 2017. [cited by applicant]
Tan, et al., “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” International Conference on Machine Learning (2019), 11 pages, retrieved from https://arxiv.org/pdf/1905.11946.pdf, 2019. 2019. [cited by applicant]
Johnson, et al., “Billion-Scale Similarity Search with GPUs,” IEEE Transactions on Big Data (2019), 12 pages, retrieved from https://arxiv.org/pdf/1702.08734.pdf, 2019. 2019. [cited by applicant]
Mu, et al., “Fast and Exact Nearest Neighbor Search in Hamming Space on Full-Text Search Engines,” International Conference on Similarity Search and Applications (2019), 14 pages, retrieved from https://arxiv.org/pdf/19… [cited by applicant]
Yang, et al., “Visual Search at eBay,” Proceedings of the SIGKDD International Conference on Knowledge Discovery and Data Mining (2017), 10 pages, retrieved from https://arxiv.org/abs/1706.03154, 2017. 2017. [cited by applicant]
Mu, et al., “Towards Practical Visual Search Engine Within Elasticsearch,” SIGIR eCom (2018), 8 pages, retrieved from http://ceur-ws.org/Vol-2319/paper7.pdf, 2018. 2018. [cited by applicant]
Zhang, et al., “Visual Search at Alibaba,” Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (2018), 9 pages, retrieved from https://arxiv.org/pdf/2102.04674.pdf, 2018. 201… [cited by applicant]
Li, et al., “The Design and Implementation of a Real Time Visual Search System on JD E-Commerce Platform,” Proceedings of the 19th International Middleware Conference Industry (2018), 7 pages, retrieved from https://arx… [cited by applicant]
Deng, et al., “Imagenet: A Large-Scale Hierarchical Image Database,” 2009 IEEE Conference on Computer Vision and Pattern Recognition (2009), 8 pages, retrieved from https://www-cs.stanford.edu/groups/vision/documents/Im… [cited by applicant]