IP Library Granted Patent US 12,412,149
Granted Patent B2
US 12,412,149 · App. 18/161,788 · Granted Sep 9, 2025

Systems and methods for analyzing and labeling images in a retail facility

Inventors: Raghava Balusu (Achanta, IN); Siddhartha Chakraborty (Kolkata, IN); Ashlin Ghosh (Ernakulam, IN); Avinash M. Jade (Bangalore, IN); Lingfeng Zhang (Dallas, TX); Amit Jhunjhunwala (Bangalore, IN)
Assignee: Walmart Apollo, LLC
G06Q10/087G06V10/762
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,149
App. No.
18/161,788
Granted
Sep 9, 2025
Kind
B2
Abstract

In some embodiments, apparatuses and methods are provided herein useful to processing captured images. In some embodiments, there is provided a system for processing captured images of objects including a memory and a control circuit executing a trained machine learning model. The memory may be configured to store a plurality of images comprising first images and second images. The control circuit may be configured to: allocate each of the first images into one of a plurality of datasets; cluster each image in the dataset into one of a plurality of groups; select a sample from at least one of the plurality of groups; cluster each of the second images into one of dominant product identifier group and a non-dominant product identifier group; select a sample from the dominant product identifier group and a sample from the non-dominant product identifier group; and output the selected sample.

Claims (46)

1. A system for processing captured images of objects at a product storage facility, the system comprising:

a memory configured to store a plurality of images, the plurality of images comprising first images and second images, wherein each of the first images contain items not detected by a trained machine learning model as being associated with a recognized product identifier, and wherein each of the second images contain items detected by the trained machine learning model as being associated with multiple recognized product identifiers; and

a control circuit executing the trained machine learning model configured to:

allocate each of the first images into one of a plurality of datasets based on areas in the product storage facility the first images were captured;

for each dataset of the plurality of datasets, cluster each image in the dataset into one of a plurality of groups based on a degree of resemblance of items depicted in the image relative to those items depicted in other images in the dataset;

select a sample from at least one of the plurality of groups;

cluster each of the second images into one of a dominant product identifier group and a non-dominant product identifier group;

select a sample from the dominant product identifier group;

select a sample from the non-dominant product identifier group; and

output at least one of the selected sample from the at least one of the plurality of groups, the selected sample from the dominant product identifier group, and the selected sample from the non-dominant product identifier group to be used to retrain the trained machine learning model.

2. The system of claim 1 , wherein the plurality of groups comprise a homogenous group, a heterogeneous group, a low similarity group, and an individual group.

3. The system of claim 2 , wherein the homogenous group includes images where the degree of resemblance of items depicted in each image is greater than 70%.

4. The system of claim 2 , wherein the heterogeneous group includes images where the degree of resemblance of items depicted in each image is at least 30% but no more than 70%.

5. The system of claim 2 , wherein the low similarity group includes images where the degree of resemblance of items depicted in each image is less than 30%.

6. The system of claim 2 , wherein the individual group includes images that do not have resemblance with any other image in the dataset.

7. The system of claim 2 , wherein the control circuit executing the trained machine learning model is further configured to:

process each image included in the low similarity group by determining the degree of resemblance of items depicted in the image, and clustering the image into one of the homogenous group, the heterogeneous group, the low similarity group, or the individual group in accordance with the determined degree of resemblance of the items; and

repeat processing of subsequent images included in the low similarity group until each of the subsequent images is clustered into one of the homogenous group, the heterogeneous group, or the individual group in accordance with the determined degree of resemblance of the items.

8. The system of claim 1 , wherein the control circuit executing the trained machine learning model is further configured to determine the degree of resemblance of items based on at least one of a textual similarity or a visual similarity of the items.

9. The system of claim 1 , wherein the selected sample from the at least one of the plurality of groups comprises a sample of one or more images based on a rule from each of the plurality of groups.

10. The system of claim 9 , wherein the control circuit executing the trained machine learning model is further configured to group the sample of the one or more images from each of the plurality of groups into the selected sample, wherein the selected sample is output in response to a determination that an image count of the selected sample is less than a threshold.

11. The system of claim 10 , wherein the control circuit executing the trained machine learning model is further configured to:

in response to a determination that the image count of the selected sample is at least the threshold, allocate each image in the selected sample into one of the plurality of datasets based on the areas in the product storage facility images in the selected sample were captured;

for each dataset of the plurality of datasets, cluster each image in the dataset into one of the plurality of groups based on the degree of resemblance of items depicted in the image relative to those items depicted in other images in the dataset;

select a second sample from at least one of the plurality of groups; and

determine whether a second image count of the second sample is less than the threshold, wherein the allocation of each image in the selected sample into one of the plurality of datasets, the clustering of each image in the dataset into one of the plurality of groups, and the selection of a subsequent sample from at least one of the plurality of groups are repeated until the subsequent sample is less than the threshold.

12. The system of claim 9 , wherein the plurality of groups comprise a homogenous group, a heterogeneous group, a low similarity group, and an individual group, and wherein the rule comprises:

selecting, by the control circuit executing the trained machine learning model and for images included in the homogenous group, at least two or three images based on a highest number of at least one of depicted bounding boxes, boundary box aspect ratio, or Optical Character Recognition (OCR) count;

selecting, by the control circuit executing the trained machine learning model and for images included in the heterogeneous group, a predetermined percentage of the images included in the heterogeneous group; and

selecting, by the control circuit executing the trained machine learning model and for images included in the individual group, all images included in the individual group.

13. The system of claim 1 , wherein the selected sample from the dominant product identifier group comprises an image having a threshold range of depicted items being associated with a dominant product identifier, wherein the dominant product identifier is a product identifier that is most identified with the depicted items relative to other identified product identifiers of the depicted items in the image.

14. The system of claim 1 , wherein the selected sample from the non-dominant product identifier group comprises all images associated with the non-dominant product identifier group.

15. The system of claim 1 , wherein the areas comprise one of an aisle, a bin, a pallet, or a rack storing one or more office supply products, grocery products, electronic products, and household supply products.

16. A method for processing captured images of objects at a product storage facility, the method comprising:

storing, by a memory, a plurality of images, the plurality of images comprising first images and second images, wherein each of the first images contain items not detected by a trained machine learning model as being associated with a recognized product identifier, and wherein each of the second images contain items detected by the trained machine learning model as being associated with multiple recognized product identifiers;

allocating, by a control circuit executing the trained machine learning model, each of the first images into one of a plurality of datasets based on areas in the product storage facility the first images were captured;

for each dataset of the plurality of datasets, clustering, by the control circuit executing the trained machine learning model, each image in the dataset into one of a plurality of groups based on a degree of resemblance of items depicted in the image relative to those items depicted in other images in the dataset;

selecting, by the control circuit executing the trained machine learning model, a sample from at least one of the plurality of groups;

clustering, by the control circuit executing the trained machine learning model, each of the second images into one of a dominant product identifier group and a non-dominant product identifier group;

selecting, by the control circuit executing the trained machine learning model, a sample from the dominant product identifier group;

selecting, by the control circuit executing the trained machine learning model, a sample from the non-dominant product identifier group; and

outputting, by the control circuit executing the trained machine learning model, at least one of the selected sample from the at least one of the plurality of groups, the selected sample from the dominant product identifier group, and the selected sample from the non-dominant product identifier group to be used to retrain the trained machine learning model.

17. The method of claim 16 , wherein the plurality of groups comprise a homogenous group, a heterogeneous group, a low similarity group, and an individual group, wherein the homogenous group includes images where the degree of resemblance of items depicted in each image is greater than 70%, wherein the low similarity group includes images where the degree of resemblance of items depicted in each image is less than 30%, wherein the heterogeneous group includes images where the degree of resemblance of items depicted in each image is at least 30% but no more than 70%, and wherein the individual group includes images that do not have resemblance with any other image in the dataset.

18. The method of claim 16 , further comprising determining, by the control circuit executing the trained machine learning model, the degree of resemblance of items based on at least one of a textual similarity or a visual similarity of the items.

19. The method of claim 16 , wherein the selected sample from the at least one of the plurality of groups comprises a sample of one or more images based on a rule from each of the plurality of groups.

20. The method of claim 19 , further comprising grouping, by the control circuit executing the trained machine learning model, the sample of the one or more images from each of the plurality of groups into the selected sample, wherein the selected sample is output in response to a determination that an image count of the selected sample is less than a threshold.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2023
From: WM GLOBAL TECHNOLOGY SERVICES INDIA PRIVATE LIMITED
To: WALMART APOLLO, LLC
Reel/Frame 064567/0745 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2023
From: BALUSU, RAGHAVA; CHAKRABORTY, SIDDHARTHA; GHOSH, ASHLIN; JADE, AVINASH M.; JHUNJHUNWALA, AMIT
To: WM GLOBAL TECHNOLOGY SERVICES INDIA PRIVATE LIMITED
Reel/Frame 062548/0938 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2023
From: ZHANG, LINGFENG
To: WALMART APOLLO, LLC
Reel/Frame 062549/0039 →
Continuity (1)
Related Publication 20240257047A1 · Aug 1, 2024
References Cited (141)
US 5074594A · Laganowski · 1991 [cited by applicant]
US 6570492B1 · Peratoner · 2003 [cited by applicant]
US 8700494B2 · Carlson · 2014 [cited by applicant]
US 8923650B2 · Wexler · 2014 [cited by applicant]
US 8965104B1 · Hickman · 2015 [cited by applicant]
US 9275308B2 · Szegedy · 2016 [cited by applicant]
US 9477955B2 · Goncalves · 2016 [cited by applicant]
US 9526127B1 · Taubman · 2016 [cited by applicant]
US 9576310B2 · Cancro · 2017 [cited by applicant]
US 9659204B2 · Wu · 2017 [cited by applicant]
US 9811754B2 · Schwartz · 2017 [cited by applicant]
US 10002344B2 · Wu · 2018 [cited by applicant]
US 10019803B2 · Venable · 2018 [cited by applicant]
US 10032072B1 · Tran · 2018 [cited by applicant]
US 10129524B2 · Ng · 2018 [cited by applicant]
US 10210432B2 · Pisoni · 2019 [cited by applicant]
US 10373116B2 · Medina · 2019 [cited by applicant]
US 10572757B2 · Graham · 2020 [cited by applicant]
US 10592854B2 · Schwartz · 2020 [cited by applicant]
US 10796352B2 · Chechuy · 2020 [cited by applicant]
US 10839452B1 · Guo · 2020 [cited by applicant]
US 10922574B1 · Tariq · 2021 [cited by applicant]
US 10943278B2 · Benkreira · 2021 [cited by applicant]
US 10956711B2 · Adato · 2021 [cited by applicant]
US 10990950B2 · Garner · 2021 [cited by applicant]
US 10991036B1 · Bergstrom · 2021 [cited by applicant]
US 11036949B2 · Powell · 2021 [cited by applicant]
US 11055905B2 · Tagra · 2021 [cited by applicant]
US 11087272B2 · Skaff · 2021 [cited by applicant]
US 11151426B2 · Dutta · 2021 [cited by applicant]
US 11163805B2 · Arocho · 2021 [cited by applicant]
US 11276034B2 · Shah · 2022 [cited by applicant]
US 11282287B2 · Gausebeck · 2022 [cited by applicant]
US 11295163B1 · Schoner · 2022 [cited by applicant]
US 11308775B1 · Sinha · 2022 [cited by applicant]
US 11409977B1 · Glaser · 2022 [cited by applicant]
US 20050238465A1 · Razumov · 2005 [cited by applicant]
US 20100188580A1 · Paschalaks · 2010 [cited by applicant]
US 20110040427A1 · Ben-Tzvi · 2011 [cited by applicant]
US 20120303412A1 · Etzioni · 2012 [cited by applicant]
US 20140002239A1 · Rayner · 2014 [cited by applicant]
US 20140247116A1 · Davidson · 2014 [cited by applicant]
US 20140307938A1 · Doi · 2014 [cited by applicant]
US 20150363660A1 · Vidal · 2015 [cited by applicant]
US 20160203525A1 · Hara · 2016 [cited by applicant]
US 20170106738A1 · Gillett · 2017 [cited by applicant]
US 20170286773A1 · Skaff · 2017 [cited by applicant]
US 20180005176A1 · Williams · 2018 [cited by applicant]
US 20180018788A1 · Olmstead · 2018 [cited by applicant]
US 20180108134A1 · Venable · 2018 [cited by applicant]
US 20180197223A1 · Grossman · 2018 [cited by applicant]
US 20180260772A1 · Chaubard · 2018 [cited by applicant]
US 20190025849A1 · Dean · 2019 [cited by applicant]
US 20190043003A1 · Fisher · 2019 [cited by applicant]
US 20190050427A1 · Wiesel et al. · 2019 [cited by applicant]
US 20190050932A1 · Dey · 2019 [cited by applicant]
US 20190087772A1 · Medina · 2019 [cited by applicant]
US 20190163698A1 · Kwon · 2019 [cited by applicant]
US 20190180150A1 · Taylor et al. · 2019 [cited by applicant]
US 20190197561A1 · Adato · 2019 [cited by applicant]
US 20190220482A1 · Crosby · 2019 [cited by applicant]
US 20190236531A1 · Adato · 2019 [cited by applicant]
US 20200118063A1 · Fu · 2020 [cited by applicant]
US 20200246977A1 · Swietojanski · 2020 [cited by applicant]
US 20200265494A1 · Glaser · 2020 [cited by applicant]
US 20200324976A1 · Diehr · 2020 [cited by applicant]
US 20200356813A1 · Sharma · 2020 [cited by applicant]
US 20200380226A1 · Rodriguez · 2020 [cited by applicant]
US 20200387858A1 · Hasan · 2020 [cited by applicant]
US 20210049541A1 · Gong · 2021 [cited by applicant]
US 20210049542A1 · Dalal · 2021 [cited by applicant]
US 20210142105A1 · Siskind · 2021 [cited by applicant]
US 20210150231A1 · Kehl · 2021 [cited by applicant]
US 20210192780A1 · Kulkarni · 2021 [cited by applicant]
US 20210216954A1 · Chaubard · 2021 [cited by applicant]
US 20210272269A1 · Suzuki · 2021 [cited by applicant]
US 20210319684A1 · Ma · 2021 [cited by applicant]
US 20210342914A1 · Dalal · 2021 [cited by applicant]
US 20210400195A1 · Adato · 2021 [cited by applicant]
US 20210406812A1 · Deshmukh · 2021 [cited by applicant]
US 20220043547A1 · Jahjah · 2022 [cited by applicant]
US 20220051177A1 · Savvides · 2022 [cited by examiner]
US 20220051179A1 · Savvides · 2022 [cited by applicant]
US 20220058425A1 · Savvides · 2022 [cited by applicant]
US 20220067085A1 · Nihas · 2022 [cited by applicant]
US 20220114403A1 · Shaw · 2022 [cited by applicant]
US 20220114821A1 · Arroyo · 2022 [cited by applicant]
US 20220138914A1 · Wang · 2022 [cited by applicant]
US 20220165074A1 · Srivastava · 2022 [cited by applicant]
US 20220222924A1 · Pan · 2022 [cited by examiner]
US 20220262008A1 · Kidd · 2022 [cited by applicant]
US 20220406030A1 · Zheng et al. · 2022 [cited by applicant]
US 20230004736A1 · Glaser · 2023 [cited by examiner]
US 20230252343A1 · McDaniel et al. · 2023 [cited by applicant]
CN 106347550B · 2019 [cited by applicant]
CN 110348439B · 2019 [cited by applicant]
CN 110443298B · 2022 [cited by applicant]
CN 114898358A · 2022 [cited by applicant]
CN 115205584A · 2022 [cited by applicant]
EP 3217324A1 · 2017 [cited by applicant]
EP 3437031 · 2019 [cited by applicant]
EP 3479298 · 2019 [cited by applicant]
WO 2006113281A2 · 2006 [cited by applicant]
WO 2017201490A1 · 2017 [cited by applicant]
WO 2018093796 · 2018 [cited by applicant]
WO 2020051213A1 · 2020 [cited by applicant]
WO 2021186176A1 · 2021 [cited by applicant]
WO 2021247420A2 · 2021 [cited by applicant]
U.S. Appl. No. 17/963,751, filed Oct. 11, 2022, Yilun Chen. [cited by applicant]
U.S. Appl. No. 17/963,787, filed Oct. 11, 2022, Lingfeng Zhang. [cited by applicant]
U.S. Appl. No. 17/963,802, filed Oct. 11, 2022, Lingfeng Zhang. [cited by applicant]
U.S. Appl. No. 17/963,903, filed Oct. 11, 2022, Raghava Balusu. [cited by applicant]
U.S. Appl. No. 17/966,580, filed Oct. 14, 2022, Paarvendhan Puviyarasu. [cited by applicant]
U.S. Appl. No. 17/971,350, filed Oct. 21, 2022, Jing Wang. [cited by applicant]
U.S. Appl. No. 17/983,773, filed Nov. 9, 2022, Lingfeng Zhang. [cited by applicant]
U.S. Appl. No. 18/102,999, filed Jan. 30, 2023, Han Zhang. [cited by applicant]
U.S. Appl. No. 18/103,338, filed Jan. 30, 2023, Wei Wang. [cited by applicant]
U.S. Appl. No. 18/106,269, filed Feb. 6, 2023, Zhaoliang Duan. [cited by applicant]
U.S. Appl. No. 18/158,925, filed Jan. 24, 2023, Raghava Balusu. [cited by applicant]
U.S. Appl. No. 18/158,950, filed Jan. 24, 2023, Ishan Arora. [cited by applicant]
U.S. Appl. No. 18/158,969, filed Jan. 24, 2023, Zhaoliang Duan. [cited by applicant]
U.S. Appl. No. 18/158,983, filed Jan. 24, 2023, Ashlin Ghosh. [cited by applicant]
U.S. Appl. No. 18/165,152, filed Feb. 6, 2023, Han Zhang. [cited by applicant]
U.S. Appl. No. 18/168,174, filed Feb. 13, 2023, Abhinav Pachauri. [cited by applicant]
U.S. Appl. No. 18/168,198, filed Feb. 13, 2023, Ashlin Ghosh. [cited by applicant]
Chaudhuri, Abon et al.; “A Smart System for Selection of Optimal Product Images in E-Commerce”; 2018 IEEE Conference on Big Data (Big Data); Dec. 10-13, 2018; IEEE; <https://ieeexplore.ieee.org/document/8622259>; pp. 17… [cited by applicant]
Chenze, Brandon et al.; “Iterative Approach for Novel Entity Recognition of Foods in Social Media Messages”; 2022 IEEE 23rd International Conference on Information Reuse and Integration for Data Science (IRI); Aug. 9-11… [cited by applicant]
Kaur, Ramanpreet et al.; “A Brief Review on Image Stitching and Panorama Creation Methods”; International Journal of Control Theory and Applications; 2017; vol. 10, No. 28; International Science Press; Gurgaon, India; <… [cited by applicant]
Naver Engineering Team; “Auto-classification of NAVER Shopping Product Categories using TensorFlow”; <https://blog.tensorflow.org/2019/05/auto-classification-of-naver-shopping.html>; May 20, 2019; pp. 1-15. [cited by applicant]
Paolanti, Marine et al.; “Mobile robot for retail surveying and inventory using visual and textual analysis of monocular pictures based on deep learning”; European Conference on Mobile Robots; Sep. 2017, 6 pages. [cited by applicant]
Refills; “Final 3D object perception and localization”; European Commision, Dec. 31, 2016, 16 pages. [cited by applicant]
Retech Labs; “Storx | RetechLabs”; <https://retechlabs.com/storx/>; available at least as early as Jun. 22, 2019; retrieved from Internet Archive Wayback Machine <https://web.archive.org/web/20190622012152/https://retec… [cited by applicant]
Schroff, Florian et al.; “Facenet: a unified embedding for face recognition and clustering”; 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); Jun. 7-12, 2015; IEEE; <https://ieeexplore.ieee.org/do… [cited by applicant]
Tan, Mingxing et al.; “EfficientDet: Scalable and Efficient Object Detection”; 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); Jun. 13-19, 2020; IEEE; <https://ieeexplore.ieee.org/document/91… [cited by applicant]
Tan, Mingxing et al.; “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks”; Proceedings of the 36th International Conference on Machine Learning; 2019; vol. 97; PLMR; <http://proceedings.mlr.press/… [cited by applicant]
Technology Robotix Society; “Colour Detection”; <https://medium.com/image-processing-in-robotics/colour-detection-e15bc03b3f61>; Jul. 2, 2019; pp. 1-6. [cited by applicant]
Tonioni, Alessio et al.; “A deep learning pipeline for product recognition on store shelves”; 2018 IEEE International Conference on Image Processing, Applications and Systems (IPAS); Dec. 12-14, 2018; IEEE; <https://iee… [cited by applicant]
Trax Retail; “Image Recognition Technology for Retail | Trax”; <https://traxretail.com/retail/>; available at least as early as Apr. 20, 2021; retrieved from Internet Wayback Machine <https://web.archive.org/web/2021042… [cited by applicant]
Verma, Nishchal, et al.; “Object identification for inventory management using convolutional neural network”; IEEE Applied Imagery Pattern Recognition Workshop (AIPR); Oct. 2016, 6 pages. [cited by applicant]
Zhang, Jicun, et al.; “An Improved Louvain Algorithm for Community Detection”; Advanced Pattern and Structure Discovery from Complex Multimedia Data Environments 2021; Nov. 23, 2021; Mathematical Problems in Engineering… [cited by applicant]
Rodriquez, Kari, “International Search Report & Written Opinion”, International Application No. PCT/US2024/012702, mailed Apr. 19, 2024, 7 pages. [cited by applicant]