IP Library › Granted Patent US 12,591,847
Granted Patent B2
US 12,591,847 · App. 17/963,751 · Granted Mar 31, 2026

Systems and methods of transforming image data to product storage facility location information

Inventors: Yilun Chen (Dallas, TX); Ravi Kumar Dalal (Dallas, TX); Adam Cantor (Austin, TX); Monisha Elumalai (Dallas, TX); Benjamin R. Ellison (San Francisco, CA)
Assignee: Walmart Apollo, LLC
G06Q10/087G06T5/73G06T7/11G06V10/75G06V20/36G06V20/62G06V30/153G06T2207/20081G06T2207/20132
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,847
App. No.
17/963,751
Filed
Oct 11, 2022
Granted
Mar 31, 2026
Kind
B2
Art Unit
2667
USPC
382/155
Abstract

Some embodiments provide systems comprising: a machine learning model database; a deblur system configured to receive at least a portion of an image comprising a presumed location label captured by the image capture device, and apply at least a deblurring machine learning framework to generate a deblurred label image comprising the presumed location label; a rectification system configured to apply an machine learning transform algorithm to the deblurred label image to generate a rectified label image; an optical character recognition (OCR) system configured to apply a recognition machine learning model to the rectified label image to estimate text; and a location estimation system configured to estimate a location of the presumed location label as a function of the estimated text of the presumed location label relative to known text on known location labels position at respective different known locations within the product storage facility.

Claims (55)

1 . An image based retail location confirmation system, comprising:

a machine learning model database storing a set of two or more machine learning models;

a deblur system communicatively coupled over a distributed communication network with the machine learning model database, wherein the deblur system is configured to receive a portion of a first image comprising a first presumed location label and a second presumed location label captured by an image capture device configured to capture images within a product storage facility, and apply at least a deblurring machine learning framework to the portion of the first image and generate a first deblurred label image comprising the first presumed location label and the second presumed location label;

a rectification system communicatively coupled over the distributed communication network with the machine learning model database, wherein the rectification system is configured to apply a rectification machine learning transform algorithm to justify the first deblurred label image according to a known shape of location labels to thereby generate a first rectified label image;

an optical character recognition (OCR) system communicatively coupled over the distributed communication network with the machine learning model database, wherein the OCR system is configured to:

apply a recognition machine learning model to the first rectified label image to estimate text of the first presumed location label and the second presumed location label; and

verify that identified patterns of the first presumed location label and the second presumed location label conform with a known location label pattern for location labels; and

a location estimation system configured to;

select one of the first presumed location label and the second presumed location label as a selected presumed location label based on which one of the first presumed location label and the second presumed location label is assigned a higher prioritized value for determining location; and

estimate a location within the product storage facility of the selected presumed location label as a function of the estimated text of the selected presumed location label relative to known text on known location labels positioned at respective different known locations within the product storage facility.

2 . The system of claim 1 , further comprising:

an image cropping system communicatively coupled over the distributed communication network with the machine learning model database, wherein the image cropping system is configured to apply a trained cropping machine learning model to identify within and extract from the first image the portion of the first image comprising the first presumed location label the second presumed location label.

3 . The system of claim 2 , further comprising:

a blur evaluation system configured to estimate a level of blur of the portion of the first image and enable the deblur system to generate the first deblurred label image when the estimated level of blur has a predefined relationship with a blur threshold.

4 . The system of claim 2 , further comprising:

a confidence evaluation system configured to apply at least one confidence rule to determine a confidence score of an accuracy of the estimated text of the first presumed location label and the second presumed location label; and

enable the location estimation system configured to estimate the location of the image capture device when the confidence score has a predefined relationship with a confidence threshold, and prevent the location estimation system from estimating locations of the image capture device when other confidence scores do not have the predefined relationship with the confidence threshold.

5 . The system of claim 4 , wherein the confidence evaluation system in applying the at least one confidence rule is configured to confirm that the estimated text complies with a predefined alphanumeric pattern of multiple alphanumeric characters.

6 . The system of claim 4 , further comprising:

a mobile task system comprising the image capture device, a task system, and a movement system communicatively coupled with the task system, wherein the task system is configured to implement code to control the movement system to control movement of the task system through the product storage facility;

wherein the image capture device is configured to capture the images as the task system moves through one or more portions of the product storage facility.

7 . The system of claim 4 , further comprising:

a portable user computing device comprising the image capture device, wherein the portable user computing device is configured to be transported by a user associated with the portable user computing device as the user moves through the product storage facility, and wherein the image capture device is configured to capture the images as the portable user computing device is transported through one or more portions of the product storage facility.

8 . The system of claim 4 , wherein the deblur system applying the deblurring machine learning framework is configured to apply a generative adversarial network (GAN) to the portion of the first image and generate the first deblurred label image comprising the first presumed location label and the second presumed location label.

9 . The system of claim 4 , further comprising:

a blur training database storing numerous training sets of actual images and numerous training sets of artificial images, wherein the numerous training sets of actual images each comprise: an actual image of one of the known location labels and at least one artificially blurred version of the one of the known location labels; and wherein the numerous training sets of artificial images each comprises: an artificially generated image of a representative location label based on known format, font and size of alphanumeric characters included on the known location labels, and at least one artificially blurred version of the artificially generated image; and

a machine learning training system communicatively coupled over the distributed communication network with the machine learning model database and the blur training database, wherein the machine learning training system is configured to repeatedly train over time the deblurring machine learning framework utilizing the numerous training sets of actual images and the numerous training sets of artificial images.

10 . A method of confirming locations within a product storage facility, comprising:

storing, in a machine learning model database, a set of two or more machine learning models;

receiving images captured by an image capture device configured to capture the images within the product storage facility;

receiving a portion of a first image, of the images, wherein the portion of the first image comprises a first presumed location label and a second presumed location label captured by the image capture device;

applying a deblurring machine learning framework to the portion of the first image and generating a first deblurred label image comprising the first presumed location label and the second presumed location label;

applying a rectification machine learning transform algorithm to justify the first deblurred label image according to a known shape of location labels and thereby generating a first rectified label image;

applying a recognition machine learning model to the first rectified label image to:

estimate text of the first presumed location label and the second presumed location label; and

verify that identified patterns of the first presumed location label and the second presumed location label conform with a known location label pattern for location labels;

selecting one of the first presumed location label and the second presumed location label as a selected presumed location label based on which one of the first presumed location label and the second presumed location label is assigned a higher prioritized value for determining location; and

estimating a location within the product storage facility of the selected presumed location label as a function of the estimated text of the selected presumed location label relative to known text on known location labels position at respective different known locations within the product storage facility.

11 . The method of claim 10 , further comprising:

applying a trained cropping machine learning model to identify within and extract from the first image the portion of the first image comprising the first presumed location label and the second presumed location label.

12 . The method of claim 11 , further comprising:

estimating, by a blur evaluation system, a level of blur of the portion of the first image and enabling the generating of the first deblurred label image when the estimated level of blur has a predefined relationship with a blur threshold.

13 . The method of claim 11 , further comprising:

applying at least one confidence rule and determining a confidence score of an accuracy of the estimated text of the first presumed location label and the second presumed location label; and

enabling the estimating of the location of the image capture device when the confidence score has a predefined relationship with a confidence threshold, and preventing the estimation of locations of the image capture device when other confidence scores do not have the predefined relationship with the confidence threshold.

14 . The method of claim 13 , wherein the applying the at least one confidence rule comprises confirming that the estimated text complies with a predefined alphanumeric pattern of multiple alphanumeric characters.

15 . The method of claim 13 , further comprising:

controlling a movement system of a mobile task system in controlling movement of the mobile task system through the product storage facility, wherein the mobile task system comprises the image capture device; and

controlling the image capture device to capture the images as the mobile task system moves through one or more portions of the product storage facility.

16 . The method of claim 13 , further comprising:

controlling the image capture device of a portable user computing device to capture the images as the portable user computing device is transported through one or more portions of the product storage facility by a user associated with the portable user computing device as the user moves through the product storage facility.

17 . The method of claim 13 , wherein the applying the deblurring machine learning framework comprises applying a generative adversarial network (GAN) to the portion of the first image and generating the first deblurred label image comprising the first presumed location label and the second presumed location label.

18 . The method of claim 13 , further comprising:

generating and storing in a blur training database storing numerous training sets of actual images and numerous training sets of artificial images, wherein the numerous training sets of actual images each comprise: an actual image of one of the known location labels and at least one artificially blurred version of the one of the known location labels; and wherein the numerous training sets of artificial images each comprises: an artificially generated image of a representative location label based on known format, font and size of alphanumeric characters included on the known location labels, and at least one artificially blurred version of the artificially generated image; and

repeatedly training over time the deblurring machine learning framework utilizing the numerous training sets of actual images and the numerous training sets of artificial images.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2022
From: CHEN, YILUN; DALAL, RAVI KUMAR; CANTOR, ADAM; ELUMALAI, MONISHA; ELLISON, BENJAMIN R.
To: WALMART APOLLO, LLC
Reel/Frame 061401/0567 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2022
From: CHEN, YILUN; DALAL, RAVI KUMAR; CANTOR, ADAM; ELUMALAI, MONISHA; ELLISON, BENJAMIN R.
To: WALMART APOLLO, LLC
Reel/Frame 061401/0933 →
Continuity (1)
Related Publication 20240119408A1 · Apr 11, 2024
References Cited (128)
US 5074594A · Laganowski · 1991 [cited by applicant]
US 6570492B1 · Peratoner · 2003 [cited by applicant]
US 8923650B2 · Wexler · 2014 [cited by applicant]
US 8965104B1 · Hickman · 2015 [cited by applicant]
US 9275308B2 · Szegedy · 2016 [cited by applicant]
US 9477955B2 · Goncalves · 2016 [cited by applicant]
US 9526127B1 · Taubman · 2016 [cited by applicant]
US 9576310B2 · Cancro · 2017 [cited by applicant]
US 9659204B2 · Wu · 2017 [cited by applicant]
US 9811754B2 · Schwartz · 2017 [cited by applicant]
US 10002344B2 · Wu · 2018 [cited by applicant]
US 10019803B2 · Venable · 2018 [cited by applicant]
US 10032072B1 · Tran · 2018 [cited by applicant]
US 10129524B2 · Ng · 2018 [cited by applicant]
US 10210432B2 · Pisoni · 2019 [cited by applicant]
US 10373116B2 · Medina · 2019 [cited by applicant]
US 10572757B2 · Graham · 2020 [cited by applicant]
US 10592854B2 · Schwartz · 2020 [cited by applicant]
US 10839452B1 · Guo · 2020 [cited by applicant]
US 10922574B1 · Tariq · 2021 [cited by applicant]
US 10943278B2 · Benkreira · 2021 [cited by applicant]
US 10956711B2 · Adato · 2021 [cited by applicant]
US 10990950B2 · Garner · 2021 [cited by applicant]
US 10991036B1 · Bergstrom · 2021 [cited by applicant]
US 11036949B2 · Powell · 2021 [cited by applicant]
US 11055905B2 · Tagra · 2021 [cited by applicant]
US 11087272B2 · Skaff · 2021 [cited by applicant]
US 11151426B2 · Dutta · 2021 [cited by applicant]
US 11163805B2 · Arocho · 2021 [cited by applicant]
US 11276034B2 · Shah · 2022 [cited by applicant]
US 11282287B2 · Gausebeck · 2022 [cited by applicant]
US 11295163B1 · Schoner · 2022 [cited by examiner]
US 11308775B1 · Sinha · 2022 [cited by applicant]
US 11409977B1 · Glaser · 2022 [cited by applicant]
US 11481683B1 · Singh · 2022 [cited by examiner]
US 20050238465A1 · Razumov · 2005 [cited by applicant]
US 20110040427A1 · Ben-Tzvi · 2011 [cited by applicant]
US 20140002239A1 · Rayner · 2014 [cited by applicant]
US 20140247116A1 · Davidson · 2014 [cited by applicant]
US 20140307938A1 · Doi · 2014 [cited by applicant]
US 20150363660A1 · Vidal · 2015 [cited by applicant]
US 20160203525A1 · Hara · 2016 [cited by applicant]
US 20170106738A1 · Gillett · 2017 [cited by applicant]
US 20170286773A1 · Skaff · 2017 [cited by applicant]
US 20180005176A1 · Williams · 2018 [cited by applicant]
US 20180018788A1 · Olmstead · 2018 [cited by applicant]
US 20180197223A1 · Grossman · 2018 [cited by applicant]
US 20180260772A1 · Chaubard · 2018 [cited by applicant]
US 20190025849A1 · Dean · 2019 [cited by applicant]
US 20190043003A1 · Fisher · 2019 [cited by applicant]
US 20190050932A1 · Dey · 2019 [cited by applicant]
US 20190087772A1 · Medina · 2019 [cited by applicant]
US 20190163698A1 · Kwon · 2019 [cited by applicant]
US 20190197561A1 · Adato · 2019 [cited by applicant]
US 20190220482A1 · Crosby · 2019 [cited by applicant]
US 20190236531A1 · Adato · 2019 [cited by applicant]
US 20200246977A1 · Swietojanski · 2020 [cited by applicant]
US 20200265494A1 · Glaser · 2020 [cited by applicant]
US 20200324976A1 · Diehr · 2020 [cited by applicant]
US 20200356813A1 · Sharma · 2020 [cited by applicant]
US 20200380226A1 · Rodriguez · 2020 [cited by applicant]
US 20200387858A1 · Hasan · 2020 [cited by applicant]
US 20210049541A1 · Gong · 2021 [cited by applicant]
US 20210049542A1 · Dalal · 2021 [cited by applicant]
US 20210142105A1 · Siskind · 2021 [cited by applicant]
US 20210150231A1 · Kehl · 2021 [cited by applicant]
US 20210192780A1 · Kulkarni · 2021 [cited by applicant]
US 20210216954A1 · Chaubard · 2021 [cited by applicant]
US 20210272269A1 · Suzuki · 2021 [cited by applicant]
US 20210319684A1 · Ma · 2021 [cited by applicant]
US 20210342914A1 · Dalal · 2021 [cited by applicant]
US 20210400195A1 · Adato · 2021 [cited by applicant]
US 20220043547A1 · Jahjah · 2022 [cited by applicant]
US 20220051179A1 · Savvides · 2022 [cited by applicant]
US 20220058425A1 · Savvides · 2022 [cited by applicant]
US 20220067085A1 · Nihas · 2022 [cited by applicant]
US 20220114403A1 · Shaw · 2022 [cited by applicant]
US 20220114821A1 · Arroyo · 2022 [cited by applicant]
US 20220138914A1 · Wang · 2022 [cited by applicant]
US 20220165074A1 · Srivastava · 2022 [cited by applicant]
US 20220222924A1 · Pan · 2022 [cited by applicant]
US 20220262008A1 · Kidd · 2022 [cited by applicant]
US 20230005108A1 · Gopalkrishna · 2023 [cited by examiner]
US 20240394854A1 · Koo · 2024 [cited by examiner]
CN 106347550B · 2019 [cited by applicant]
CN 110348439B · 2019 [cited by applicant]
CN 110443298B · 2022 [cited by applicant]
CN 114898358A · 2022 [cited by applicant]
EP 3217324A1 · 2017 [cited by applicant]
EP 3437031 · 2019 [cited by applicant]
EP 3479298 · 2019 [cited by applicant]
WO 2006113281A2 · 2006 [cited by applicant]
WO 2017201490A1 · 2017 [cited by applicant]
WO 2018093796 · 2018 [cited by applicant]
WO 2020051213A1 · 2020 [cited by applicant]
WO 2021186176A1 · 2021 [cited by applicant]
WO 2021247420A2 · 2021 [cited by applicant]
X. Zhang, M. Chen and D. Wu, “Industrial Character Image Motion Deblurring and Target Region Dynamic Location Method,” 2022 5th International Conference on Pattern Recognition and Artificial Intelligence (PRAI), Chengdu… [cited by examiner]
P. Jadhav et al., “Pix2Pix Generative Adversarial Network with ResNet for Document Image Denoising,” 2022 4th International Conference on Inventive Research in Computing Applications (ICIRCA), Coimbatore, India, 2022, p… [cited by examiner]
N. Jaiswal and Y. K. Meghrajani, “Saliency based automatic image cropping using support vector machine classifier,” 2015 International Conference on Innovations in Information, Embedded and Communication Systems (ICIIEC… [cited by examiner]
U.S. Appl. No. 17/963,787, filed Oct. 11, 2022, Lingfeng Zhang. [cited by applicant]
U.S. Appl. No. 17/963,802, filed Oct. 11, 2022, Lingfeng Zhang. [cited by applicant]
U.S. Appl. No. 17/963,903, filed Oct. 11, 2022, Raghava Balusu. [cited by applicant]
U.S. Appl. No. 17/966,580, filed Oct. 14, 2022, Paarvendhan Puviyarasu. [cited by applicant]
U.S. Appl. No. 17/971,350, filed Oct. 21, 2022, Jing Wang. [cited by applicant]
U.S. Appl. No. 17/983,773, filed Nov. 9, 2022, Lingfeng Zhang. [cited by applicant]
Chaudhuri, Abon et al.; “A Smart System for Selection of Optimal Product Images in E-Commerce”; 2018 IEEE Conference on Big Data (Big Data); Dec. 10-13, 2018; IEEE; <https://ieeexplore.ieee.org/document/8622259>; pp. 17… [cited by applicant]
Chenze, Brandon et al.; “Iterative Approach for Novel Entity Recognition of Foods in Social Media Messages”; 2022 IEEE 23rd International Conference on Information Reuse and Integration for Data Science (IRI); Aug. 9-11… [cited by applicant]
Naver Engineering Team; “Auto-classification of NAVER Shopping Product Categories using TensorFlow”; <https://blog.tensorflow.org/2019/05/auto-classification-of-naver-shopping.html>; May 20, 2019; pp. 1-13. [cited by applicant]
Paolanti, Marine et al.; “Mobile robot for retail surveying and inventory using visual and textual analysis of monocular pictures based on deep learning”; European Conference on Mobile Robots; Sep. 2017, 6 pages. [cited by applicant]
Ramanpreet Kaur et al.; “A Brief Review on Image Stitching and Panorama Creation Methods”; International Journal of Control Theory and Applications; 2017; vol. 10, No. 28; International Science Press; Gurgaon, India; <h… [cited by applicant]
Refills; “Final 3D object perception and localization”; European Commision, Dec. 31, 2016, 16 pages. [cited by applicant]
Retech Labs; “Storx | RetechLabs”; <https://retechlabs.com/storx/>; available at least as early as Jun. 22, 2019; retrieved from Internet Archive Wayback Machine <https://web.archive.org/web/20190622012152/https://retec… [cited by applicant]
Schroff, Florian et al.; “Facenet: a unified embedding for face recognition and clustering”; 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); Jun. 7-12, 2015; IEEE; <https://ieeexplore.ieee.org/do… [cited by applicant]
Singh, Ankit; “Automated Retail Shelf Monitoring Using AI”; <https://blog.paralleldots.com/shelf-monitoring/automated-retail-shelf-monitoring-using-ai/>; Sep. 20, 2019; pp. 1-12. [cited by applicant]
Singh, Ankit; “Image Recognition and Object Detection in Retail”; <https://blog.paralleldots.com/featured/image-recognition-and-object-detection-in-retail/>; Sep. 26, 2019; pp. 1-11. [cited by applicant]
Tan, Mingxing et al.; “EfficientDet: Scalable and Efficient Object Detection”; 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); Jun. 13-19, 2020; IEEE; <https://ieeexplore.ieee.org/document/91… [cited by applicant]
Tan, Mingxing et al.; “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks”; Proceedings of the 36th International Conference on Machine Learning; 2019; vol. 97; PLMR; <http://proceedings.mlr.press/… [cited by applicant]
Technology Robotix Society; “Colour Detection”; < https://medium.com/image-processing-in-robotics/colour-detection-e15bc03b3f61>; Jul. 2, 2019; pp. 1-8. [cited by applicant]
Tonioni, Alessio et al.; “A deep learning pipeline for product recognition on store shelves”; 2018 IEEE International Conference on Image Processing, Applications and Systems (IPAS); Dec. 12-14, 2018; IEEE; <https://iee… [cited by applicant]
Trax Retail; “Image Recognition Technology for Retail | Trax”; < https://traxretail.com/retail/>; available at least as early as Apr. 20, 2021; retrieved from Internet Wayback Machine <https://web.archive.org/web/202104… [cited by applicant]
USPTO; U.S. Appl. No. 16/991,885; Final Rejection mailed Mar. 7, 2022; (pp. 1-13). [cited by applicant]
USPTO; U.S. Appl. No. 16/991,885; Notice of Allowance and Fees Due (PTOL-85) mailed Aug. 24, 2022; (pp. 1-13). [cited by applicant]
USPTO; U.S. Appl. No. 16/991,885; Notice of Allowance and Fees Due (PTOL-85) mailed Dec. 7, 2022; (pp. 1-13). [cited by applicant]
USPTO; U.S. Appl. No. 16/991,885; Notice of Allowance and Fees Due (PTOL-85) mailed Dec. 23, 2022; (pp. 1-2). [cited by applicant]
USPTO; U.S. Appl. No. 16/991,885; Office Action mailed Sep. 20, 2021; (pp. 1-13). [cited by applicant]
USPTO; U.S. Appl. No. 16/991,980; Non-Final Rejection mailed Sep. 14, 2022; (pp. 1-19). [cited by applicant]
Verma, Nishchal et al.; “Object identification for inventory management using convolutional neural network”; IEEE Applied Imagery Pattern Recognition Workshop (AIPR); Oct. 2016, 6 pages. [cited by applicant]