IP Library Granted Patent US 12,412,209
Granted Patent B2
US 12,412,209 · App. 18/176,700 · Granted Sep 9, 2025

Image processing arrangements

Inventors: Ravi K. Sharma (Portland, OR); Vahid Sedighianaraki (Portland, OR); William Y. Conwell (Portland, OR)
Assignee: Digimarc Corporation
G06Q30/0641G06F18/214G06F18/22G06F18/2431G06V10/17G06V10/462G06V10/764G06V10/774G06V10/776G06V10/82G06V20/00G06F16/906G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,209
App. No.
18/176,700
Granted
Sep 9, 2025
Kind
B2
Abstract

Aspects of the detailed technologies concern training and use of neural networks for fine-grained classification of large numbers of items, e.g., as may be encountered in a supermarket. Mitigating false positive errors is an exemplary area of emphasis. Novel network topologies are also detailed—some employing recognition technologies in addition to neural networks. A great number of other features and arrangements are also detailed.

Claims (52)

1. A method of processing imagery captured from items in a retail store to compile an item checkout tally, the method including the acts:

for a first particular item:

capturing, via one or more cameras, first image data depicting the first particular item;

applying, in real-time during checkout, the first image data depicting the first particular item to first and second recognition processes executing concurrently on one or more computer processors,

wherein the first recognition process applies neural network-based item identification by processing the first image data through multiple neural network layers of a trained convolutional neural network to generate probability outputs across multiple output neurons, and the second recognition process applies at least one of fingerprint-, histogram-, OCR-, barcode- or watermark-based item identification;

receiving first result data for said first particular item from the first recognition process, and receiving second result data for said first particular item from the second recognition process;

disregarding said first result data for the first particular item received from the first recognition process and adding said first particular item to said checkout tally in accordance with said second result data received from the second recognition process;

for a second particular item:

capturing, via one or more cameras, second image data depicting the second particular item;

applying, in real-time during checkout, the second image data depicting the second particular item to said first and second recognition processes;

receiving first result data for said second particular item from the first recognition process, and second result identification data for said second particular item from the second recognition process; and

adding said second particular item to said checkout tally in accordance with said first result data received from the first recognition process;

whereby mis-identification and false-positive errors are minimized when compiling the item checkout tally.

2. The method of claim 1 wherein said first result data for the first particular item is disregarded based on said first result data received for said first particular item.

3. The method of claim 1 wherein said first result data for the first particular item is disregarded based on second result data received from the second recognition process applied to the first image data.

4. The method of claim 1 wherein said first result data for the first particular item is disregarded based on second result data received from a fingerprint-based recognition process applied to the first image data.

5. The method of claim 1 wherein said first result data for the first particular item is disregarded based on second result data received from an OCR-based recognition process applied to the first image data.

6. The method of claim 1 wherein said first result data for the first particular item is disregarded based on second result data received from a color histogram-based recognition process applied to the first image data.

7. The method of claim 1 wherein second result data received from the second recognition process indicates that the first image data depicts an item susceptible to being incorrectly indicated by the first recognition process.

8. The method of claim 1 that includes applying fingerprint-based item identification to the second image data, and not disregarding the first result data for the second particular item from the first recognition process due to said fingerprint-based item identification failing to identify a match for said second image data.

9. The method of claim 1 in which the second recognition process applies fingerprint-based item identification.

10. The method of claim 1 in which the second recognition process applies color histogram-based item identification.

11. The method of claim 1 in which the second recognition process applies OCR-based item identification.

12. The method of claim 1 in which the second recognition process applies barcode-based item identification.

13. The method of claim 1 in which the second recognition process applies watermark-based item identification.

14. A system to compile a checkout tally of items in a retail store, comprising:

one or more cameras;

one or more processors configured in accordance with software and/or parameters stored in one or more memories, said software and/or parameters configuring the one or more processors to perform acts including:

for a first particular item:

capturing, via the one or more cameras, first image data depicting the first particular item;

applying, in real-time during checkout, the first image data depicting the first particular item to first and second recognition processes, executing concurrently on the one or more processors,

wherein the first recognition process applies neural network-based item identification by processing the first image data through multiple neural network layers of a trained convolutional neural network to generate probability outputs across multiple output neurons, and the second recognition process applies at least one of fingerprint-, histogram-, OCR-, barcode- or watermark-based item identification;

receiving first result data for said first particular item from the first recognition process, and receiving second result data for said first particular item from the second recognition process;

disregarding said first result data for the first particular item from the first recognition process and adding said first particular item to said checkout tally in accordance with said second result data received from the second recognition process;

for a second particular item:

capturing, via the one or more cameras, second image data depicting the second particular item;

applying second image data depicting the second particular item to said first and second recognition processes;

receiving first result data for said second particular item from the first recognition process, and receiving second result data for said second particular item from the second recognition process; and

adding said second particular item to said checkout tally in accordance with said first result data received from the first recognition process;

whereby mis-identification and false-positive errors are minimized when compiling the checkout tally.

15. The method of claim 1 that includes, for said first particular item:

applying the first image data initially to the second recognition module, said second recognition module applying at least fingerprint-based item identification;

determining that no fingerprint match can be found from said first image data; and

following said determining, applying the first image data to the first recognition module;

wherein the trained convolutional neural network is not relied on to identify item classes for which the second recognition module can find a fingerprint match.

16. The method of claim 1 in which applying the first image data to the first recognition process comprises initially applying the first image data to a neural network- or fingerprint-based classifier that examines the first image data for the presence of one or several logo markings indicating a particular retail brand family, and upon detection of such a logo marking indicating a particular retail brand family, applying related image data to a specialized classifier that has been trained on retail items in said particular retail brand family, to thereby reduce false-positive item mis-identification due to logo confusion.

17. The system of claim 14 in which the second recognition process applies OCR-based item identification.

18. The system of claim 14 in which the second recognition process applies barcode-based item identification.

19. The system of claim 14 in which the second recognition process applies watermark-based item identification.

20. The system of claim 14 , wherein:

the first recognition process produces a probability output indicating a likelihood of correct item identification that is below a predetermined threshold; and

the second recognition process applies digital watermark-based item identification that provides an item identification with greater accuracy than the neural network-based item identification, whereby the digital watermark-based identification is used instead of the less reliable neural network identification for the first particular item.

Assignments (3)
ARTICLES OF CONVERSION Recorded Jun 19, 2026
From: DIGIMARC CORPORATION
To: DIGIMARC LLC
Reel/Frame 075863/0211 →
ARTICLES OF AMENDMENT OFTHE ARTICLES OF ORGANIZATION OF DIGIMARC LLC Recorded Jun 19, 2026
From: DIGIMARC LLC
To: DMRC LLC
Reel/Frame 075863/0266 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2025
From: SHARMA, RAVI K.; FILLER, TOMAS; DESHMUKH, UTKARSH; SEDIGHIANARAKI, VAHID; CONWELL, WILLIAM Y.
To: DIGIMARC CORPORATION
Reel/Frame 071554/0245 →
Continuity (9)
Division 16880778 · May 21, 2020
Division 15726290 · Oct 5, 2017
Provisional Application 62556276 · Sep 8, 2017
Provisional Application 62456446 · Feb 8, 2017
Provisional Application 62426148 · Nov 23, 2016
Provisional Application 62418047 · Nov 4, 2016
Provisional Application 62414368 · Oct 28, 2016
Provisional Application 62404721 · Oct 5, 2016
Related Publication 20230316386A1 · Oct 5, 2023
References Cited (92)
US 5285523A · Takahashi · 1994 [cited by examiner]
US 5524178A · Yokono · 1996 [cited by applicant]
US 5903884A · Lyon · 1999 [cited by applicant]
US 6311174B1 · Kato · 2001 [cited by examiner]
US 9799098B2 · Seung · 2017 [cited by applicant]
US 10007863B1 · Pereira · 2018 [cited by applicant]
US 20100312730A1 · Weng · 2010 [cited by applicant]
US 20130308045A1 · Rhoads · 2013 [cited by applicant]
US 20140032461A1 · Weng · 2014 [cited by applicant]
US 20140108309A1 · Frank · 2014 [cited by applicant]
US 20140201111A1 · Kasravi · 2014 [cited by applicant]
US 20150055855A1 · Rodriguez · 2015 [cited by applicant]
US 20150100529A1 · Sarah · 2015 [cited by applicant]
US 20150161482A1 · Preetham · 2015 [cited by applicant]
US 20150170001A1 · Rabinovich · 2015 [cited by applicant]
US 20150242707A1 · Wilf · 2015 [cited by applicant]
US 20150278224A1 · Jaber · 2015 [cited by applicant]
US 20160110642A1 · Matsuda · 2016 [cited by applicant]
US 20160140425A1 · Kulkarni · 2016 [cited by applicant]
US 20160224892A1 · Sawada · 2016 [cited by applicant]
US 20160247061A1 · Trask · 2016 [cited by applicant]
US 20160267358A1 · Shoaib · 2016 [cited by applicant]
US 20160284346A1 · Visser · 2016 [cited by applicant]
US 20160379091A1 · Lin · 2016 [cited by applicant]
US 20170068887A1 · Kwon · 2017 [cited by applicant]
US 20170229117A1 · van der Made · 2017 [cited by examiner]
US 20170294010A1 · Shen · 2017 [cited by applicant]
US 20170316281A1 · Criminisi · 2017 [cited by applicant]
US 20180032844A1 · Yao · 2018 [cited by applicant]
US 20180046619A1 · Shi · 2018 [cited by applicant]
US 20180165548A1 · Wang · 2018 [cited by examiner]
US 20180247107A1 · Murthy · 2018 [cited by examiner]
US 20190057314A1 · Julian · 2019 [cited by examiner]
US 20190087677A1 · Wolf · 2019 [cited by applicant]
US 20200401851A1 · Mau · 2020 [cited by applicant]
US 20210103818A1 · Du · 2021 [cited by applicant]
Babenko, et al, Neural Codes for Image Retrieval, arXiv preprint, arXiv:1404.1777 (2014), 16 pages. [cited by applicant]
Baz, et al, Context-aware hybrid classification system for fine-grained retail product recognition, 2016 IEEE 12th Image, Video, and Multidimensional Signal Processing Workshop (IVMSP) Jul. 11, 2016 (pp. 1-5). [cited by applicant]
Bazzani, et al, Self-Taught Object Localization with Deep Networks, arXiv preprint 1409:3964v7, 2016, 13 pages. [cited by applicant]
Bengio, et al, Curriculum Learning, Proceedings of the 26th ACM Annual International Conference on Machine Learning, pp. 41-48, 2009. [cited by applicant]
Cinbis, et al, Weakly supervised object localization with multi-fold multiple instance learning, arXiv:1503.00949, 2015, 15 pages. [cited by applicant]
Dai, et al, Convolutional Feature Masking for Joint Object and Stuff Segmentation, IEEE Conference on Computer Vision and Pattern Recognition, pp. 3992-4000, 2015. [cited by applicant]
Dai, R-fcn: Object detection via region-based fully convolutional networks, Advances in Neural Information Processing Systems, pp. 379-387, 2016. [cited by applicant]
Felzenszwalb, et al, Object Detection with Discriminatively Trained Part-Based Models, IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1627-1645. 2010. [cited by applicant]
Floerkemeier, Recognizing products: a per-exemplar multi-label image classification approach, European Conference on Computer Vision Sep. 6, 2014 (pp. 440-455). [cited by applicant]
Girshick, Fast R-CNN, Proceedings of the IEEE International Conference on Computer Vision, pp. 1440-1448, 2015. [cited by applicant]
Goodfellow et al, Explaining and Harnessing Adversarial Examples, arXiv preprint, arXiv:1412.6572v3, 2015, 11 pages. [cited by applicant]
Gupta, et al, Learning rich features from RGB-D images for object detection and segmentation, European Conference on Computer Vision, pp. 345-360, 2014. [cited by applicant]
Hariharan, et al, Simultaneous Detection and Segmentation, arXiv preprint 1407.1808v1, 2014, 16 pages. [cited by applicant]
Held, et al, Robust single-view instance recognition, 2016 IEEE International Conference on Robotics and Automation (ICRA) May 16, 2016 (pp. 2152-2159). [cited by applicant]
Ji, Weng, and Prokhorov, Where-what network 1:“Where” and “What” assist each other through top-down connections, 7th IEEE International Conference on Development and Learning, pp. 61-66, 2008. [cited by applicant]
Jiang, et al, Salient Object Detection: a Discriminative Regional Feature Integration Approach, IEEE Conference on Computer Vision and Pattern Recognition, pp. 2083-2090, 2013. [cited by applicant]
Karlinsky, et al, Fine-grained recognition of thousands of object categories with single-example training, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2017 (pp. 4113-4122). [cited by applicant]
Kolesnikov, et al, Seed, Expand and Constrain: Three Principles for Weakly-Supervised Image Segmentation, arXiv preprint 1603:06098v3, 2016, 20 pages. [cited by applicant]
Krizhevsky, et al, Imagenet Classification with Deep Convolutional Neural Networks, Advances in Neural Information Processing Systems, 2012, 9 pages. [cited by applicant]
Krizhevsky, et al, Imagenet classification with deep convolutional neural networks, Advances in Neural Information Processing Systems, pp. 1097-1105, 2012. [cited by applicant]
Kuncheva LI. Combining pattern classifiers: methods and algorithms. John Wiley & Sons, 2014. [cited by applicant]
Kurakin, Adversarial Examples in the Physical World, arXiv preprint arXiv: 1607.02533v3, Nov. 4, 2016, 13 pages. [cited by applicant]
LeCun et al, Handwritten Digit Recognition with a Back-Propagation Network, Advances in Neural Information Processing Systems, pp. 396-404, 1990. [cited by applicant]
Merler, et al, Recognizing groceries in situ using in vitro training data, 2007 IEEE Conference on Computer Vision and Pattern Recognition Jun. 17, 2007 (pp. 1-8). [cited by applicant]
Mircic, et al, Fine-grained product class recognition for assisted shopping, Proceedings of the IEEE International Conference on Computer Vision Workshops 2015 (pp. 154-162). [cited by applicant]
Oquab, et al, Is Object Localization for Free?—Weakly-Supervised Learning with Convolutional Neural Networks, IEEE Conference on Computer Vision and Pattern Recognition, pp. 685-694, 2015. [cited by applicant]
Pathak, et al, Context Encoders: Feature Learning by Inpainting, IEEE Conference on Computer Vision and Pattern Recognition, pp. 2536-2544, 2016. [cited by applicant]
Polikar, Ensemble learning. Scholarpedia, 4(1):2776, revision #186077, 2009. [cited by applicant]
Ponti, Combining classifiers: from the creation of ensembles to the decision fusion, 24th IEEE SIBGRAPI Conference on Graphics, Patterns, and Images Tutorials, pp. 1-10, 2011. [cited by applicant]
Rumelhart et al, Learning Representations by Back-Propagating Errors, Nature, 323, 533-536, 1986. [cited by applicant]
Selvaraju et al, Visual Explanations from Deep Networks via Gradient-based Localization, arXiv preprint, arXiv:1610.02391v3, Mar. 21, 2017, 24 pages. [cited by applicant]
Sermanet, et al, Overfeat: Integrated recognition, localization and detection using convolutional networks, arXiv preprint arXiv: 1312.6229, 2013, 16 pages. [cited by applicant]
Shrivastava et al, Training Region-based Object Detectors with Online Hard Example Mining, arXiv preprint, arXiv:1604.03540v1, Apr. 12, 2016, 9 pages. [cited by applicant]
Shrivastava, et al, Training region-based object detectors with online hard example mining, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 761-769, 2016. [cited by applicant]
Simonyan, et al, Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps, arXiv preprint, arXiv:1312.6034v2, 2014, 8 pages. [cited by applicant]
Softmax function, archived Wikipedia article, Sep. 28, 2016. [cited by applicant]
Solgi and Weng, WWN-8: incremental online stereo with shape-from-x using life-long big data from multiple modalities, Procedia Computer Science 53 (2015): 316-326. [cited by applicant]
Song, et al, On Learning to Localize Objects with Minimal Supervision, International Conference on Machine Learning, 2014, 9 pages. [cited by applicant]
Springenber, et al, Striving for Simplicity: the All Convolutional Net, arXiv preprint, arXiv:1412.6806v3, 2015, 14 pages. [cited by applicant]
Sung et al, Example-based learning for view-based human face detection, IEEE Transactions on Pattern Analysis and Machine Intelligence, 20(1):39-51, 1998. [cited by applicant]
Szegedy, et al, Going Deeper with Convolutions. IEEE Conference on Computer Vision and Pattern Recognition, pp. 1-9, 2015. [cited by applicant]
Szegedy, et al, Intriguing Properties of Neural Networks, arXiv preprint, arXiv:1312.6199v4, 2014, 10 pages. [cited by applicant]
Tonioni, et al, Product recognition in store shelves as a sub-graph isomorphism problem, International Conference on Image Analysis and Processing Sep. 11, 2017 (pp. 682-693). [cited by applicant]
Tsai, et al, Mobile product recognition, Proceedings of the 18th ACM international conference on Multimedia Oct. 25, 2010 (pp. 1587-1590). [cited by applicant]
Wang, et al, A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection, arXiv preprint, arXiv:1704.03414v1, Apr. 11, 2017, 10 pages. [cited by applicant]
Wang, Wu, and Weng, Brain-like learning directly from dynamic cluttered natural video, Proc. International Conference on Brain-Mind, pages, pp. 1-8. 2012. [cited by applicant]
Wei, et al, A Simple to Complex Framework for Weakly-Supervised Semantic Segmentation, arXiv preprint 1509:03150v2, 2016, 8 pages. [cited by applicant]
Wei, et al, Object Region Mining with Adversarial Erasing: a Simple Classification to Semantic Segmentation Approach, arXiv preprint, arXiv:1703.08448v2, Mar. 27, 2017, 10 pages. [cited by applicant]
Weng, Computational introduction to the biological brain-mind, Natural Intelligence, vol. 1, No. 3, 2012. [cited by applicant]
Weng, et al, Brain-inspired concept networks: Learning concepts from cluttered scenes, IEEE Intelligent Systems, vol. 29, No. 6, Oct. 7, 2014, pp. 14-22. [cited by applicant]
Weng, et al, Brain-like emergent spatial processing, IEEE Transactions on Autonomous Mental Development 4, No. 2 (2011): 161-185. [cited by applicant]
Weng, et al, Online learning for attention, recognition, and tracking by a single developmental framework., IEEE Computer Society Conference on Computer Vision and Pattern Recognition-Workshops, pp. 7-14, 2010. [cited by applicant]
Weng, et al, Skull-closed autonomous development: Wwn-6 using natural video, International Joint Conference on Neural Networks (IJCNN), pp. 1-8. IEEE, 2012. [cited by applicant]
Wu, Guo, and Weng, Skull-closed autonomous development: WWN-7 dealing with scales, Proc. International Conference on Brain-Mind. East Lansing, Michigan: BMI Press, pp. 1-8. 2013. [cited by applicant]
Zeiler et al, Visualizing and Understanding Convolutional Networks, European Conference on Computer Vision, 2014, 16 pages. [cited by applicant]
Zhou, et al, Learning Deep Features for Discriminative Localization, IEEE Conference on Computer Vision and Pattern Recognition, pp. 2921-2929, Jun. 26, 2016. [cited by applicant]