IP Library Granted Patent US 11,393,082
Granted Patent B2
US 11,393,082 · App. 16/521,741 · Granted Jul 19, 2022

System and method for produce detection and classification

Inventors: Issac Mathew (Bangalore, IN); Pushkar Pushp (Bangalore, IN); Viraj Patel (Bangalore, IN); Emily Xavier (Bangalore, IN); Gaurav Savlani (Bangalore, IN); Venkataraja Nellore (Bangalore, IN); Rahul Agarwal (Bangalore, IN); Girish Thiruvenkadam (Bangalore, IN); Shivani Naik (Bangalore, IN)
Assignee: Walmart Apollo, LLC
G06T7/0004G06K9/6267G06N3/0454G06N3/08G06N20/20G06T2207/20081G06T2207/20084G06T2207/30128G06V20/68
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,393,082
App. No.
16/521,741
Granted
Jul 19, 2022
Kind
B2
Abstract

Systems, methods, and computer-readable storage media for object detection and classification, and particularly produce detection and classification. A system configured according to this disclosure can receiving, at a processor, an image of an item. The system can then perform, across multiple pre-trained neural networks, feature detection on the image, resulting in feature maps of the image. These feature maps can be concatenated and combined, then input into an additional neural network for feature detection on the combined feature map, resulting in tiered neural network features. The system then classifies, via the processor, the item based on the tiered neural network features.

Claims (45)

1. A method comprising:

receiving, at a processor, an image of an item;

performing, via the processor using a first pre-trained neural network, feature detection on the image, resulting in a first feature map of the image;

concatenating the first feature map, resulting in a first concatenated feature map;

performing, via the processor using a second pre-trained neural network, feature detection on the image, resulting in a second feature map of the image;

concatenating the second feature map, resulting in a second concatenated feature map;

combining the first concatenated feature map and the second concatenated feature map, resulting in a combined feature map;

performing, via the processor using a third pre-trained neural network, feature detection on the combined feature map, resulting in tiered neural network features; and

classifying, via the processor, the item based on the tiered neural network features, the classifying including implementing a set of pre-trained neural networks, the set of pre-trained neural networks having been produced based on the tiered neural network features, the classification being a combination of results of the set of pre-trained neural networks, and a result of each pre-trained neural network of the set of pre-trained neural networks being weighted based on a corresponding accuracy.

2. The method of claim 1 , wherein the item is produce.

3. The method of claim 2 , wherein the feature detection identifies defects within the produce.

4. The method of claim 1 , wherein at least one of the first pre-trained neural network, the second pre-trained neural network, and the third pre-trained neural network is a Faster Regional Convolutional Neural Network.

5. The method of claim 4 , wherein the Faster Regional Convolutional Neural Network identifies a top-left coordinate of a rectangular region for each item within the image and a bottom-right coordinate of the rectangular region.

6. The method of claim 1 , wherein the third pre-trained neural network uses distinct neural links than the neural links of the first pre-trained neural network and the second pre-trained neural network.

7. The method of claim 1 , wherein the processor is a Graphical Processing Unit.

8. A system, comprising:

a processor; and a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

receiving an image of an item;

performing, using a first pre-trained neural network, feature detection on the image, resulting in a first feature map of the image;

concatenating the first feature map, resulting in a first concatenated feature map;

performing, using a second pre-trained neural network, feature detection on the image, resulting in a second feature map of the image;

concatenating the second feature map, resulting in a second concatenated feature map;

combining the first concatenated feature map and the second concatenated feature map, resulting in a combined feature map;

performing, using a third pre-trained neural network, feature detection on the combined feature map, resulting in tiered neural network features; and

classifying the item based on the tiered neural network features, the classifying including implementing a set of pre-trained neural networks, the set of pre-trained neural networks having been produced based on the tiered neural network features, the classification being a combination of results of the set of pre-trained neural networks, and a result of each pre-trained neural network of the set of pre-trained neural networks being weighted based on a corresponding accuracy.

9. The system of claim 8 , wherein the item is produce.

10. The system of claim 9 , wherein the feature detection identifies defects within the produce.

11. The system of claim 8 , wherein at least one of the first pre-trained neural network, the second pre-trained neural network, and the third pre-trained neural network is a Faster Regional Convolutional Neural Network.

12. The system of claim 11 , wherein the Faster Regional Convolutional Neural Network identifies a top-left coordinate of a rectangular region for each item within the image and a bottom-right coordinate of the rectangular region.

13. The system of claim 8 , wherein the third pre-trained neural network uses distinct neural links than the neural links of the first pre-trained neural network and the second pre-trained neural network.

14. The system of claim 8 , wherein the processor is a Graphical Processing Unit.

15. A non-transitory computer-readable storage medium having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

receiving an image of an item;

performing, using a first pre-trained neural network, feature detection on the image, resulting in a first feature map of the image;

concatenating the first feature map, resulting in a first concatenated feature map;

performing, using a second pre-trained neural network, feature detection on the image, resulting in a second feature map of the image;

concatenating the second feature map, resulting in a second concatenated feature map;

combining the first concatenated feature map and the second concatenated feature map, resulting in a combined feature map;

performing, using a third pre-trained neural network, feature detection on the combined feature map, resulting in tiered neural network features; and

classifying the item based on the tiered neural network features, the classifying including implementing a set of pre-trained neural networks, the set of pre-trained neural networks having been produced based on the tiered neural network features, the classification being a combination of results of the set of pre-trained neural networks, and a result of each pre-trained neural network of the set of pre-trained neural networks being weighted based on a corresponding accuracy.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the item is produce.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the feature detection identifies defects within the produce.

18. The non-transitory computer-readable storage medium of claim 15 , wherein at least one of the first pre-trained neural network, the second pre-trained neural network, and the third pre-trained neural network is a Faster Regional Convolutional Neural Network.

19. The non-transitory computer-readable storage medium of claim 18 , wherein the Faster Regional Convolutional Neural Network identifies a top-left coordinate of a rectangular region for each item within the image and a bottom-right coordinate of the rectangular region.

20. The non-transitory computer-readable storage medium of claim 15 , wherein the third pre-trained neural network uses distinct neural links than the neural links of the first pre-trained neural network and the second pre-trained neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2021
From: MATHEW, ISSAC; PUSHP, PUSHKAR; PATEL, VIRAJ; XAVIER, EMILY; SAVLANI, GAURAV; NELLORE, VENKATARAJA; AGARWAL, RAHUL; THIRUVENKADAM, GIRISH; NAIK, SHIVANI
To: WALMART APOLLO, LLC
Reel/Frame 058312/0439 →
Priority Claims (1)
IN 201811028178 · Jul 26, 2018 · national
Continuity (2)
Provisional Application 62773756 · Nov 30, 2018
Related Publication 20200034962A1 · Jan 30, 2020
Cited By (1)
US 12,450,564