System and method for pose tolerant feature extraction using generated pose-altered images
Disclosed herein is a system and method for augmenting data by generating a plurality of pose-altered images of an item from one or more 2D images of the item and using the augmented data to train a train a feature extractor. In other aspects of the invention, the trained feature extractor is used to enroll features extracted from images of new products in a library database of known products or to identify images of unknown products by matching features of an image of the unknown product with features stored in the library database.
1. A method comprising:
generating a plurality of synthetic pose-altered images of a plurality of items based on one or more 2D images of each of the plurality of items; and
training a feature extractor to extract features of the synthetic pose-altered images;
wherein the synthetic pose-altered images include views of the items which are modified from the one or more 2D images of the items, the modifying including presenting the items at arbitrary angles, rotating the items, translating the items, crumpling the items, modifying the colors of the items, over-exposing the items and under-exposing the items.
2. The method of claim 1 further comprising:
generating a plurality of synthetic pose-altered images of a new item based on one or more 2D images of the new item;
inputting the synthetic pose-altered images of the new item to the feature extractor to obtain features of the synthetic pose-altered images; and
enrolling the extracted features in a library database.
3. The method of claim 2 wherein the one or more 2D images of the new item are also input to the feature extractor.
4. The method of claim 1 further comprising:
identifying an unknown item in a testing image by:
exposing the testing image of the unknown item to the feature extractor to obtain features of the testing image; and
matching the features extracted from the testing image with features enrolled in the library database to determine a match to identify the unknown item.
5. The method of claim 4 wherein the matching is performed by a trained classifier.
6. The method of claim 1 wherein the step of generating a plurality of synthetic pose-altered images comprises:
generating a 3D model of the item from the one or more 2D images of the item;
generating images showing different viewpoints of the item by rotating the 3D model along one or more axes.
7. The method of claim 1 wherein the step of generating a plurality of synthetic pose-altered images comprises:
fixing the one or more 2D images of the item to a plane with a known depth; and
generating images showing different viewpoints of the item by rotating the 2D images along one or more of the axes of the plane.
8. The method of claim 1 wherein the step of generating a plurality of synthetic pose-altered images comprises:
using planar homography by fitting corner endpoints of the one or more 2D images to different configurations to simulate different views of the item.
9. The method of claim 1 wherein the step of generating a plurality of synthetic pose-altered images comprises:
training a machine learning model to generate the synthetic pose-altered images; and
inputting the one or more 2D images of the item to the machine learning model.
10. The method of claim 1 further comprising:
pose correcting the one or more 2D images of the item to eliminate portions of the one or more 2D containing views of sides of the item other than the frontal side.
11. A system comprising:
a processor; and
a memory storing software that, when executed by the processor, causes the system to:
generate a plurality of synthetic pose-altered images of a plurality of items based on one or more 2D images of each of the plurality of items; and
train a feature extractor to extract features of the pose-altered images;
wherein the synthetic pose-altered images include views of the items which are modified from the one or more 2D images of the items, the modifying including presenting the item at arbitrary angles, rotating the item, translating the item, crumpling the item, modifying the colors of the item, over-exposing the item and under-exposing the item.
12. The system of claim 11 the software further causing the system to:
generate a plurality of synthetic pose-altered images of a new item based on one or more 2D images of the new item;
input the one or more synthetic pose-altered images of the new item to the feature extractor to obtain features of the synthetic pose-altered images; and
enroll the extracted features in a library database.
13. The system of claim 12 , the software further causing the system to:
input the one or more 2D images of the new item into the feature extractor.
14. The system of claim 1 , the software identifying an unknown item in a testing image by causing the system to:
expose the testing image of the unknown item to the feature extractor to obtain features of the testing image; and
match the features extracted from the testing image with features enrolled in the library database to determine a match to identify the unknown item.
15. The system of claim 14 wherein the matching is performed by a trained classifier.
16. The system of claim 11 wherein the software causes the generation of the plurality of synthetic pose-altered images by causing the system to:
generate a 3D model of the item from the one or more 2D images of the item; and
generate images showing different viewpoints of the item by rotating the 3D model along one or more of the axes.
17. The system of claim 11 wherein the software causes the generation of the plurality of synthetic pose-altered images by causing the system to:
fix the one or more 2D images of the item to a plane with a known depth; and
generate images showing different viewpoints of the item by rotating the 2D images along one or more of the axes of the plane.
18. The system of claim 11 wherein the software causes the generation of the plurality of synthetic pose-altered images by causing the system to:
use planar homography by fitting corner endpoints of the one or more 2D images to different configurations to simulate different views of the item.
19. The system of claim 11 wherein the software causes the generation of the plurality of synthetic pose-altered images by causing the system to:
train a machine learning model to generate the synthetic pose-altered images; and
input the one or more 2D images of the item to the machine learning model.
20. The system of claim 11 wherein the software further causes the system to:
pose correct the one or more 2D images of the item to eliminate portions of the one or more 2D images containing views of sides of the item other than the frontal side.