Multiview estimation of 6D pose
View Patent ↗This disclosure describes a method and system to perform object detection and 6D pose estimation. The system comprises a database of 3D models, a CNN-based object detector, multiview pose verification, and a hard example generator for CNN training. The accuracy of that detection and estimation can be iteratively improved by retraining the CNN with increasingly hard ground truth examples. The additional images are detected and annotated by an automatic process of pose estimation and verification.
1. A system for automatically training a CNN (Convolutional Neural Network), the system comprising:
a computer processor and one or more non-transitory computer readable storage media configured to execute the steps of:
initializing a set of training data with a plurality of annotated images of an object;
initializing a CNN;
proceeding, starting with the initial set of training data to iteratively enlarge the set of training data by repeating the steps of:
training the CNN with the set of training data;
using the CNN to process a first image and make a first estimate of the position of the object in physical space;
using the CNN to process a second image to make a second estimate of the position of the object in physical space;
computing, by using geometric constraints between the first estimate and the second estimate, a more accurate position of the object in three translational dimensions;
computing, by minimizing measurements of feature distance between features of the first and second images and features of images of the object in known rotational positions, a more accurate position of the object in three rotational dimensions;
back-projecting the object into a third image and annotating the third image with the more accurate position in at least two translational dimensions and with the more accurate position in at least two rotational dimensions; and
adding the annotated third image to the set of training data.