IP Library › Granted Patent US 10,977,827
Granted Patent B2
US 10,977,827 · App. 15/937,424 · Granted Apr 13, 2021

Multiview estimation of 6D pose

Inventors: J. William Mauchly (Berwyn, PA); Joseph T. Friel (Ardmore, PA)
G06T7/75G06N3/08G06T15/20G06T2207/20081G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,977,827
App. No.
15/937,424
Granted
Apr 13, 2021
Kind
B2
Abstract

This disclosure describes a method and system to perform object detection and 6D pose estimation. The system comprises a database of 3D models, a CNN-based object detector, multiview pose verification, and a hard example generator for CNN training. The accuracy of that detection and estimation can be iteratively improved by retraining the CNN with increasingly hard ground truth examples. The additional images are detected and annotated by an automatic process of pose estimation and verification.

Claims (12)

1. A system for automatically training a CNN (Convolutional Neural Network), the system comprising:

a computer processor and one or more non-transitory computer readable storage media configured to execute the steps of:

initializing a set of training data with a plurality of annotated images of an object;

initializing a CNN;

proceeding, starting with the initial set of training data to iteratively enlarge the set of training data by repeating the steps of:

training the CNN with the set of training data;

using the CNN to process a first image and make a first estimate of the position of the object in physical space;

using the CNN to process a second image to make a second estimate of the position of the object in physical space;

computing, by using geometric constraints between the first estimate and the second estimate, a more accurate position of the object in three translational dimensions;

computing, by minimizing measurements of feature distance between features of the first and second images and features of images of the object in known rotational positions, a more accurate position of the object in three rotational dimensions;

back-projecting the object into a third image and annotating the third image with the more accurate position in at least two translational dimensions and with the more accurate position in at least two rotational dimensions; and

adding the annotated third image to the set of training data.

Continuity (1)
Related Publication 20190304134A1 · Oct 3, 2019
Cited By (2)
US 12,277,745 US 12,555,342