IP Library Granted Patent US 11,182,640
Granted Patent B2
US 11,182,640 · App. 16/530,778 · Granted Nov 23, 2021

Analyzing content of digital images

Inventors: Huguens Jean (Woodstock, MD); Yoriyasu Yano (Oakland, CA); Hui Peng Hu (Berkeley, CA); Kuang Chen (Oakland, CA)
Assignee: DST Technologies, Inc.
G06K9/4671G06F16/5838G06F16/5846G06F16/5854G06K9/00483G06K9/4642G06K9/4676
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,182,640
App. No.
16/530,778
Granted
Nov 23, 2021
Kind
B2
Abstract

Methods, apparatuses, and embodiments related to analyzing the content of digital images. A computer extracts multiple sets of visual features, which can be keypoints, based on an image of a selected object. Each of the multiple sets of visual features is extracted by a different visual feature extractor. The computer further extracts a visual word count vector based on the image of the selected object. An image query is executed based on the extracted visual features and the extracted visual word count vector to identify one or more candidate template objects of which the selected object may be an instance. When multiple candidate template objects are identified, a matching algorithm compares the selected object with the candidate template objects to determine a particular candidate template of which the selected object is an instance.

Claims (45)

1. A method comprising:

assigning a first template prediction based on a multi-dimensional descriptor vector, the multi-dimensional descriptor vector being associated with a keypoint of a first plurality of keypoints identified based on a digital image of a physical object and by use of a computer vision algorithm;

assigning a second template predictions based on a multi-dimensional bag of visual words (BoVW) vector, the BoVW vector is based on a keypoint of a second plurality of keypoints identified based on the digital image by use of a second computer vision algorithm;

reconciling each respective template prediction wherein the template predictions are merged where each respective template prediction is in agreement and eliminating each fewest vote receiving prediction where each respective template prediction is not in agreement; and

in response to said reconciling, classifying the physical object based on a remaining template prediction.

2. The method of claim 1 , wherein the assigning the first template prediction includes:

generating, by a processor, the first plurality of keypoints,

wherein the keypoint of the first plurality of keypoints includes a multi-dimensional descriptor vector that represents an image gradient at a location of the digital image associated with the keypoint of the first plurality of keypoints;

comparing, by the processor, the multi-dimensional descriptor vector that represents the image gradient to a nearest neighbor multi-dimensional descriptor vector found in a set of previously classified example digital images; and

based on the comparison, assigning, by the processor, a template prediction to the multi-dimensional descriptor vector that represents the image gradient.

3. The method of claim 1 , further comprising:

assigning, a third template prediction based on a keypoint of a third plurality of keypoints identified based on the digital image by use of a third computer vision algorithm, wherein the third computer vision algorithm is different from the first and the second computer vision algorithms.

4. The method of claim 3 , wherein the first and the second computer vision algorithms are a same computer vision algorithm.

5. The method of claim 1 , wherein the assigning the second template prediction includes:

partitioning, by a processor, an instance of the digital image into a plurality of partitions;

generating, by the processor, the second plurality of keypoints based on a partition of the plurality of partitions,

wherein the keypoint of the second plurality of keypoints includes a multi-dimensional descriptor vector that represents an image gradient at a location of the digital image associated with the keypoint of the second plurality of keypoints;

comparing, by the processor, the multi-dimensional descriptor vector that represents the image gradient to a plurality of clusters of multi-dimensional descriptor vectors associated with keypoints in previously classified example digital images;

mapping, by the processor, the multi-dimensional descriptor vector that represents the image gradient to a nearest cluster of the plurality of clusters, based on the comparison;

generating, by the processor, a BoVW vector that represents a histogram of most frequently mapped clusters; and

assigning a template prediction to the partition of the plurality of partitions based on a most frequently mapped cluster of the BoVW vector.

6. The method of claim 1 , further comprising:

receiving, via an image capture device, the digital image of the physical object.

7. The method of claim 1 , wherein the keypoints identified based on the digital image are associated with scale and rotation invariant features of the physical object depicted in the digital image.

8. The method of claim 1 , wherein the first and the second computer vision algorithms are each any one of: Scale Invariant Feature Transform (SIFT), Speeded Up Robust Features (SURF), and Oriented Features from Accelerated Segment Test and Rotated Binary Robust Independent Elementary Features (ORB).

9. The method of claim 1 , wherein the physical object is a paper form document.

10. The method of claim 1 , wherein the plurality of identified keypoints includes between 64 and 512 keypoints.

11. The method of claim 1 , wherein the multi-dimensional descriptor vector is a 128-dimension descriptor vector that represents an image gradient at a location of the digital image associated with the keypoint of the first plurality of keypoints.

12. A system comprising:

a processor;

a communication interface, coupled to the processor, through which to communicate over a network with remote devices; and

a memory coupled to the processor, the memory storing instructions which when executed by the processor cause the system to perform operations including:

assign a first template prediction based on a multi-dimensional descriptor vector, the multi-dimensional descriptor vector being associated with a keypoint of a first plurality of keypoints identified based on a digital image of a physical object and by use of a computer vision algorithm;

assign a second template predictions based on a multi-dimensional bag of visual words (BoVW) vector, the BoVW vector is based on a keypoint of a second plurality of keypoints identified based on the digital image by use of a second computer vision algorithm;

reconcile each respective template prediction wherein the template predictions are merged where each respective template prediction is in agreement and eliminating each fewest vote receiving prediction where each respective template prediction is not in agreement; and

in response to said reconciliation, classify the physical object based on a remaining template prediction.

13. The system of claim 12 , wherein the first plurality of keypoints and the second plurality of keypoints are a same plurality of keypoints, and wherein the first computer vision algorithm and the second computer vision algorithm are a same computer vision algorithm.

14. The system of claim 12 , wherein the operations further include:

extract a first plurality of visual features of the digital image;

extract a second plurality of visual features of the digital image;

identify a plurality of clusters of visual features based on any of the first plurality or the second plurality of visual features; and

calculate the multi-dimensional BoVW vector based on the plurality of clusters of visual features.

15. The system of claim 12 , wherein the operations further include:

receive the digital image from a mobile device after the mobile device captured the digital image of the object.

16. The system of claim 12 , wherein the digital image is a frame of a video that includes the object.

Assignments (2)
MERGER AND CHANGE OF NAME Recorded Jan 15, 2021
From: CAPTRICITY, INC.; DST TECHNOLOGIES, INC.
To: DST TECHNOLOGIES, INC.
Reel/Frame 054933/0472 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2019
From: JEAN, HUGUENS; YANO, YORIYASU; HU, HUI PENG; CHEN, KUANG
To: CAPTRICITY, INC.
Reel/Frame 049947/0087 →
Continuity (4)
Continuation 15483291 · Apr 10, 2017
Division 14713863 · May 15, 2015
Provisional Application 62085237 · Nov 26, 2014
Related Publication 20210049401A1 · Feb 18, 2021