IP Library › Granted Patent US 9,342,758
Granted Patent B2
US 9,342,758 · App. 13/711,500 · Granted May 17, 2016

Image classification based on visual words

Inventor: Hui Xue (Hangzhou, CN)
Assignee: Alibaba Group Holding Limited
G06K9/6256G06K9/00496G06K9/00536G06K9/00624G06K9/4642G06K9/4676G06K9/62G06K9/6201G06K9/6212G06K9/6215G06K9/6226G06K9/6267
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,342,758
App. No.
13/711,500
Granted
May 17, 2016
Kind
B2
Abstract

The present disclosure introduces a method and an apparatus for classifying images. Classification image features of an image for classification are extracted. Based on a similarity relationship between each classification image feature and one or more visual words in a pre-generated visual dictionary, each classification image feature is quantified by multiple visual words in the visual dictionary and a similarity coefficient between each classification image feature and each of the visual words is determined. Based on the similarity coefficient of each visual word that corresponds to different classification image features, a weight of each visual word is determined to establish a classification visual word histogram of the image for classification. The classification visual word histogram is input into an image classifier that is trained by sample visual word histograms arising from multiple sample images. An output result is used to determine a classification of the image for classification.

Claims (59)

1. A method performed by one or more processors configured with computer-executable instructions, the method comprising:

extracting one or more classification image features from an image for classification;

quantifying, based on a similarity relationship between each classification image feature and each visual word in a pre-generated visual dictionary, each classification image feature by multiple visual words in the visual dictionary and determining a similarity coefficient between each classification image feature and each visual word after the quantifying;

determining, based on one or more similarity coefficients of each visual word corresponding to different classification image features, a weight of each visual word to establish a classification visual word histogram;

inputting the classification visual word histogram into an image classifier; and

using an output of the inputting to determine a classification of the image for classification,

wherein the quantifying each classification image feature by multiple visual words in the visual dictionary and determining the similarity coefficient between each classification image feature and each visual word after the quantifying includes:

calculating, based on the similarity relationship between each classification image feature and the visual words in the pre-generated visual dictionary, a Euclidean distance between each classification image feature and each visual word,

determining a smallest Euclidean distance among calculated Euclidean distances,

determining, with respect to each classification image feature, one or more visual words of which Euclidean distances are within a preset times range of the smallest Euclidean distance as the visual words for quantification of the respective classification image feature, and

calculating, based on the Euclidean distance between the respective classification image feature and each of the visual words for quantification, the one or more similarity coefficients between the respective classification image feature and the visual words, the one or more similarity coefficients being calculated respectively as a percentage relationship of each of the one or more visual words for quantification,

wherein the percentage relationship of a particular visual word for quantification is calculated by dividing the respective Euclidean distance of the particular visual word for quantification by a sum total of the Euclidean distances of the one or more visual words for quantification.

2. The method as recited in claim 1 , wherein the image classifier is generated through training by sample visual word histograms from multiple sample images.

3. The method as recited in claim 2 , further comprising

comparing, after inputting the classification visual word histogram into an image classifier, the inputted classification visual word histogram to a pre-generated classification visual word histogram in the image classifier to determine a classification of the image for classification.

4. The method as recited in claim 1 , wherein the determining the weight of each visual word to establish the classification visual word histogram comprises:

adding up the one or more coefficients of a respective visual word that correspond to different classification image features to calculate the weight of the respective visual word; and

establishing the classification visual word histogram.

5. The method as recited in claim 1 , wherein the determining the weight of each visual word to establish the classification visual word histogram comprises:

dividing the image for classification into multiple layer images based on a pyramid image algorithm; and

dividing each layer image to form multiple child images.

6. The method as recited in claim 1 , wherein the visual dictionary is generated through clustering of multiple sample image features extracted from multiple sample images.

7. The method as recited in claim 1 , wherein the classification visual word histogram increases space information of the image for classification.

8. A system comprising:

one or more processors; and

memory containing computer readable media storing one or more modules including instructions, which when executed by the processors, cause the modules to perform:

extracting one or more classification image features from an image for classification,

quantifying, based on a similarity relationship between each classification image feature and each visual word in a pre-generated visual dictionary, each classification image feature by multiple visual words in the visual dictionary and determining a similarity coefficient between each classification image feature and each visual word after the quantifying,

determining, based on one or more similarity coefficients of each visual word corresponding to different classification image features, a weight of each visual word to establish a classification visual word histogram, and

inputting the classification visual word histogram into an image classifier and using an output to determine a classification of the image for classification,

wherein the quantifying includes:

calculating, based on the similarity relationship between each classification image feature and the visual words in the pre-generated visual dictionary, a Euclidean distance between each classification image feature and each visual word,

determining a smallest Euclidean distance among calculated Euclidean distances and, with respect to each classification image feature, determining one or more visual words of which Euclidean distances are within a preset times range of the smallest Euclidean distance as the visual words for quantification of the respective classification image feature, and

calculating, based on the Euclidean distance between the respective classification image feature and each of the visual words for quantification, the one or more similarity coefficients between the respective classification image feature and the visual words, the one or more similarity coefficients being calculated respectively as a percentage relationship of each of the one or more visual words for quantification, and

wherein the percentage relationship of a particular visual word for quantification is calculated by dividing the respective Euclidean distance of the particular visual word for quantification by a sum total of the Euclidean distances of the one or more visual words for quantification.

9. The apparatus as recited in claim 8 , wherein the image classifier is generated through training by sample visual word histograms from multiple sample images.

10. The apparatus as recited in claim 9 , wherein the image classifier compares the inputted classification visual word histogram with a pre-generated classification visual word histogram to determine a classification of the image for classification.

11. The apparatus as recited in claim 8 , wherein the determining a weight of each visual word includes:

dividing the image for classification into multiple child images based on a pyramid image algorithm,

determining classification image features of each child image,

adding the coefficients of a respective visual word that correspond to each classification image features in a respective child image,

calculating the weight of the respective visual word corresponding to the respective child image to establish a child classification visual word histogram of the respective child image, and

combining each child classification visual word histogram of each child image to establish the classification visual word histogram.

12. The apparatus as recited in claim 8 , wherein the classification visual word histogram increases space information of the image for classification.

13. One or more computer storage media including processor-executable instructions that, when executed by one or more processors, direct the one or more processors to perform a method comprising:

extracting one or more classification image features from an image for classification;

quantifying, based on a similarity relationship between each classification image feature and each visual word in a pre-generated visual dictionary, each classification image feature by multiple visual words in the visual dictionary and determining a similarity coefficient between each classification image feature and each visual word after the quantifying;

determining, based on one or more similarity coefficients of each visual word corresponding to different classification image features, a weight of each visual word to establish a classification visual word histogram;

inputting the classification visual word histogram into an image classifier that is generated through training by sample visual word histograms from multiple sample images; and

using an output of the inputting to determine a classification of the image for classification,

wherein the quantifying each classification image feature by multiple visual words in the visual dictionary and determining the similarity coefficient between each classification image feature and each visual word after the quantifying includes:

calculating, based on the similarity relationship between each classification image feature and the visual words in the pre-generated visual dictionary, a Euclidean distance between each classification image feature and each visual word,

determining a smallest Euclidean distance among calculated Euclidean distances,

determining, with respect to each classification image feature, one or more visual words of which Euclidean distances are within a preset times range of the smallest Euclidean distance as the visual words for quantification of the respective classification image feature, and

calculating, based on the Euclidean distance between the respective classification image feature and each of the visual words for quantification, the one or more similarity coefficients between the respective classification image feature and the visual words, the one or more similarity coefficients being calculated respectively as a percentage relationship of each of the one or more visual words for quantification,

wherein the percentage relationship of a particular visual word for quantification is calculated by dividing the respective Euclidean distance of the particular visual word for quantification by a sum total of the Euclidean distances of the one or more visual words for quantification.

14. The one or more computer storage media as recited in claim 13 , wherein the method further comprises

comparing, after inputting the classification visual word histogram into an image classifier, the inputted classification visual word histogram to a pre-generated classification visual word histogram in the image classifier to determine a classification of the image for classification.

15. The one or more computer storage media as recited in claim 13 , wherein the classification visual word histogram increases space information of the image for classification.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2013
From: XUE, HUI
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 029726/0482 →
Priority Claims (1)
CN 2011 1 0412537 · Dec 12, 2011 · national
Continuity (1)
Related Publication 20130148881A1 · Jun 13, 2013