IP Library Granted Patent US 8,224,849
Granted Patent B2
US 8,224,849 · App. 13/092,083 · Granted Jul 17, 2012

Object similarity search in high-dimensional vector spaces

Assignee: Microsoft Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,224,849
App. No.
13/092,083
Granted
Jul 17, 2012
Kind
B2
Abstract

An object search system generates a hierarchical clustering of objects of a collection based on similarity of the objects. The object search system generates a separate hierarchical clustering of objects for multiple features of the objects. To identify objects similar to a target object, the object search system first generates a feature vector for the target object. For each feature of the feature vector, the object search system uses the hierarchical clustering of objects to identify the cluster of objects that is most “feature similar” to that feature of the target object. The object search system indicates the similarity of each candidate object based on the features for which the candidate object is similar.

Claims (28)

1. A computer-readable storage device storing computer-executable instructions for controlling a computing device to identify images of a collection that are similar to a target image, by a method comprising:

for each of a plurality of features, providing a cluster index data structure for the collection of images, the cluster index data structure defining clusters of images that are feature similar based on the values of that feature, such that for each feature, the images in the collection are clustered differently based on the values for that feature;

for each of the plurality of features, identifying, from the cluster index data structure for that feature, candidate images that are feature similar to the target image based on that feature, the cluster index data structure defining, for each the plurality of features of images, clusters of images that are feature similar based on that feature; and

for each of the candidate images, indicating similarity of that candidate image to the target image based on the features for which that candidate image is feature similar to the target image.

2. The computer-readable storage device of claim 1 wherein the cluster index data structure for a feature stores, for each image in the collection, a hash code representing the feature for that image, and wherein images are clustered that are feature similar using the hash code to represent the feature of an image.

3. The computer-readable storage device of claim 1 wherein the target image is provided in a search request and including ranking the candidate images based on the indicated similarity of the images.

4. The computer-readable storage device of claim 1 wherein each feature representing a characteristic of the image.

5. The computer-readable storage device of claim 1 wherein the candidate image is feature similar to the target object based on the values of the features.

6. The computer-readable storage device of claim 1 wherein the candidate image is feature similar to the target object based on the number of identified clusters containing the candidate object.

7. A method performed by a computing device with a processor and a memory for identifying objects of a collection that are similar to a target object, the method comprising:

for each of a plurality of features, identifying, from a cluster index data structure for that feature, candidate objects that are feature similar to the target object based on that feature, the cluster index data structure defining, for each of the plurality of features of objects, clusters of objects that are feature similar based on that feature, wherein each cluster index data structure provides a separate clustering of the objects in the collection based on a different feature; and

for candidate objects, indicating by the processor similarity of the candidate object to the target object based on the features for which the candidate object is feature similar to the target object.

8. The method of claim 7 wherein the cluster index data structure for a feature stores, for each object in the collection, a hash code representing the feature for that object, and wherein objects are clustered that are feature similar using the hash code to represent the feature of an object.

9. The method of claim 8 includes generating the cluster index data structure.

10. The method of claim 7 wherein the target object is provided in a search request and including ranking the candidate objects based on the indicated similarity of the objects.

11. The method of claim 7 wherein each feature represents a characteristic of the object.

12. The method of claim 7 including generating, for each feature, a cluster index data structure for the collection of objects, the cluster index data structure defining clusters of objects that are feature similar based on the values of feature, such that for each feature, the objects in the collection are clustered differently based on the values for that feature.

13. A computing device for identifying images of a collection that are similar to a target image, the computing device comprising:

a memory storing computer-executable instructions of:

a component that, for each of a plurality of features, identifies, from a cluster index data structure for that one feature, candidate images that are feature similar to the target image based on that one feature, wherein images are feature similar to the target image based on that one feature when the images have similar values for that one feature; and

a component that, for each of the candidate images, indicates similarity of that candidate image to the target image based on the features for which that candidate image is feature similar to the target image

wherein each of the plurality of cluster index data structures provides a mapping of values for one feature to clusters of images that have similar values for that one feature; and

a processor that executes the computer-executable instructions stored in the memory.

14. The computing device of claim 13 wherein the cluster index data structure for a feature stores, for each object in the collection, a hash code representing that feature for that object, and wherein objects are clustered that are feature similar using the hash code to represent the feature of an object.

15. The computing device of claim 14 wherein feature similarity for a feature is based on a Hamming distance between hash codes of that feature.

16. The computing device of claim 13 includes a component that generates the cluster index data structure for each of the plurality of features.

17. The computing device of claim 13 wherein the target object is provided in a search request and including a component that ranks the candidate objects based on the indicated similarity of the objects.

18. The computing device of claim 13 wherein each feature represents a characteristic of the object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034544/0001 →
Continuity (2)
Continuation 11737075 · Apr 18, 2007
Related Publication 20110194780A1 · Aug 11, 2011