Image search result summarization with informative priors
An informative priors image search result summarization system and method that summarizes image search results based on the image relevance (as determined by a search engine's initial ranking) and the image quality. Embodiments of the system and method cluster the image search results, rank images within each cluster based on a computed image score, and then select a summary image for the cluster. Each cluster is analyzed and an image in the cluster having the maximum image score is included in a selected summary collection. The image score is computed using the image relevance and the image quality, as well as a cluster coherence, a density, and a diversity. The selection of images from a collection of candidate images generates an image search result summarization, which is presented to a user. The summaries are presented to the user in a ranked order based on their image scores.
1. A method implemented on a computing device having a processor for summarizing image search results, comprising:
using the computing device having the processor to perform the following:
estimating an image relevance for each image in the image search results as a ranking of each image given by a search engine that provided the image search results in response to a query from a user to the search engine;
computing an image quality for each image based on one or more image quality measures;
clustering images in the image search results using a clustering technique that has as a first informative prior the image quality for each image and as a second informative prior the image relevance for each image to obtain a summary candidate collection containing image clusters and an exemplar image for each cluster;
selecting and ranking each image in the summary candidate collection to obtain an image search results summarization;
selecting a number representing a desired number of summaries; and
presenting the image search results summarization to a user based on whether a number of images contained in the image search results summarization is less than the number representing a desired number of summaries.
2. The method of claim 1 , further comprising:
computing an image score for each image in the summary candidate collection; and
ranking each image in the summary candidate collection based on its image score.
3. The method of claim 2 , further comprising:
computing a similarity between a selected image and another image in a image cluster;
using the similarity to compute a cluster coherence, density, and diversity for the selected image.
4. The method of claim 3 , further comprising computing an image score for the selected image using the cluster coherence, density, diversity, image quality, and image relevance.
5. The method of claim 4 , further comprising:
computing an image score for each image in a the image cluster; and
identifying an image in the image cluster having a maximum image score.
6. The method of claim 5 , further comprising removing the image in the image cluster having the maximum image score from the summary candidate collection:
adding the image in the image cluster having the maximum image score to the selected summaries collection; and
removing the image in the image cluster having the maximum image score from the summary candidate collection.
7. The method of claim 6 , further comprising presenting images in the selected summaries collection to the user in an order according to a ranking of the images based on the image score of each image.
8. The method of claim 1 , further comprising computing the image quality for each image using color entropy of the image as an image quality measure.
9. The method of claim 1 , further comprising clustering images in the image search results using an Affinity Propagation clustering technique that has as a first preference the image quality and as a second preference the image relevance for each image.
10. The method of claim 9 , further comprising computing the image quality for each image using at least one of the following as an image quality measure: (a) dynamic range, which denotes a luminance range of a scene being photographed and is a ratio between a maximum and a minimum measurable light intensities; (b) color entropy, which describes a colorlessness of image content; (c) brightness, which describes an amount of light in an image; (d) blur, which describes a sharpness of an image; and, (e) contrast, where good-quality images are generally under strong contrast between a subject and a background.
11. A method implemented on a computing device having a processor for performing image search results summarization on a plurality of initially-ranked search results ranked by a search engine, comprising:
using the computing device having the processor to perform the following:
setting as a first preference an image relevance for each image corresponding to an initial ranking by the search engine;
selecting one or more image quality measures;
computing an image quality for each image using the selected image quality measures;
clustering images in the plurality of initially-ranked search results using a clustering algorithm using the first preference and a second preference to obtain a plurality of clusters;
selecting an exemplar image for each of the plurality of clusters and including the exemplar image and images in the plurality of clusters in a summary candidate collection;
selecting a cluster and an image from the summary candidate collection;
selecting a number representing a desired number of summaries;
determining whether a number of images contained in a selected summaries collection is less than the number representing a desired number of summaries;
if so, then setting the summary candidate collection equal to images in the plurality of initially-ranked search results minus the number of images contained in the selected summaries collection;
computing an image score for each image in the selected cluster using the image relevance, the image quality, a cluster coherence, a density, and a diversity;
identifying an image in the selected cluster having a maximum image score as compared to other images in the selected cluster;
adding the image having a maximum image score to a selected summaries collection and removing the image having a maximum image score from the summary candidate collection; and
displaying to a user images in the selected summary collection in a ranked order based on the image score of the image.
12. The method of claim 11 , further comprising:
determining whether a number of images contained in the summary candidate collection is greater than zero; and
if not, then displaying to the user the images in the selected summaries collection.
13. The method of claim 11 , further comprising:
computing a visual distance between a selected image and other images in the selected cluster;
defining a scaling parameter;
computing a similarity between the selected image and the other images in the selected cluster using the visual distance and the scaling parameter.
14. The method of claim 13 , further comprising computing the cluster coherence, density, and diversity for the selected image using the computed similarity.
15. A computer-implemented method for generating summary images for an image search result containing a plurality of initially-ranked images, comprising:
setting as a first preference an image relevance of the plurality of initially-ranked images, denoted as R(i), which is an image relevance of an image in the i th position of the image search result;
computing an image quality of each of the plurality of initially-ranked images using a quality measure based on color entropy, denoted as Q(i), which is an image quality of an image in the i th position of the image search result;
clustering the plurality of initially-ranked images using an Affinity Propagation clustering technique having the image relevance as the first preference and the image quality as the second preference to obtain a plurality of clusters;
selecting an exemplar image from each of the plurality of clusters and including the exemplars and images in the plurality of clusters in a summary candidate collection;
selecting a cluster from the summary candidate collection and an image from the selected cluster;
obtaining a cluster coherence, Coh(i), a density, Dens(i), a diversity, Div(i), the image quality, Q(i), and the image relevance, R(i), for an i th image in the selected cluster;
computing an image score, S i , for the i th image using the following equation:
S i =W 1 ×Coh( i )+ W 2 ×Dens( i )+ W 3 ×Div(i)+α× R ( i )+β× Q ( i ),
where W 1 is a first weight, W 2 is a second weight, W 3 is a third weight, α is a first parameter, and β is a second parameter, until each image in the selected cluster has an image score;
identifying an image in the selected cluster having a maximum image score, removing the identified image from the summary candidate collection, and adding the identified image to a selected summaries collection; and
displaying to a user images in the selected summaries collection that are ranked accordingly to a respective image score.
16. The computer-implemented method of claim 15 , further comprising clustering the plurality of initially-ranked images using the Affinity Propagation clustering technique to find an overall preference, P(i), for the image in the i th position of the image search result using the equation:
P ( i )=α× R ( i )+β× Q ( i )+ c,
where c is a constant.
17. The computer-implemented method of claim 15 , further comprising computing the cluster coherence, Coh(i), for an i th image in the selected cluster, C i , using the equation:
Coh( i )=Σ I i ,I j εC i S ( I i ,I j ),
where S(I i ,I j ) is a similarity based on a distance between the i th image, I i , and a j th image, I j , and the distance, Dis(i,j), given by the equation:
S
(
I
i
,
I
j
)
=
exp
(
-
Dis
(
i
,
j
)
α
2
)
.
18. The computer-implemented method of claim 17 , further comprising computing the density, Dens(i), for the i th image in the selected cluster, using the equation:
Dens( i )=Σ I j S ( I i ,I j ),
where I j is an image in the image search result other than I i .
19. The computer-implemented method of claim 18 , further comprising computing the diversity, Div(i), for the i th image in the selected cluster, using the equation:
Div( i )=max I j εA S ( I i ,I j ),
where A is a set of selected images.