DETECTING DUPLICATE AND NEAR-DUPLICATE FILES
Improved duplicate and near-duplicate detection techniques may assign a number of fingerprints to a given document by (i) extracting parts from the document, (ii) assigning the extracted parts to one or more of a predetermined number of lists, and (iii) generating a fingerprint from each of the populated lists. Two documents may be considered to be near-duplicates if any one of their fingerprints match.
1 . A method for filtering a plurality of candidate search results to remove near-duplicates, the method comprising:
a) for one or more of the plurality of candidate search results, determining that one candidate search result of the one or more candidate search results is a near-duplicate of another of the plurality of candidate search results by
1) determining that a cluster identifier of the one candidate search result matches a cluster identifier of the other candidate search result; and
2) in response to determining that a cluster identifier of the one candidate search result matches a cluster identifier of the other candidate search result, concluding that the one candidate search is a near-duplicate of the other candidate search result; and
b) in response to the determination that the one candidate search result is a near-duplicate of the other candidate search result, rejecting the one candidate search result thereby defining a filtered set of search results including only those of the plurality of candidate search results that have not been rejected.
2 . The method of claim 1 , wherein determining that the one candidate search result is a near-duplicate of another of the plurality of candidate search results is preceded by:
determining that the one candidate search result is a near-duplicate of a previously processed search result; and
associating the one candidate search result with a cluster identifier of the previously processed search result.
3 . The method of claim 2 , wherein associating the one candidate search result with the cluster identifier of the previously processed search result comprises:
determining that the one candidate search result is associated with a different cluster identifier, wherein the different cluster identifier identifies a cluster of related search results; and
associating each of the cluster of related search results with the cluster identifier of the previously processed search result.
4 . The method of claim 1 , comprising:
receiving the plurality of candidate search results from a search engine.
5 . The method of claim 1 , wherein rejecting the one candidate search result comprises:
determining, for both of the one candidate search result and the other candidate search result, a quality measure associated with the search result;
determining that the one candidate search result has a lower quality measure than the other candidate search result; and
adding the other candidate search result to the filtered set of search results.
6 . The method of claim 1 , wherein rejecting the one candidate search result comprises:
determining that the document referenced by the other candidate search result is more recent than the document associated with the one candidate search result; and
adding the other candidate search result to the filtered set of search results.
7 . The method of claim 1 , comprising:
providing the filtered set of search results for display.
8 . An apparatus, comprising:
at least one processor; and
at least one storage device storing a processor executable program which, when executed by the at least one processor, causes the at least one processor to perform operations comprising:
a) for one or more of the plurality of candidate search results, determining that one candidate search result of the one or more candidate search results is a near-duplicate of another of the plurality of candidate search results by
1) determining that a cluster identifier of the one candidate search result matches a cluster identifier of the other candidate search result; and
2) in response to determining that a cluster identifier of the one candidate search result matches a cluster identifier of the other candidate search result, concluding that the one candidate search is a near-duplicate of the other candidate search result; and
b) in response to the determination that the one candidate search result is a near-duplicate of the other candidate search result, rejecting the one candidate search result thereby defining a filtered set of search results including only those of the plurality of candidate search results that have not been rejected.
9 . The apparatus of claim 8 , wherein performing the operation of determining that the one candidate search result is a near-duplicate of another of the plurality of candidate search results is preceded by performing operations comprising:
determining that the one candidate search result is a near-duplicate of a previously processed search result; and
associating the one candidate search result with a cluster identifier of the previously processed search result.
10 . The apparatus of claim 9 , wherein performing the operation of associating the one candidate search result with the cluster identifier of the previously processed search result comprises:
determining that the one candidate search result is associated with a different cluster identifier, wherein the different cluster identifier identifies a cluster of related search results; and
associating each of the cluster of related search results with the cluster identifier of the previously processed search result.
11 . The apparatus of claim 8 , wherein the operations further comprise:
receiving the plurality of candidate search results from a search engine.
12 . The apparatus of claim 8 , wherein performing the operations rejecting the one candidate search result comprises:
determining, for both of the one candidate search result and the other candidate search result, a quality measure associated with the search result;
determining that the one candidate search result has a lower quality measure than the other candidate search result; and
adding the other candidate search result to the filtered set of search results.
13 . The apparatus of claim 8 , wherein performing the operations rejecting the one candidate search result comprises:
determining that the document referenced by the other candidate search result is more recent than the document associated with the one candidate search result; and
adding the other candidate search result to the filtered set of search results.
14 . The apparatus of claim 8 , wherein the operations further comprise:
providing the filtered set of search results for display.
15 . A computer storage device encoded with a computer program, the computer program comprising instructions that, when executed by data processing apparatus, cause the data processing apparatus to perform operations comprising:
a) for one or more of the plurality of candidate search results, determining that one candidate search result of the one or more candidate search results is a near-duplicate of another of the plurality of candidate search results by
1) determining that a cluster identifier of the one candidate search result matches a cluster identifier of the other candidate search result; and
2) in response to determining that a cluster identifier of the one candidate search result matches a cluster identifier of the other candidate search result, concluding that the one candidate search is a near-duplicate of the other candidate search result; and
b) in response to the determination that the one candidate search result is a near-duplicate of the other candidate search result, rejecting the one candidate search result thereby defining a filtered set of search results including only those of the plurality of candidate search results that have not been rejected.
16 . The computer storage device of claim 15 , wherein performing the operation of determining that the one candidate search result is a near-duplicate of another of the plurality of candidate search results is preceded by performing operations comprising:
determining that the one candidate search result is a near-duplicate of a previously processed search result; and
associating the one candidate search result with a cluster identifier of the previously processed search result.
17 . The computer storage device of claim 16 , wherein performing the operation of associating the one candidate search result with the cluster identifier of the previously processed search result comprises:
determining that the one candidate search result is associated with a different cluster identifier, wherein the different cluster identifier identifies a cluster of related search results; and
associating each of the cluster of related search results with the cluster identifier of the previously processed search result.
18 . The computer storage device of claim 15 , wherein the operations further comprise:
receiving the plurality of candidate search results from a search engine.
19 . The computer storage device of claim 15 , wherein performing the operations rejecting the one candidate search result comprises:
determining, for both of the one candidate search result and the other candidate search result, a quality measure associated with the search result;
determining that the one candidate search result has a lower quality measure than the other candidate search result; and
adding the other candidate search result to the filtered set of search results.
20 . The computer storage device of claim 15 , wherein performing the operations rejecting the one candidate search result comprises:
determining that the document referenced by the other candidate search result is more recent than the document associated with the one candidate search result; and
adding the other candidate search result to the filtered set of search results.
21 . The computer storage device of claim 15 , wherein the operations further comprise:
providing the filtered set of search results for display.