IP Library Patent Application 13313913
Patent Application
App. No. 13/313,913

DETECTING DUPLICATE AND NEAR-DUPLICATE FILES

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
13/313,913
Abstract

Improved duplicate and near-duplicate detection techniques may assign a number of fingerprints to a given document by (i) extracting parts from the document, (ii) assigning the extracted parts to one or more of a predetermined number of lists, and (iii) generating a fingerprint from each of the populated lists. Two documents may be considered to be near-duplicates if any one of their fingerprints match.

Claims (68)

1 . A method for filtering a plurality of candidate search results to remove near-duplicates, the method comprising:

a) for one or more of the plurality of candidate search results, determining that one candidate search result of the one or more candidate search results is a near-duplicate of another of the plurality of candidate search results by

1) determining that a cluster identifier of the one candidate search result matches a cluster identifier of the other candidate search result; and

2) in response to determining that a cluster identifier of the one candidate search result matches a cluster identifier of the other candidate search result, concluding that the one candidate search is a near-duplicate of the other candidate search result; and

b) in response to the determination that the one candidate search result is a near-duplicate of the other candidate search result, rejecting the one candidate search result thereby defining a filtered set of search results including only those of the plurality of candidate search results that have not been rejected.

2 . The method of claim 1 , wherein determining that the one candidate search result is a near-duplicate of another of the plurality of candidate search results is preceded by:

determining that the one candidate search result is a near-duplicate of a previously processed search result; and

associating the one candidate search result with a cluster identifier of the previously processed search result.

3 . The method of claim 2 , wherein associating the one candidate search result with the cluster identifier of the previously processed search result comprises:

determining that the one candidate search result is associated with a different cluster identifier, wherein the different cluster identifier identifies a cluster of related search results; and

associating each of the cluster of related search results with the cluster identifier of the previously processed search result.

4 . The method of claim 1 , comprising:

receiving the plurality of candidate search results from a search engine.

5 . The method of claim 1 , wherein rejecting the one candidate search result comprises:

determining, for both of the one candidate search result and the other candidate search result, a quality measure associated with the search result;

determining that the one candidate search result has a lower quality measure than the other candidate search result; and

adding the other candidate search result to the filtered set of search results.

6 . The method of claim 1 , wherein rejecting the one candidate search result comprises:

determining that the document referenced by the other candidate search result is more recent than the document associated with the one candidate search result; and

adding the other candidate search result to the filtered set of search results.

7 . The method of claim 1 , comprising:

providing the filtered set of search results for display.

8 . An apparatus, comprising:

at least one processor; and

at least one storage device storing a processor executable program which, when executed by the at least one processor, causes the at least one processor to perform operations comprising:

a) for one or more of the plurality of candidate search results, determining that one candidate search result of the one or more candidate search results is a near-duplicate of another of the plurality of candidate search results by

1) determining that a cluster identifier of the one candidate search result matches a cluster identifier of the other candidate search result; and

2) in response to determining that a cluster identifier of the one candidate search result matches a cluster identifier of the other candidate search result, concluding that the one candidate search is a near-duplicate of the other candidate search result; and

b) in response to the determination that the one candidate search result is a near-duplicate of the other candidate search result, rejecting the one candidate search result thereby defining a filtered set of search results including only those of the plurality of candidate search results that have not been rejected.

9 . The apparatus of claim 8 , wherein performing the operation of determining that the one candidate search result is a near-duplicate of another of the plurality of candidate search results is preceded by performing operations comprising:

determining that the one candidate search result is a near-duplicate of a previously processed search result; and

associating the one candidate search result with a cluster identifier of the previously processed search result.

10 . The apparatus of claim 9 , wherein performing the operation of associating the one candidate search result with the cluster identifier of the previously processed search result comprises:

determining that the one candidate search result is associated with a different cluster identifier, wherein the different cluster identifier identifies a cluster of related search results; and

associating each of the cluster of related search results with the cluster identifier of the previously processed search result.

11 . The apparatus of claim 8 , wherein the operations further comprise:

receiving the plurality of candidate search results from a search engine.

12 . The apparatus of claim 8 , wherein performing the operations rejecting the one candidate search result comprises:

determining, for both of the one candidate search result and the other candidate search result, a quality measure associated with the search result;

determining that the one candidate search result has a lower quality measure than the other candidate search result; and

adding the other candidate search result to the filtered set of search results.

13 . The apparatus of claim 8 , wherein performing the operations rejecting the one candidate search result comprises:

determining that the document referenced by the other candidate search result is more recent than the document associated with the one candidate search result; and

adding the other candidate search result to the filtered set of search results.

14 . The apparatus of claim 8 , wherein the operations further comprise:

providing the filtered set of search results for display.

15 . A computer storage device encoded with a computer program, the computer program comprising instructions that, when executed by data processing apparatus, cause the data processing apparatus to perform operations comprising:

a) for one or more of the plurality of candidate search results, determining that one candidate search result of the one or more candidate search results is a near-duplicate of another of the plurality of candidate search results by

1) determining that a cluster identifier of the one candidate search result matches a cluster identifier of the other candidate search result; and

2) in response to determining that a cluster identifier of the one candidate search result matches a cluster identifier of the other candidate search result, concluding that the one candidate search is a near-duplicate of the other candidate search result; and

b) in response to the determination that the one candidate search result is a near-duplicate of the other candidate search result, rejecting the one candidate search result thereby defining a filtered set of search results including only those of the plurality of candidate search results that have not been rejected.

16 . The computer storage device of claim 15 , wherein performing the operation of determining that the one candidate search result is a near-duplicate of another of the plurality of candidate search results is preceded by performing operations comprising:

determining that the one candidate search result is a near-duplicate of a previously processed search result; and

associating the one candidate search result with a cluster identifier of the previously processed search result.

17 . The computer storage device of claim 16 , wherein performing the operation of associating the one candidate search result with the cluster identifier of the previously processed search result comprises:

determining that the one candidate search result is associated with a different cluster identifier, wherein the different cluster identifier identifies a cluster of related search results; and

associating each of the cluster of related search results with the cluster identifier of the previously processed search result.

18 . The computer storage device of claim 15 , wherein the operations further comprise:

receiving the plurality of candidate search results from a search engine.

19 . The computer storage device of claim 15 , wherein performing the operations rejecting the one candidate search result comprises:

determining, for both of the one candidate search result and the other candidate search result, a quality measure associated with the search result;

determining that the one candidate search result has a lower quality measure than the other candidate search result; and

adding the other candidate search result to the filtered set of search results.

20 . The computer storage device of claim 15 , wherein performing the operations rejecting the one candidate search result comprises:

determining that the document referenced by the other candidate search result is more recent than the document associated with the one candidate search result; and

adding the other candidate search result to the filtered set of search results.

21 . The computer storage device of claim 15 , wherein the operations further comprise:

providing the filtered set of search results for display.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2012
From: PUGH, WILLIAM; HENZINGER, MONIKA H.
To: GOOGLE INC.
Reel/Frame 028376/0317 →