IP Library Granted Patent US 7,366,718
Granted Patent B1
US 7,366,718 · App. 10/608,468 · Granted Apr 29, 2008

Detecting duplicate and near-duplicate files

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,366,718
App. No.
10/608,468
Granted
Apr 29, 2008
Kind
B1
Abstract

Improved duplicate and near-duplicate detection techniques may assign a number of fingerprints to a given document by (i) extracting parts from the document, (ii) assigning the extracted parts to one or more of a predetermined number of lists, and (iii) generating a fingerprint from each of the populated lists. Two documents may be considered to be near-duplicates if any one of their fingerprints match.

Claims (8)

1. A method for filtering a plurality of candidate search results to remove near-duplicates, the method comprising:

a) for one of the plurality of candidate search results, determining whether the one candidate search result is a near-duplicate of another of the plurality of candidate search results by

1) comparing a cluster identifier of the one candidate search result with a cluster identifier of the other candidate search result, and

2) if the cluster identifiers of the one and the other candidate search results match, then concluding that the one candidate search is a near-duplicate of the other candidate search result; and

b) in response to a determination that the one candidate search result is a near-duplicate of the other candidate search result, rejecting the one candidate search result thereby defining a filtered set of search results including only those of the plurality of candidate search results that have not been rejected.

2. A search filter for processing a plurality of search results to remove near-duplicates, the search filter comprising:

a) a near-duplicate determination facility for determining, for one of the plurality of candidate search results, whether the one candidate search result is a near-duplicate of another of the plurality of candidate search results, and wherein the near-duplicate determination facility includes a comparison facility for comparing a cluster identifier of the one candidate search result with a cluster identifier of the other candidate search result, and wherein if the cluster identifiers of the one candidate search result and the other candidate search result match, then it is concluded that the one candidate search result and the other candidate search result are near-duplicates; and

b) a filter for rejecting the one candidate search result if it is determined that the one candidate search result is a near-duplicate of the other candidate search result and passing the one candidate search result if it is not determined that the one candidate search result is a near-duplicate of the other candidate search result, thereby defining a filtered set of search results including only those of the plurality of candidate results that have not been rejected by the filter.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044101/0610 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 18, 2009
From: PUGH, WILLIAM; HENZINGER, MONIKA H.
To: GOOGLE INC.
Reel/Frame 022414/0155 →