IP Library Granted Patent US 8,156,092
Granted Patent B2
US 8,156,092 · App. 12/022,144 · Granted Apr 10, 2012

Document de-duplication and modification detection

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,156,092
App. No.
12/022,144
Granted
Apr 10, 2012
Kind
B2
Abstract

Provided is a system and method for the de-duplication and modification detection of documents collected during document production. The disclosed technology provides a simple, legally defensible, rapid and cost-efficient system for collecting responsive electronic document sets, identifying and eliminating unnecessary documents by comparing a collected document to previously collected documents and copying only information that has not been duplicated. The disclosed technology provides a method for copying the unduplicated information without transmitting or storing the duplicated portions. In addition, the claimed subject matter provides a system for detecting whether or not a document being submitted to a project archive is a modification of a previously submitted document. A document being submitted that represents a modification of a previously submitted document is prevented from being added to the project document archive.

Claims (56)

1. A method for the organization and collection of documents, comprising:

collecting a first document, including a metadata portion and a non-metadata portion, in response to a document collection request;

generating a first hash code corresponding to the non-metadata portion of the first document;

comparing the first hash code to a plurality of hash codes, each hash code of the plurality of hash codes corresponding to a non-metadata portion of a corresponding document of a plurality of documents, each document collected in response to the document collection request;

if the first hash code does not match any hash code of the plurality of hash codes, storing the first document, including the metadata portion and the non-metadata portion, on a data storage device; and

if the first hash code matches any hash code of the plurality of hash codes, extracting metadata corresponding to the first document; and

storing on the data storage device the extracted metadata, but not the non-metadata portion of the first document, in conjunction with the particular document corresponding to the hash code that matches the first hash code.

2. The method of claim 1 , further comprising:

generating collection event information concerning the collection of the first document; and

storing on the data storage device the collection event information in conjunction with the first document or the extracted metadata, depending upon whether the first document or the extracted metadata is stored, respectively.

3. The method of claim 1 , further comprising:

generating chain of custody information corresponding to a party collecting the first document; and

storing on the data storage the chain of custody information in conjunction with the with the first document if the first hash code and second hash code do not match.

4. The method of claim 1 , further comprising:

determining whether or not the first document and a second document of the plurality of documents represent the same document; and

preventing the first document from being stored on the data storage device if the first document and the second document represent the same document and the first hash code does not match the hash code corresponding to the second document.

5. The method of claim 1 , wherein the hash code is generated using a Message-digest algorithm 5 (MD5) hash function.

6. A system for the organization and collection of documents, comprising:

a processor;

a memory coupled to the processor; and

logic, stored on the memory for execution on the processor, for

collecting a first document, including a metadata portion and a non-metadata portion, in response to a document collection request;

generating a first hash code corresponding to the non-metadata portion of the first document;

comparing the first hash code to a plurality of hash codes, each hash code of the plurality of hash codes corresponding to a non-metadata portion of a corresponding document of a plurality of documents, each document collected in response to the document collection request;

if the first hash code does not match any hash code of the plurality of hash codes, storing the first document, including the metadata portion and the non-metadata portion, on a data storage; and

if the first hash code matches any hash code of the plurality of hash codes, extracting metadata corresponding to the first document: and

storing on the data storage the extracted metadata, but not the non-metadata portion of the first document, in conjunction with the particular document corresponding to the hash code that matches the first hash code.

7. The system of claim 6 , the logic further comprising logic for:

generating collection event information concerning the collection of the first document; and

storing on the data storage the collection event information in conjunction with the first document or the extracted metadata, depending upon whether the first document or the extracted metadata is stored, respectively.

8. The system of claim 6 , the logic further comprising logic for:

generating chain of custody information corresponding to a party collecting the first document; and

storing on the data storage the chain of custody information in conjunction with the first document.

9. The system of claim 6 , the logic further comprising logic for:

determining whether or not the first document and a second document of the plurality of documents represent the same document; and

preventing the first document from being stored on the data storage in conjunction with the second document if the first document and the second document represent the same document and the generated hash code does not match the second hash code.

10. The system of claim 6 , wherein the hash code is generated using a Message-digest algorithm 5 (MD5) hash function.

11. A computer programming product for the organization and collection of documents, comprising:

a memory; and

logic, stored on the memory for execution on a processor, for:

collecting a first document, including a metadata portion and a non-metadata portion, in response to a document collection request;

generating a first hash code corresponding to the non-metadata portion of the first document;

comparing the first hash code to a plurality of hash codes, each hash code of the plurality of hash codes corresponding to a non-metadata portion of a corresponding document of a plurality of documents, each document collected in response to the document collection request;

the first hash code does not match any hash code of the plurality of hash codes, storing the first document, including the metadata portion and the non-metadata portion, on a data storage; and

if the first hash code matches any hash code of the plurality of hash codes, extracting metadata corresponding to the first document; and

storing on the data storage the extracted metadata, but not the non-metadata portion of the first document, in conjunction with the particular document corresponding to the hash code that matches the first hash code.

12. The computer programming product of claim 11 , the logic further comprising logic for:

generating collection event information concerning the collection of the first document; and

storing on the data storage the collection event information in conjunction with the first document or the extracted metadata, depending upon whether the first document or the extracted metadata is stored, respectively.

13. The computer programming product of claim 11 , the logic further comprising logic for:

generating chain of custody information corresponding to a party collecting the first document; and

storing on the data storage the chain of custody information in conjunction with the first document.

14. The computer programming product of claim 11 , the logic further comprising logic for:

determining whether or not the first document and a second document of the plurality of documents represent the same document; and

preventing the first document from being stored on the data storage in conjunction with the second document if the first document and the second document represent the same document and the generated hash code does not match the second hash code.

15. The computer programming product of claim 11 , wherein the hash code is generated using a Message-digest algorithm 5 (MD5) hash function.

Continuity (2)
Continuation 12022137 · Jan 29, 2008
Related Publication 20090192978A1 · Jul 30, 2009