IP Library Granted Patent US 8,276,065
Granted Patent B2
US 8,276,065 · App. 11/526,470 · Granted Sep 25, 2012

System and method for classifying electronically posted documents

Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,276,065
App. No.
11/526,470
Granted
Sep 25, 2012
Kind
B2
Abstract

A method for classifying electronically posted documents includes receiving two posted documents and generating corresponding metadata summaries for each, wherein each of the metadata summaries includes at least one sub-tree structure. The structures of the two summary sub-trees within the respective metadata summaries are subsequently compared. If the two summary sub-trees are different, the two documents are deemed distinct. If the two summary sub-trees are the same, attribute values and text content of the metadata summaries are compared over a portion of the metadata summaries. If the compared attribute values and text content are determined to be the same, the documents are deemed duplicative.

Claims (58)

1. A method for classifying electronically posted documents, the method comprising:

receiving a first document and a second document;

generating a first metadata summary for the first document and a second metadata summary for the second document, wherein the first metadata summary includes a first plurality of sub-trees and the second metadata summary includes a second plurality of sub-trees, wherein each of the sub-trees includes a plurality of nodes, wherein each of the sub-trees in the first and second pluralities of sub-trees comprises a structural format defined by a plurality of hierarchical structural constructs, and wherein subsets of the metadata summaries are grouped into a set of different summary groups based on a mime-type designation associated with each subset of the metadata summaries;

comparing the first and second metadata summaries on a structural level by comparing the entire structural format defined by the plurality of hierarchical structural constructs of the sub-trees of the first metadata summary with the entire structural format defined by the plurality of hierarchical structural constructs of the sub-trees of the second metadata summary;

identifying the first and second documents as distinct if the entire structural formats of the sub-trees of the first and second metadata summaries are not equivalent;

if the entire structural formats of the sub-trees of the first and second metadata summaries are equivalent, performing a further comparison of the first and second metadata summaries,

wherein the further comparison of the first and second metadata summaries includes the sub-steps of:

comparing the first and second metadata summaries on a textual level by comparing textual content from the first document that is contained in the sub-trees of the first metadata summary with textual content from the second document that is contained in the sub-trees of the second metadata summary;

identifying the first and second documents as distinct if the textual content within the sub-trees of the first and second metadata summaries are not equivalent;

before comparing the first and second metadata summaries on a textual level, comparing the first and second metadata summaries on an attribute level by comparing attribute values within the sub-trees of the first metadata summary with attribute values within the sub-trees of the second metadata summary; and

identifying the first and second documents as distinct if the attribute values within the sub-trees of the first and second metadata summaries are not equivalent.

2. The method of claim 1 , further comprising identifying the first and second documents as duplicates if the textual content within the sub-trees of the first and second metadata summaries are equivalent.

3. The method of claim 2 , further comprising removing the second metadata summary if the first and second documents are identified as duplicates.

4. The method of claim 1 , further comprising:

defining a first equivalence metadata table comprising:

a first row corresponding to the first metadata summary;

a second row corresponding to the second metadata summary;

a first column corresponding to the first metadata summary; and

a second column corresponding to the second metadata summary,

wherein the step of identifying the first and second documents as distinct if the structural formats of the sub-trees of the first and second metadata summaries are not equivalent comprises storing a zero value in the first row and second column position of the first equivalence metadata table.

5. The method of claim 4 , wherein the step of identifying the first and second documents as distinct if the textual content within the sub-trees of the first and second metadata summaries are not equivalent comprises storing a zero value in the first row and second column position of the first equivalence metadata table.

6. The method of claim 1 , further comprising:

defining a first equivalence metadata table comprising:

a first row corresponding to the first metadata summary;

a second row corresponding to the second metadata summary;

a first column corresponding to the first metadata summary; and

a second column corresponding to the second metadata summary,

wherein the step of identifying the first and second documents as distinct if the attribute values within the sub-trees of the first and second metadata summaries are not equivalent comprises storing a zero value in the first row and second column position of the first equivalence metadata table.

7. A system for classifying electronically posted documents, the system comprising:

a memory;

a processor communicatively coupled to the memory;

a metadata parser module coupled to receive electronically posted documents, the metadata parser configured to output a metadata summary for each of the posted documents, wherein each of the metadata summaries comprises a plurality of sub-trees and each of the sub-trees includes a plurality of nodes, wherein each of the sub-trees in the pluralities of sub-trees comprises a structural format defined by a plurality of hierarchical structural constructs, and wherein subsets of the metadata summaries are grouped into a set of different summary groups based on a mime-type designation associated with each subset of the metadata summaries;

a summary repository coupled to receive and store the metadata summaries; and

a summary consolidator coupled to the summary repository, the summary consolidator configured to:

select a first metadata summary and a second metadata summary from one of the summary groups in the set of different summary groups;

compare the first metadata summary and the second metadata summary on a structural level by comparing the entire structural format defined by the plurality of hierarchical structural constructs of the sub-trees of the first metadata summary with the entire structural format defined by the plurality of hierarchical structural constructs of the sub-trees of the second metadata summary;

identify the first and second documents corresponding to the first and second metadata summaries as distinct if the entire structural formats of the sub-trees of the first and second metadata summaries are not equivalent; and

if the entire structural formats of the sub-trees of the first and second metadata summaries are equivalent, further compare the first and second metadata summaries,

wherein the further comparison of the first and second metadata summaries includes:

comparing the first and second metadata summaries on a textual level by comparing textual content from the first document that is contained in the sub-trees of the first metadata summary with textual content from the second document that is contained in the sub-trees of the second metadata summary;

identifying the first and second documents corresponding to the first and second metadata summaries as distinct if the textual content within the sub-trees of the first and second metadata summaries are not equivalent;

before comparing the first and second metadata summaries on a textual level, comparing the first and second metadata summaries on an attribute level by comparing attribute values within the sub-trees of the first metadata summary with attribute values within the sub-trees of the second metadata summary; and

identifying the first and second documents corresponding to the first and second metadata summaries as distinct if the attribute values within the sub-trees of the first and second metadata summaries are not equivalent.

8. A program product for use in a computer system that executes program steps recorded in a computer-readable media to perform a method for classifying electronically posted documents, the program product comprising:

a record-able media;

a program of computer-readable instructions executable by the computer system to perform processes comprising the steps of:

receiving a first document and a second document;

generating a first metadata summary for the first document and a second metadata summary for the second document, wherein the first metadata summary includes a first plurality of sub-trees and the second metadata summary includes a second plurality of sub-trees, and wherein each of the sub-trees includes a plurality of nodes, wherein each of the sub-trees in the first and second pluralities of sub-trees comprises a structural format defined by a plurality of hierarchical structural constructs, and wherein subsets of the metadata summaries are grouped into a set of different summary groups based on a mime-type designation associated with each subset of the metadata summaries;

comparing the first and second metadata summaries on a structural level by comparing the entire structural format defined by the plurality of hierarchical structural constructs of the sub-trees of the first metadata summary with the entire structural format defined by the plurality of hierarchical structural constructs of the sub-trees of the second metadata summary;

identifying the first and second documents as distinct if the entire structural formats of the sub-trees of the first and second metadata summaries are not equivalent;

if the entire structural formats of the sub-trees of the first and second metadata summaries are equivalent, performing a further comparison of the first and second metadata summaries,

wherein the further comparison of the first and second metadata summaries includes the sub-steps of:

comparing the first and second metadata summaries on a textual level by comparing textual content from the first document that is contained in the sub-trees of the first metadata summary with textual content from the second document that is contained in the sub-trees of the second metadata summary;

identifying the first and second documents as distinct if the textual content within the sub-trees of the first and second metadata summaries are not equivalent;

before comparing the first and second metadata summaries on a textual level, comparing the first and second metadata summaries on an attribute level by comparing attribute values within the sub-trees of the first metadata summary with attribute values within the sub-trees of the second metadata summary; and

identifying the first and second documents as distinct if the attribute values within the sub-trees of the first and second metadata summaries are not equivalent.

9. The program product of claim 8 , further comprising the step of identifying the first and second documents as duplicates if the textual content within the sub-trees of the first and second metadata summaries are equivalent.

10. The program product of claim 9 , further comprising the step of removing the second metadata summary if the first and second documents are identified as duplicates.

Continuity (2)
Continuation 09513058 · Feb 24, 2000
Related Publication 20070022374A1 · Jan 25, 2007