IP Library Granted Patent US 7,519,607
Granted Patent B2
US 7,519,607 · App. 10/640,460 · Granted Apr 14, 2009

Computer-based system and method for generating, classifying, searching, and analyzing standardized text templates and deviations from standardized text templates

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,519,607
App. No.
10/640,460
Granted
Apr 14, 2009
Kind
B2
Abstract

A method for generating, classifying, searching, and analyzing standardized text templates drawn from a plurality of text documents and for identifying standardized text deviations from standardized text templates. Semi-standardized documents may be represented as standardized templates and deviations from standardized templates, with such templates themselves automatically generated by a computer-implemented method from a plurality of similar text documents. The method enables enhanced analysis of semi-standardized documents and automatic extraction of information from standardized text templates.

Claims (33)

1. A method for representing similar text documents and similar segments of text documents as standardized text templates and deviations from standardized text templates, the method comprising:

partitioning documents into segments;

generating standardized text segment templates from segments of documents;

determining the deviations of individual text segments from the standardized text segment templates so generated;

representing the individual text segments as the combination of standardized text segment templates and deviations from standardized text segment templates;

generating standardized text templates representing text documents as sequences of standardized text segment templates;

determining the deviations of the individual text documents from the standardized text templates so generated; and

representing the individual text documents as a combination of the standardized text templates and the deviations from the standardized text templates.

2. The method of claim 1 , where the documents are partitioned into segments corresponding to paragraphs.

3. The method of claim 1 , where the standardized text template generated is the document or portion of a document with the least average edit distance with respect to each other text document in a family of similar text documents.

4. The method of claim 1 , further comprising:

identifying elements or segments of the generated template that frequently vary with respect to the other similar text documents; and

coding those elements or segments as “blanks.”

5. The method of claim 1 , where the documents are of the file types used in a computer word processor.

6. The method of claim 1 , where the similar text documents are business or legal documents.

7. The method of claim 1 , further comprising:

clustering text documents into families of similar text inputs or portions of text inputs;

comparing each text input in the family to each other text input in the family; and

selecting the text input with the least average distance from all other text inputs in the family as the standardized text template.

8. The method of claim 7 , where the documents are partitioned into segments corresponding to paragraphs.

9. The method of claim 7 , where the template selected is the document or portion of a document with the least average edit distance with respect to each other text document in the family of similar text documents.

10. The method of claim 7 , further comprising:

identifying elements or segments of the generated template that frequently vary with respect to other matching documents; and

coding those elements or segments as “blanks.”

11. The method of claim 7 , where the documents are of the file types used in a computer word processor.

12. Computer readable media having computer readable program code thereon which, when loaded into and executed by a computer, performs the steps of:

partitioning documents into segments;

generating a standardized text segment templates from segments of documents;

determining the deviations of individual text segments from the standardized text templates so generated;

representing the individual text segments as the combination of standardized text segment templates and the deviations from the standardized text segment templates;

generating a standardized text templates representing text documents as sequences of standardized text segment templates;

determining the deviations of individual text documents from the standardized text templates so generated; and

representing the individual text documents as a combination of the standardized text templates and the deviations from the standardized text templates.

Assignments (3)
SECURITY INTEREST Recorded Oct 16, 2017
From: BLOOMBERG FINANCE L.P.
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 044217/0047 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2014
From: DOCUMENT ANALYTIC TECHNOLOGIES, LLC, D/B/A/ EXEMPLIFY
To: THE BUREAU OF NATIONAL AFFAIRS, INC.
Reel/Frame 034165/0650 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2012
From: ANDERSON, ROBERT, IV
To: DOCUMENT ANALYTIC TECHNOLOGIES, LLC
Reel/Frame 028209/0547 →