IP Library Granted Patent US 8,285,750
Granted Patent B2
US 8,285,750 · App. 12/398,728 · Granted Oct 9, 2012

Computer-based system and method for generating, classifying, searching, and analyzing standardized text templates and deviations from standardized text templates

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,285,750
App. No.
12/398,728
Granted
Oct 9, 2012
Kind
B2
Abstract

A method for generating, classifying, searching, and analyzing standardized text templates drawn from a plurality of text documents and for identifying standardized text deviations from standardized text templates. Semi-standardized documents may be represented as standardized templates and deviations from standardized templates, with such templates themselves automatically generated by a computer-implemented method from a plurality of similar text documents. The method enables enhanced analysis of semi-standardized documents and automatic extraction of information from standardized text templates.

Claims (42)

1. A computer system configured to automatically analyze text documents by performing the following steps in the order recited:

receiving a subject text document;

comparing the text of the subject text document or the text of a subset of the subject text document to the text of a plurality of given text templates, each text template containing at least one paragraph of text;

determining which given text template or text templates has text that matches the text of the subject text document or the text of the subset of the subject text document to a given degree of correspondence; and

generating a report of the differences between the text of the subject text document or the text of the subset of the subject text document and the text of the matching text template or text templates.

2. The computer system of claim 1 also configured to perform the steps of:

comparing a family of specimen text documents;

identifying one paragraph of text within one of the family of specimen text documents that most closely matches a paragraph of text in all of the other specimen text documents, as compared to all of the other paragraphs in the one specimen text document; and

generating one of the text templates containing at least the one identified paragraph of text.

3. The computer system of claim 2 wherein the identifying the paragraph that most closely matches uses an edit distance algorithm.

4. The computer system of claim 2 wherein the identifying the paragraph that most closely matches uses a longest common subsequence algorithm.

5. The computer system of claim 1 wherein the determining matches to a degree of correspondence uses an edit distance algorithm.

6. The computer system of claim 1 wherein the determining matches to a degree of correspondence uses a longest common subsequence algorithm.

7. The computer system of claim 1 wherein the differences in the report are computed according to an edit distance algorithm.

8. Non-transitory, tangible, computer readable media containing computer programming instructions that, when loaded in a computer system having a memory, cause the computer system, in response to user commands, to automatically analyze text documents by performing the following steps in the order recited:

receiving a subject text document;

comparing the text of the subject text document or the text of a subset of the subject text document to the text of a plurality of given text templates, each text template containing at least one paragraph of text;

determining which given text template or text templates has text that matches the text of the subject text document or the text of the subset of the subject text document to a given degree of correspondence; and

generating a report of the differences between the text of the subject text document or the text of the subset of the subject text document and the text of the matching text template or text templates.

9. The media of claim 8 wherein the programming instructions, when loaded in the computer system, also cause the computer system, in response to user commands, to perform the following steps:

comparing a family of specimen text documents;

identifying one paragraph of text within one of the family of specimen text documents that most closely matches a paragraph of text in all of the other specimen text documents, as compared to all of the other paragraphs in the one specimen text document; and

generating one of the text templates containing at least the one identified paragraph of text.

10. The media of claim 9 wherein the identifying the paragraph that most closely matches uses an edit distance algorithm.

11. The media of claim 9 wherein the identifying the paragraph that most closely matches uses a longest common subsequence algorithm.

12. The media of claim 8 wherein the determining matches to a degree of correspondence uses an edit distance algorithm.

13. The media of claim 8 wherein the determining matches to a degree of correspondence uses a longest common subsequence algorithm.

14. The media of claim 8 wherein the differences in the report are computed according to an edit distance algorithm.

15. A computer-implemented method of automatically analyzing text documents in which a computer performs the following steps in the order recited:

receiving a subject text document;

comparing the text of the subject text document or the text of a subset of the subject text document to the text of a plurality of given text templates, each text template containing at least one paragraph of text;

determining which given text template or text templates has text that matches the text of the subject text document or the text of the subset of the subject test document to a given degree of correspondence; and

generating a report of the differences between the text of the subject text document or the text of the subset of the subject text document and the text of the matching text template or text templates.

16. The computer-implemented method of claim 15 further comprising the steps of:

comparing a family of specimen text documents;

identifying one paragraph of text within one of the family of specimen text documents that most closely matches a paragraph of text in all of the other specimen text documents, as compared to all of the other paragraphs in the one specimen text document; and

generating one of the text templates containing at least the one identified paragraph of text.

17. The computer-implemented method of claim 16 wherein the identifying the paragraph that most closely matches uses an edit distance algorithm.

18. The computer-implemented method of claim 16 wherein the identifying the paragraph that most closely matches uses a longest common subsequence algorithm.

19. The computer-implemented method of claim 15 wherein the determining matches to a degree of correspondence uses an edit distance algorithm.

20. The computer-implemented method of claim 15 wherein the determining matches to a degree of correspondence uses a longest common subsequence algorithm.

21. The computer-implemented method of claim 15 wherein the differences in the report are computed according to an edit distance algorithm.

Assignments (3)
SECURITY INTEREST Recorded Oct 16, 2017
From: BLOOMBERG FINANCE L.P.
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 044217/0047 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2014
From: DOCUMENT ANALYTIC TECHNOLOGIES, LLC, D/B/A/ EXEMPLIFY
To: THE BUREAU OF NATIONAL AFFAIRS, INC.
Reel/Frame 034165/0650 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2012
From: ANDERSON, ROBERT, IV
To: DOCUMENT ANALYTIC TECHNOLOGIES, LLC
Reel/Frame 028209/0547 →