IP Library Granted Patent US 10,467,252
Granted Patent B1
US 10,467,252 · App. 13/754,780 · Granted Nov 5, 2019

Document classification and characterization using human judgment, tiered similarity analysis and language/concept analysis

Inventors: Stephen John Barsony (King of Prussia, PA); Yerachmiel Tzvi Messing (Baltimore, MD); David Matthew Shub (Cranford, NJ); Philip L. Richards (Charlotte, NC); Stephen H. Schreiber (Warren, MI)
G06F16/285
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,467,252
App. No.
13/754,780
Granted
Nov 5, 2019
Kind
B1
Abstract

Systems, methods, and articles are provided for characterizing and defining groups within large corpuses of documents using a combination of one or more of human judgment, tiered similarity analysis techniques, and language/concept analysis. Related apparatus, systems, techniques and articles are also described.

Claims (51)

1. A method comprising:

receiving a corpus of documents;

characterizing similarities among the corpus of documents using at least three similarity algorithms having different similarity criteria, the characterizing comprising:

obtaining contextual characteristics for each of the corpus of documents and associating the contextual characteristics with the corresponding document, the contextual characteristics selected from a group consisting of: similarity score, type of similarity algorithm used to characterize the document, document family, document type, and metadata describing properties of the document;

first removing a first portion of the corpus of documents based on applying a first similarity algorithm to the corpus of documents;

second removing, after the first removing, a second portion of the corpus of documents based on applying a second similarity algorithm to the corpus of documents; and

third removing, after the first removing and the second removing, a third portion of the corpus of documents based on applying a third similarity algorithm to the corpus of documents, the third similarity algorithm based on a criteria other than that implemented by the first similarity algorithm and the second similarity algorithm, wherein the third similarity algorithm identifies conceptually similar documents in the corpus of documents based on content of each respective document, and wherein the conceptually similar documents are neither exact duplicates nor substantial duplicates;

defining stacks of documents based on pre-defined grouping criteria as applied to the characterized similarities among the corpus of documents, the characterized similarities based on the first removing, the second removing, or the third removing;

identifying, within each stack, a prime document; and

initiating provision of each prime document to at least one human reviewer via a computer-implemented document review and characterization system.

2. The method as in claim 1 , wherein at least one of the similarity algorithms identifies exact duplicates in the corpus of documents.

3. The method as in claim 1 , wherein at least one of the similarity algorithms identifies substantial duplicates in the corpus of documents.

4. The method as in claim 1 , wherein the pre-defined grouping criteria is adjusted based on quality control review of documents within the stack other than the corresponding prime document.

5. The method as in claim 1 , further comprising: receiving data characterizing quality control review of at least a portion of the corpus of documents, the received data being used to modify the pre-defined grouping criteria to either increase or decrease one or more similarity metrics.

6. The method as in claim 1 , wherein at least a portion of the corpus of documents comprise families of documents, the families of documents having a pre-defined logical interrelation.

7. A method comprising:

receiving a corpus of documents;

obtaining contextual characteristics for each of the corpus of documents and associating the contextual characteristics with the corresponding document, the contextual characteristics selected from a group consisting of: similarity score, type of similarity algorithm used to characterize the document, document family, document type, and metadata describing properties of the document;

generating a first subset of the corpus of documents by identifying and characterizing similarities among the corpus of documents based on applying a first similarity algorithm to the corpus of documents;

generating a second subset of the corpus of documents by identifying and characterizing similarities among the first subset of the corpus of documents based on applying a second similarity algorithm to the first subset of the corpus of documents, the second similarity algorithm having a relaxed similarity standard as compared to the first similarity algorithm; and

generating a third subset of the corpus of documents by identifying and characterizing similarities among the corpus of documents based on applying a third similarity algorithm to the second subset of the corpus of documents, the third similarity algorithm having a similarity standard as other than that implemented by the first similarity algorithm and the second similarity algorithm, wherein the third similarity algorithm identifies conceptually similar documents in the corpus of documents based on content of each respective document, and wherein the conceptually similar documents are neither exact duplicates nor substantial duplicates;

defining stacks of documents based on pre-defined grouping criteria as applied to the second subset of the corpus of documents and the third subset of the corpus of documents;

identifying, within each stack, a prime document; and

initiating provision of each prime document to at least one human reviewer via a computer-implemented document review and characterization system.

8. A non-transitory computer program product storing instructions, which when executed by at least one data processor of at least one computing system, result in operations comprising:

receiving a corpus of documents;

characterizing similarities among the corpus of documents using at least three similarity algorithms having different similarity criteria, the characterizing comprising:

obtaining contextual characteristics for each of the corpus of documents and associating the contextual characteristics with the corresponding document, the contextual characteristics selected from a group consisting of: similarity score, type of similarity algorithm used to characterize the document, document family, document type, and metadata describing properties of the document;

first removing a first portion of the corpus of documents based on applying a first similarity algorithm to the corpus of documents; and

second removing, after the first removing, a second portion of the corpus of documents based on applying a second similarity algorithm to the corpus of documents; and

third removing, after the first removing and the second removing, a third portion of the corpus of documents based on applying a third similarity algorithm to the corpus of documents, the third similarity algorithm based on a criteria other than that implemented by the first similarity algorithm and the second similarity algorithm, wherein the third similarity algorithm identifies conceptually similar documents in the corpus of documents based on content of each respective document, and wherein the conceptually similar documents are neither exact duplicates nor substantial duplicates;

defining stacks of documents based on pre-defined grouping criteria as applied to the characterized similarities among the corpus of documents, the characterized similarities based on the first removing, the second removing, or the third removing;

identifying, within each stack, a prime document; and

initiating provision of each prime document to at least one human reviewer via a computer-implemented document review and characterization system.

9. The non-transitory computer program product as in claim 8 , wherein at least one of the similarity algorithms identifies exact duplicates in the corpus of documents.

10. The non-transitory computer program product as in claim 8 , wherein at least one of the similarity algorithms identifies substantial duplicates in the corpus of documents.

11. The non-transitory computer program product as in claim 8 , wherein the pre-defined grouping criteria is adjusted based on quality control review of documents within the stack other than the corresponding prime document.

12. The non-transitory computer program product as in claim 8 , wherein the operations further comprise: receiving data characterizing quality control review of at least a portion of the corpus of documents, the received data being used to modify the pre-defined grouping criteria to either increase or decrease one or more similarity metrics.

13. The non-transitory computer program product as in claim 8 , wherein at least a portion of the corpus of documents comprise families of documents, the families of documents having a pre-defined logical interrelation.

14. A system comprising:

at least one data processor; and

memory storing instructions, which when executed by the at least one data processor, result in operations comprising:

receiving a corpus of documents;

characterizing similarities among the corpus of documents using at least three similarity algorithms having different similarity criteria, the characterizing comprising:

obtaining contextual characteristics for each of the corpus of documents and associating the contextual characteristics with the corresponding document, the contextual characteristics selected from a group consisting of: similarity score, type of similarity algorithm used to characterize the document, document family, document type, and metadata describing properties of the document;

first removing a first portion of the corpus of documents based on applying a first similarity algorithm to the corpus of documents;

second removing, after the first removing, a second portion of the corpus of documents based on applying a second similarity algorithm to the corpus of documents; and

third removing, after the first removing and the second removing, a third portion of the corpus of documents based on applying a third similarity algorithm to the corpus of documents, the third similarity algorithm based on a criteria other than that implemented by the first similarity algorithm and the second similarity algorithm, wherein the third similarity algorithm identifies conceptually similar documents in the corpus of documents based on content of each respective document, and wherein the conceptually similar documents are neither exact duplicates nor substantial duplicates;

defining stacks of documents based on pre-defined grouping criteria as applied to the characterized similarities among the corpus of documents, the characterized similarities based on the first removing, the second removing, or the third removing;

identifying, within each stack, a prime document; and

initiating provision of each prime document to at least one human reviewer via a computer-implemented document review and characterization system.

Assignments (10)
RELEASE OF SECOND LIEN SECURITY INTEREST IN PATENTS, RECORDED AT REEL 056329, FRAME 0610 Recorded Jan 9, 2025
From: UBS AG, STAMFORD BRANCH, AS COLLATERAL AGENT
To: XCELLENCE, LLC; CONSILIO, LLC
Reel/Frame 069856/0006 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2022
From: DISCOVERREADY LLC
To: CONSILIO, LLC
Reel/Frame 059080/0227 →
SECURITY INTEREST Recorded May 24, 2021
From: DISCOVERREADY LLC; FASTLINE TECHNOLOGIES, LLC; CONSILIO, LLC; XCELLENCE, INC.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS FIRST LIEN COLLATERAL AGENT
Reel/Frame 056329/0599 →
SECURITY INTEREST Recorded May 24, 2021
From: DISCOVERREADY LLC; FASTLINE TECHNOLOGIES, LLC; CONSILIO, LLC; XCELLENCE, INC.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS SECOND LIEN COLLATERAL AGENT
Reel/Frame 056329/0610 →
RELEASE OF SECURITY INTEREST Recorded May 20, 2021
From: JEFFERIES FINANCE LLC (RELEASING REEL/FRAME 47311/0298)
To: DISCOVERREADY LLC
Reel/Frame 056302/0799 →
RELEASE OF SECURITY INTEREST Recorded May 20, 2021
From: JEFFERIES FINANCE LLC (RELEASING REEL/FRAME 47311/0193)
To: DISCOVERREADY LLC
Reel/Frame 056302/0772 →
SECURITY INTEREST Recorded Oct 25, 2018
From: DISCOVERREADY LLC
To: JEFFERIES FINANCE LLC
Reel/Frame 047311/0193 →
SECURITY INTEREST Recorded Oct 25, 2018
From: DISCOVERREADY LLC
To: JEFFERIES FINANCE LLC
Reel/Frame 047311/0298 →
RELEASE OF SECURITY INTEREST Recorded Jun 25, 2018
From: BAYSIDE CAPITAL, INC.
To: DISCOVERREADY LLC
Reel/Frame 046191/0124 →
SECURITY INTEREST Recorded Nov 20, 2014
From: DISCOVERREADY LLC
To: BAYSIDE CAPITAL, INC.
Reel/Frame 034222/0146 →
Continuity (1)
Provisional Application 61592487 · Jan 30, 2012