IP Library Granted Patent US 11,023,828
Granted Patent B2
US 11,023,828 · App. 15/406,542 · Granted Jun 1, 2021

Systems and methods for predictive coding

Inventors: Jan Puzicha (Bonn, DE); Steve Vranas (Ashburn, VA)
Assignee: Open Text Holdings, Inc.
G06N20/10G06F16/93G06N5/04G06N5/048G06N7/005G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,023,828
App. No.
15/406,542
Filed
Jan 13, 2017
Granted
Jun 1, 2021
Kind
B2
Art Unit
2126
USPC
706/52
Abstract

Systems and methods for analyzing documents are provided herein. A plurality of documents and user input are received via a computing device. The user input includes hard coding of a subset of the plurality of documents, based on an identified subject or category. Instructions stored in memory are executed by a processor to generate an initial control set, analyze the initial control set to determine at least one seed set parameter, automatically code a first portion of the plurality of documents based on the initial control set and the seed set parameter associated with the identified subject or category, analyze the first portion of the plurality of documents by applying an adaptive identification cycle, and retrieve a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle test on the first portion of the plurality of documents.

Claims (51)

1. A method for analyzing a plurality of documents, comprising:

receiving the plurality of documents via a computing device;

filtering the plurality of documents to produce a subset of the plurality of documents;

executing instructions stored in memory, wherein execution of the instructions by a processor generates an initial control set based on random sampling of the subset of the plurality of documents;

receiving user input from the computing device, the user input based on an identified subject or category; and

executing instructions stored in memory, wherein execution of the instructions by a processor:

reviews the initial control set to determine at least one seed set parameter associated with the identified subject or category,

automatically codes a first portion of the plurality of documents, based on the initial control set and the at least one seed set parameter associated with the identified subject or category, the automatic coding being performed by a predictive coding system,

analyzes the first portion of the plurality of documents by applying an adaptive identification cycle, the adaptive identification cycle being based on at least one of the initial control set, user validation of soft coding of the first portion of the plurality of documents, and a confidence threshold validation,

automatically codes a second portion of the plurality of documents with the initial control set using the predictive coding system,

presents at least one document from the second portion of the plurality of documents to a human reviewer, the at least one document having automated coding from the predictive coding system,

allows the human reviewer to correct at least a portion of the automated coding by performing a hard coding correction, the hard coding correction comprising changing of the at least a portion of the automated coding from a first coding to a second coding,

receives the hard coding correction to the at least a portion of the automated coding of the at least one document,

updates the initial control set with the at least one document having the hard coding correction, and

applies the updated coded control set to additional documents to automatically code the additional documents.

2. The method according to claim 1 , wherein the plurality of documents is partially coded.

3. The method according to claim 1 , wherein the hard coding correction further comprises tagging based on the identified subject or category.

4. The method according to claim 3 , wherein the identified subject or category comprises any of relevancy, issue, and privilege.

5. The method according to claim 1 , further comprising:

querying for the additional documents that correspond to the hard coding of the plurality of documents using the at least one seed set parameter; and

retrieving the additional documents that correspond to the hard coding of the plurality of documents.

6. The method according to claim 5 , further comprising generating a confidence score for the additional documents that comprises a computer-generated judgment on soft coding of the additional documents.

7. The method according to claim 6 , further comprising updating the initial control set with the additional documents that correspond to the hard coding of the plurality of documents to create a new seed set.

8. The method according to claim 7 , further comprising:

identifying non-relevant documents as a set of negative examples;

enriching the set of negative examples to a given total of documents by randomly sampling the non-relevant documents; and

batching out and reviewing all non-relevant documents that correspond to a category or identified subject.

9. The method according to claim 8 , further comprising analyzing the non-relevant documents using a support vector machine that maps the non-relevant documents into an internal high-dimensional representation.

10. The method according to claim 9 , further comprising computing linear functions on the internal high-dimensional representation to model training examples.

11. The method according to claim 1 , further comprising retrieving the second portion of the plurality of documents based on a result of the application of the adaptive identification cycle on the first portion of the plurality of documents.

12. The method according to claim 1 , wherein generating the initial control set comprises receiving any of a search request or a request for content, the initial control set being selected based on the search request or the request for content.

13. A method for analyzing a plurality of documents, comprising:

executing instructions stored in memory, wherein execution of the instructions by a processor generates an initial control set of documents based on random sampling of a subset of the plurality of documents on a static basis and on a rolling load basis;

receiving hard coding of the subset of the plurality of documents;

reviewing the initial control set and a user input to determine at least one seed set parameter associated with the hard coding;

predictively coding a first portion of the plurality of documents, based on the initial control set and the at least one seed set parameter associated with an identified subject or category, the automatic coding being performed by a predictive coding system,

analyzing the first portion of the plurality of documents by applying an adaptive identification cycle, the adaptive identification cycle being based on at least one of the initial control set, user validation of soft coding of the first portion of the plurality of documents, and a confidence threshold validation;

automatically coding a second portion of the plurality of documents with the initial control set using the predictive coding system;

presenting at least one document from the second portion of the plurality of documents to a human reviewer, the at least one document having automated coding from the predictive coding system;

allowing the human reviewer to correct at least a portion of the automated coding by performing a hard coding correction, the hard coding correction comprising changing of the at least a portion of the automated coding from a first coding to a second coding;

receiving the hard coding correction to the at least a portion of the automated coding of the at least one document;

updating the initial control set with the at least one document having the hard coding correction; and

applying the updated coded control set to additional documents to automatically code the additional documents.

14. The method according to claim 13 , further comprising applying confidence threshold validation testing on one or more of the plurality of documents.

15. The method according to claim 14 , wherein applying confidence threshold validation testing comprises: setting a size of a quality control (QC) sample set at a size of the initial control set, creating the QC sample set by random sampling from unreviewed documents, and reviewing the QC sample set.

16. The method according to claim 13 , further comprising adding further documents to the plurality of documents on a rolling load basis by:

querying for additional documents that correspond to the hard coding of the plurality of documents using the at least one seed set parameter; and

retrieving the additional documents that correspond to the hard coding of the plurality of documents.

17. The method according to claim 13 , further comprising: predictively coding the additional documents; and adding the additional documents that have been predictively coded to the initial control set.

18. The method according to claim 13 , wherein the first portion of the plurality of documents comprise electronically stored information.

19. The method according to claim 13 , further comprising generating a pre-populated coding form to a reviewer based on the predictive coding the first portion of the plurality of documents.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2026
From: OPEN TEXT HOLDINGS, INC.
To: OPEN TEXT INC.
Reel/Frame 074362/0745 →
MERGER Recorded Jul 20, 2018
From: RECOMMIND, INC.
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 046418/0220 →
CHANGE OF ADDRESS Recorded Jul 13, 2017
From: RECOMMIND, INC.
To: RECOMMIND, INC.
Reel/Frame 043187/0308 →
CHANGE OF ADDRESS Recorded Jun 13, 2017
From: RECOMMIND, INC.
To: RECOMMIND, INC.
Reel/Frame 042801/0885 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 2, 2017
From: PUZICHA, JAN; VRANAS, STEVE
To: RECOMMIND, INC.
Reel/Frame 041161/0335 →
Continuity (5)
Continuation 13848023 · Mar 20, 2013
Continuation 13624854 · Sep 21, 2012
Continuation 13074005 · Mar 28, 2011
Continuation 12787354 · May 25, 2010
Related Publication 20170132530A1 · May 11, 2017
Cited By (5)
US 12,299,051 US 12,306,700 US 12,547,944 US 12,572,857 US 12,664,481