IP Library Granted Patent US 12664481
Granted Patent B2
US 12664481 · App. 17/220,445 · Granted Jun 23, 2026

Systems and methods for predictive coding utilizing confidence levels

Inventors: Jan Puzicha (Bonn, DE); Steve Vranas (Ashburn, VA)
Assignee: Open Text Inc.
G06N20/10G06F16/93G06N5/04G06N5/048G06N7/01G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12664481
App. No.
17/220,445
Granted
Jun 23, 2026
Kind
B2
Abstract

Systems and methods for analyzing documents are provided herein. A plurality of documents and user input are received via a computing device. The user input includes hard coding of a subset of the plurality of documents, based on an identified subject or category. Instructions stored in memory are executed by a processor to generate an initial control set, analyze the initial control set to determine at least one seed set parameter, automatically code a first portion of the plurality of documents based on the initial control set and the seed set parameter associated with the identified subject or category, analyze the first portion of the plurality of documents by applying an adaptive identification cycle, and retrieve a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle test on the first portion of the plurality of documents.

Claims (42)

1 . A system for document review in electronic discovery, comprising:

a processor; and

memory storing instructions that, when executed by the processor, cause the system to perform a set of operations for document review in electronic discovery, the set of operations comprising:

receiving a plurality of documents via a computing device;

generating an initial set based on random sampling of a subset of the plurality of documents;

receiving a user input from a computing device, the user input based on an identified subject or category;

reviewing the initial set and the user input to determine at least one seed set parameter associated with the identified subject or category;

coding a first portion of the plurality of documents, based on the initial set and the at least one seed set parameter associated with the identified subject or category;

analyzing the first portion by applying an adaptive identification cycle based on the initial set and user validation of the coding of the first portion;

receiving a second user input via the computing device, the second user input corresponding to a confidence level;

calculating a statistic regarding machine-only coding accuracy rate;

comparing the statistic regarding the machine-only coding accuracy rate against the second user input based on a defined confidence interval; and

retrieving a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle on the first portion.

2 . The system of claim 1 , further comprising coding the second portion of the plurality of documents resulting from an application of user analysis and the adaptive identification cycle.

3 . The system of claim 2 , further comprising adding the coded second portion of the plurality of documents to the initial set.

4 . The system of claim 1 , wherein receiving the user input includes a validation of the initial set.

5 . The system of claim 1 , wherein the random sampling is on a static basis, and wherein generating the initial set is further based on random sampling of the subset of the plurality of documents on a rolling load basis.

6 . The system of claim 1 , wherein receiving the user input from the computing device comprises a designation corresponding to key documents of the initial set.

7 . The system of claim 1 , wherein the adaptive identification cycle is further based on confidence threshold validation.

8 . The system of claim 7 , further comprising retrieving a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle on the first portion of the plurality of documents.

9 . The system of claim 1 , wherein coding the first portion of the plurality of documents further comprises coding based on probabilistic latent semantic analysis and support vector machine analysis of the first portion of the plurality of documents.

10 . The system of claim 1 , wherein the confidence level determines a likelihood that the defined confidence interval contains a parameter.

11 . A computer product storing instructions that, when executed by a processor, are capable of performing a method for document review in electronic discovery, the method comprising:

receiving a plurality of documents via a computing device;

generating an initial set based on random sampling of a subset of the plurality of documents;

receiving a user input from a computing device, the user input based on an identified subject or category;

reviewing the initial set and the user input to determine at least one seed set parameter associated with the identified subject or category;

coding a first portion of the plurality of documents, based on the initial set and the at least one seed set parameter associated with the identified subject or category;

analyzing the first portion by applying an adaptive identification cycle based on the initial set and user validation of the coding of the first portion;

receiving a second user input via the computing device, the second user input corresponding to a confidence level;

calculating a statistic regarding machine-only coding accuracy rate;

comparing the statistic regarding the machine-only coding accuracy rate against the second user input based on a defined confidence interval; and

retrieving a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle on the first portion.

12 . The computer product of claim 11 , further comprising coding the second portion of the plurality of documents resulting from an application of user analysis and the adaptive identification cycle.

13 . The computer product of claim 12 , further comprising adding the coded second portion of the plurality of documents to the initial set.

14 . The computer product of claim 11 , wherein receiving the user input includes a validation of the initial set.

15 . The computer product of claim 11 , wherein the random sampling is on a static basis, and wherein generating the initial set is further based on random sampling of the subset of the plurality of documents on a rolling load basis.

16 . The computer product of claim 11 , wherein receiving the user input from the computing device comprises a designation corresponding to key documents of the initial set.

17 . The computer product of claim 11 , wherein the adaptive identification cycle is further based on confidence threshold validation.

18 . The computer product of claim 17 , further comprising retrieving a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle on the first portion of the plurality of documents.

19 . The computer product of claim 11 , wherein the coding the first portion of the plurality of documents further comprises coding based on probabilistic latent semantic analysis and support vector machine analysis of the first portion of the plurality of documents.

20 . The computer product of claim 11 , wherein the confidence level determines a likelihood that the defined confidence interval contains a parameter.