IP Library Granted Patent US 12,547,944
Granted Patent B2
US 12,547,944 · App. 17/222,283 · Granted Feb 10, 2026

Systems and methods for predictive coding analysis for electronic discovery

Inventors: Jan Puzicha (Bonn, DE); Steve Vranas (Ashburn, VA)
Assignee: Open Text Holdings, Inc.
G06N20/10G06F16/93G06N5/04G06N5/048G06N7/01G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,547,944
App. No.
17/222,283
Filed
Apr 5, 2021
Granted
Feb 10, 2026
Kind
B2
Art Unit
2126
USPC
706/12
Abstract

Systems and methods for analyzing documents are provided herein. A plurality of documents and user input are received via a computing device. The user input includes hard coding of a subset of the plurality of documents, based on an identified subject or category. Instructions stored in memory are executed by a processor to generate an initial control set, analyze the initial control set to determine at least one seed set parameter, automatically code a first portion of the plurality of documents based on the initial control set and the seed set parameter associated with the identified subject or category, analyze the first portion of the plurality of documents by applying an adaptive identification cycle, and retrieve a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle test on the first portion of the plurality of documents.

Claims (59)

1 . A system for analyzing a plurality of documents for electronic discovery using an adaptive identification cycle, comprising:

a processor for training the adaptive identification cycle; and

memory storing instructions that, when executed by the processor, cause the system to perform a set of training operations for training the adaptive identification cycle, the set of training operations comprising:

receiving the plurality of documents via a computing device to be used for training the adaptive identification cycle;

filtering the plurality of documents to produce a subset of the plurality of documents used in training the adaptive identification cycle;

generating an initial control set based on random sampling of the subset of the plurality of documents;

receiving user input corresponding to an identified subject or category;

reviewing the initial control set to determine at least one seed set parameter associated with the identified subject or category;

soft coding a first portion of the plurality of documents, based on the initial control set and the at least one seed set parameter associated with the identified subject or category;

analyzing the first portion of the plurality of documents based on at least one of: the initial control set, user validation of the soft coding of the first portion of the plurality of documents, and a confidence threshold validation, the analyzing the first portion of the plurality of documents comprising:

soft coding a second portion of the plurality of documents with the initial control set;

presenting at least one document from the second portion of the plurality of documents to a human reviewer and allowing the human reviewer to correct at least a portion of the soft coding by performing a hard coding correction, the hard coding correction comprising changing of the at least a portion of the soft coding from the soft coding to a hard coding;

receiving the hard coding correction to the at least a portion of the soft coding of the at least one document;

updating the initial control set with the at least one document having the hard coding correction; and

applying the updated coded control set to additional documents to automatically code the additional documents; and

validating the plurality of documents by determining that the plurality of documents meets the confidence threshold.

2 . The system according to claim 1 , wherein the plurality of documents are partially coded.

3 . The system according to claim 1 , further comprising:

querying for the additional documents that correspond to the hard coding of the plurality of documents using the at least one seed set parameter; and

retrieving the additional documents that correspond to the hard coding of the plurality of documents.

4 . The system according to claim 3 , further comprising generating a confidence score for the additional documents that comprises a computer-generated judgment on soft coding of the additional documents.

5 . The system according to claim 4 , further comprising updating the initial set with the additional documents that correspond to the hard coding of the plurality of documents to create a new seed set.

6 . The system according to claim 5 , further comprising:

identifying non-relevant documents as a set of negative examples;

enriching the set of negative examples to a given total of documents by randomly sampling the non-relevant documents; and

reviewing the non-relevant documents that correspond to the identified subject or category.

7 . The system according to claim 6 , further comprising analyzing the non-relevant documents using a support vector machine that maps the non-relevant documents into an internal high-dimensional representation.

8 . The system according to claim 7 , further comprising computing linear functions on the internal high-dimensional representation to model training examples.

9 . The system according to claim 1 , further comprising retrieving the second portion of the plurality of documents based on a result of application of the adaptive identification cycle on the first portion of the plurality of documents.

10 . The system according to claim 1 , wherein generating the initial set comprises receiving a request for content, the initial set being selected based on the request for content.

11 . A computer product storing instructions that, when executed by a processor, are capable of performing a method of training an adaptive identification cycle for document review in electronic discovery, the method of training an adaptive identification cycle comprising:

receiving a plurality of documents via a computing device to be used for training the adaptive identification cycle;

filtering the plurality of documents to produce a subset of the plurality of documents used in training the adaptive identification cycle;

generating an initial control set based on random sampling of the subset of the plurality of documents;

receiving user input corresponding to an identified subject or category;

reviewing the initial control set to determine at least one seed set parameter associated with the identified subject or category;

soft coding a first portion of the plurality of documents, based on the initial control set and the at least one seed set parameter associated with the identified subject or category;

analyzing the first portion of the plurality of documents based on at least one of: the initial control set, user validation of the soft coding of the first portion of the plurality of documents, and a confidence threshold validation, the analyzing the first portion of the plurality of documents comprising:

soft coding a second portion of the plurality of documents with the initial control set;

presenting at least one document from the second portion of the plurality of documents to a human reviewer and allowing the human reviewer to correct at least a portion of the soft coding by performing a hard coding correction, the hard coding correction comprising changing of the at least a portion of the soft coding from the soft coding to a hard coding;

receiving the hard coding correction to the at least a portion of the soft coding of the at least one document;

updating the initial control set with the at least one document having the hard coding correction; and

applying the updated coded control set to additional documents to automatically code the additional documents; and

validating the plurality of documents by determining that the plurality of documents meets the confidence threshold.

12 . The computer product according to claim 11 , wherein the plurality of documents are partially coded.

13 . The computer product according to claim 11 , further comprising:

querying for the additional documents that correspond to the hard coding of the plurality of documents using the at least one seed set parameter; and

retrieving the additional documents that correspond to the hard coding of the plurality of documents.

14 . The computer product according to claim 13 , further comprising generating a confidence score for the additional documents that comprises a computer-generated judgment on soft coding of the additional documents.

15 . The computer product according to claim 14 , further comprising updating the initial set with the additional documents that correspond to the hard coding of the plurality of documents to create a new seed set.

16 . The computer product according to claim 15 , further comprising:

identifying non-relevant documents as a set of negative examples;

enriching the set of negative examples to a given total of documents by randomly sampling the non-relevant documents; and

reviewing the non-relevant documents that correspond to the identified subject or category.

17 . The computer product according to claim 16 , further comprising analyzing the non-relevant documents using a support vector machine that maps the non-relevant documents into an internal high-dimensional representation.

18 . The computer product according to claim 17 , further comprising computing linear functions on the internal high-dimensional representation to model training examples.

19 . The computer product according to claim 11 , further comprising retrieving the second portion of the plurality of documents based on a result of application of the adaptive identification cycle on the first portion of the plurality of documents.

20 . The computer product according to claim 11 , wherein generating the initial set comprises receiving a request for content, the initial set being selected based on the request for content.

21 . The system according to claim 1 , wherein the confidence threshold is based on a confidence score of the plurality of documents, with the confidence score relating to one or more of relevancy, responsiveness, and privileged nature of the plurality of documents.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2026
From: OPEN TEXT HOLDINGS, INC.
To: OPEN TEXT INC.
Reel/Frame 074362/0745 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2021
From: PUZICHA, JAN; VRANAS, STEVE
To: RECOMMIND, INC.
Reel/Frame 055934/0864 →
MERGER Recorded Apr 15, 2021
From: RECOMMIND, INC.
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 055935/0158 →
Continuity (6)
Continuation 15406542 · Jan 13, 2017
Continuation 13848023 · Mar 20, 2013
Continuation 13624854 · Sep 21, 2012
Continuation 13074005 · Mar 28, 2011
Continuation 12787354 · May 25, 2010
Related Publication 20210224694A1 · Jul 22, 2021
References Cited (128)
US 4839853A · Deerwester et al. · 1989 [cited by applicant]
US 6687696B2 · Hofmann et al. · 2004 [cited by applicant]
US 7051017B2 · Marchisio · 2006 [cited by applicant]
US 7089238B1 · Davis et al. · 2006 [cited by applicant]
US 7107266B1 · Breyman et al. · 2006 [cited by applicant]
US 7328216B2 · Hofmann et al. · 2008 [cited by applicant]
US 7376635B1 · Porcari et al. · 2008 [cited by applicant]
US 7428541B2 · Houle · 2008 [cited by applicant]
US 7454407B2 · Chaudhuri et al. · 2008 [cited by applicant]
US 7519589B2 · Charnock et al. · 2009 [cited by applicant]
US 7558778B2 · Carus et al. · 2009 [cited by applicant]
US 7657522B1 · Puzicha et al. · 2010 [cited by applicant]
US 7933859B1 · Puzicha et al. · 2011 [cited by applicant]
US 7945600B1 · Thomas et al. · 2011 [cited by applicant]
US 8015124B2 · Milo · 2011 [cited by applicant]
US 8196030B1 · Wang et al. · 2012 [cited by applicant]
US 8250008B1 · Cao et al. · 2012 [cited by applicant]
US 8296309B2 · Brassil et al. · 2012 [cited by applicant]
US 8433705B1 · Dredze et al. · 2013 [cited by applicant]
US 8489538B1 · Puzicha et al. · 2013 [cited by applicant]
US 8527523B1 · Ravid · 2013 [cited by applicant]
US 8554716B1 · Puzicha et al. · 2013 [cited by applicant]
US 8577866B1 · Osinga et al. · 2013 [cited by applicant]
US 8620842B1 · Cormack · 2013 [cited by applicant]
US 9058327B1 · Lehrman et al. · 2015 [cited by applicant]
US 9223858B1 · Gummaregula et al. · 2015 [cited by applicant]
US 9269053B2 · Naslund et al. · 2016 [cited by applicant]
US 9558265B1 · Tacchi et al. · 2017 [cited by applicant]
US 9595005B1 · Puzicha et al. · 2017 [cited by applicant]
US 9607272B1 · Yu · 2017 [cited by applicant]
US 9785634B2 · Puzicha · 2017 [cited by applicant]
US 10062039B1 · Lockett · 2018 [cited by applicant]
US 10691760B2 · Pattabiraman et al. · 2020 [cited by applicant]
US 10902066B2 · Puzicha et al. · 2021 [cited by applicant]
US 11023828B2 · Puzicha et al. · 2021 [cited by applicant]
US 11282000B2 · Puzicha et al. · 2022 [cited by applicant]
US 12299051B2 · Puzicha et al. · 2025 [cited by applicant]
US 20010037324A1 · Agrawal et al. · 2001 [cited by applicant]
US 20020032564A1 · Ehsani et al. · 2002 [cited by applicant]
US 20020080170A1 · Goldberg et al. · 2002 [cited by applicant]
US 20020164070A1 · Kuhner et al. · 2002 [cited by applicant]
US 20030120653A1 · Brady et al. · 2003 [cited by applicant]
US 20030135818A1 · Goodwin et al. · 2003 [cited by applicant]
US 20040167877A1 · Thompson, III · 2004 [cited by applicant]
US 20040210834A1 · Duncan et al. · 2004 [cited by applicant]
US 20050021397A1 · Cui et al. · 2005 [cited by applicant]
US 20050027664A1 · Johnson et al. · 2005 [cited by applicant]
US 20050262039A1 · Kreulen et al. · 2005 [cited by applicant]
US 20060020571A1 · Patterson · 2006 [cited by applicant]
US 20060161423A1 · Scott et al. · 2006 [cited by applicant]
US 20060242190A1 · Wnek · 2006 [cited by applicant]
US 20060259475A1 · Dehlinger · 2006 [cited by applicant]
US 20060294101A1 · Wnek · 2006 [cited by applicant]
US 20070226211A1 · Heinze et al. · 2007 [cited by applicant]
US 20080069456A1 · Perronnin · 2008 [cited by applicant]
US 20080086433A1 · Schmidtler · 2008 [cited by examiner]
US 20090012984A1 · Ravid et al. · 2009 [cited by applicant]
US 20090043797A1 · Dorie et al. · 2009 [cited by applicant]
US 20090083200A1 · Pollara et al. · 2009 [cited by applicant]
US 20090106239A1 · Getner et al. · 2009 [cited by applicant]
US 20090119343A1 · Jiao et al. · 2009 [cited by applicant]
US 20090164416A1 · Guha · 2009 [cited by applicant]
US 20090306933A1 · Chan et al. · 2009 [cited by applicant]
US 20100014762A1 · Renders et al. · 2010 [cited by applicant]
US 20100030798A1 · Kumar et al. · 2010 [cited by applicant]
US 20100097634A1 · Meyers et al. · 2010 [cited by applicant]
US 20100118025A1 · Smith et al. · 2010 [cited by applicant]
US 20100250474A1 · Richards et al. · 2010 [cited by applicant]
US 20100250541A1 · Richards et al. · 2010 [cited by applicant]
US 20100257127A1 · Owens · 2010 [cited by examiner]
US 20100293117A1 · Xu · 2010 [cited by applicant]
US 20100312725A1 · Privault et al. · 2010 [cited by applicant]
US 20100325102A1 · Maze · 2010 [cited by applicant]
US 20110023034A1 · Nelson et al. · 2011 [cited by applicant]
US 20110029536A1 · Knight et al. · 2011 [cited by applicant]
US 20110047156A1 · Knight et al. · 2011 [cited by applicant]
US 20110135209A1 · Oba · 2011 [cited by applicant]
US 20120101965A1 · Hennig et al. · 2012 [cited by applicant]
US 20120191708A1 · Barsony et al. · 2012 [cited by applicant]
US 20120278266A1 · Naslund et al. · 2012 [cited by applicant]
US 20120296891A1 · Rangan · 2012 [cited by applicant]
US 20120310930A1 · Kumar et al. · 2012 [cited by applicant]
US 20120310935A1 · Puzicha · 2012 [cited by applicant]
US 20130006996A1 · Kadarkarai · 2013 [cited by applicant]
US 20130124552A1 · Stevenson et al. · 2013 [cited by applicant]
US 20130132394A1 · Puzicha · 2013 [cited by applicant]
US 20140059038A1 · McPherson et al. · 2014 [cited by applicant]
US 20140059069A1 · Taft et al. · 2014 [cited by applicant]
US 20140156567A1 · Scholtes · 2014 [cited by applicant]
US 20140207786A1 · Tal-Rothschild et al. · 2014 [cited by applicant]
US 20140310588A1 · Bhogal et al. · 2014 [cited by applicant]
US 20150347576A1 · Endert et al. · 2015 [cited by applicant]
US 20160019282A1 · Lewis et al. · 2016 [cited by applicant]
US 20160110826A1 · Morimoto et al. · 2016 [cited by applicant]
US 20170132530A1 · Puzicha et al. · 2017 [cited by applicant]
US 20170270115A1 · Cormack et al. · 2017 [cited by applicant]
US 20170322931A1 · Puzicha · 2017 [cited by applicant]
US 20180121831A1 · Puzicha et al. · 2018 [cited by applicant]
US 20180341875A1 · Carr · 2018 [cited by applicant]
US 20190138615A1 · Huh et al. · 2019 [cited by applicant]
US 20190205400A1 · Puzicha · 2019 [cited by applicant]
US 20190325031A1 · Puzicha · 2019 [cited by applicant]
US 20200005218A1 · Cheung et al. · 2020 [cited by applicant]
US 20200026768A1 · Puzicha et al. · 2020 [cited by applicant]
US 20210133255A1 · Puzicha et al. · 2021 [cited by applicant]
US 20210216915A1 · Puzicha et al. · 2021 [cited by applicant]
US 20210224693A1 · Puzicha et al. · 2021 [cited by applicant]
US 20220036244A1 · Puzicha et al. · 2022 [cited by applicant]
US 20220188708A1 · Puzicha et al. · 2022 [cited by applicant]
US 20250156485A1 · Puzicha et al. · 2025 [cited by applicant]
EP 2718803A1 · 2014 [cited by applicant]
WO WO2012170048A1 · 2012 [cited by applicant]
Machine Learning in Automated Text Categorization Fabrizio Sebastiani ACM Computing Surveys, vol. 34, No. 1, Mar. 2002, pp. 1-47 (Year: 2002). [cited by examiner]
Joachims, Thorsten, “Transductive Inference for Text Classification Using Support Vector Machines”, Proceedings of the Sixteenth International Conference on Machine Learning, 1999, 10 pages. [cited by applicant]
Webber et al. “Assessor Error in Stratified Evaluation,” Proceedings of the 19th ACM International Conference on Information and Knowledge Management, 2010. p. 539-548. [Accessed Jun. 2, 2011—ACM Digital Library] http:/… [cited by applicant]
Webber et al. “Score Adjustment for Correction of Pooling Bias,” Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, 2009. p. 444-451. [Accessed Jun. 2, 2011—… [cited by applicant]
Buckley et al. “Bias and the Limits of Pooling for Large Collections,” Journal of Information Retrieval, Dec. 2007. vol. 10, No. 6, pp. 1-16 [Accessed Jun. 2, 2011—Google, via ACM Digital Library] http://www.cs.umbc.edu… [cited by applicant]
Carpenter, “E-Discovery: Predictive Tagging To Reduce Cost and Error”, The Metropolitan Corporate Counsel, 2009, p. 40. [cited by applicant]
Zad et al. “Collaborative Movie Annotation”, Handbook of Multimedia for Digital Entertainment and Arts, 2009, pp. 265-288. [cited by applicant]
“Extended European Search Report”, European Patent Application No. 11867283.1, Feb. 24, 2015, 6 pages. [cited by applicant]
“Axcelerate 5 Case Manager Guide”, Recommind, Inc., [online], 2016 [retrieved Aug. 12, 2020], retrieved from the Internet: <URL: http://axcelerate-docs.opentext.com/help/axc-main/5.15/en_us/content/resources/pdf%20guide… [cited by applicant]
“Axcelerate 5 Reviewer User Guide”, Recommind, Inc., [online], 2016 [retrieved Aug. 12, 2020], retrieved from the Internet: <URL: http://axcelerate-docs.opentext.com/help/axc-main/5.15/en_us/content/resources/pdf%20guid… [cited by applicant]
“Discovery-Assistant—Near Duplicates”, ImageMAKER Development Inc. [online], 2010, [retrieved Aug. 12, 2020], retrieved from the Internet: <URL:www.discovery-assistant.com > Download > Near-Duplicates.pdf>, 14 pages. [cited by applicant]
Doherty, Sean, “Recornrnind's Axcelerate: An E-Discovery Speedway”, Legal Technology News, Sep. 20, 2014, 3 pages. [cited by applicant]
YouTube, “Introduction to Axcelerate 5”, OpenText Discovery, [online], uploaded Apr. 17, 2014, [retrieved Jul. 10, 2020], retrieved from the Internet: <URL:www.youtube.com/watch?v=KBzbZL9Uxyw>, 41 pages. [cited by applicant]
“Axcelerate 5.9.0 Release Notes”, Recommind, Inc., [online], Aug. 17, 2016 [retrieved Aug. 12, 2020], retrieved from the Internet: <URL: http://axcelerate-docs.opentext.com/help/axc-user/5.9/en_us/content/resources/pdf%… [cited by applicant]
“Axcelerate 5.7.2 Release Notes”, Recommind, Inc., [online], Mar. 3, 2016 [retrieved Aug. 12, 2020], retrieved from the Internet: <URL: http://axcelerate-docs.opentext.com/help/axc-user/5.9/en_us/content/resources/pdf%2… [cited by applicant]
Hofmann, “Unsupervised Learning by Probabilistic Latent Semantic Analysis,” Machine Learning, vol. 42, 2001, pp. 177-196. [cited by applicant]
Cited By (1)
US 12,664,481