IP Library › Granted Patent US 12,725,091
Granted Patent B2
US 12,725,091 · App. 17/684,186 · Granted Sep 1, 2026

Systems and methods for updating predictive coding for document category review

Inventors: Jan Puzicha (Bonn, DE); Steve Vranas (Ashburn, VA)
Assignee: Open Text Inc.
G06N20/10G06F16/93G06N5/04G06N5/048G06N7/01G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,091
App. No.
17/684,186
Filed
Mar 1, 2022
Granted
Sep 1, 2026
Kind
B2
Art Unit
2126
USPC
706/52
Abstract

Systems and methods for analyzing documents are provided herein. A plurality of documents and user input are received via a computing device. The user input includes hard coding of a subset of the plurality of documents, based on an identified subject or category. Instructions stored in memory are executed by a processor to generate an initial control set, analyze the initial control set to determine at least one seed set parameter, automatically code a first portion of the plurality of documents based on the initial control set and the seed set parameter associated with the identified subject or category, analyze the first portion of the plurality of documents by applying an adaptive identification cycle, and retrieve a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle test on the first portion of the plurality of documents.

Claims (55)

1 . A method of predictive coding for document category review comprising:

receiving a coding category that includes one of responsiveness, issues, and privileges;

receiving training documents that comprise an initial set of review documents;

generating a coded control set of data based on the training documents using coding determinations for the coded control set, the coding determinations performed using a predictive coding of the initial set of review documents;

training machine learning for the coding category utilizing an adaptive identification cycle trained with the initial set of review documents;

coding additional review documents using the coded control set;

presenting a subject document from the additional review documents to a human reviewer;

receiving a correction from the human reviewer to correct at least a portion of the predictive coding by performing a hard coding correction on the subject document, the hard coding correction comprising updating the predictive coding;

receiving the hard coding correction and updating the coded control set with the subject document having the hard coding correction;

applying the updated coded control set to another set of additional documents to predictively code the another set of additional documents;

identifying similar documents to the initial set of review documents utilizing the machine learning to detect concepts within the initial set of review documents;

updating the machine learning by utilizing the adaptive identification cycle with the additional review documents until a confidence threshold validation is met, the confidence threshold validation calculated using the correction from the human reviewer of the subject document and the identifying of the similar documents by the machine learning; and

supplementing the initial set of review documents with the similar documents.

2 . The method according to claim 1 , further comprising applying the coding determinations of the coded control set to contextually similar data in a corpus of documents.

3 . The method according to claim 1 , further comprising receiving a review of each document in the updated coded control set to ensure that proper coding has been implemented.

4 . The method according to claim 1 , wherein generating the coded control set of data comprises generating the coded control set of data based on a plurality of uncoded documents from a corpus, the plurality of uncoded documents comprising documents that are relevant or selected randomly.

5 . The method according to claim 4 , wherein a portion of the initial set of review documents are randomly sampled from an un-reviewed document population.

6 . The method according to claim 1 , further comprising receiving a determination if the subject document coded using the coded control set was miscoded, prior to the step of receiving the hard coding correction to the subject document.

7 . The method according to claim 1 , further comprising updating the training documents with each new coded document created using the coded control set.

8 . A system for predictive coding for document category review, the system comprising:

a processor; and

a memory for storing instructions, the processor executing the instructions to:

receive a coding category that includes one of responsiveness, issues, and privileges;

receive training documents that comprise an initial set of review documents;

generate a coded control set of data based on the training documents using coding determinations for the coded control set, the coding determinations performed using a predictive coding of the initial set of review documents;

train machine learning for the coding category utilizing an adaptive identification cycle trained with the initial set of review documents;

code additional review documents using the coded control set;

present a subject document from the additional review documents to a human reviewer;

receive a correction from the human reviewer to correct at least a portion of the predictive coding by performing a hard coding correction on the subject document, the hard coding correction comprising updating the predictive coding;

receive the hard coding correction and updating the coded control set with the subject document having the hard coding correction;

apply the updated coded control set to another set of additional documents to predictively code the another set of additional documents;

identify similar documents to the initial set of review documents utilizing the machine learning to detect concepts within the initial set of review documents;

update the machine learning by utilizing the adaptive identification cycle with the additional review documents until a confidence threshold validation is met, the confidence threshold validation calculated using the correction from the human reviewer of the subject document and the identifying of the similar documents by the machine learning; and

supplement the initial set of review documents with the similar documents.

9 . The system according to claim 8 , further comprising applying the coding determinations of the coded control set to contextually similar data in a corpus of documents.

10 . The system according to claim 8 , wherein the processor is configured to receive a review of each document in the updated coded control set to ensure that proper coding has been implemented.

11 . The system according to claim 8 , wherein the processor is configured to generate the coded control set of data by generating the coded control set of data based on a plurality of uncoded documents from a corpus, the plurality of uncoded documents comprising documents that are relevant or selected randomly.

12 . The system according to claim 11 , wherein a portion of the initial set of review documents are randomly sampled from an un-reviewed document population.

13 . The system according to claim 8 , wherein the processor is configured to receive a determination if the subject document coded using the coded control set was miscoded, prior to the step of receiving the hard coding correction to the subject document.

14 . The system according to claim 8 , wherein the processor is configured to update the training documents with each new coded document created using the coded control set.

15 . A method for document category review, the method comprising:

receiving a coding category that includes one of responsiveness, issues, and privileges;

generating a coded control set of data based on received training documents using coding determinations for the coded control set, the coding determinations performed using a predictive coding of an initial set of review documents;

training machine learning for the coding category utilizing an adaptive identification cycle trained with the initial set of review documents;

coding additional review documents using the coded control set;

presenting a subject document from additional review documents to a human reviewer, the additional review documents being generated from the coded control set;

receiving a correction from the human reviewer to correct at least a portion of the predictive coding, the correction comprising a hard coding correction on the subject document, the hard coding correction being used to update the predictive coding by updating the coded control set with the subject document having the hard coding correction;

applying the updated coded control set to another set of additional documents to predictively code the another set of additional documents;

identifying similar documents to the initial set of review documents utilizing the machine learning to detect concepts within the initial set of review documents; and

updating the machine learning by utilizing the adaptive identification cycle with the additional review documents until a confidence threshold validation is met, the confidence threshold validation calculated using the correction from the human reviewer of the subject document and the identifying of the similar documents by the machine learning.

16 . The method according to claim 15 , further comprising receiving a determination if the subject document coded using the coded control set was miscoded, prior to the step of receiving the correction.

17 . The method according to claim 16 , further comprising supplementing the initial set of review documents with the similar documents.

18 . The method according to claim 15 , wherein the additional review documents are selected based on comparison of documents with an initial set of relevant documents.

19 . The method according to claim 15 , further comprising automatically coding a document using the coded control set of data.

20 . The method according to claim 19 , further comprising providing the document to a reviewer if the document is incorrectly coded, prior to the step of receiving a hard coding correction to the document.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYING PARTY EXECUTION DATE PREVIOUSLY RECORDED AT REEL: 59228 FRAME: 990. ASSIGNOR(S) HEREBY CONFIRMS THE MERGER . Recorded Jul 15, 2026
From: RECOMMIND, INC.
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 076011/0622 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2026
From: OPEN TEXT HOLDINGS, INC.
To: OPEN TEXT INC.
Reel/Frame 074362/0745 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2022
From: PUZICHA, JAN; VRANAS, STEVE
To: RECOMMIND, INC.
Reel/Frame 059228/0382 →
MERGER Recorded Mar 10, 2022
From: RECOMMIND, INC.
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 059228/0990 →
Continuity (7)
Continuation 15851645 · Dec 21, 2017
Continuation 15406542 · Jan 13, 2017
Continuation 13848023 · Mar 20, 2013
Continuation 13624854 · Sep 21, 2012
Continuation 13074005 · Mar 28, 2011
Continuation 12787354 · May 25, 2010
Related Publication 20220188708A1 · Jun 16, 2022
References Cited (148)
US 4839853A · Deerwester et al. · 1989 [cited by applicant]
US 6473753B1 · Katariya et al. · 2002 [cited by applicant]
US 6687696B2 · Hofmann et al. · 2004 [cited by applicant]
US 7051017B2 · Marchisio · 2006 [cited by applicant]
US 7089238B1 · Davis et al. · 2006 [cited by applicant]
US 7107266B1 · Breyman et al. · 2006 [cited by applicant]
US 7251637B1 · Caid et al. · 2007 [cited by applicant]
US 7328216B2 · Hofmann et al. · 2008 [cited by applicant]
US 7376635B1 · Porcari et al. · 2008 [cited by applicant]
US 7428541B2 · Houle · 2008 [cited by applicant]
US 7454407B2 · Chaudhuri et al. · 2008 [cited by applicant]
US 7519589B2 · Charnock et al. · 2009 [cited by applicant]
US 7558778B2 · Carus et al. · 2009 [cited by applicant]
US 7657522B1 · Puzicha et al. · 2010 [cited by applicant]
US 7933859B1 · Puzicha et al. · 2011 [cited by applicant]
US 7945600B1 · Thomas et al. · 2011 [cited by applicant]
US 8015124B2 · Milo · 2011 [cited by applicant]
US 8196030B1 · Wang et al. · 2012 [cited by applicant]
US 8250008B1 · Cao et al. · 2012 [cited by applicant]
US 8296309B2 · Brassil et al. · 2012 [cited by applicant]
US 8433705B1 · Dredze et al. · 2013 [cited by applicant]
US 8489538B1 · Puzicha et al. · 2013 [cited by applicant]
US 8527523B1 · Ravid · 2013 [cited by applicant]
US 8533148B1 · Feuersanger et al. · 2013 [cited by applicant]
US 8554716B1 · Puzicha et al. · 2013 [cited by applicant]
US 8577866B1 · Osinga et al. · 2013 [cited by applicant]
US 8620842B1 · Cormack · 2013 [cited by applicant]
US 9058327B1 · Lehrman et al. · 2015 [cited by applicant]
US 9223858B1 · Gummaregula et al. · 2015 [cited by applicant]
US 9269053B2 · Naslund et al. · 2016 [cited by applicant]
US 9558265B1 · Tacchi et al. · 2017 [cited by applicant]
US 9595005B1 · Puzicha et al. · 2017 [cited by applicant]
US 9607272B1 · Yu · 2017 [cited by applicant]
US 9785634B2 · Puzicha · 2017 [cited by applicant]
US 10062039B1 · Lockett · 2018 [cited by applicant]
US 10324936B2 · Feuersänger et al. · 2019 [cited by applicant]
US 10691760B2 · Pattabiraman et al. · 2020 [cited by applicant]
US 10902066B2 · Puzicha et al. · 2021 [cited by applicant]
US 11023828B2 · Puzicha et al. · 2021 [cited by applicant]
US 11282000B2 · Puzicha et al. · 2022 [cited by applicant]
US 12299051B2 · Puzicha et al. · 2025 [cited by applicant]
US 12547944B2 · Puzicha et al. · 2026 [cited by applicant]
US 12572857B2 · Puzicha et al. · 2026 [cited by applicant]
US 12664481B2 · Puzicha · 2026 [cited by applicant]
US 20010037324A1 · Agrawal et al. · 2001 [cited by applicant]
US 20020032564A1 · Ehsani et al. · 2002 [cited by applicant]
US 20020080170A1 · Goldberg et al. · 2002 [cited by applicant]
US 20020164070A1 · Kuhner et al. · 2002 [cited by applicant]
US 20030120653A1 · Brady et al. · 2003 [cited by applicant]
US 20030135818A1 · Goodwin et al. · 2003 [cited by applicant]
US 20040167877A1 · Thompson, III · 2004 [cited by applicant]
US 20040210834A1 · Duncan et al. · 2004 [cited by applicant]
US 20050021397A1 · Cui et al. · 2005 [cited by applicant]
US 20050027664A1 · Johnson et al. · 2005 [cited by applicant]
US 20050262039A1 · Kreulen et al. · 2005 [cited by applicant]
US 20060020571A1 · Patterson · 2006 [cited by applicant]
US 20060161423A1 · Scott et al. · 2006 [cited by applicant]
US 20060242190A1 · Wnek · 2006 [cited by applicant]
US 20060259475A1 · Dehlinger · 2006 [cited by applicant]
US 20060294101A1 · Wnek · 2006 [cited by applicant]
US 20070043774A1 · Davis et al. · 2007 [cited by applicant]
US 20070226211A1 · Heinze et al. · 2007 [cited by applicant]
US 20080069456A1 · Perronnin · 2008 [cited by applicant]
US 20080086433A1 · Schmidtler et al. · 2008 [cited by applicant]
US 20090012984A1 · Ravid et al. · 2009 [cited by applicant]
US 20090024554A1 · Murdock et al. · 2009 [cited by applicant]
US 20090043797A1 · Dorie et al. · 2009 [cited by applicant]
US 20090083200A1 · Pollara et al. · 2009 [cited by applicant]
US 20090106239A1 · Getner et al. · 2009 [cited by applicant]
US 20090119343A1 · Jiao et al. · 2009 [cited by applicant]
US 20090164416A1 · Guha · 2009 [cited by applicant]
US 20090210406A1 · Freire et al. · 2009 [cited by applicant]
US 20090300007A1 · Hiraoka · 2009 [cited by applicant]
US 20090306933A1 · Chan et al. · 2009 [cited by applicant]
US 20100014762A1 · Renders et al. · 2010 [cited by applicant]
US 20100030798A1 · Kumar et al. · 2010 [cited by applicant]
US 20100097634A1 · Meyers et al. · 2010 [cited by applicant]
US 20100118025A1 · Smith et al. · 2010 [cited by applicant]
US 20100145900A1 · Zheng et al. · 2010 [cited by applicant]
US 20100179933A1 · Bai et al. · 2010 [cited by applicant]
US 20100198837A1 · Wu et al. · 2010 [cited by applicant]
US 20100250474A1 · Richards et al. · 2010 [cited by applicant]
US 20100250541A1 · Richards et al. · 2010 [cited by applicant]
US 20100257127A1 · Owens · 2010 [cited by applicant]
US 20100293117A1 · Xu · 2010 [cited by applicant]
US 20100312725A1 · Privault et al. · 2010 [cited by applicant]
US 20100325102A1 · Maze · 2010 [cited by applicant]
US 20110010363A1 · Homma et al. · 2011 [cited by applicant]
US 20110023034A1 · Nelson et al. · 2011 [cited by applicant]
US 20110029536A1 · Knight et al. · 2011 [cited by applicant]
US 20110047156A1 · Knight et al. · 2011 [cited by applicant]
US 20110135209A1 · Oba · 2011 [cited by applicant]
US 20120101965A1 · Hennig et al. · 2012 [cited by applicant]
US 20120191708A1 · Barsony et al. · 2012 [cited by applicant]
US 20120278266A1 · Naslund et al. · 2012 [cited by applicant]
US 20120296891A1 · Rangan · 2012 [cited by applicant]
US 20120310930A1 · Kumar et al. · 2012 [cited by applicant]
US 20120310935A1 · Puzicha · 2012 [cited by applicant]
US 20130006996A1 · Kadarkarai · 2013 [cited by applicant]
US 20130124552A1 · Stevenson et al. · 2013 [cited by applicant]
US 20130132394A1 · Puzicha · 2013 [cited by applicant]
US 20140059038A1 · McPherson et al. · 2014 [cited by applicant]
US 20140059069A1 · Taft et al. · 2014 [cited by applicant]
US 20140095493A1 · Feuersanger et al. · 2014 [cited by applicant]
US 20140156567A1 · Scholtes · 2014 [cited by applicant]
US 20140207786A1 · Tal-Rothschild et al. · 2014 [cited by applicant]
US 20140310588A1 · Bhogal et al. · 2014 [cited by applicant]
US 20150347576A1 · Endert et al. · 2015 [cited by applicant]
US 20160019282A1 · Lewis et al. · 2016 [cited by applicant]
US 20160110826A1 · Morimoto et al. · 2016 [cited by applicant]
US 20170132530A1 · Puzicha et al. · 2017 [cited by applicant]
US 20170270115A1 · Cormack et al. · 2017 [cited by applicant]
US 20170322931A1 · Puzicha · 2017 [cited by applicant]
US 20180121831A1 · Puzicha et al. · 2018 [cited by applicant]
US 20180341875A1 · Carr · 2018 [cited by applicant]
US 20190138615A1 · Huh et al. · 2019 [cited by applicant]
US 20190205400A1 · Puzicha · 2019 [cited by applicant]
US 20190213197A1 · Feuersanger et al. · 2019 [cited by applicant]
US 20190325031A1 · Puzicha · 2019 [cited by applicant]
US 20200005218A1 · Cheung et al. · 2020 [cited by applicant]
US 20200026768A1 · Puzicha et al. · 2020 [cited by applicant]
US 20210133255A1 · Puzicha et al. · 2021 [cited by applicant]
US 20210216915A1 · Puzicha et al. · 2021 [cited by applicant]
US 20210224693A1 · Puzicha et al. · 2021 [cited by applicant]
US 20210224694A1 · Puzicha et al. · 2021 [cited by applicant]
US 20220036244A1 · Puzicha et al. · 2022 [cited by applicant]
US 20250156485A1 · Puzicha et al. · 2025 [cited by applicant]
Roitblat, Herbert L., Anne Kershaw, and Patrick Oot. “Document categorization in legal electronic discovery: computer classification vs. manual review.” Journal of the American Society for Information Science and Techno… [cited by examiner]
Joachims, Thorsten, “Transductive Inference for Text Classification Using Support Vector Machines”, Proceedings of the Sixteenth International Conference on Machine Learning, 1999, 10 pages. [cited by applicant]
Webber et al. “Assessor Error in Stratified Evaluation,” Proceedings of the 19th ACM International Conference on Information and Knowledge Management, 2010. p. 539-548. [Accessed Jun. 2, 2011—ACM Digital Library] http:/… [cited by applicant]
Webber et al. “Score Adjustment for Correction of Pooling Bias,” Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, 2009. p. 444-451. [Accessed Jun. 2, 2011—… [cited by applicant]
Buckley et al. “Bias and the Limits of Pooling for Large Collections,” Journal of Information Retrieval, Dec. 2007. vol. 10, No. 6, pp. 1-16 [Accessed Jun. 2, 2011—Google, via ACM Digital Library] http://www.cs.umbc.edu… [cited by applicant]
Carpenter, “E-Discovery: Predictive Tagging To Reduce Cost and Error”, The Metropolitan Corporate Counsel, 2009, p. 40. [cited by applicant]
Zad et al. “Collaborative Movie Annotation”, Handbook of Multimedia for Digital Entertainment and Arts, 2009, pp. 265-288. [cited by applicant]
“Axcelerate 5 Case Manager Guide”, Recommind, Inc., [online], 2016 [retrieved Aug. 12, 2020], retrieved from the Internet: <URL: http://axcelerate-docs.opentext.com/help/axc-main/5.15/en_us/content/resources/pdf%20guide… [cited by applicant]
“Axcelerate 5 Reviewer User Guide”, Recommind, Inc., [online], 2016 [retrieved Aug. 12, 2020], retrieved from the Internet: <URL: http://axcelerate-docs.opentext.com/help/axc-main/5.15/en_us/content/resources/pdf%20guid… [cited by applicant]
“Discovery-Assistant—Near Duplicates”, ImageMAKER Development Inc. [online], 2010, [retrieved Aug. 12, 2020], retrieved from the Internet: <URL:www.discovery-assistant.com > Download > Near-Duplicates.pdf>, 14 pages. [cited by applicant]
Doherty, Sean, “Recornrnind's Axcelerate: An E-Discovery Speedway”, Legal Technology News, Sep. 20, 2014, 3 pages. [cited by applicant]
YouTube, “Introduction to Axcelerate 5”, OpenText Discovery, [online], uploaded Apr. 17, 2014, [retrieved Jul. 10, 2020], retrieved from the Internet: <URL:www.youtube.com/watch?v=KBzbZL9Uxyw>, 41 pages. [cited by applicant]
“Axcelerate 5.9.0 Release Notes”, Recommind, Inc., [online], Aug. 17, 2016 [retrieved Aug. 12, 2020], retrieved from the Internet: <URL: http://axcelerate-docs.opentext.com/help/axc-user/5.9/en_us/content/resources/pdf%… [cited by applicant]
“Axcelerate 5.7.2 Release Notes”, Recommind, Inc., [online], Mar. 3, 2016 [retrieved Aug. 12, 2020], retrieved from the Internet: <URL: http://axcelerate-docs.opentext.com/help/axo-user/5.9/en_us/content/resources/pdf%2… [cited by applicant]
Onoda, Takashi et al., “Relevance Feedback Document Retrieval Using Support Vector Machines”, 2005, Springer Verlag, AM 2003, LNAI 3430, pp. 59-73. [cited by applicant]
Piaseczny, Wojciech, Adaptive Document Discovery for Vertical Search Engines, Thesis Submitted in Partial Fulfillment of the Requirements for the Degree of Master of Applied Science, Simon Fraser University School of En… [cited by applicant]
Pazzani, Michael et al., “Content-Based Recommendation Systems”, Apr. 24, 2007, The Adaptive Web (Lecture Notes in Computer Science), Springer Berlin Heidelberg, pp. 325-341. [cited by applicant]
Lee, Dik et al., “Document Ranking and the Vector-Space Model,” IEEE Software, Institute of Electrical and Electronics Engineers, US, vol. 14 No. 2, Mar. 1, 1997, pp. 67-75. [cited by applicant]
Sebastiani, F., Machine Learning in Automated Text Categorization:, ACM Computing Surveys, vol. 34, No. 1, Mar. 1, 2002, pp. 1-47. [cited by applicant]
Hofmann, “Unsupervised Learning by Probabilistic Latent Semantic Analysis,” Machine Learning, vol. 42, 2001, pp. 177-196. [cited by applicant]
Bot et al., “A Hybrid Classifier Approach for Web Retrieved Documents Classification” International Conference on Information Technology: Coding and Computing, Proceedings, ITCC 2004, Las Vegas, NV, USA, 2004, vol. 1doi… [cited by applicant]