IP Library Granted Patent US 12,572,857
Granted Patent B2
US 12,572,857 · App. 17/216,362 · Granted Mar 10, 2026

Adaptive probabilistic latent semantic analysis system for automated document coding and review in electronic discovery

Inventors: Jan Puzicha (Bonn, DE); Steve Vranas (Ashburn, VA)
Assignee: Open Text Inc.
G06N20/10G06F16/93G06N5/04G06N5/048G06N7/01G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,857
App. No.
17/216,362
Filed
Mar 29, 2021
Granted
Mar 10, 2026
Kind
B2
Art Unit
2126
USPC
706/12
Abstract

Systems and methods for analyzing documents are provided herein. A plurality of documents and user input are received via a computing device. The user input includes hard coding of a subset of the plurality of documents, based on an identified subject or category. Instructions stored in memory are executed by a processor to generate an initial control set, analyze the initial control set to determine at least one seed set parameter, automatically code a first portion of the plurality of documents based on the initial control set and the seed set parameter associated with the identified subject or category, analyze the first portion of the plurality of documents by applying an adaptive identification cycle, and retrieve a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle test on the first portion of the plurality of documents.

Claims (61)

1 . A method for document review in electronic discovery, comprising:

receiving from user input, a document review coding category;

generating, by machine assisted automatic coding, based on the document review coding category, an initial control set comprising documents from a corpus of documents;

training probabilistic latent semantic analysis (PLSA) on the initial control set to automatically detect concepts within the documents of the corpus via a statistical analysis of word contexts;

performing the machine assisted automatic coding on the documents of the corpus, by executing the PLSA trained on the initial control set to assign coding determinations to the documents of the corpus according to the document review coding category, based on the PLSA's automatic detection of concepts within the documents of the corpus via the statistical analysis of the word contexts;

generating targeted documents based on an analysis of the initial control set;

reviewing the generated targeted documents based on the machine assisted automatic coding for the initial control set and human-only assisted hard coding of documents;

based on the reviewing of the generated targeted documents, modifying the machine assisted automatic coding and reiterating the generating of the targeted documents and reviewing of the generated targeted documents;

generating a confidence validation by comparing the machine assisted automatic coding to a human-only assisted hard coding of the documents;

based on the confidence validation, determining whether the machine assisted automatic coding is more accurate than the human-only assisted hard coding; and

when the machine assisted automatic coding is not more accurate than the human-only assisted hard coding, evaluating the machine assisted automatic coding with a randomly selected document and modifying the machine assisted automatic coding based on the evaluating.

2 . The method of claim 1 , wherein the document review coding category corresponds to electronic discovery categories comprising responsive, issues, or privilege.

3 . The method of claim 1 , wherein the initial control set is further generated based on the received user input on the documents.

4 . The method of claim 3 , wherein generating the initial control set comprises:

receiving the human-only assisted hard coding of a selected subset of the documents based on the document review coding category; and

analyzing the selected subset of the documents to generate the machine assisted automatic coding.

5 . The method of claim 1 , wherein reviewing the generated targeted documents based on the machine assisted automatic coding for the initial control set and human-only assisted hard coding of the documents comprises:

receiving human assisted validation of the generated targeted documents.

6 . The method of claim 5 , wherein the randomly selected document is a plurality of randomly selected documents used to reseed the initial control set for further evaluation.

7 . The method of claim 1 , wherein when the machine assisted automatic coding is more accurate than the human-only assisted hard coding, generating a selected subset of the documents from the corpus of the documents and corresponding to the document review coding category.

8 . A system for document review in electronic discovery, comprising:

a processor; and

memory storing instructions that, when executed by the processor, causes the system to perform a set of operations, the set of operations comprising:

receiving from user input, a document review coding category;

generating, by machine assisted automatic coding, based on the document review coding category, an initial control set comprising documents from a corpus of documents;

training probabilistic latent semantic analysis (PLSA) on the initial control set to automatically detect concepts within the documents of the corpus via a statistical analysis of word contexts;

performing the machine assisted automatic coding on the documents of the corpus, by executing the PLSA trained on the initial control set to assign coding determinations to the documents of the corpus according to the document review coding category, based on the PLSA's automatic detection of concepts within the documents of the corpus via the statistical analysis of the word contexts;

generating targeted documents based on an analysis of the initial control set;

reviewing the generated targeted documents based on the machine assisted automatic coding for the initial control set and human-only assisted hard coding of documents;

based on the reviewing of the generated targeted documents, modifying the machine assisted automatic coding and reiterating the generating of the targeted documents and reviewing of the generated targeted documents;

generating a confidence validation by comparing the machine assisted automatic coding to the human-only assisted hard coding of the documents;

based on the confidence validation, determining whether the machine assisted automatic coding is more accurate than the human-only assisted hard coding; and

when the machine assisted automatic coding is not more accurate than the human-only assisted hard coding, evaluating the machine assisted automatic coding with a randomly selected document and modifying the machine assisted automatic coding based on the evaluating.

9 . The system of claim 8 , wherein the document review coding category corresponds to electronic discovery categories comprising responsive, issues, or privilege.

10 . The system of claim 8 , wherein the initial control set is further generated based on the received user input on the documents.

11 . The system of claim 10 , wherein generating the initial control set comprises:

receiving the human-only assisted hard coding of a selected subset of the documents based on the document review coding category; and

analyzing the selected subset of the documents to generate the machine assisted automatic coding.

12 . The system of claim 8 , wherein reviewing the generated targeted documents based on the machine assisted automatic coding for the initial control set and human-only assisted hard coding of the documents comprises:

receiving human assisted validation of the generated targeted documents.

13 . The system of claim 12 , wherein the randomly selected document is a plurality of randomly selected documents used to reseed the initial control set for further evaluation.

14 . The system of claim 8 , wherein when the machine assisted automatic coding is more accurate than the human-only assisted hard coding, generating a selected subset of the documents from the corpus of the documents and corresponding to the document review coding category.

15 . A computer product storing instructions that, when executed by a processor, are capable of performing a method for document review in electronic discovery, the method comprising:

receiving from user input, a document review coding category;

generating, by machine assisted automatic coding, based on the document review coding category, an initial control set comprising documents from a corpus of documents;

training probabilistic latent semantic analysis (PLSA) on the initial control set to automatically detect concepts within the documents of the corpus via a statistical analysis of word contexts;

performing the machine assisted automatic coding on the documents of the corpus, by executing the PLSA trained on the initial control set to assign coding determinations to the documents of the corpus according to the document review coding category, based on the PLSA's automatic detection of concepts within the documents of the corpus via the statistical analysis of the word contexts;

generating targeted documents based on an analysis of the initial control set;

reviewing the generated targeted documents based on the machine assisted automatic coding for the initial control set and human-only assisted hard coding of documents;

based on the reviewing of the generated targeted documents, modifying the machine assisted automatic coding and reiterating the generating of the targeted documents and reviewing of the generated targeted documents;

generating a confidence validation by comparing the machine assisted automatic coding to a human-only assisted hard coding of the documents;

based on the confidence validation, determining whether the machine assisted automatic coding is more accurate than the human-only assisted hard coding; and

when the machine assisted automatic coding is not more accurate than the human-only assisted hard coding, evaluating the machine assisted automatic coding with a randomly selected document and modifying the machine assisted automatic coding based on the evaluating.

16 . The computer product of claim 15 , wherein the initial control set is further generated based on the received user input on the documents.

17 . The computer product of claim 16 , wherein generating the initial control set comprises:

receiving the human-only assisted hard coding of a selected subset of the documents based on the document review coding category; and

analyzing the selected subset of the documents to generate the machine assisted automatic coding.

18 . The computer product of claim 15 , wherein reviewing the generated targeted documents based on the machine assisted automatic coding for the initial control set and human-only assisted hard coding of the documents comprises:

receiving human assisted validation of the generated targeted documents.

19 . The computer product of claim 18 , wherein the randomly selected document is a plurality of randomly selected documents used to reseed the initial control set for further evaluation.

20 . The computer product of claim 15 , wherein when the machine assisted automatic coding is more accurate than the human-only assisted hard coding, generating a selected subset of documents from the corpus of documents and corresponding to the document review coding category.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2026
From: OPEN TEXT HOLDINGS, INC.
To: OPEN TEXT INC.
Reel/Frame 074362/0745 →
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 055854 FRAME: 0633. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 26, 2021
From: PUZICHA, JAN; VRANAS, STEVE
To: RECOMMIND, INC.
Reel/Frame 056047/0481 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2021
From: PUZICHA, JAN; VRANAS, STEVE
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 055854/0633 →
MERGER Recorded Apr 7, 2021
From: RECOMMIND, INC.
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 055855/0699 →
Continuity (6)
Continuation 15406542 · Jan 13, 2017
Continuation 13848023 · Mar 20, 2013
Continuation 13624854 · Sep 21, 2012
Continuation 13074005 · Mar 28, 2011
Continuation 12787354 · May 25, 2010
Related Publication 20210216915A1 · Jul 15, 2021
References Cited (129)
US 4839853A · Deerwester et al. · 1989 [cited by applicant]
US 6687696B2 · Hofmann et al. · 2004 [cited by applicant]
US 7051017B2 · Marchisio · 2006 [cited by applicant]
US 7089238B1 · Davis et al. · 2006 [cited by applicant]
US 7107266B1 · Breyman et al. · 2006 [cited by applicant]
US 7328216B2 · Hofmann et al. · 2008 [cited by applicant]
US 7376635B1 · Porcari et al. · 2008 [cited by applicant]
US 7428541B2 · Houle · 2008 [cited by applicant]
US 7454407B2 · Chaudhuri et al. · 2008 [cited by applicant]
US 7519589B2 · Charnock et al. · 2009 [cited by applicant]
US 7558778B2 · Carus et al. · 2009 [cited by applicant]
US 7657522B1 · Puzicha et al. · 2010 [cited by applicant]
US 7933859B1 · Puzicha et al. · 2011 [cited by applicant]
US 7945600B1 · Thomas et al. · 2011 [cited by applicant]
US 8015124B2 · Milo · 2011 [cited by applicant]
US 8196030B1 · Wang et al. · 2012 [cited by applicant]
US 8250008B1 · Cao et al. · 2012 [cited by applicant]
US 8296309B2 · Brassil et al. · 2012 [cited by applicant]
US 8433705B1 · Dredze et al. · 2013 [cited by applicant]
US 8489538B1 · Puzicha et al. · 2013 [cited by applicant]
US 8527523B1 · Ravid · 2013 [cited by applicant]
US 8554716B1 · Puzicha et al. · 2013 [cited by applicant]
US 8577866B1 · Osinga et al. · 2013 [cited by applicant]
US 8620842B1 · Cormack · 2013 [cited by applicant]
US 9058327B1 · Lehrman et al. · 2015 [cited by applicant]
US 9223858B1 · Gummaregula et al. · 2015 [cited by applicant]
US 9269053B2 · Naslund et al. · 2016 [cited by applicant]
US 9558265B1 · Tacchi et al. · 2017 [cited by applicant]
US 9595005B1 · Puzicha et al. · 2017 [cited by applicant]
US 9607272B1 · Yu · 2017 [cited by applicant]
US 9785634B2 · Puzicha · 2017 [cited by applicant]
US 10062039B1 · Lockett · 2018 [cited by applicant]
US 10691760B2 · Pattabiraman et al. · 2020 [cited by applicant]
US 10902066B2 · Puzicha et al. · 2021 [cited by applicant]
US 11023828B2 · Puzicha et al. · 2021 [cited by applicant]
US 11282000B2 · Puzicha et al. · 2022 [cited by applicant]
US 12299051B2 · Puzicha et al. · 2025 [cited by applicant]
US 20010037324A1 · Agrawal et al. · 2001 [cited by applicant]
US 20020032564A1 · Ehsani et al. · 2002 [cited by applicant]
US 20020080170A1 · Goldberg et al. · 2002 [cited by applicant]
US 20020164070A1 · Kuhner et al. · 2002 [cited by applicant]
US 20030120653A1 · Brady et al. · 2003 [cited by applicant]
US 20030135818A1 · Goodwin et al. · 2003 [cited by applicant]
US 20040167877A1 · Thompson, III · 2004 [cited by applicant]
US 20040210834A1 · Duncan et al. · 2004 [cited by applicant]
US 20050021397A1 · Cui et al. · 2005 [cited by applicant]
US 20050027664A1 · Johnson et al. · 2005 [cited by applicant]
US 20050262039A1 · Kreulen et al. · 2005 [cited by applicant]
US 20060020571A1 · Patterson · 2006 [cited by applicant]
US 20060161423A1 · Scott et al. · 2006 [cited by applicant]
US 20060242190A1 · Wnek · 2006 [cited by applicant]
US 20060259475A1 · Dehlinger · 2006 [cited by applicant]
US 20060294101A1 · Wnek · 2006 [cited by applicant]
US 20070226211A1 · Heinze et al. · 2007 [cited by applicant]
US 20080069456A1 · Perronnin · 2008 [cited by examiner]
US 20080086433A1 · Schmidtler et al. · 2008 [cited by applicant]
US 20090012984A1 · Ravid et al. · 2009 [cited by applicant]
US 20090043797A1 · Dorie et al. · 2009 [cited by applicant]
US 20090083200A1 · Pollara et al. · 2009 [cited by applicant]
US 20090106239A1 · Getner et al. · 2009 [cited by applicant]
US 20090119343A1 · Jiao et al. · 2009 [cited by applicant]
US 20090164416A1 · Guha · 2009 [cited by applicant]
US 20090306933A1 · Chan et al. · 2009 [cited by applicant]
US 20100014762A1 · Renders et al. · 2010 [cited by applicant]
US 20100030798A1 · Kumar et al. · 2010 [cited by applicant]
US 20100097634A1 · Meyers et al. · 2010 [cited by applicant]
US 20100118025A1 · Smith et al. · 2010 [cited by applicant]
US 20100250474A1 · Richards et al. · 2010 [cited by applicant]
US 20100250541A1 · Richards et al. · 2010 [cited by applicant]
US 20100257127A1 · Owens · 2010 [cited by examiner]
US 20100293117A1 · Xu · 2010 [cited by applicant]
US 20100312725A1 · Privault et al. · 2010 [cited by applicant]
US 20100325102A1 · Maze · 2010 [cited by applicant]
US 20110023034A1 · Nelson et al. · 2011 [cited by applicant]
US 20110029536A1 · Knight et al. · 2011 [cited by applicant]
US 20110047156A1 · Knight et al. · 2011 [cited by applicant]
US 20110135209A1 · Oba · 2011 [cited by applicant]
US 20120101965A1 · Hennig et al. · 2012 [cited by applicant]
US 20120191708A1 · Barsony et al. · 2012 [cited by applicant]
US 20120278266A1 · Naslund et al. · 2012 [cited by applicant]
US 20120296891A1 · Rangan · 2012 [cited by applicant]
US 20120310930A1 · Kumar et al. · 2012 [cited by applicant]
US 20120310935A1 · Puzicha · 2012 [cited by applicant]
US 20130006996A1 · Kadarkarai · 2013 [cited by applicant]
US 20130124552A1 · Stevenson et al. · 2013 [cited by applicant]
US 20130132394A1 · Puzicha · 2013 [cited by applicant]
US 20140059038A1 · McPherson et al. · 2014 [cited by applicant]
US 20140059069A1 · Taft et al. · 2014 [cited by applicant]
US 20140156567A1 · Scholtes · 2014 [cited by applicant]
US 20140207786A1 · Tal-Rothschild et al. · 2014 [cited by applicant]
US 20140310588A1 · Bhogal et al. · 2014 [cited by applicant]
US 20150347576A1 · Endert et al. · 2015 [cited by applicant]
US 20160019282A1 · Lewis et al. · 2016 [cited by applicant]
US 20160110826A1 · Morimoto et al. · 2016 [cited by applicant]
US 20170132530A1 · Puzicha et al. · 2017 [cited by applicant]
US 20170270115A1 · Cormack et al. · 2017 [cited by applicant]
US 20170322931A1 · Puzicha · 2017 [cited by applicant]
US 20180121831A1 · Puzicha et al. · 2018 [cited by applicant]
US 20180341875A1 · Carr · 2018 [cited by applicant]
US 20190138615A1 · Huh et al. · 2019 [cited by applicant]
US 20190205400A1 · Puzicha · 2019 [cited by applicant]
US 20190325031A1 · Puzicha · 2019 [cited by applicant]
US 20200005218A1 · Cheung et al. · 2020 [cited by applicant]
US 20200026768A1 · Puzicha et al. · 2020 [cited by applicant]
US 20210133255A1 · Puzicha et al. · 2021 [cited by applicant]
US 20210224693A1 · Puzicha et al. · 2021 [cited by applicant]
US 20210224694A1 · Puzicha et al. · 2021 [cited by applicant]
US 20220036244A1 · Puzicha et al. · 2022 [cited by applicant]
US 20220188708A1 · Puzicha et al. · 2022 [cited by applicant]
US 20250156485A1 · Puzicha et al. · 2025 [cited by applicant]
EP 2718803A1 · 2014 [cited by applicant]
WO WO2012170048A1 · 2012 [cited by applicant]
Machine Learning in Automated Text Categorization Fabrizio Sebastiani ACM Computing Surveys, vol. 34, No. 1, Mar. 2002, pp. 1-47. (Year: 2002). [cited by examiner]
“Unsupervised Learning by Probabilistic Latent Semantic Analysis” Thomas Hofmann [email protected] Department of Computer Science, Brown University, Providence, RI 02912, USA (Year: 2001). [cited by examiner]
Joachims, Thorsten, “Transductive Inference for Text Classification Using Support Vector Machines”, Proceedings of the Sixteenth International Conference on Machine Learning, 1999, 10 pages. [cited by applicant]
Webber et al. “Assessor Error in Stratified Evaluation, Proceedings of the 19th ACM International Conference on Information and Knowledge Management,” 2010. p. 539-548. [Accessed Jun. 2, 2011—ACM Digital Library] http:/… [cited by applicant]
Webber et al. “Score Adjustment for Correction of Pooling Bias,” Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, 2009. p. 444-451. [Accessed Jun. 2, 2011—… [cited by applicant]
Buckley et al. “Bias and the Limits of Pooling for Large Collections,” Journal of Information Retrieval, Dec. 2007. vol. 10, No. 6, pp. 1-16 [Accessed Jun. 2, 2011—Google, via ACM Digital Library] http://www.cs.umbc.edu… [cited by applicant]
Carpenter, “E-Discovery: Predictive Tagging to Reduce Cost and Error”, The Metropolitan Corporate Counsel, 2009, p. 40. [cited by applicant]
Zad et al. “Collaborative Movie Annotation, Handbook of Multimedia for Digital Entertainment and Arts”, 2009, pp. 265-288. [cited by applicant]
“Extended European Search Report”, European Patent Application No. 11867283.1, Feb. 24, 2015, 6 pages. [cited by applicant]
“Axcelerate 5 Case Manager Guide”, Recommind, Inc., [online], 2016 [retrieved Aug. 12, 2020], retrieved from the Internet: <URL: http://axcelerate-docs.opentext.com/help/axc-main/5.15/en_us/content/resources/pdf%20guide… [cited by applicant]
“Axcelerate 5 Reviewer User Guide”, Recommind, Inc., [online], 2016 [retrieved Aug. 12, 2020], retrieved from the Internet: <URL: http://axcelerate-docs.opentext.com/help/axc-main/5.15/en_us/content/resources/pdf%20guid… [cited by applicant]
“Discovery-Assistant—Near Duplicates”, ImageMAKER Development Inc. [online], 2010, [retrieved Aug. 12, 2020], retrieved from the Internet: <URL:www.discovery-assistant.com > Download > Near-Duplicates.pdf>, 14 pages. [cited by applicant]
Doherty, Sean, “Recornrnind's Axcelerate: An E-Discovery Speedway”, Legal Technology News, Sep. 20, 2014, 3 pages. [cited by applicant]
YouTube, “Introduction to Axcelerate 5”, OpenText Discovery, [online], uploaded Apr. 17, 2014, [retrieved Jul. 10, 2020], retrieved from the Internet: <URL:www.youtube.com/watch?v=KBzbZL9Uxyw>, 41 pages. [cited by applicant]
“Axcelerate 5.9.0 Release Notes”, Recommind, Inc., [online], Aug. 17, 2016 [retrieved Aug. 12, 2020], retrieved from the Internet: <URL: http://axcelerate-docs.opentext.com/help/axc-user/5.9/en_us/content/resources/pdf%… [cited by applicant]
“Axcelerate 5.7.2 Release Notes”, Recommind, Inc., [online], Mar. 3, 2016 [retrieved Aug. 12, 2020], retrieved from the Internet: <URL: http://axcelerate-docs.opentext.com/help/axc-user/5.9/en_us/content/resources/pdf%2… [cited by applicant]
Roitblat et al., “Document Categorization in Legal Electronic Discovery: Computer Classification vs. Manual Review,” Journal of the American Society for Information Science and Technology, vol. 61, No. 1, Dec. 9, 2009, … [cited by applicant]
Cited By (1)
US 12,664,481