IP Library Granted Patent US 10,229,117
Granted Patent B2
US 10,229,117 · App. 15/186,382 · Granted Mar 12, 2019

Systems and methods for conducting a highly autonomous technology-assisted review classification

Inventors: Gordon V. Cormack (Waterloo, CA); Maura R. Grossman (New York, NY)
G06F17/30011G06F17/30345G06F17/30705G06N99/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,229,117
App. No.
15/186,382
Granted
Mar 12, 2019
Kind
B2
Abstract

Systems and methods for classifying electronic information are provided by way of a Technology-Assisted Review (“TAR”) process, specifically an “Auto-TAR” process that limits discretionary choices in an information classification effort, while still achieving superior results. In certain embodiments, Auto-TAR selects an initial relevant document from a document collection, selects a number of other documents from the document collection and assigns them a default classification, trains a classifier using a training set made up of the selected relevant document and the documents assigned a default classification, scores documents in the document collection and determines if a stopping criteria is met. If a stopping criteria has not been met, the process sorts the documents according to scores, selects a batch of documents from the collection for further review, receives user coding decisions for them, and re-trains a classifier using the received user coding decisions and an adjusted training set.

Claims (61)

1. A system for classifying information, the system comprising:

at least one computing device having a processor and physical memory, the physical memory storing instructions that cause the processor to:

receive an identification of a relevant document;

select a first set of documents from a document collection, wherein the document collection is stored on a non-transitory storage medium;

assign a first set of default classifications to documents in the first set of documents to be used as a training set along with the relevant document;

train a classifier using the training set;

score one or more documents in the document collection using the classifier;

upon determining that a stopping criteria has been reached, classify one or more documents in the document collection using the classifier;

upon determining that a stopping criteria has not been reached, select a second set of documents having a batch size for presenting to a reviewer for review prior to repeating the step of training the classifier;

present one or more documents in the second set of documents to the reviewer;

receive from the reviewer user coding decisions associated with the presented documents;

add one or more of the documents presented to the reviewer for which user coding decisions were received to the training set;

remove one or more documents in the first set of documents from the training set;

add a third set of documents from the document collection to the training set;

assign a second set of default classifications to one or more documents in the third set of documents;

update the classifier using one or more documents in the training set;

increase the batch size of documents selected for the second set of documents; and

repeat the steps of training, scoring and determining whether a stopping criteria has been reached;

wherein the first and second set of default classifications are presumptively assigned classifications used for the purpose of training the classifier in order to form a decision boundary, the presumptively assigned classifications not being based on a review; and wherein the one or more documents in the first set of documents removed from the training set are documents previously assigned a presumptively assigned classification.

2. The system of claim 1 , wherein the number of documents presented for review is increased between iterations.

3. The system of claim 2 , wherein the increase is 10%.

4. The system of claim 2 , wherein the batch size is increased exponentially between iterations.

5. The system of claim 1 , wherein the number of documents presented for review is varied between iterations or is selected to achieve an effectiveness target.

6. The system of claim 1 , wherein the stopping criteria is the exhaustion of the first set of documents.

7. The system of claim 1 , wherein the stopping criteria is a targeted level of recall.

8. The system of claim 1 , wherein the stopping criteria is a targeted level of F 1 , where F 1 is a measure that combines recall and precision.

9. The system of claim 1 , further comprising instructions that cause the processor to sort the scored documents and present to the reviewer the highest scored documents for review.

10. The system of claim 1 , wherein the presumptively assigned classification assigned to documents in the first set or second set of default classifications is “non-relevant”.

11. The system of claim 1 , wherein the documents in the first or third sets are selected randomly.

12. The system of claim 1 , wherein the identified relevant document is a synthetic document.

13. The system of claim 1 , wherein the presumptively assigned classification is a temporary label that is used to train the classifier, and the presumptively assigned classification is not based on actual determination of relevance by the reviewer.

14. A computerized method for classifying information, the method comprising:

receiving an identification of a relevant document;

selecting a first set of documents from a document collection, wherein the document collection is stored on a non-transitory storage medium;

assigning a first set of default classifications to documents in the first set of documents to be used as a training set along with the relevant document;

training a classifier using the training set;

scoring one or more documents in the document collection using the classifier;

upon determining that a stopping criteria has been reached, classifying one or more documents in the document collection using the classifier;

upon determining that a stopping criteria has not been reached, selecting a second set of documents having a batch size for presenting to a reviewer for review prior to repeating the step of training the classifier;

presenting one or more documents in the second set of documents to the reviewer;

receiving from the reviewer user coding decisions associated with the presented documents;

adding one or more of the documents presented to the reviewer for which user coding decisions were received to the training set;

removing one or more documents in the first set of documents from the training set;

adding a third set of documents from the document collection to the training set;

assigning a second set of default classifications to one or more documents in the third set of documents;

updating the classifier using one or more documents in the training set;

increasing the batch size of documents selected for the second set of documents; and

repeating the steps of training, scoring and determining whether a stopping criteria has been reached;

wherein the first and second set of default classifications are presumptively assigned classifications used for the purpose of training the classifier in order to form a decision boundary, the presumptively assigned classifications not being based on a review; and wherein the one or more documents in the first set of documents removed from the training set are documents previously assigned a presumptively assigned classification.

15. The method of claim 14 , wherein the number of documents presented for review is increased between iterations.

16. The method of claim 15 , wherein the increase is 10%.

17. The method of claim 14 , wherein the batch size is increased exponentially between iterations.

18. The method of claim 14 , wherein the number of documents presented for review is varied between iterations or is selected to achieve an effectiveness target.

19. The method of claim 14 , wherein the stopping criteria is the exhaustion of the first set of documents.

20. The method of claim 14 , wherein the stopping criteria is a targeted level of recall.

21. The method of claim 14 , wherein the stopping criteria is a targeted level of F 1 , where F 1 is a measure that combines recall and precision.

22. The method of claim 14 , further comprising sorting the scored documents and presenting to the reviewer the highest scored documents for review.

23. The method of claim 14 , wherein the presumptively assigned classification assigned to documents in the first set or second set of default classifications is “non-relevant”.

24. The method of claim 14 , wherein the documents in the first or third sets are selected randomly.

25. The method of claim 14 , wherein the identified relevant document is a synthetic document.

26. The method of claim 14 , wherein the presumptively assigned classification is a temporary label that is used to train the classifier, and the presumptively assigned classification is not based on actual determination of relevance by the reviewer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 24, 2025
From: GROSSMAN, MAURA R, MS.; CORMACK, GORDON V., MR.
To: ADAPTIVE CLASSIFICATION TECHNOLOGIES LLC
Reel/Frame 070310/0938 →
Continuity (3)
Provisional Application 62182028 · Jun 19, 2015
Provisional Application 62182072 · Jun 19, 2015
Related Publication 20160371261A1 · Dec 22, 2016
Cited By (2)
US 12,572,746 US 12,639,108