IP Library Granted Patent US 8,713,023
Granted Patent B1
US 8,713,023 · App. 13/921,046 · Granted Apr 29, 2014

Systems and methods for classifying electronic information using advanced active learning techniques

Inventors: Gordon Villy Cormack (Waterloo, CA); Maura Robin Grossman (New York, NY)
G06F17/30601G06F17/3071G06F17/30648G06F17/30705
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,713,023
App. No.
13/921,046
Granted
Apr 29, 2014
Kind
B1
Abstract

Systems and methods for classifying electronic information or documents into a number of classes and subclasses are provided through an active learning algorithm. Such document classification systems are easily scalable for large document collections, require less manpower and can be employed on a single computer, thus requiring fewer resources. Furthermore, the classification systems and methods described can be used for any pattern recognition or classification effort in a wide variety of fields, including electronic discovery in legal proceedings.

Claims (35)

1. A system for classifying documents in a document collection as relevant or non-relevant in connection with conducting e-discovery in a legal proceeding, the system comprising:

a memory configured to store the document collection;

a computing device coupled to the memory, the computing device comprising:

a display;

a physical input interface;

a processor coupled to the display and the input interface, the processor being configured to:

generate a document information profile for the documents in the collection, each document information profile corresponding to a particular document and representing features and related metadata of that document and no other document;

select a document from the collection to present to a human reviewer;

display a portion of the selected document on the display;

receive, through the input interface, one or more user coding decisions associated with the selected document;

update a classifier using at least one received user coding decision and the document information profile for the document associated with the at least one received user coding decision, wherein the classifier is updated using an incremental learning technique;

compute a set of scores for the documents in the collection by applying the updated classifier to the document information profile associated with each document to be scored;

estimate a number of relevant documents in the document collection by (i) fitting scores computed for documents for which user coding decisions were received to a standard distribution curve, and (ii) calculating an area beneath the curve in order to determine whether review is complete by comparing the estimate to a number of documents in the document collection that the user coded as relevant and that were used to update the classifier;

indicate on the display statistics pertaining to the extent to which review is complete;

in response to determining that review is not complete, repeat the steps of selecting a document, displaying a portion of the selected document, receiving one or more user coding decisions associated with the selected document, updating a classifier, computing a set of scores, and estimating a number of relevant documents; and

classify documents in the document collection as relevant or non-relevant to the legal proceeding using the computed scores or the received user coding decisions.

2. The system of claim 1 , wherein the selected document is one whose score is not within a range of scores associated with a previously selected document.

3. The system of claim 1 , wherein the processor is further configured to receive or provide relevance rankings, wherein the relevance rankings are generated by one or more keyword searching algorithms or by a comparison with one or more exemplary documents.

4. The system of claim 3 , wherein the relevance rankings are generated using more than one technique.

5. The system of claim 4 , wherein an output of the relevance rankings techniques are combined using reciprocal rank fusion.

6. The system of claim 4 , wherein an output of the relevance rankings techniques are combined using stacking.

7. The system of claim 1 , wherein the processor is further configured to manage priority queues for ranking the one or more documents of the document collection.

8. The system of claim 7 , wherein the selected document is one whose score or priority queue ranking is not within a range of scores or rankings associated with a previously selected document.

9. The system of claim 7 , wherein the processor is further configured to receive or provide relevance rankings, wherein the relevance rankings are generated by one or more keyword searching algorithms or by a comparison with one or more exemplary documents.

10. The system of claim 9 , wherein the ranking of the priority queues is derived from a subset of the set of scores and the relevance rankings.

11. The system of claim 1 , wherein the selected document is identified as not being similar to a previously selected document using unsupervised learning techniques.

12. The system of claim 1 , wherein the processor is further configured to pre-process documents from the document collection to reduce dimensionalities of document information profiles.

13. The system of claim 12 , wherein pre-processing the documents includes converting characters of the document to a common case or compressing strings of non-alphanumeric characters to a single character.

14. The system of claim 1 , wherein selecting the document comprises choosing among a set of techniques for selecting the document.

15. The system of claim 14 , wherein the choosing among the set of techniques is prioritized using move-to-front pooling.

16. The system of claim 1 , wherein the standard distribution is a Gaussian distribution.

17. The system of claim 1 , wherein the document information profile of the document to be scored is generated using an N-gram technique.

18. The system of claim 1 , wherein the document information profile of the document to be scored is generated from at least a portion of the contents of that particular document and related metadata.

19. The system of claim 1 , wherein the document information profile of the document to be scored is generated from multiple overlapping portions of the contents of that particular document and related metadata.

20. The system of claim 1 , wherein the incremental learning technique is a gradient ascent or descent technique.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 24, 2025
From: GROSSMAN, MAURA R, MS.; CORMACK, GORDON V., MR.
To: ADAPTIVE CLASSIFICATION TECHNOLOGIES LLC
Reel/Frame 070310/0938 →
Continuity (1)
Continuation 13840029 · Mar 15, 2013