IP Library Granted Patent US 11,080,340
Granted Patent B2
US 11,080,340 · App. 15/607,022 · Granted Aug 3, 2021

Systems and methods for classifying electronic information using advanced active learning techniques

Inventors: Gordon Villy Cormack (Waterloo, CA); Maura Robin Grossman (New York, NY)
G06F16/93G06F16/285G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,080,340
App. No.
15/607,022
Granted
Aug 3, 2021
Kind
B2
Abstract

Systems and methods for classifying electronic information or documents into a number of classes and subclasses are provided through an active learning algorithm. In certain embodiments, seed sets may be eliminated by merging relevance feedback and machine learning phases. In certain embodiments, the active learning algorithm forks a number of classification paths corresponding to predicted user coding decisions for a selected document. The active learning algorithm determines an order in which the documents of the collection may be processed and scored by the forked classification paths. Such document classification systems are easily scalable for large document collections, require less manpower and can be employed on a single computer, thus requiring fewer resources. Furthermore, the classification systems and methods described can be used for any pattern recognition or classification effort in a wide variety of fields.

Claims (36)

1. A method for classifying documents in a document collection as a member of one or more classes or subclasses, the method comprising:

selecting a document from the document collection;

calculating at least two predicted classifiers for a particular class of the one or more classes or subclasses, each predicted classifier being calculated using a document information profile for the selected document, a current classifier associated with the particular classes, and a different coding decision selected from a set of possible user coding decisions to be received from a user, thereby resulting in a plurality of predicted classifiers each one corresponding to a different user coding decision having been applied to update the same current classifier, wherein the possible user coding decisions indicate whether the selected document is a member of the particular class;

determining a processing order for a subset of documents in the document collection that indicates an order in which the documents of the subset are to be scored;

for each one of the predicted classifiers, calculating a set of scores for one or more documents in the document collection, at least in part, according to the processing order, wherein each score is generated for a document by utilizing the corresponding predicted classifier and a document information profile of the document to be scored; and

in response to receiving a user coding decision, identifying and propagating forward the predicted classifier that corresponds to the received user coding decision as the current classifier and terminating or suspending use of at least one predicted classifier that does not correspond to the received user coding decision.

2. The method of claim 1 , further comprising in response to determining whether one or more stopping criteria have been met, classifying a set of documents in the document collection into one or more of the one or more classes or subclasses using the propagated predicted classifier.

3. The method of claim 2 , wherein in order to determine whether one or more stopping criteria have been met, the method comprises:

selecting and presenting documents from the document collection to a user; and

calculating an estimate of system effectiveness using the user coding decisions received for the selected and presented documents.

4. The method of claim 2 , wherein in order to determine whether one or more stopping criteria have been met, the method comprises:

calculating an estimate of the number of documents in the document collection that are relevant to one of the classes or subclasses by fitting calculated scores to a distribution; and

validating the estimate using a sequence of user coding decisions.

5. The method of claim 1 , further comprising receiving or providing relevance rankings, wherein the relevance rankings are generated by one or more keyword searching algorithms or by a comparison with one or more exemplary documents.

6. The method of claim 5 , wherein the processing order is derived, at least in part, from scores calculated using a classifier and the relevance rankings.

7. The method of claim 5 , wherein the document is selected by choosing between utilizing document scores calculated using a classifier, the relevance rankings or a combination of document scores and the relevance rankings.

8. A system for classifying documents in a document collection as a member of one or more classes or subclasses, the system comprising:

a processor adapted to:

select a document from the document collection;

calculate at least two predicted classifiers for a particular class of the one or more classes or subclasses, each predicted classifier being calculated using a document information profile for the selected document, a current classifier associated with the particular, and a different coding decision selected from a set of possible user coding decisions to be received from a user, thereby resulting in a plurality of predicted classifiers each one corresponding to a different user coding decision having been applied to update the same current classifier, wherein the possible user coding decisions indicate whether the selected document is a member of the particular class;

determine a processing order for a subset of documents in the document collection that indicates an order in which the documents of the subset are to be scored;

for each one of the predicted classifiers, calculate a set of scores for one or more documents in the document collection, at least in part, according to the processing order, wherein each score is generated for a document by utilizing the corresponding predicted classifier and a document information profile of the document to be scored; and

identify and propagate forward the predicted classifier that corresponds to a received user coding decision as the current classifier and terminating or suspending use of at least one predicted classifier that does not correspond to the received user coding decision.

9. The system of claim 8 , wherein the processor is further adapted to:

classify a set of documents in the document collection into one or more of the one or more classes or subclasses using the propagated predicted classifier in response to a determination that one or more stopping criteria have been met.

10. The system of claim 9 , wherein in order to determine whether one or more stopping criteria have been met, the processor is further adapted to:

select and present documents from the document collection to a user; and

calculate an estimate of system effectiveness using user coding decisions received for the selected and presented documents.

11. The system of claim 9 , wherein in order to determine whether one or more stopping criteria have been met, the processor is further adapted to:

calculate an estimate of the number of documents in the document collection that are relevant to one of the classes or subclasses by fitting calculated scores to a distribution; and

validate the estimate using a sequence of user coding decisions.

12. The system of claim 8 , wherein the processor is further adapted to:

receive or provide relevance rankings, wherein the relevance rankings are generated by one or more keyword searching algorithms or by a comparison with one or more exemplary documents.

13. The system of claim 12 , wherein the processor is further adapted to determine the processing order for the subset of documents, at least in part, from scores calculated using a classifier and the relevance rankings.

14. The system of claim 12 , wherein the processor is further adapted to

select the document from the document collection by choosing between utilizing document scores calculated using a classifier, the relevance rankings or a combination of document scores and the relevance rankings.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 24, 2025
From: GROSSMAN, MAURA R, MS.; CORMACK, GORDON V., MR.
To: ADAPTIVE CLASSIFICATION TECHNOLOGIES LLC
Reel/Frame 070310/0938 →
Continuity (3)
Continuation 14806029 · Jul 22, 2015
Continuation 13840029 · Mar 15, 2013
Related Publication 20170270115A1 · Sep 21, 2017
Cited By (2)
US 12,572,746 US 12,639,108