Systems and methods for machine learning-based document classification
In some aspects, the disclosure is directed to methods and systems for machine learning-based document classification using multiple classifiers. Various classifiers may be employed during different iterations of the method to advance the classification of a document. The document may be classified and labeled in response to a predetermined number of classifiers agreeing upon a meaningful label. Further, the meaningful label may only be applied to the document in the event that the classifiers predicted the document label with a confidence score in excess of a threshold value.
1 . A method for machine learning-based document sorting, comprising:
receiving, by a computing device, a plurality of candidate documents for sorting, the received candidate documents stored in a memory of the computing device, each candidate document lacking an identifier of a related candidate document;
for each candidate document of the plurality of documents:
iteratively, by the computing device:
(a) selecting a subset of classifiers from a plurality of classifiers executable by a processor of the computing device,
(b) extracting a corresponding set of feature characteristics from the candidate document stored in the memory of the computing device, responsive to the selected subset of classifiers, the feature characteristics comprising text or visual content extracted from the candidate document;
(c) classifying, by the processor of the computing device, the candidate document according to each classifier of the selected subset of classifiers, and
(d) repeating steps (a)-(c) until a predetermined number of the selected subset of classifiers at each iteration agrees on a classification, wherein a number of classifiers in the selected subset of classifiers in a first iteration is different from a number of classifiers in the selected subset of classifiers in a second iteration; and
classifying, by the computing device, the candidate document according to the agreed-upon classification; and
collating a subset of the plurality of candidate documents, by the computing device, into a single multi-page document, responsive to each document of the subset having the same classification.
2 . The method of claim 1 , wherein each classifier in a selected subset utilizes different feature characteristics of the candidate document.
3 . The method of claim 1 , wherein in a final iteration, a first number of the selected subset of classifiers classify the candidate document with a first classification, and a second number of the selected subset of classifiers classify the candidate document with a second classification.
4 . The method of claim 1 , wherein classifying the candidate document according to the agreed-upon classification is responsive to a confidence score of the classification exceeding a threshold.
5 . The method of claim 1 , wherein step (d) further comprises repeating steps (a)-(c) responsive to a classifier of the selected subset of classifiers returning an unknown classification.
6 . The method of claim 1 , wherein during at least one iteration, step (d) further comprises repeating steps (a)-(c) responsive to all of the selected subset of classifiers not agreeing on a classification.
7 . The method of claim 1 , wherein extracting the corresponding set of feature characteristics from the candidate document further comprises at least one of extracting text of the candidate document, identifying coordinates of text within the candidate document, or identifying vertical or horizontal edges of an image the candidate document.
8 . The method of claim 1 , wherein the plurality of classifiers comprise a gradient boosting classifier, a neural network, a time series analysis, a regular expression parser, or one or more image comparators.
9 . The method of claim 1 , wherein the predetermined number of the selected subset of classifiers in at least one iteration is equal to a majority of the classifiers in the at least one iteration.
10 . The method of claim 1 , wherein the predetermined number of the selected subset of classifiers in at least one iteration is equal to a minority of the classifiers in the at least one iteration.
11 . The method of claim 1 , wherein during at least one iteration,
step (b) further comprises extracting feature characteristics of a parent document of the candidate document, the parent document comprising a document related to the candidate document that has been previously classified by the computing device, the feature characteristics of the parent document comprising text or visual content extracted from the parent document; and
step (c) further comprises classifying the candidate document according to the extracted feature characteristics of the parent document of the candidate document.
12 . A system for machine learning-based sorting, comprising:
a computing device comprising a storage device, processing circuitry, and a receiver;
wherein the receiver is configured to receive a plurality of candidate documents for sorting, each candidate document lacking an identifier of a related candidate document;
wherein the storage device is configured to store the received candidate documents; and
wherein the processing circuitry is configured to:
for each candidate document of the plurality of documents:
iteratively:
select a subset of classifiers from a plurality of classifiers,
extract a set of feature characteristics from the candidate document, the extracted set of feature characteristics based on the selected subset of classifiers, the feature characteristics comprising text or visual content extracted from the candidate document, and
classify the candidate document according to each classifier of the selected subset of classifiers,
until determining that a predetermined number of the selected subset of classifiers agrees on a classification, wherein a number of classifiers in the selected subset of classifiers in a first iteration is different from a number of classifiers in the selected subset of classifiers in a second iteration;
compare a confidence score to a threshold based on the selected subset of classifiers, the confidence score calculated based on the classification of the candidate document by each of the selected subset of classifiers agreeing upon the classification; and
classify the candidate document according to the agreed-upon classification, responsive to the confidence score exceeding the threshold; and
collate a subset of the candidate documents into a single multi-page document, responsive to the classification, responsive to each document of the subset having the same classification.
13 . The system of claim 12 , wherein each classifier in a selected subset utilizes different feature characteristics of the candidate document.
14 . The system of claim 12 , wherein the processing circuitry is further configured to extract the set of feature characteristics from the candidate document by at least one of extracting text of the candidate document, identifying coordinates of text within the candidate document, or identifying vertical or horizontal edges of an image of the candidate document.
15 . The system of claim 12 , wherein the plurality of classifiers comprise an elastic search model, a gradient boosting classifier, a neural network, a time series analysis, a regular expression parser, or one or more image comparators.
16 . The system of claim 12 , wherein the predetermined number of the selected subset of classifiers is equal to a majority of the selected subset of classifiers.
17 . The system of claim 12 , wherein the predetermined number of the selected subset of classifiers is equal to a minority of the selected subset of classifiers.
18 . The system of claim 12 , wherein the processing circuitry is further configured to return an unknown classification.
19 . The system of claim 12 , wherein during at least one iteration, the processing circuitry is further configured to:
extract feature characteristics of a parent document of the candidate document, the parent document comprising a document related to the candidate document that has been previously classified by the computing device, the feature characteristics of the parent document comprising text or visual content extracted from the parent document, and
classify the candidate document according to the extracted feature characteristics of the parent document of the candidate document.
20 . A method for collating data, comprising:
receiving, by a computing system, a plurality of documents, each lacking an identifier of a related document of the plurality of documents;
for each document, iteratively applying a plurality of different machine learning classifiers, by the computing system, to feature characteristics comprising text or visual content extracted from the document, wherein a number of applied classifiers in a first iteration is different from a number of applied classifiers in a second iteration, wherein each classifier generates an identification of a corresponding document type, the iterations being repeated until a predetermined number of generated identifications from different classifiers during an iteration match, the matching identification applied to the document;
selecting, by the computing system, a first subset of the plurality documents having matching identifications; and
collating, by the computing system, the first subset of the plurality of documents into a single document.