Efficient multilabel classification by chaining ordered classifiers and optimizing on uncorrelated labels
Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and/or the like for converting a multilabel classification model into a sequence of a plurality of binary classification models based on a plurality of label subgroups associated with the multilabel classification model, where the label subgroups comprise an optimal subgroup size, the optimal subgroup size is generated by optimizing an optimization measure defined by a subgroup size variable and a total inner group correlation measure, and identifying label membership to a particular subgroup by using a mixed integer linear program model.
1 . A computer-implemented method comprising:
generating, by one or more processors and during a machine learning model training process, a plurality of label groups by:
(i) generating a plurality of correlation values based on a training dataset and a plurality of training classification labels,
(ii) generating a plurality of candidate label groups of different group count ranges based on a total inner group correlation measure,
(iii) determining an optimal label group count from a plurality of candidate label group counts by minimizing (a) a quantity of label groups and (b) the total inner group correlation measure,
(iv) generating, based on the optimal label group count, the plurality of label groups, wherein (a) a label group of the plurality of label groups includes a grouped subset of the plurality of label groups and (b) the grouped subset comprises an inner group correlation measure that satisfies an inner group correlation measure condition, and
(v) generating, using a hyperparameter optimization algorithm, an optimal sequence of the plurality of label groups;
receiving, by the one or more processors, a prediction input data object associated with a plurality of classification labels assigned to the plurality of label groups;
inputting, by the one or more processors, the prediction input data object to a multi-label classification machine learning model to receive a plurality of classification scores corresponding to the plurality of classification labels; and
initiating, by the one or more processors, one or more prediction-based actions based on the plurality of classification scores.
2 . The computer-implemented method of claim 1 , wherein:
determining the optimal label group count comprises determining a candidate label group count g of the plurality of candidate label group counts that is associated with a largest decline of a group count selection relationship between the plurality of candidate label group counts and corresponding total inner group correlation measures.
3 . The computer-implemented method of claim 2 , wherein generating the plurality of candidate label groups for the candidate label group count g comprises:
generating, using a mixed integer linear program machine learning model and based on the plurality of correlation values and the candidate label group count g, g group inclusion indicators corresponding to the plurality of candidate label groups.
4 . The computer-implemented method of claim 2 , wherein generating the plurality of candidate label groups comprises:
generating a total absolute correlation summation measure based on one or more correlation values associated with a training classification label of the plurality of training classification labels;
generating the plurality of candidate label groups comprising at least two candidate label groups, the at least two candidate label groups including at least two initially-grouped classification labels from the plurality of training classification labels;
updating a candidate label group having a highest total inner group correlation measure to include a current classification label among a subgroup of the plurality of training classification labels that is initially generated by excluding all initially-grouped classification labels starting from the training classification label having a highest total absolute correlation summation measure, and
excluding the current classification label from the subgroup.
5 . The computer-implemented method of claim 4 , wherein generating the plurality of label groups comprises:
generating a plurality of classification label mappings corresponding to the plurality of training classification labels in a multi-dimensional clustering space;
generating, based on the plurality of classification label mappings and using a clustering machine learning model, a defined number of label clusters; and
generating the plurality of label groups based on the defined number of label clusters.
6 . The computer-implemented method of claim 5 , wherein a classification label mapping of the plurality of classification label mappings comprises one or more correlation values associated with the training classification label.
7 . The computer-implemented method of claim 2 wherein the total inner group correlation measure comprises an absolute relative spatial distance between the plurality of training classification labels.
8 . The computer-implemented method of claim 1 further comprising assigning a given one of the plurality of training classification labels to a given one of the plurality of candidate label groups based on a lowest total inner group correlation measure.
9 . The computer-implemented method of claim 1 wherein a per-label classifier of a plurality of preceding per-label classifier sets is associated with a respective label from the plurality of label groups and is configured to generate a classification score for the respective label.
10 . The computer-implemented method of claim 6 , wherein generating the classification label mapping further comprises generating the classification label mapping based on an inverse K-means clustering.
11 . The computer-implemented method of claim 1 , wherein a training classification label is associated with one of a plurality of predefined label groups.
12 . The computer-implemented method of claim 1 , wherein:
(i) the multi-label classification machine learning model comprises a sequence of classifier groups ordered according to the optimal sequence and comprising an initial classifier group and a plurality of non-initial classifier groups,
(ii) the initial classifier group comprises an initial per-label classifier set associated with an initial label group of the plurality of label groups, and
(iii) a non-initial classifier group of the plurality of non-initial classifier groups comprises a subsequent per-label classifier set associated with a subsequent label group, subsequent to the initial classifier group, of the plurality of label groups and a plurality of preceding per-label classifier sets from a plurality of preceding classifier groups, including the initial label group, in the sequence of classifier groups.
13 . A system comprising:
one or more processors; and
at least one memory storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to:
generate, during a machine learning model training process, a plurality of label groups by:
(i) generating a plurality of correlation values based on a training dataset and a plurality of training classification labels,
(ii) generating a plurality of candidate label groups of different group count ranges based on a total inner group correlation measure,
(iii) determining an optimal label group count from a plurality of candidate label group counts by minimizing (a) a quantity of label groups and (b) the total inner group correlation measure,
(iv) generating, based on the optimal label group count, the plurality of label groups, wherein (a) a label group of the plurality of label groups includes a grouped subset of the plurality of label groups and (b) the grouped subset comprises an inner group correlation measure that satisfies an inner group correlation measure condition, and
(v) generating, using a hyperparameter optimization algorithm, an optimal sequence of the plurality of label groups;
receive a prediction input data object associated with a plurality of classification labels assigned to the plurality of label groups;
input the prediction input data object to a multi-label classification machine learning model to receive a plurality of classification scores corresponding to the plurality of classification labels;
initiate one or more prediction-based actions based on the plurality of classification scores.
14 . The system of claim 13 , wherein:
determining the optimal label group count comprises determining a candidate label group count g of the plurality of candidate label group counts that is associated with a largest decline of a group count selection relationship between the plurality of candidate label group counts and corresponding total inner group correlation measures.
15 . The system of claim 14 , wherein generating the plurality of candidate label groups for the candidate label group count g comprises:
generating, using a mixed integer linear program machine learning model and based on the plurality of correlation values and the candidate label group count g, g group inclusion indicators corresponding to the plurality of candidate label groups.
16 . The system of claim 13 , wherein generating the plurality of candidate label groups comprises:
generating a total absolute correlation summation measure based on one or more correlation values associated with a training classification label of the plurality of training classification labels;
generating the plurality of candidate label groups comprising at least two candidate label groups, the at least two candidate label groups including at least two initially-grouped classification labels from the plurality of training classification labels;
updating a candidate label group having a highest total inner group correlation measure to include a current classification label among a subgroup of the plurality of training classification labels that is initially generated by excluding all initially-grouped classification labels starting from the training classification label having a highest total absolute correlation summation measure, and
excluding the current classification label from the subgroup.
17 . The system of claim 13 , wherein the total inner group correlation measure comprises an absolute relative spatial distance between the plurality of training classification labels.
18 . The system of claim 13 wherein the one or more processors are further caused to:
assign a given one of the plurality of training classification labels to a given one of the plurality of candidate label groups based on a lowest total inner group correlation measure.
19 . The system of claim 13 , wherein generating the plurality of label groups comprises:
generating a plurality of classification label mappings corresponding to the plurality of training classification labels in a multi-dimensional clustering space;
generating, based on the plurality of classification label mappings and using a clustering machine learning model, a defined number of label clusters; and
generating the plurality of label groups based on the defined number of label clusters.
20 . One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to:
generate, during a machine learning model training process, a plurality of label groups by:
(i) generating a plurality of correlation values based on a training dataset and a plurality of training classification labels,
(ii) generating a plurality of candidate label groups of different group count ranges based on a total inner group correlation measure,
(iii) determining an optimal label group count from a plurality of candidate label group counts by minimizing (a) a quantity of label groups and (b) the total inner group correlation measure,
(iv) generating, based on the optimal label group count, the plurality of label groups, wherein (a) a label group of the plurality of label groups includes a grouped subset of the plurality of label groups and (b) the grouped subset comprises an inner group correlation measure that satisfies an inner group correlation measure condition, and
(v) generating, using a hyperparameter optimization algorithm, an optimal sequence of the plurality of label groups;
receive a prediction input data object associated with a plurality of classification labels assigned to the plurality of label groups;
input the prediction input data object to a multi-label classification machine learning model to receive a plurality of classification scores corresponding to the plurality of classification labels;
initiate one or more prediction-based actions based on the plurality of classification scores.