IP Library › Granted Patent US 12,632,782
Granted Patent B2
US 12,632,782 · App. 17/805,742 · Granted May 19, 2026

Efficient multilabel classification by chaining ordered classifiers and optimizing on uncorrelated labels

Inventors: Neill Michael Byrne (Dublin, IE); Kieran O'Donoghue (Dublin, IE); Michael J. McCarthy (Dublin, IE)
Assignee: Optum Services (Ireland) Limited
G06N20/00G06F18/2113G06F18/23213G06F18/2411G06F18/2431
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,782
App. No.
17/805,742
Granted
May 19, 2026
Kind
B2
Abstract

Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and/or the like for converting a multilabel classification model into a sequence of a plurality of binary classification models based on a plurality of label subgroups associated with the multilabel classification model, where the label subgroups comprise an optimal subgroup size, the optimal subgroup size is generated by optimizing an optimization measure defined by a subgroup size variable and a total inner group correlation measure, and identifying label membership to a particular subgroup by using a mixed integer linear program model.

Claims (71)

1 . A computer-implemented method comprising:

generating, by one or more processors and during a machine learning model training process, a plurality of label groups by:

(i) generating a plurality of correlation values based on a training dataset and a plurality of training classification labels,

(ii) generating a plurality of candidate label groups of different group count ranges based on a total inner group correlation measure,

(iii) determining an optimal label group count from a plurality of candidate label group counts by minimizing (a) a quantity of label groups and (b) the total inner group correlation measure,

(iv) generating, based on the optimal label group count, the plurality of label groups, wherein (a) a label group of the plurality of label groups includes a grouped subset of the plurality of label groups and (b) the grouped subset comprises an inner group correlation measure that satisfies an inner group correlation measure condition, and

(v) generating, using a hyperparameter optimization algorithm, an optimal sequence of the plurality of label groups;

receiving, by the one or more processors, a prediction input data object associated with a plurality of classification labels assigned to the plurality of label groups;

inputting, by the one or more processors, the prediction input data object to a multi-label classification machine learning model to receive a plurality of classification scores corresponding to the plurality of classification labels; and

initiating, by the one or more processors, one or more prediction-based actions based on the plurality of classification scores.

2 . The computer-implemented method of claim 1 , wherein:

determining the optimal label group count comprises determining a candidate label group count g of the plurality of candidate label group counts that is associated with a largest decline of a group count selection relationship between the plurality of candidate label group counts and corresponding total inner group correlation measures.

3 . The computer-implemented method of claim 2 , wherein generating the plurality of candidate label groups for the candidate label group count g comprises:

generating, using a mixed integer linear program machine learning model and based on the plurality of correlation values and the candidate label group count g, g group inclusion indicators corresponding to the plurality of candidate label groups.

4 . The computer-implemented method of claim 2 , wherein generating the plurality of candidate label groups comprises:

generating a total absolute correlation summation measure based on one or more correlation values associated with a training classification label of the plurality of training classification labels;

generating the plurality of candidate label groups comprising at least two candidate label groups, the at least two candidate label groups including at least two initially-grouped classification labels from the plurality of training classification labels;

updating a candidate label group having a highest total inner group correlation measure to include a current classification label among a subgroup of the plurality of training classification labels that is initially generated by excluding all initially-grouped classification labels starting from the training classification label having a highest total absolute correlation summation measure, and

excluding the current classification label from the subgroup.

5 . The computer-implemented method of claim 4 , wherein generating the plurality of label groups comprises:

generating a plurality of classification label mappings corresponding to the plurality of training classification labels in a multi-dimensional clustering space;

generating, based on the plurality of classification label mappings and using a clustering machine learning model, a defined number of label clusters; and

generating the plurality of label groups based on the defined number of label clusters.

6 . The computer-implemented method of claim 5 , wherein a classification label mapping of the plurality of classification label mappings comprises one or more correlation values associated with the training classification label.

7 . The computer-implemented method of claim 2 wherein the total inner group correlation measure comprises an absolute relative spatial distance between the plurality of training classification labels.

8 . The computer-implemented method of claim 1 further comprising assigning a given one of the plurality of training classification labels to a given one of the plurality of candidate label groups based on a lowest total inner group correlation measure.

9 . The computer-implemented method of claim 1 wherein a per-label classifier of a plurality of preceding per-label classifier sets is associated with a respective label from the plurality of label groups and is configured to generate a classification score for the respective label.

10 . The computer-implemented method of claim 6 , wherein generating the classification label mapping further comprises generating the classification label mapping based on an inverse K-means clustering.

11 . The computer-implemented method of claim 1 , wherein a training classification label is associated with one of a plurality of predefined label groups.

12 . The computer-implemented method of claim 1 , wherein:

(i) the multi-label classification machine learning model comprises a sequence of classifier groups ordered according to the optimal sequence and comprising an initial classifier group and a plurality of non-initial classifier groups,

(ii) the initial classifier group comprises an initial per-label classifier set associated with an initial label group of the plurality of label groups, and

(iii) a non-initial classifier group of the plurality of non-initial classifier groups comprises a subsequent per-label classifier set associated with a subsequent label group, subsequent to the initial classifier group, of the plurality of label groups and a plurality of preceding per-label classifier sets from a plurality of preceding classifier groups, including the initial label group, in the sequence of classifier groups.

13 . A system comprising:

one or more processors; and

at least one memory storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to:

generate, during a machine learning model training process, a plurality of label groups by:

(i) generating a plurality of correlation values based on a training dataset and a plurality of training classification labels,

(ii) generating a plurality of candidate label groups of different group count ranges based on a total inner group correlation measure,

(iii) determining an optimal label group count from a plurality of candidate label group counts by minimizing (a) a quantity of label groups and (b) the total inner group correlation measure,

(iv) generating, based on the optimal label group count, the plurality of label groups, wherein (a) a label group of the plurality of label groups includes a grouped subset of the plurality of label groups and (b) the grouped subset comprises an inner group correlation measure that satisfies an inner group correlation measure condition, and

(v) generating, using a hyperparameter optimization algorithm, an optimal sequence of the plurality of label groups;

receive a prediction input data object associated with a plurality of classification labels assigned to the plurality of label groups;

input the prediction input data object to a multi-label classification machine learning model to receive a plurality of classification scores corresponding to the plurality of classification labels;

initiate one or more prediction-based actions based on the plurality of classification scores.

14 . The system of claim 13 , wherein:

determining the optimal label group count comprises determining a candidate label group count g of the plurality of candidate label group counts that is associated with a largest decline of a group count selection relationship between the plurality of candidate label group counts and corresponding total inner group correlation measures.

15 . The system of claim 14 , wherein generating the plurality of candidate label groups for the candidate label group count g comprises:

generating, using a mixed integer linear program machine learning model and based on the plurality of correlation values and the candidate label group count g, g group inclusion indicators corresponding to the plurality of candidate label groups.

16 . The system of claim 13 , wherein generating the plurality of candidate label groups comprises:

generating a total absolute correlation summation measure based on one or more correlation values associated with a training classification label of the plurality of training classification labels;

generating the plurality of candidate label groups comprising at least two candidate label groups, the at least two candidate label groups including at least two initially-grouped classification labels from the plurality of training classification labels;

updating a candidate label group having a highest total inner group correlation measure to include a current classification label among a subgroup of the plurality of training classification labels that is initially generated by excluding all initially-grouped classification labels starting from the training classification label having a highest total absolute correlation summation measure, and

excluding the current classification label from the subgroup.

17 . The system of claim 13 , wherein the total inner group correlation measure comprises an absolute relative spatial distance between the plurality of training classification labels.

18 . The system of claim 13 wherein the one or more processors are further caused to:

assign a given one of the plurality of training classification labels to a given one of the plurality of candidate label groups based on a lowest total inner group correlation measure.

19 . The system of claim 13 , wherein generating the plurality of label groups comprises:

generating a plurality of classification label mappings corresponding to the plurality of training classification labels in a multi-dimensional clustering space;

generating, based on the plurality of classification label mappings and using a clustering machine learning model, a defined number of label clusters; and

generating the plurality of label groups based on the defined number of label clusters.

20 . One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to:

generate, during a machine learning model training process, a plurality of label groups by:

(i) generating a plurality of correlation values based on a training dataset and a plurality of training classification labels,

(ii) generating a plurality of candidate label groups of different group count ranges based on a total inner group correlation measure,

(iii) determining an optimal label group count from a plurality of candidate label group counts by minimizing (a) a quantity of label groups and (b) the total inner group correlation measure,

(iv) generating, based on the optimal label group count, the plurality of label groups, wherein (a) a label group of the plurality of label groups includes a grouped subset of the plurality of label groups and (b) the grouped subset comprises an inner group correlation measure that satisfies an inner group correlation measure condition, and

(v) generating, using a hyperparameter optimization algorithm, an optimal sequence of the plurality of label groups;

receive a prediction input data object associated with a plurality of classification labels assigned to the plurality of label groups;

input the prediction input data object to a multi-label classification machine learning model to receive a plurality of classification scores corresponding to the plurality of classification labels;

initiate one or more prediction-based actions based on the plurality of classification scores.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2022
From: BYRNE, NEILL MICHAEL; O'DONOGHUE, KIERAN; MCCARTHY, MICHAEL J.
To: OPTUM SERVICES (IRELAND) LIMITED
Reel/Frame 060122/0920 →
Continuity (1)
Related Publication 20230394352A1 · Dec 7, 2023
References Cited (32)
US 7725329B2 · Kil et al. · 2010 [cited by applicant]
US 10650927B2 · Are et al. · 2020 [cited by applicant]
US 10930395B2 · Mowery · 2021 [cited by applicant]
US 20110173018A1 · Hoffner et al. · 2011 [cited by applicant]
US 20110302111A1 · Chidlovskii · 2011 [cited by examiner]
US 20160350846A1 · Dintenfass et al. · 2016 [cited by applicant]
US 20200118691A1 · Kiljanek · 2020 [cited by applicant]
US 20200365268A1 · Michuda · 2020 [cited by examiner]
US 20210118571A1 · Hsu et al. · 2021 [cited by applicant]
US 20220059224A1 · Tulley et al. · 2022 [cited by applicant]
US 20220092338A1 · Yang · 2022 [cited by examiner]
CN 110688482A · 2020 [cited by applicant]
CN 111553127A · 2020 [cited by examiner]
CN 113989607A · 2022 [cited by examiner]
CN 114298191A · 2022 [cited by examiner]
WO WO2020095408A1 · 2020 [cited by examiner]
Gao et al., “A Three-phase Augmented Classifiers Chain Approach Based on Co-occurrence Analysis for Multi-Label Classification”, ARXIV ID: 2204.06138, Apr. 12, 2022. (Year: 2022). [cited by examiner]
Niaz et al., “Leveraging from group classification for video concept detection”, 2013 11th International Workshop on Content-Based Multimedia Indexing (CBMI), Jun. 2013, pp. 173-178. (Year: 2013). [cited by examiner]
Rokach et al., “Ensemble Methods for Multi-label Classification”, ARXIV ID: 1307.1769, Jul. 6, 2013, pp. 1-32. (Year: 2013). [cited by examiner]
Kommu et al., “A novel approach for multi-label classification using probabilistic classifiers”, International Conference on Advances in Engineering & Technology Research (ICAETR—2014), Aug. 2014, pp. 1-8. (Year: 2014). [cited by examiner]
Huang et al., “Group sensitive Classifier Chains for multi-label classification”, 2015 IEEE International Conference on Multimedia and Expo (ICME), Jun. 2015, pp. 1-6. (Year: 2015). [cited by examiner]
Rastin et al, “Multi-label classification systems by the use of supervised clustering”, 2017 Artificial Intelligence and Signal Processing Conference (AISP), Oct. 2017, pp. 246-249. (Year: 2017). [cited by examiner]
Biswas et al., “Parallelization of Multi-label classification for large data sets”, 2018 IEEE Symposium Series on Computational Intelligence (SSCI), Nov. 2018, pp. 2005-2010. (Year: 2018). [cited by examiner]
Bidgoli et al., “A Novel Multi-objective Binary Differential Evolution Algorithm for Multi-label Feature Selection”, 2019 IEEE Congress on Evolutionary Computation (CEC), Jun. 10-13, 2019, pp. 1588-1595. (Year: 2019). [cited by examiner]
Weng et al., “An Efficient Stacking Model of Multi-Label Classification Based on Pareto Optimum”, IEEE Access, vol. 7, Jul. 26, 2019, pp. 127427-127437. (Year: 2019). [cited by examiner]
Zhang et al., “Multi-Label Feature Selection Based on High-Order Label Correlation Assumption”, Entropy, 22(7), 797, Jul. 21, 2020, pp. 1-24. (Year: 2020). [cited by examiner]
Jiaman et al., “Association Rules-Based Classifier Chains Method”, IEEE Access, vol. 10, Feb. 4, 2022, pp. 18210-18221. (Year: 2022). [cited by examiner]
Gao et al., “A Three-phase Augmented Classifiers Chain Approach Based on Co-occurrence Analysis for Multi-Label Classification”, ARXIV ID: 2204.06138, Apr. 12, 2022, pp. 1-24. (Year: 2022). [cited by examiner]
Fissler, Tobias et al. “Model Comparison and Calibration Assessment,” arXiv PrePrint arXiv:2202.12780v1 [stat.ML], Feb. 25, 2022, pp. 1-68. [cited by applicant]
Fu, Bin. “Learning Label Dependency for Multi-Label Classification,” Doctoral Dissertation, Faculty of Engineering and Information Technology, University of Technology, Sydney, Feb. 2018, (158 pages). [cited by applicant]
Luo, Li et al. “Using Machine Learning Approaches to Predict High-Cost Chronic Obstructive Pulmonary Disease Patients in China,” Health Informatics Journal, vol. 26, No. 3, Sep. 2020, pp. 1577-1598, DOI: 10.1177/1460458… [cited by applicant]
Wang, Zhenwu et al. “Partial Classifier Chains with Feature Selection by Exploiting Label Correlation in Multi-Label Classification,” Entropy, vol. 22, No. 10:1143, pp. 1-22, Oct. 10, 2020, DOI: 10.3390/e22101143. [cited by applicant]