IP Library › Granted Patent US 12,548,303
Granted Patent B2
US 12,548,303 · App. 18/528,880 · Granted Feb 10, 2026

Diversity-aware weighted majority vote classifier for decision making on imbalanced datasets

Inventors: Anil Goyal (Heidelberg, DE); Jihed Khiari (Heidelberg, DE)
Assignee: NEC Corporation
G06V10/7747G06N20/20G06V10/817G06V10/87
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,303
App. No.
18/528,880
Granted
Feb 10, 2026
Kind
B2
Abstract

An ensemble learning based method is for a binary classification on an imbalanced dataset. The imbalanced dataset has a minority class comprising positive samples and a majority class comprising negative samples. The method includes: generatively oversampling the imbalanced dataset by synthetically generating minority class examples, thereby generating a generated dataset; using the generated dataset to generate subsamples, and learning a base classifier on each of the subsamples to determine a plurality of base classifiers; and learning a weighted majority vote classifier by combining outputs of the base classifiers. Each of the base classifiers is assigned a weight in such a way that a diversity between the base classifiers on the positive samples is minimized.

Claims (42)

1 . An ensemble learning based method, executed by one or more computer processors, for a binary classification task on an imbalanced dataset, wherein the imbalanced dataset comprises a minority class comprising positive samples and a majority class comprising negative samples, the ensemble learning based method comprising:

generatively oversampling minority class examples included in the imbalanced dataset by synthetically generating the minority class examples using computer-implemented algorithms, thereby generating a modified dataset D′ of the minority class;

using the modified dataset D′ of the minority class to generate a plurality of subsamples in a random and stratified way such that each subsample, of the subsamples, contains random examples while preserving a class-distribution ratio of the modified dataset D′, and learning a base classifier on each of the subsamples using machine learning algorithms executed on the one or more computer processors to determine a plurality of base classifiers;

learning a diversity-aware weighted majority vote classifier by using PAC-Bayesian theory to control operation of the one or more computer processors, the learning comprising minimizing an objective function including:

a classification-error term of the diversity-aware weighted majority vote classifier, and

a pairwise diversity term computed over positive examples only; and

executing the binary classification task by applying the diversity-aware weighted majority vote classifier to the imbalanced dataset,

wherein the binary classification task outputs a final prediction indicating samples as ones of the positive samples and the negative samples, and

wherein the diversity-aware weighted majority vote classifier is configured to process input data and generate classification outputs with low false positive rates.

2 . The ensemble learning based method according to claim 1 , wherein a minimization of a diversity between the base classifiers on the positive samples is determined by an optimization problem.

3 . The ensemble learning based method according to claim 2 , wherein the optimization problem is based on a pairwise comparison of the base classifiers and on a classification error of the base classifiers.

4 . The ensemble learning based method according to claim 1 , wherein synthetically generating the minority class examples comprises:

selecting, from the positive samples of the imbalanced dataset, safe examples using k-nearest neighbors, wherein a positive sample, of the positive samples, is considered to be a safe example, of the safe examples, based upon determining that a majority vote of k-nearest neighbors for that example is positive.

5 . The ensemble learning based method according to claim 4 , wherein synthetically generating the minority class examples further comprises:

learning, for the selected safe examples, parameters for a multivariate probability distribution, and

generating the minority class examples based on the multivariate probability distribution using the parameters learned from the safe examples.

6 . The ensemble learning based method according to claim 1 , wherein the random examples are of the modified dataset D′.

7 . The ensemble learning based method according to claim 1 , wherein the subsamples are generated in such a way that the examples contained in each of the subsamples have the same ratio of class distribution as the examples contained in the modified dataset D′.

8 . The ensemble learning based method according to claim 1 , wherein the base classifiers learned on the subsamples are based Decision Trees, Logistic Regression, or Support Vector Machines.

9 . The ensemble learning based method according to claim 1 , wherein the base classifier is learned on each of the subsamples.

10 . The ensemble learning based method according to claim 1 , wherein heterogeneous base classifiers are learned on different subsamples.

11 . The ensemble learning based method according to claim 1 , wherein the imbalanced dataset comprises medical diagnosis data, comprising samples of photographic, computer-graphic, or medical-tissue images, and wherein the binary classification task is a differentiation between the positive samples, as diagnosed with a condition in question, and the negative samples as not diagnosed with the condition in question.

12 . The ensemble learning based method according to claim 1 , wherein the imbalanced dataset comprises predictive maintenance data, comprising equipment status samples of a set of assets in operation, and wherein the binary classification task is a differentiation between the positive samples, as related to assets that are qualified to be defective or to become defective in the near future, and the negative samples as related to assets that are qualified to function properly without becoming defective in the near future.

13 . The ensemble learning based method according to claim 1 , wherein the imbalanced dataset comprises credit cards transactions data, and wherein the binary classification task is a differentiation between the positive samples, as related to transactions that are qualified to be fraudulent, and the negative samples as related to transactions that are qualified to be legitimate.

14 . A system for configured to execute an ensemble learning based method, by one or more computer processors of the system, for a binary classification task on an imbalanced dataset, the imbalanced dataset comprising a minority class comprising positive samples and a majority class comprising negative samples, the ensemble learning based method comprising:

generatively oversampling minority class examples included in the imbalanced dataset by synthetically generating the minority class examples using computer-implemented algorithms, thereby generating a modified dataset D′ of the minority class;

using the modified dataset D′ of the minority class to generate a plurality of subsamples in a random and stratified way such that each subsample, of the subsamples, contains random examples while preserving a class-distribution ratio of the modified dataset D′, and learning a base classifier on each of the subsamples using machine learning algorithms executed on the one or more computer processors to determine a plurality of base classifiers;

learning a diversity-aware weighted majority vote classifier by using PAC-Bayesian theory to control operation of the one or more computer processors, the learning comprising minimizing an objective function including:

a classification-error term of the diversity-aware weighted majority vote classifier, and

a pairwise diversity term computed over positive examples only; and

executing the binary classification task by applying the diversity-aware weighted majority vote classifier to the imbalanced dataset,

wherein the binary classification task outputs a final prediction indicating samples as ones of the positive samples and the negative samples, and

wherein the diversity-aware weighted majority vote classifier is configured to process input data and generate classification outputs with low false positive rates.

15 . A non-transitory computer readable medium having, stored thereon, instructions for performing an ensemble learning based method, executed by one or more computer processors, for a binary classification task on an imbalanced dataset, the imbalanced dataset comprising a minority class comprising positive samples and a majority class comprising negative samples, the ensemble learning based method comprising:

generatively oversampling minority class examples included in the imbalanced dataset by synthetically generating the minority class examples using computer-implemented algorithms, thereby generating a modified dataset D′ of the minority class;

using the modified dataset D′ of the minority class to generate a plurality of subsamples in a random and stratified way such that each subsample, of the subsamples, contains random examples while preserving a class-distribution ratio of the modified dataset D′, and learning a base classifier on each of the subsamples using machine learning algorithms executed on the one or more computer processors to determine a plurality of base classifiers;

learning a diversity-aware weighted majority vote classifier by using PAC-Bayesian theory to control operation of the one or more computer processors, the learning comprising minimizing an objective function including:

a classification-error term of the diversity-aware weighted majority vote classifier, and

a pairwise diversity term computed over positive examples only; and

executing the binary classification task by applying the diversity-aware weighted majority vote classifier to the imbalanced dataset,

wherein the binary classification task outputs a final prediction indicating samples as ones of the positive samples and the negative samples, and

wherein the diversity-aware weighted majority vote classifier is configured to process input data and generate classification outputs with low false positive rates.

Continuity (2)
Continuation 17612281
Related Publication 20240112451A1 · Apr 4, 2024
References Cited (47)
US 9224104B2 · Lin et al. · 2015 [cited by applicant]
US 11544570B2 · Roy · 2023 [cited by applicant]
US 11676719B2 · Feczko et al. · 2023 [cited by applicant]
US 20050069936A1 · Diamond et al. · 2005 [cited by applicant]
US 20140257122A1 · Ong et al. · 2014 [cited by applicant]
US 20180144352A1 · Ram et al. · 2018 [cited by applicant]
US 20190130215A1 · Kaestle et al. · 2019 [cited by applicant]
US 20190213605A1 · Patel et al. · 2019 [cited by applicant]
US 20190370384A1 · Dalek et al. · 2019 [cited by applicant]
US 20200005901A1 · Cohen et al. · 2020 [cited by applicant]
US 20200050964A1 · Vassilev · 2020 [cited by applicant]
US 20200334744A1 · Gupta et al. · 2020 [cited by applicant]
WO 2018167404A1 · 2018 [cited by applicant]
Paass, G., & Kindermann, J. (Apr. 1998). Bayesian classification trees with overlapping leaves applied to credit-scoring. In Pacific-Asia Conference on Knowledge Discovery and Data Mining (pp. 234-245). Berlin, Heidelbe… [cited by examiner]
Zhang Yongqing, Zhu Min, Zhang Danling, Mi Gang and Ma Daichuan, “Improved SMOTEBagging and its application in imbalanced data classification,” IEEE Conference Anthology, China, 2013, pp. 1-5, doi: 10.1109/ANTHOLOGY.201… [cited by examiner]
B. Das, N. C. Krishnan and D. J. Cook, “RACOG and wRACOG: Two Probabilistic Oversampling Techniques,” in IEEE Transactions on Knowledge and Data Engineering, vol. 27, No. 1, pp. 222-234, Jan. 1, 2015, doi: 10.1109/TKDE.… [cited by examiner]
Wang H, Xu Q, Zhou L. Large unbalanced credit scoring using Lasso-logistic regression ensemble. PLoS One. 2015;10(2):e0117844. Published Feb. 23, 2015. doi:10.1371/journal.pone.0117844 (Year: 2015). [cited by examiner]
Lim, Pin, Chi Keong Goh, and Kay Chen Tan. “Evolutionary cluster-based synthetic oversampling ensemble (eco-ensemble) for imbalance learning.” IEEE transactions on cybernetics 47.9 (2016): 2850-2861. (Year: 2016). [cited by examiner]
K. E. Bennin et al., “MAHAKIL: Diversity Based Oversampling Approach to Alleviate the Class Imbalance Issue in Software Defect Prediction,” in IEEE Transactions on Software Engineering, vol. 44, No. 6, pp. 534-550, Jun.… [cited by examiner]
Douzas, Georgios, and Fernando Bacao. “Effective data generation for imbalanced learning using conditional generative adversarial networks.” Expert Systems with applications 91 (2018): 464-471. (Year: 2018). [cited by examiner]
Morvant et al., (2014). Majority vote of diverse classifiers for late fusion. In Structural, Syntactic, and Statistical Pattern Recognition: Joint IAPR International Workshop, S+ SSPR 2014, Joensuu, Finland, Aug. 20-22,… [cited by examiner]
I. Kononenko, “Machine learning for medical diagnosis: history, state of the art and perspective”, Artificial Intelligence in medicine, 23.1 (2001):89-109. [cited by applicant]
Z. Ding, “Diversified ensemble classifiers for highly imbalanced data learning and their application in bioinformatics” (2011). [cited by applicant]
G. N. Hortobagyi, “Treatment of breast cancer”, New England Journal of Medicine 339.14( 1998), pp. 974-984. [cited by applicant]
Kou, Yufeng, et al., “Survey of fraud detection techniques”, IEEE International Conference on Networking, Sensing and Control, 2004. vol. 2. IEEE, 2004, pp. 749-754. [cited by applicant]
A report from Javelin Strategy and Research firm published in 2015, retrieved from the Internet on May 2, 2025, https://www.javelinstrategy.com/press-release/false-positive-card-declines-push-consumersabandon-issuers-an… [cited by applicant]
“Credit Card Fraud Detection” dataset, retrieved from the Internet on May 14, 2025, https://urldefense.com/v3/_https://www.kaggle.com/datasets/mig-ulb/creditcardfraud_:!!OhYLZkit9p47d2A!u8z5fhdR7sx7Gm-rspQAuBvwBk6WUHDWT… [cited by applicant]
R. Wedge et al., “Solving the False Positives Problem in Fraud Prediction Using Automated Feature Engineering”, Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Springer, Cham, 2018. [cited by applicant]
“Breast Cancer Wisconsin diagnostic”, retrieved from the internet on May 14, 2025, https://urldefense.com/v3/_https://archive.ics.uci.edu/dataset/17/breast*cancer*wisconsin*diagnostic_; Kysr!!OhYLZkit9p47d2A!t28VFZtlw0N… [cited by applicant]
“Pima Indians diabetes database”, retrieved from the Internet on May 14, 2025, https://urldefense.com/v3/_https://www.kaggle.com/datasets/ucimi/pima-indians-diabetes-database_;!!OhYLZkit9p47d2A!u8z5fhdR7sx7Gm-rspQAuBvwB… [cited by applicant]
Chawla, et al., “Exploiting Diversity in Ensembles: Improving the Performance on Unbalanced Datasets,” [cited by applicant]
Wang, et al., “Diversity Analysis on Imbalanced Data Sets by Using Ensemble Models,” [cited by applicant]
Wang, et al., “Relationships Between Diversity of Classification Ensembles and Single-Class Performance Measures,” [cited by applicant]
Galar, et al., “A Review on Ensembles for the Class Imbalance Problem: Bagging-, Boosting-, and Hybrid-Based Approaches,” [cited by applicant]
He, et al., “Learning from Imbalanced Data.” [cited by applicant]
US Office Action for U.S. Appl. No. 17/612,281, mailed on Feb. 11, 2025. [cited by applicant]
C.-L. Liu and P.-Y. Hsieh, “Model-Based Synthetic Sampling for Imbalanced Data,” in IEEE Transactions on Knowledge and Data Engineering, vol. 32, No. 8, pp. 1543-1556, Aug. 1, 2020, doi: 10.1109/TKDE.2019.2905559. (Year… [cited by applicant]
T Ryan Hoens et al., “Imbalanced Datasets: From Sampling to Classifiers,” in Imbalanced Learning: Foundations, Algorithms, and Applications, IEEE, 2013, pp. 43-59, doi: 10.1002/9781118646106.ch3. (Year: 2013). [cited by applicant]
M. Y. Arafat et al., “Cluster-based under-sampling with random forest for multiclass imbalanced classification,” 2017 11th International Conference on Software, Knowledge, Information Management and Applications (SKIMA)… [cited by applicant]
A-D. Lipitakis and S. Kotsiantis, “A hybrid Machine Learning methodology for imbalanced datasets,” IISA 2014, The 5th International Conference on Information, Intelligence, Systems and Applications, Chania, Greece, 2014… [cited by applicant]
Felix Last et al., “Oversampling for Imbalanced Learning Based on K-Means and SMOTE,” arXiv:1711.00837v2 submitted on 2017, 19 pages. [cited by applicant]
US Office Action for U.S. Appl. No. 18/530,960, mailed on Jun. 17, 2025. [cited by applicant]
Chandra, Arjun, and Xin Yao, “Ensemble learning using multi-objective evolutionary algorithms.” Journal of Mathematical Modelling and Algorithms 5 (2006): 417-445. (Year: 2006). [cited by applicant]
Parvin, Hamid, et al. “A scalable method for improving the performance of classifiers in multiclass applications by pairwise classifiers and GA” 2008 fourth international conference on networked computing and advanced i… [cited by applicant]
Fernandes, Everlandio RQ, Andre CPLF de Carvalho, and Andre LV Coelho. “An evolutionary sampling approach for classification with imbalanced data.” 2015 international joint conference on neural networks (IJCN N). IEEE, … [cited by applicant]
Onan, Aytug, Serdar Korukoglu, and Hasan Bulut. “A multiobjective weighted voting ensemble classifier based on differential evolution algorithm for text sentiment classification.” Expert Systems with Applications 62 (20… [cited by applicant]
Saleena, Nabizath. “An ensemble classification system for twitter sentiment analysis.” Procedia computer science 132 (2018): 937-946. (Year: 2018). [cited by applicant]