IP Library Granted Patent US 10,970,650
Granted Patent B1
US 10,970,650 · App. 16/876,574 · Granted Apr 6, 2021

AUC-maximized high-accuracy classifier for imbalanced datasets

Inventors: Abdullah Abusorrah (Jeddah, SA); Yusuf Al-Turki (Jeddah, SA); Mengchu Zhou (Newark, NJ); Siya Yao (Shanghai, CN)
Assignee: King Abdulaziz University
G06N20/00G06F16/285
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,970,650
App. No.
16/876,574
Granted
Apr 6, 2021
Kind
B1
Abstract

An AUC-maximized high-accuracy classification method and system for imbalanced datasets integrates an under-sampling-and-ensemble strategy, a true-outliers-removing strategy and a fake-outliers-concealing strategy, with the hope to effectively and robustly enhance both the AUC and the accuracy metrics in imbalanced classification. Applying under-sampling to construct multiple sub-datasets and assembling classification results of multiple classifiers greatly decline the risk of misclassification and lead to highly accurate and robust results in imbalanced classification task. Moreover, this invention pays attention to detect and identify extremely hidden outliers in a sub-dataset which includes a sub-majority dataset and the entire minority dataset. In this way, more hidden outliers can be located and thus exert less influence on the decision boundary, which contributes to both high AUC and accuracy. Furthermore, this invention proposes to conceal fake outliers when building decision boundary, which can achieve a higher classification accuracy of the majority class without changing that of the minority class.

Claims (22)

1. An AUC-maximized high-accuracy classification method for imbalanced datasets composed of a majority samples dataset and a minority samples dataset, comprising the steps of:

under-sampling of the majority samples dataset to produce k clusters of sub-majority datasets,

combining the minority samples dataset with each of the k clusters of sub-majority datasets to transform an original imbalanced dataset into multiple balanced sub-datasets,

detecting outliers in each balanced sub-dataset,

categorizing the detected outliers into a first category consisting of samples which user input has identified as outliers and a second category of outliers consisting of samples which the user input has not identified as outliers;

removing outliers of the first category from training data and validation data of each balanced sub-dataset,

removing outliers of the second category from only the training data of each balanced sub-dataset,

building an AUC-maximized decision boundary between majority samples and minority samples of each sub-dataset, and

assembling above k MaxAUC classifiers as the final classifier which achieves a higher classification accuracy of the majority samples dataset without changing that of the minority samples dataset.

2. The AUC-maximized high-accuracy classification method for imbalanced datasets of claim 1 , wherein the steps of under-sampling and combining produces multiple balanced sub-datasets in which the numbers of positive and negative samples are about the same.

3. The AUC-maximized high-accuracy classification method for imbalanced datasets of claim 1 , wherein each classifier for the sub-dataset is based on the criterion of maximizing AUC.

4. A machine learning based system, comprising:

a computer or computer system trained by an AUC-maximized high-accuracy classification method for imbalanced datasets composed of a majority samples dataset and a minority samples dataset, wherein training comprises the steps of:

under-sampling of the majority samples dataset to produce k clusters of sub-majority datasets,

combining the minority samples dataset with each of the k clusters of sub-majority datasets to transform an original imbalanced dataset into multiple balanced sub-datasets,

detecting outliers in each balanced sub-dataset,

categorizing the detected outliers into a first category consisting of samples which user input has identified as outliers and a second category of outliers consisting of samples which the user input has not identified as outliers;

removing outliers of the first category from training data and validation data of each balanced sub-dataset,

removing outliers of the second category from only the training data of each balanced sub-dataset,

building an AUC-maximized decision boundary between majority samples and minority samples of each sub-dataset, and

assembling above k MaxAUC classifiers as the final classifier which achieves a higher classification accuracy of the majority samples dataset without changing that of the minority samples dataset; and

inputs to and outputs from the computer or computer system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2020
From: ABUSORRAH, ABDULLAH; AL-TURKI, YUSUF; ZHOU, MENGCHU; YAO, SIYA
To: KING ABDULAZIZ UNIVERSITY
Reel/Frame 052687/0741 →
Cited By (3)
US 12,229,222 US 12,645,799 US 12,657,515