IP Library › Granted Patent US 12,481,918
Granted Patent B2
US 12,481,918 · App. 17/702,156 · Granted Nov 25, 2025

Method and apparatus for improving performance of classification on the basis of mixed sampling

Inventor: Won Joo Park (Daejeon, KR)
Assignee: Electronics and Telecommunication Research Institute
G06N20/00G06F16/285G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,481,918
App. No.
17/702,156
Granted
Nov 25, 2025
Kind
B2
Abstract

A method of improving performance of classification on the basis of mixed sampling is applied. The present invention is directed to providing a method and apparatus for improving the performance of classification on the basis of mixed sampling that are capable of, when learning a model that classifies types using deep learning by dividing the entire data set and using the divided data set for training, validation, and testing, applying different sampling techniques by types of data according to the characteristics of the training data in order to improve the classification performance.

Claims (45)

1 . A method of improving performance of classification on the basis of mixed sampling, which is performed by a computer, the method comprising:

dividing an entire data set requiring classification into training data and testing data;

training an N th classification model on the basis of the training data;

setting the testing data to an input of the N th classification model to provide a type inference result;

applying a mixed sampling technique to the type inference result based on characteristic information for each type group of the type inference result;

reconstructing the training data based on mixed sampling for an N+1 th classification model in which mixed sampling-based data according to a result of the applying of the mixed sampling technique is concatenated;

training the N+1th classification model on the basis of the mixed sampling-based training data; and

setting the testing data to an input of the N+1th classification model to provide the type inference result;

wherein the setting of the testing data to the input of the N th classification model to provide the type inference result includes providing a type inference result including a plurality of type groups distinguished by grades according to the characteristic information for each type group including a number of pieces of data to be learned and a classification inference performance of the N th classification model; and

wherein the plurality of type groups include a first type group, a second type group, a third type group, and a fourth type group,

wherein the first type group is a group in which the number of pieces of the data to be learned is relatively large, and a type inference performance is relatively high,

the second type group is a group in which the number of pieces of the data to be learned is relatively large, and the type inference performance is relatively low,

the third type group is a group in which the number of pieces of the data to be learned is relatively small, and the type inference performance is relatively high, and

the fourth type group is a group in which the number of pieces of the data to be learned is relatively small, and the type inference performance is relatively low.

2 . The method of claim 1 , wherein the dividing of the entire data set, which is a target for the classification, into the training data and the testing data includes dividing numbers of pieces of the training data according to types and numbers of pieces of the testing data according to types in proportion to a distribution of numbers of pieces of data in the entire data set according to types.

3 . The method of claim 1 , further comprising applying the training data to a predetermined model learning technique to learn a plurality of classification models.

4 . The method of claim 1 , wherein the applying of the mixed sampling technique to the type inference result on the basis of the characteristic information for each type group of the type inference result includes selectively applying a technique of under-sampling, a technique of data augmentation, and a technique of re-labeling to each of the type groups on the basis of the characteristic information for each type group of the plurality of type groups.

5 . The method of claim 4 , wherein the applying of the mixed sampling technique to the type inference result on the basis of the characteristic information for each type group of the type inference result includes:

applying the technique of under-sampling to the first type group to randomly under-sample training data;

applying the technique of re-labeling to the second type group;

allowing the third type group to be included in the mixed sampling-based training data without sampling; and

applying the technique of data augmentation to the fourth type group.

6 . The method of claim 1 , further providing each of the pieces of type inference result inferred by the N th classification model and the N+1 th classification model as a visualization graph type.

7 . An apparatus for improving performance of training data classification on the basis of mixed sampling, the apparatus comprising:

a communication module;

a memory; and

a processor for executing a program stored in the memory,

wherein the processor is configured to, according to execution of the program:

divide an entire data set requiring classification into training data and testing data through a data pre-processing unit, train an N th classification model on the basis of the training data through a model training unit, set the testing data to an input of the N th classification model to provide a type inference result through a model testing unit, and evaluate characteristic information for each type group based on the type inference result to determine type groups through a model evaluation unit; and

apply a mixed sampling technique to the training data and reconstruct the training data based on mixed sampling for an N+1th classification model, in which mixed sampling-based data according to a result of the application of the mixed sampling technique is concatenated, through the data pre-processing unit, train the N+1th classification model on the basis of the mixed sampling-based training data through the model training unit, and set the testing data to an input of the N+1th classification model to provide the type inference result through the model testing unit;

wherein the model testing unit provides a type inference result including a plurality of type groups distinguished by grades according to the characteristic information for each type group including a number of pieces of data to be learned and a classification inference performance of the N th classification model;

wherein the plurality of type groups include a first type group, a second type group, a third type group, and a fourth type group,

wherein the first type group is a group in which the number of pieces of the data to be learned is relatively large, and the type inference performance is relatively high,

the second type group is a group in which the number of pieces of the data to be learned is relatively large, and the type inference performance is relatively low,

the third type group is a group in which the number of pieces of the data to be learned is relatively small, and the type inference performance is relatively high, and

the fourth type group is a group in which the number of pieces of the data to be learned is relatively small, and the type inference performance is relatively low.

8 . The apparatus of claim 7 , wherein the data pre-processing unit divides numbers of pieces of the training data according to types and numbers of pieces of the testing data according to types in proportion to a distribution of numbers of pieces of data in the entire data set according to types.

9 . The apparatus of claim 7 , wherein the data pre-processing unit applies the training data to a predetermined model learning technique to learn a plurality of classification models.

10 . The apparatus of claim 7 , wherein the data pre-processing unit selectively applies a technique of under-sampling, a technique of data augmentation, and a technique of re-labeling to each of the type groups on the basis of the characteristic information for each type group of the plurality of type groups.

11 . The apparatus of claim 10 , wherein the data pre-processing unit is configured to:

apply the technique of under-sampling to the first type group to randomly under-sample training data;

apply the technique of re-labeling to the second type group;

allow the third type group to be included in the mixed sampling-based training data without sampling; and

apply the technique of data augmentation to the fourth type group.

12 . The apparatus of claim 7 , further comprising a user interface configured to provide each of pieces of type inference result inferred by the N th classification model and N+1 th classification model as a visualization graph type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2022
From: PARK, WON JOO
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 059376/0914 →
Priority Claims (1)
KR 10-2021-0038141 · Mar 24, 2021 · national
Continuity (1)
Related Publication 20220309401A1 · Sep 29, 2022
References Cited (35)
US 7725408B2 · Lee et al. · 2010 [cited by applicant]
US 10089109B2 · Shukla · 2018 [cited by examiner]
US 11620558B1 · Xu · 2023 [cited by examiner]
US 11734937B1 · Pushkin · 2023 [cited by examiner]
US 11756572B2 · Shor · 2023 [cited by examiner]
US 11892897B2 · Shakarian · 2024 [cited by examiner]
US 11893772B1 · Gokalp · 2024 [cited by examiner]
US 20130097103A1 · Chari · 2013 [cited by examiner]
US 20140297570A1 · Garera · 2014 [cited by examiner]
US 20140372351A1 · Sun · 2014 [cited by examiner]
US 20200005901A1 · Cohen · 2020 [cited by examiner]
US 20200242154A1 · Haneda · 2020 [cited by examiner]
US 20210012213A1 · Kim et al. · 2021 [cited by applicant]
US 20210027143A1 · Joo · 2021 [cited by applicant]
US 20210073660A1 · Zhang · 2021 [cited by examiner]
US 20210192288A1 · Cao · 2021 [cited by examiner]
US 20210224696A1 · Nasr-Azadani · 2021 [cited by examiner]
US 20210233615A1 · Banavar · 2021 [cited by examiner]
US 20210264260A1 · Kim · 2021 [cited by examiner]
US 20210312351A1 · Pourmohammad · 2021 [cited by examiner]
US 20210319333A1 · Lee · 2021 [cited by examiner]
US 20220012535A1 · Ben-Itzhak · 2022 [cited by examiner]
US 20220019848A1 · Takayama · 2022 [cited by examiner]
US 20220083856A1 · Yaguchi · 2022 [cited by examiner]
US 20220120727A1 · Al-Dabbagh · 2022 [cited by examiner]
US 20220172342A1 · Zepeda Salvatierra · 2022 [cited by examiner]
US 20220207352A1 · Barr · 2022 [cited by examiner]
US 20220253526A1 · Sanders · 2022 [cited by examiner]
JP 2020149618A · 2020 [cited by applicant]
KR 1020170107283A · 2017 [cited by applicant]
KR 20190135329A · 2019 [cited by applicant]
KR 20190140824A · 2019 [cited by applicant]
KR 20200027834A1 · 2020 [cited by applicant]
KR 1020200110400A · 2020 [cited by applicant]
KR 102283283B1 · 2021 [cited by applicant]