IP Library Granted Patent US 12,469,608
Granted Patent B2
US 12,469,608 · App. 17/219,901 · Granted Nov 11, 2025

Mining method for sample grouping

Inventors: Guan-An Chen (Hsinchu, TW); Jhen-Yang Syu (Tainan, TW)
Assignee: Industrial Technology Research Institute
G16H50/70G16H30/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,469,608
App. No.
17/219,901
Granted
Nov 11, 2025
Kind
B2
Abstract

A mining method for sample grouping is provided. The method includes the following steps. A field dataset including multiple samples is obtained, and each sample corresponds to an actual labeled result. The samples are respectively input to an existing model, so as to obtain the estimated results. An outlier sample set in the field dataset is removed based on a difference distribution of the estimated results and the actual labeled results, and the samples that remain in the field dataset form a remaining sample set. The remaining sample set is grouped into a hard sample set and an easy sample set based on the estimated results of the remaining sample set.

Claims (25)

1 . A mining method for sample grouping, comprising:

performing following steps through a processor:

(a) obtaining an existing model that has been trained and a field dataset that is different from a training data set of the existing model, wherein the field dataset comprises a plurality of samples collected based on a specified field, and the plurality of samples comprises a plurality of corresponding actual labeled results, and a sample number of the field dataset is smaller than a sample number of the training data set of the existing model;

(b) inputting the plurality of samples respectively into the existing model, so as to obtain a plurality of estimated results;

(c) removing an outlier sample set from the field dataset based on a difference distribution of the plurality of estimated results and the plurality of actual labeled results, wherein the plurality of samples that remain in the field dataset after the outlier sample set is removed form a remaining sample set;

(d) grouping the remaining sample set into a hard sample set and an easy sample set based on the plurality of estimated results of the remaining sample set; and

(e) after obtaining the hard sample set and the easy sample set, sending the hard sample set and the easy sample set into an incremental learning framework to retrain the existing model, so that a retrained existing model becomes a model directed to the specified field corresponding to the field dataset to interpret data obtained in the specified field.

2 . The mining method for sample grouping according to claim 1 , wherein after step (b), the method further comprises:

calculating the difference distribution of the plurality of estimated results and the plurality of actual labeled results;

checking whether the difference distribution is a normal distribution through a normal distribution testing method;

performing step (c) and step (d) in sequence when the difference distribution is determined to be a normal distribution; and

selecting another existing model when the difference distribution is determined to be not a normal distribution, and performing step (b) again.

3 . The mining method for sample grouping according to claim 2 , wherein calculating the difference distribution of the plurality of estimated results and the plurality of actual labeled results comprises:

calculating a difference between an estimated result of each of the plurality of samples and a corresponding actual labeled result through a loss function to thereby obtain the difference distribution.

4 . The mining method for sample grouping according to claim 3 , wherein the loss function adopts cross entropy.

5 . The mining method for sample grouping according to claim 3 , wherein step (c) comprises:

determining a sample with the difference greater than a first setting value or a sample with the difference smaller than a second setting value as an outlier sample set.

6 . The mining method for sample grouping according to claim 3 , wherein step (d) comprises:

calculating an absolute value of the difference corresponding to each of the plurality of samples in the remaining sample set, so as to obtain an absolute-value distribution;

performing a normalization conversion on the absolute-value distribution, so as to obtain a normalized distribution; and

based on the normalized distribution, grouping the remaining sample set into the hard sample set and the easy sample set.

7 . The mining method for sample grouping according to claim 6 , wherein based on the normalized distribution, grouping the remaining sample set into the hard sample set and the easy sample set comprises:

grouping samples that meet a first threshold number into the easy sample set, starting from a sample with an absolute value of a normalized difference in the normalized distribution being 0; and

grouping samples that meet a second threshold number into the hard sample set, starting from a sample with an absolute value of a normalized difference in the normalized distribution being 1.

8 . The mining method for sample grouping according to claim 1 , wherein a sample number of the easy sample set is greater than a sample number of the hard sample set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2021
From: CHEN, GUAN-AN; SYU, JHEN-YANG
To: INDUSTRIAL TECHNOLOGY RESEARCH INSTITUTE
Reel/Frame 055859/0112 →
Priority Claims (1)
TW 110100290 · Jan 5, 2021 · national
Continuity (1)
Related Publication 20220215966A1 · Jul 7, 2022
References Cited (20)
US 10484399B1 · Curtin · 2019 [cited by applicant]
US 20200160178A1 · Kar · 2020 [cited by examiner]
US 20200211185A1 · Hu · 2020 [cited by examiner]
US 20200405148A1 · Tran · 2020 [cited by applicant]
US 20210374518A1 · Zhu · 2021 [cited by examiner]
US 20220114444A1 · Weinzaepfel · 2022 [cited by examiner]
US 20230289592A1 · Badri · 2023 [cited by examiner]
CN 105631037 · 2016 [cited by applicant]
CN 108804470 · 2018 [cited by applicant]
CN 109919928 · 2019 [cited by applicant]
CN 107330074 · 2020 [cited by applicant]
CN 111709485 · 2020 [cited by applicant]
CN 112164447 · 2021 [cited by applicant]
TW 201942769 · 2019 [cited by applicant]
TW I689875 · 2020 [cited by applicant]
TW 202020885 · 2020 [cited by applicant]
WO WO2020061489A1 · 2020 [cited by examiner]
Garg, Prachi, et al. “Memorization and Generalization in Deep CNNS Using Soft Gating Mechanisms.” (2019). ( Year: 2019). [cited by examiner]
Tsai-Yuan Chou, “Generating Virtual Attributes by Fuzzy Clustering Algorithm for Small Datasets Learning,” with English translation of the Abstract thereof, Postgraduate Programs in Management, I-Shou University, May 20… [cited by applicant]
“Office Action of Taiwan Counterpart Application”, issued on May 18, 2022, p. 1-p. 11. [cited by applicant]