IP Library Granted Patent US 12,725,080
Granted Patent B2
US 12,725,080 · App. 18/218,970 · Granted Sep 1, 2026

Simultaneous data sampling and feature selection via weak learners

Inventors: Moein Owhadi Kareshk (Burnaby, CA); Ali Seyfi (Vancouver, CA); Hesam Fathi Moghadam (Sunnyvale, CA); Sungpack Hong (Palo Alto, CA); Hassan Chafi (San Mateo, CA)
Assignee: Oracle International Corporation
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,080
App. No.
18/218,970
Granted
Sep 1, 2026
Kind
B2
Abstract

From many features and many multidimensional points, a computer generates exploratory training configurations. Each point contains a value for each of the features. Each exploratory training configuration identifies a random subset of the features and a random subset of the points. A performance score is generated for each of the exploratory training configurations. A feature weight is generated for each of the features that is based on the performance scores of the exploratory training configurations whose random subset of features contains the feature. A point weight is generated for each of the points that is based on the performance scores of the exploratory training configurations whose random subset of the many points contains the point. A machine learning model is trained using an optimized training corpus that consists of a subset of the many features based on feature weight and a subset of the many points based on point weight.

Claims (49)

1 . A method comprising:

generating a training corpus by:

a) generating, from a plurality of features and a plurality of multidimensional points, a plurality of training configurations wherein:

each point in the plurality of multidimensional points contains a value for each feature in the plurality of features, and

each configuration in the plurality of training configurations identifies a random subset of the plurality of features and a random subset of the plurality of multidimensional points;

b) generating a performance score for each configuration in the plurality of training configurations by training and evaluating a weak learner based on the random subset of the plurality of features of the configuration and the random subset of the plurality of multidimensional points of the configuration;

c) generating a feature weight for each feature in the plurality of features that is based on the performance scores of the plurality of training configurations whose random subset of the plurality of features contains the feature;

d) generating a point weight for each point in the plurality of multidimensional points that is based on the performance scores of the plurality of training configurations whose random subset of the plurality of multidimensional points contains the point; and

e) including in the training corpus:

a subset of the plurality of features having feature weights that exceed a first threshold, and

a subset of the plurality of multidimensional points having point weights that exceed a second threshold; and

training, with the training corpus, a machine learning model that does not comprise a weak learner;

wherein the method is performed by one or more computers.

2 . The method of claim 1 wherein:

the weak learner comprises a decision tree; and

the machine learning model does not comprise a decision tree.

3 . The method of claim 1 further comprising for each feature in the plurality of features, counting how many of the plurality of training configurations whose random subset of the plurality of features of the configuration contains the feature.

4 . The method of claim 1 further comprising for each point in the plurality of multidimensional points, counting how many of the plurality of training configurations whose random subset of the plurality of multidimensional points of the configuration contains the point.

5 . The method of claim 1 wherein said generating the plurality of training configurations comprises generating a fixed count of training configurations.

6 . The method of claim 1 further comprising for each feature in the plurality of features, summing performance scores of the plurality of training configurations whose random subset of the plurality of features of the configuration contains the feature.

7 . The method of claim 1 further comprising for each point in the plurality of multidimensional points, summing performance scores of the plurality of training configurations whose random subset of the plurality of multidimensional points of the configuration contains the point.

8 . The method of claim 1 wherein at least one selected from a group consisting of:

said generating the plurality of training configurations occurs before said generating the performance score for each configuration in the plurality of training configurations, and

said generating the plurality of training configurations does not depend on said generating the performance score for each configuration in the plurality of training configurations.

9 . The method of claim 1 wherein said training the machine learning model comprises the machine learning model accepting as input a point weight of a point in the plurality of multidimensional points.

10 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause:

generating a training corpus by:

a) generating, from a plurality of features and a plurality of multidimensional points, a plurality of training configurations wherein:

each point in the plurality of multidimensional points contains a value for each feature in the plurality of features, and

each configuration in the plurality of training configurations identifies a random subset of the plurality of features and a random subset of the plurality of multidimensional points;

b) generating a performance score for each configuration in the plurality of training configurations by training and evaluating a weak learner based on the random subset of the plurality of features of the configuration and the random subset of the plurality of multidimensional points of the configuration;

c) generating a feature weight for each feature in the plurality of features that is based on the performance scores of the plurality of training configurations whose random subset of the plurality of features contains the feature;

d) generating a point weight for each point in the plurality of multidimensional points that is based on the performance scores of the plurality of training configurations whose random subset of the plurality of multidimensional points contains the point; and

e) including in the training corpus:

a subset of the plurality of features having feature weights that exceed a first threshold, and

a subset of the plurality of multidimensional points having point weights that exceed a second threshold; and

training, with the training corpus, a machine learning model that does not comprise a weak learner.

11 . The one or more non-transitory computer-readable media of claim 10 wherein:

the weak learner comprises a decision tree; and

the machine learning model does not comprise a decision tree.

12 . The one or more non-transitory computer-readable media of claim 10 wherein the instructions further cause for each feature in the plurality of features, counting how many of the plurality of training configurations whose random subset of the plurality of features of the configuration contains the feature.

13 . The one or more non-transitory computer-readable media of claim 10 wherein the instructions further cause for each point in the plurality of multidimensional points, counting how many of the plurality of training configurations whose random subset of the plurality of multidimensional points of the configuration contains the point.

14 . The one or more non-transitory computer-readable media of claim 10 wherein said generating the plurality of training configurations comprises generating a fixed count of training configurations.

15 . The one or more non-transitory computer-readable media of claim 10 wherein the instructions further cause for each feature in the plurality of features, summing performance scores of the plurality of training configurations whose random subset of the plurality of features of the configuration contains the feature.

16 . The one or more non-transitory computer-readable media of claim 10 wherein the instructions further cause for each point in the plurality of multidimensional points, summing performance scores of the plurality of training configurations whose random subset of the plurality of multidimensional points of the configuration contains the point.

17 . The one or more non-transitory computer-readable media of claim 10 wherein at least one selected from a group consisting of:

said generating the plurality of training configurations occurs before said generating the performance score for each configuration in the plurality of training configurations, and

said generating the plurality of training configurations does not depend on said generating the performance score for each configuration in the plurality of training configurations.

18 . The one or more non-transitory computer-readable media of claim 10 wherein said training the machine learning model comprises the machine learning model accepting as input a point weight of a point in the plurality of multidimensional points.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 6, 2023
From: OWHADI KARESHK, MOEIN; SEYFI, ALI; FATHI MOGHADAM, HESAM; HONG, SUNGPACK; CHAFI, HASSAN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 064174/0434 →
Continuity (1)
Related Publication 20250013909A1 · Jan 9, 2025
References Cited (14)
US 10318882B2 · Brueckner · 2019 [cited by examiner]
US 11783175B2 · Leskovec · 2023 [cited by examiner]
US 11868854B2 · Moharrer · 2024 [cited by examiner]
US 12406478B2 · Janousková · 2025 [cited by examiner]
US 20190370684A1 · Gunes · 2019 [cited by examiner]
US 20210357776A1 · Quader · 2021 [cited by examiner]
US 20220027680A1 · Friedland · 2022 [cited by examiner]
Yakovlev et al., “Oracle AutoML: a fast and predictive AutoML pipeline”, Proc. VLDB Endow. 13, 12, Aug. 2020, 3166-3180. [cited by applicant]
Wang et al., “Software defect prediction based on combined sampling and feature selection”, ICMLCA , 2nd International Conference on Machine Learning and Computer Application, 2021. [cited by applicant]
Tsai et al., “Combining feature selection, instance selection, and ensemble classification techniques for improved financial distress prediction.”, Journal of Business Research 130, 2021, 200-209. [cited by applicant]
Sainin et al., “Ensemble Meta Classifier with Sampling and Feature Selection for Data with Multiclass Imbalance Problem”, Journal of Information and Communication Technology, 2021. [cited by applicant]
Kaynak, “Optical Recognition of Handwritten Digits”, MSC Thesis, Institute of Graduate Studies in Science and Engineering, Bogazici University, 5 pages, 1995. [cited by applicant]
Huang et al. “On combining feature selection and over-sampling techniques for breast cancer prediction.” Applied Sciences 11.14, 2021, 6574. [cited by applicant]
Gao et al., “Combining Feature Subset Selection and Data Sampling for Coping with Highly Imbalanced Software Data.”, SEKE, 2015. [cited by applicant]