IP Library › Granted Patent US 12,493,825
Granted Patent B2
US 12,493,825 · App. 17/893,367 · Granted Dec 9, 2025

Selecting a high coverage dataset

Inventors: Shaikh Shahriar Quader (Scarborough, CA); Aindrila Basak (Edmonton, CA); Adrian Mahjour (Toronto, CA); Petr Novotny (Mount Kisco, NY); Carlo Appugliese (Seminole, FL); Berthold Reinwald (San Jose, CA); Dheeraj Arremsetty (Austin, TX)
Assignee: International Business Machines Corporation
G06N20/00G06F18/217G06F18/22G06F18/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,493,825
App. No.
17/893,367
Granted
Dec 9, 2025
Kind
B2
Abstract

Providing a representative dataset from an initial dataset by accessing a dataset associated with a machine learning model, receiving input parameters associated with the representative dataset selection, the input parameters including an evaluation metric, determining a density of a plurality of datapoints associated with the dataset, training a first iteration of a machine learning model using a first data point selected according to the density, determining a first value of the evaluation metric for the first iteration of the machine learning model, generating a representative subset based on the first value of the evaluation metric value, and providing the representative dataset and a final machine learning model trained using the representative dataset.

Claims (83)

1 . A computer implemented method for providing a representative dataset from an initial dataset, the method comprising:

accessing, by a computing device, a dataset associated with a machine learning model;

receiving, by the computing device, input parameters associated with the representative dataset selection, the input parameters including an evaluation metric;

determining, by the computing device, a density of a plurality of datapoints associated with the dataset;

training, by the computing device, a first iteration of the machine learning model using a first data point selected according to the density;

determining a first value of the evaluation metric for the first iteration of the machine learning model;

identifying, by the computing device, neighbors and non-neighbors of the first data point;

determining, by the computing device, a dissimilarity between the first data point and all non-neighbors; and

training, by the computing device, a second iteration of the machine learning model using a second point selected according to the dissimilarity;

determining, by the computing device, a second value of the evaluation metric for the second iteration of the machine learning model;

generating, by the computing device, the representative dataset based on the second value of the evaluation metric;

training a final machine learning model using the representative dataset and

providing, by the computing device, the representative dataset and the final machine learning model trained using the representative dataset.

2 . The computer implemented method according to claim 1 , further comprising changing, by the computing device, all neighbors to inactive, and determining a dissimilarity between the first data point and all active points.

3 . The computer implemented method according to claim 1 , wherein the input parameters comprise a neighborhood size.

4 . The computer implemented method according to claim 3 , further comprising evaluating, by the computing device, a coverage according to the neighborhood size, wherein determining a coverage of the representative dataset comprises determining a ratio of points directly reachable from a point in the representative dataset and all points in the initial dataset.

5 . The computer implemented method according to claim 1 , further comprising:

receiving a target column as user input;

determining a correlation between dataset columns and the target column; and

retaining data point column values according to the correlation.

6 . The computer implemented method according to claim 1 , further comprising:

defining, by the computing device, a budget;

evaluating, by the computing device, all initial dataset points without reaching the budget;

changing, by the computing device, all points not in the representative dataset to active;

training, by the computing device, a third iteration of the machine learning model using a third point selected according to the density;

determining, by the computing device, a third value of the evaluation metric for the third iteration of the machine learning model; and

generating, by the computing device, a representative subset based on the third value of the evaluation metric value.

7 . A computer program product for providing a representative dataset from an initial dataset, the computer program product comprising one or more computer readable storage media and collectively stored program instructions on the one or more computer readable storage media, the stored program instructions comprising:

program instructions to access a dataset associated with a machine learning model;

program instructions to receive input parameters associated with the representative dataset selection, the input parameters including an evaluation metric;

program instructions to determine a density of a plurality of datapoints associated with the dataset;

program instructions to train a first iteration of the machine learning model using a first data point selected according to the density;

program instructions to determine a first value of the evaluation metric for the first iteration of the machine learning model;

program instructions to identify neighbors and non-neighbors of the first data point;

program instructions to determine a dissimilarity between the first data point and all non-neighbors;

program instructions to train a second iteration of the machine learning model using a second point selected according to the dissimilarity;

program instructions to determine a second value of the evaluation metric for the second iteration of the machine learning model;

program instructions to generate the representative dataset based on the second value of the evaluation metric;

program instructions to train a final machine learning model using the representative dataset and

program instructions to provide the representative dataset and the final machine learning model trained using the representative dataset.

8 . The computer program product according to claim 7 , the stored program instructions further comprising program instructions to change all neighbors to inactive, and program instructions to determine a dissimilarity between the first data point and all active points.

9 . The computer program product according to claim 7 , wherein the input parameters comprise a neighborhood size.

10 . The computer program product according to claim 9 , the stored program instructions further comprising program instructions to evaluate a coverage according to the neighborhood size, wherein determining a coverage of the representative dataset comprises determining a ratio of points directly reachable from a point in the representative dataset and all points in the initial dataset.

11 . The computer program product according to claim 7 , the stored program instructions further comprising:

program instructions to receive a target column as user input;

program instructions to determine a correlation between dataset columns and the target column; and

program instructions to retain data point column values according to the correlation.

12 . The computer program product according to claim 7 , the stored program instructions further comprising:

program instructions to define a budget;

program instructions to evaluate all initial dataset points without reaching the budget;

program instructions to change all points not in the representative dataset to active;

program instructions to train a third iteration of the machine learning model using a third point selected according to the density;

program instructions to determine a third value of the evaluation metric for the third iteration of the machine learning model; and

program instructions to generate a representative subset based on the third value of the evaluation metric value.

13 . A computer system for providing a representative dataset from an initial dataset, the computer system comprising:

one or more computer processors;

one or more computer readable storage devices; and

stored program instructions on the one or more computer readable storage devices for execution by the one or more computer processors, the stored program instructions comprising:

program instructions to access a dataset associated with a machine learning model;

program instructions to receive input parameters associated with the representative dataset selection, the input parameters including an evaluation metric;

program instructions to determine a density of a plurality of datapoints associated with the dataset;

program instructions to train a first iteration of the machine learning model using a first data point selected according to the density;

program instructions to determine a first value of the evaluation metric for the first iteration of the machine learning model;

program instructions to identify neighbors and non-neighbors of the first data point;

program instructions to determine a dissimilarity between the first data point and all non-neighbors;

program instructions to train a second iteration of the machine learning model using a second point selected according to the dissimilarity;

program instructions to determine a second value of the evaluation metric for the second iteration of the machine learning model;

program instructions to generate the representative dataset based on the second value of the evaluation metric;

program instructions to train a final machine learning model using the representative dataset and

program instructions to provide the representative dataset and the final machine learning model trained using the representative dataset.

14 . The computer system according to claim 13 , wherein the input parameters comprise a neighborhood size.

15 . The computer system according to claim 14 , the stored program instructions further comprising program instructions to evaluate a coverage according to the neighborhood size, wherein determining a coverage of the representative dataset comprises determining a ratio of points directly reachable from a point in the representative dataset and all points in the initial dataset.

16 . The computer system according to claim 13 , the stored program instructions further comprising:

program instructions to receive a target column as user input;

program instructions to determine a correlation between dataset columns and the target column; and

program instructions to retain data point column values according to the correlation.

17 . The computer system according to claim 13 , the stored program instructions further comprising:

program instructions to define a budget;

program instructions to evaluate all initial dataset points without reaching the budget;

program instructions to change all points not in the representative dataset to active;

program instructions to train a third iteration of the machine learning model using a third point selected according to the density;

program instructions to determine a third value of the evaluation metric for the third iteration of the machine learning model; and

program instructions to generate a representative subset based on the third value of the evaluation metric value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2022
From: QUADER, SHAIKH SHAHRIAR; BASAK, AINDRILA; MAHJOUR, ADRIAN; NOVOTNY, PETR; APPUGLIESE, CARLO; REINWALD, BERTHOLD; ARREMSETTY, DHEERAJ
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 060871/0896 →
Continuity (1)
Related Publication 20240070522A1 · Feb 29, 2024
References Cited (15)
US 11295234B2 · Ardhanari · 2022 [cited by applicant]
US 20180012143A1 · Hansen · 2018 [cited by examiner]
US 20200050968A1 · Lee · 2020 [cited by examiner]
US 20200342265A1 · Cai · 2020 [cited by applicant]
US 20210342642A1 · Shabtay · 2021 [cited by applicant]
US 20230370520A1 · Shetty · 2023 [cited by examiner]
CN 111612146A · 2020 [cited by examiner]
Bachem et al., “Practical Coreset Constructions for Machine Learning”, Jun. 4, 2017, 39 pps., <https://arxiv.org/abs/1703.06476>. [cited by applicant]
Clark, “OptiSim: An Extended Dissimilarity Selection Method for Finding Diverse Representative Subsets”, J. Chem. Inf. Comput. Sci. 1997, 37, 6, 1181-1188, Publication Date:Nov. 24, 1997, <https://doi.org/10.1021/ci9702… [cited by applicant]
Kennard et al., “Computer Aided Design of Experiments”, Feb. 1969, Technometrics, vol. 11, No. I, pp. 137-148, <http://libpls.net/publication/KS_1969.pdf>. [cited by applicant]
Lai et al., “Exploring high-dimensional data through locally enhanced projections”, Journal of Visual Languages & Computing, 13 pps., <https://doi.org/10.1016/j.jvlc.2018.08.006>. [cited by applicant]
Mall et al., “FURS: Fast and Unique Representative Subset selection retaining large-scale community structure”, Published: Oct. 22, 2013, Social Network Analysis and Mining vol. 3, pp. 1075-1095, <https://doi.org/10.100… [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, National Institute of Standards and Technology, U.S. Department of Commerce, NIST Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]
Tremblay et al., “Determinantal Point Processes for Coresets”, Journal of Machine Learning Research 20 (2019) 1-69, Submitted Mar. 2018; Revised Oct. 2019; Published Nov. 2019, <https://arxiv.org/abs/1803.08700>. [cited by applicant]
Williamson et al., “Understanding Collections of Related Datasets Using Dependent MMD Coresets”, Information 12, No. 10, 2021, 27 pages. [cited by applicant]