IP Library › Granted Patent US 12,619,906
Granted Patent B2
US 12,619,906 · App. 17/216,475 · Granted May 5, 2026

Interactive machine learning optimization

Inventors: Dhavalkumar C. Patel (White Plains, NY); Si Er Han (Xi'an, CN); Jiang Bo Kang (Xi'an, CN)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N20/00G06F8/34G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,906
App. No.
17/216,475
Granted
May 5, 2026
Kind
B2
Abstract

Methods, computer program products, and systems are presented. The method, computer program products, and systems can include, for instance: examining an enterprise dataset, the enterprise dataset defined by enterprise collected data; selecting one or more synthetic dataset in dependence on the examining, the one or more synthetic dataset including data other than data collected by the enterprise; training a set of predictive models using data of the one or more synthetic dataset to provide a set of trained predictive models; testing the set of trained predictive models with use of holdout data of the one or more synthetic dataset; and presenting prompting data on a displayed user interface of a developer user in dependence on result data resulting from the testing, the prompting data prompting the developer user to direct action with respect to one or more model of the set of predictive models.

Claims (39)

1 . A computer implemented method comprising:

examining an enterprise dataset, the enterprise dataset defined by enterprise collected data;

selecting one or more synthetic dataset in dependence on the examining, the one or more synthetic dataset including data other than data collected by the enterprise, wherein the examining the enterprise dataset includes subjecting the enterprise dataset to processing for extraction of enterprise dataset characterizing parameter values and comparing the enterprise dataset characterizing parameter values to respective sets of synthetic dataset characterizing parameter values stored in a data repository, wherein the respective sets of synthetic dataset characterizing parameter values characterize respective synthetic datasets stored in the data repository, wherein the selecting one or more synthetic dataset in dependence on the examining includes identifying from the comparing at least one synthetic dataset stored in the data repository having a threshold satisfying similarity with the enterprise dataset and identifying from the respective synthetic datasets stored in the data repository a highest ranked synthetic dataset having a greatest similarity to the enterprise dataset amongst the respective synthetic datasets stored in the data repository, and wherein the selecting one or more synthetic dataset in dependence on the examining includes performing the selecting so that the selected one or more synthetic dataset have the threshold satisfying similarity with the enterprise dataset and include the highest ranked synthetic dataset having the greatest similarity to the enterprise dataset amongst the respective synthetic datasets stored in the data repository;

training a set of predictive models using data of the selected one or more synthetic dataset having the threshold satisfying similarity with the enterprise dataset and including the highest ranked synthetic dataset having the greatest similarity to the enterprise dataset amongst the respective synthetic datasets stored in the data repository to provide a set of trained predictive models;

testing the set of trained predictive models with use of holdout data of the one or more synthetic dataset; and

presenting prompting data on a displayed user interface of a developer user in dependence on result data resulting from the testing, the prompting data prompting the developer user to direct action with respect to one or more model of the set of predictive models, wherein the prompting data prompting the developer user to direct action with respect to one or more model of the set of predictive models includes prompting data prompting the developer user to select, in dependence on data of the result data, a certain synthetic dataset for training and testing a certain predictive model of the set of predictive models, wherein the user interface guides the developer user by presenting a recommended datasets text area displaying recommended synthetic datasets including a highlighted dataset indicator distinguishing a recommended dataset from other candidate datasets, and restricts the developer user from selecting an unrecommended dataset for a predictive model such that, during a subsequent iteration, different predictive models are trained on differentiated synthetic datasets.

2 . The method of claim 1 , wherein testing the set of trained predictive models includes generating a ranked listing of model specific performance metrics, and wherein the user interface presents a model identifier text area that displays the ranked listing to visually correlate each predictive model with its respective performance metric.

3 . The method of claim 1 , wherein the user interface presents a results visualization area that plots predicted signals produced by multiple predictive models alongside a ground truth signal to enable qualitative comparison of model behavior over a shared time axis.

4 . The method of claim 1 , wherein the user interface presents a model attributes text area that displays attributes including model type, capacity, and training iteration count, and dynamically updates the displayed attributes in dependence on testing result data.

5 . The method of claim 1 , wherein the presenting prompting data includes displaying classifications of candidate synthetic datasets by attributes including seasonality, trend, variance, outliers, or level shift, and wherein the classifications are generated by processing metadata of the respective synthetic datasets stored in the data repository.

6 . The method of claim 1 , wherein the presenting prompting data includes displaying integrated visual guidance that correlates model specific error signals with dataset selection information derived from the testing, enabling the developer user to interpret relationships between error behavior and training data characteristics.

7 . The method of claim 1 , wherein the presenting prompting data includes querying a machine learning trained model that is trained to predict a next dataset for training a predictive model of the set of predictive models, wherein the trained model is configured using crowdsourced developer action data identifying historical developer selections of synthetic datasets.

8 . A system comprising:

a memory;

at least one processor in communication with the memory; and

program instructions executable by one or more processor via the memory to perform a method comprising:

examining an enterprise dataset, the enterprise dataset defined by enterprise collected data;

selecting one or more synthetic dataset in dependence on the examining, the one or more synthetic dataset including data other than data collected by the enterprise, wherein the examining the enterprise dataset includes subjecting the enterprise dataset to processing for extraction of enterprise dataset characterizing parameter values and comparing the enterprise dataset characterizing parameter values to respective sets of synthetic dataset characterizing parameter values stored in a data repository, wherein the respective sets of synthetic dataset characterizing parameter values characterize respective synthetic datasets stored in the data repository, wherein the selecting one or more synthetic dataset in dependence on the examining includes identifying from the comparing at least one synthetic dataset stored in the data repository having a threshold satisfying similarity with the enterprise dataset and identifying from the respective synthetic datasets stored in the data repository a highest ranked synthetic dataset having a greatest similarity to the enterprise dataset amongst the respective synthetic datasets stored in the data repository, and wherein the selecting one or more synthetic dataset in dependence on the examining includes performing the selecting so that the selected one or more synthetic dataset have the threshold satisfying similarity with the enterprise dataset and include the highest ranked synthetic dataset having the greatest similarity to the enterprise dataset amongst the respective synthetic datasets stored in the data repository;

training a set of predictive models using data of the selected one or more synthetic dataset having the threshold satisfying similarity with the enterprise dataset and including the highest ranked synthetic dataset having the greatest similarity to the enterprise dataset amongst the respective synthetic datasets stored in the data repository to provide a set of trained predictive models;

testing the set of trained predictive models with use of holdout data of the one or more synthetic dataset; and

presenting prompting data on a displayed user interface of a developer user in dependence on result data resulting from the testing, the prompting data prompting the developer user to direct action with respect to one or more model of the set of predictive models, wherein the prompting data prompting the developer user to direct action with respect to one or more model of the set of predictive models includes prompting data prompting the developer user to select, in dependence on data of the result data, a certain synthetic dataset for training and testing a certain predictive model of the set of predictive models, wherein the user interface guides the developer user by presenting a recommended datasets text area displaying recommended synthetic datasets including a highlighted dataset indicator distinguishing a recommended dataset from other candidate datasets, and restricts the developer user from selecting an unrecommended dataset for a predictive model such that, during a subsequent iteration, different predictive models are trained on differentiated synthetic datasets.

9 . The system of claim 8 , wherein testing the set of trained predictive models includes generating a ranked listing of model specific performance metrics, and wherein the user interface presents a model identifier text area that displays the ranked listing to visually correlate each predictive model with its respective performance metric.

10 . The system of claim 8 , wherein the user interface presents a results visualization area that plots predicted signals produced by multiple predictive models alongside a ground truth signal to enable qualitative comparison of model behavior over a shared time axis.

11 . The system of claim 8 , wherein the user interface presents a model attributes text area that displays attributes including model type, capacity, and training iteration count, and dynamically updates the displayed attributes in dependence on testing result data.

12 . The system of claim 8 , wherein the presenting prompting data includes displaying classifications of candidate synthetic datasets by attributes including seasonality, trend, variance, outliers, or level shift, and wherein the classifications are generated by processing metadata of the respective synthetic datasets stored in the data repository.

13 . The system of claim 8 , wherein the presenting prompting data includes displaying integrated visual guidance that correlates model specific error signals with dataset selection information derived from the testing, enabling the developer user to interpret relationships between error behavior and training data characteristics.

14 . The system of claim 8 , wherein the presenting prompting data includes querying a machine learning trained model that is trained to predict a next dataset for training a predictive model of the set of predictive models, wherein the trained model is configured using crowdsourced developer action data identifying historical developer selections of synthetic datasets.

15 . A computer program product comprising:

a computer readable storage medium readable by one or more processing circuit and storing instructions for execution by one or more processor for performing a method comprising:

examining an enterprise dataset, the enterprise dataset defined by enterprise collected data;

selecting one or more synthetic dataset in dependence on the examining, the one or more synthetic dataset including data other than data collected by the enterprise, wherein the examining the enterprise dataset includes subjecting the enterprise dataset to processing for extraction of enterprise dataset characterizing parameter values and comparing the enterprise dataset characterizing parameter values to respective sets of synthetic dataset characterizing parameter values stored in a data repository, wherein the respective sets of synthetic dataset characterizing parameter values characterize respective synthetic datasets stored in the data repository, wherein the selecting one or more synthetic dataset in dependence on the examining includes identifying from the comparing at least one synthetic dataset stored in the data repository having a threshold satisfying similarity with the enterprise dataset and identifying from the respective synthetic datasets stored in the data repository a highest ranked synthetic dataset having a greatest similarity to the enterprise dataset amongst the respective synthetic datasets stored in the data repository, and wherein the selecting one or more synthetic dataset in dependence on the examining includes performing the selecting so that the selected one or more synthetic dataset have the threshold satisfying similarity with the enterprise dataset and include the highest ranked synthetic dataset having the greatest similarity to the enterprise dataset amongst the respective synthetic datasets stored in the data repository;

training a set of predictive models using data of the selected one or more synthetic dataset having the threshold satisfying similarity with the enterprise dataset and including the highest ranked synthetic dataset having the greatest similarity to the enterprise dataset amongst the respective synthetic datasets stored in the data repository to provide a set of trained predictive models;

testing the set of trained predictive models with use of holdout data of the one or more synthetic dataset; and

presenting prompting data on a displayed user interface of a developer user in dependence on result data resulting from the testing, the prompting data prompting the developer user to direct action with respect to one or more model of the set of predictive models, wherein the prompting data prompting the developer user to direct action with respect to one or more model of the set of predictive models includes prompting data prompting the developer user to select, in dependence on data of the result data, a certain synthetic dataset for training and testing a certain predictive model of the set of predictive models, wherein the user interface guides the developer user by presenting a recommended datasets text area displaying recommended synthetic datasets including a highlighted dataset indicator distinguishing a recommended dataset from other candidate datasets, and restricts the developer user from selecting an unrecommended dataset for a predictive model such that, during a subsequent iteration, different predictive models are trained on differentiated synthetic datasets.

16 . The computer program product of claim 15 , wherein testing the set of trained predictive models includes generating a ranked listing of model specific performance metrics, and wherein the user interface presents a model identifier text area that displays the ranked listing to visually correlate each predictive model with its respective performance metric.

17 . The computer program product of claim 15 , wherein the user interface presents a results visualization area that plots predicted signals produced by multiple predictive models alongside a ground truth signal to enable qualitative comparison of model behavior over a shared time axis.

18 . The computer program product of claim 15 , wherein the user interface presents a model attributes text area that displays attributes including model type, capacity, and training iteration count, and dynamically updates the displayed attributes in dependence on testing result data.

19 . The computer program product of claim 15 , wherein the presenting prompting data includes displaying integrated visual guidance that correlates model specific error signals with dataset selection information derived from the testing, enabling the developer user to interpret relationships between error behavior and training data characteristics.

20 . The computer program product of claim 15 , wherein the presenting prompting data includes querying a machine learning trained model that is trained to predict a next dataset for training a predictive model of the set of predictive models, wherein the trained model is configured using crowdsourced developer action data identifying historical developer selections of synthetic datasets.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2021
From: PATEL, DHAVALKUMAR C.; HAN, SI ER; KANG, JIANG BO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 055758/0857 →
Continuity (1)
Related Publication 20220309391A1 · Sep 29, 2022
References Cited (45)
US 10860466B1 · Baldwin · 2020 [cited by examiner]
US 20140115007A1 · Harvey · 2014 [cited by examiner]
US 20150149895A1 · Furman · 2015 [cited by examiner]
US 20180068221A1 · Brennan · 2018 [cited by examiner]
US 20190156196A1 · Zoldi et al. · 2019 [cited by applicant]
US 20190392252A1 · Fighel et al. · 2019 [cited by applicant]
US 20200012657A1 · Walters · 2020 [cited by examiner]
US 20200193234A1 · Pai et al. · 2020 [cited by applicant]
US 20200201727A1 · Nie et al. · 2020 [cited by applicant]
US 20200380301A1 · Siracusa · 2020 [cited by examiner]
US 20200380409A1 · Seo · 2020 [cited by examiner]
US 20210125068A1 · Kim · 2021 [cited by examiner]
US 20210256420A1 · Elisha · 2021 [cited by examiner]
US 20220101186A1 · Sharma Mittal · 2022 [cited by examiner]
US 20220292239A1 · Kahraman · 2022 [cited by examiner]
Zhang, “A neural network ensemble method with jittered training data for time series forecasting”, 2007, Information Sciences, vol. 177 No. 23, pp. 5329-5346 (Year: 2007). [cited by examiner]
Zhang et al., “Learning Optimal Data Augmentation Policies via Bayesian Optimization for Image Classification Tasks”, 2019, arXiv, v2, pp. 1-23 (Year: 2019). [cited by examiner]
Bader et al., “Getafix: Learning to Fix Bugs Automatically”, 2019, Proceedings of the ACM on Programming Languages, vol. 3 (2019 ), pp. 1-27 (Year: 2019). [cited by examiner]
Santos et al., “Visus: An Interactive System for Automatic Machine Learning Model Building and Curation”, 2019, HILDA '19: Proceedings of the Workshop on Human-In-the-Loop Data Analytics, vol. 2019, pp. 1-7 (Year: 2019). [cited by examiner]
Hohman et al., “Understanding and Visualizing Data Iteration in Machine Learning”, 2020, Proceedings of the 2020 CHI conference on human factors in computing systems, vol. 2020, pp. 1-13 (Year: 2020). [cited by examiner]
Ding et al., “DAGA: Data Augmentation with a Generation Approach for Low-resource Tagging Task”, 2020, arXIv, v1, pp. 1-13 ( Year: 2020). [cited by examiner]
Saratchandran, “How Enterprise Software is Getting Intelligent Through Machine Learning”, 2018, Medium, retrieved from https://medium.com/data-science/how-enterprise-software-is-getting-intelligent-through-machine-learn… [cited by examiner]
Raczko et al., “Comparison of support vector machine, random forest and neural network classifiers for tree species classification on airborne hyperspectral APEX images”, 2017, European Journal of Remote Sensing, vol. 5… [cited by examiner]
Fons et al., “Adaptive weighting scheme for automatic time-series data augmentation”, Feb. 16, 2021, arXiv, v1, pp. 1-10 (Year: 2021). [cited by examiner]
Xin et al., “Accelerating Human-in-the-loop Machine Learning: Challenges and Opportunities”, 2018, DEEM'18: Proceedings of the Second Workshop on Data Management for End-To-End Machine Learning, vol. 2018, pp. 1-4 (Year… [cited by examiner]
Kohli, et al., “Identifying and eliminating bugs in learned predictive models”, DeepMind, Mar. 28, 2019, 12 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://deepmind.com/blog/article/robust-and-verified-a… [cited by applicant]
“TSimulus—A realistic time series generator”, 2016, 2 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://simulus.readthedocs.io/en/latest/>. [cited by applicant]
Herlambang, “Data Generator for Supervised Time Series”, kerasgenerator, Oct. 9, 2019, 29 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://kerasgenerator.bagasbgy.com/articles/timeseries.html>. [cited by applicant]
Yldrm, “Time Series Analysis: Creating Synthetic Datasets”, Towards Data Science, May 16, 2020, 9 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://towardsdatascience.com/time-series-analysis-creating-synt… [cited by applicant]
Kang et al., “GRATIS: GeneRAting Time Series with diverse and controllable characteristics”, 2019, 2 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://robjhyndman.com/publications/tsgeneration/>. [cited by applicant]
Kang et al., “GRATIS: GeneRAting Time Series with diverse and controllable characteristics”, 2019, 38 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://arxiv.org/ftp/arxiv/papers/1903/1903.02787.pdf>. [cited by applicant]
Goswami, “tsBNgen, a Python Library to Generate Synthetic Data From an Arbitrary Bayesian Network”, MarketTechPost, Sep. 15, 2020, 3 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://www.marktechpost.com/2… [cited by applicant]
Wen, “TSAUG: An Open-Source Python Package for Time Series Augmentation”, Arundo, Nov. 20, 2019, 5 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://www.arundo.com/arundo_tech_blog/tsaug-an-open-source-pyt… [cited by applicant]
Tadayon, et al., “tsBNgen: A Python Library to Generate Time Series Data from an Arbitrary Dynamic Bayesian Network Structure”, Sep. 9, 2020, 4 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://arxiv.org/p… [cited by applicant]
Schwartz, B., “Generating Realistic Time Series Data”, Jan. 24, 2014, 2 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://www.xaprb.com/blog/2014/01/24/methods-generate-realistic-time-series-data/>. [cited by applicant]
“Getting Started”, SDV, MIT Data to A1 Lab, 2018, 2 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://sdv.dev/SDV/getting_started/index.html>. [cited by applicant]
“Keras TimeseriesGenerator tutorial”, 15 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://github.com/Tony607/Keras_TimeseriesGenerator/blob/master/TimeseriesGenerator.ipynb>. [cited by applicant]
“TSimulus—A realistic time series generator”, Cetic A.S.B.L., 2016, 2 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://tsimulus.readthedocs.io/en/latest/>. [cited by applicant]
Herrera, et al., “A Time Series Simulator Processor for NiFi”, Apr. 9, 2018, 4 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://github.com/hashmapinc/nifi-simulator-bundle>. [cited by applicant]
Lichtenwalner, “Modeling Seasonal Data”, Ocean Data Labs, Mar. 24, 2020, 19 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://datalab.marine.rutgers.edu/2020/03/modeling-seasonal-data/>. [cited by applicant]
Uchidalab, “time series augmentation”, 6 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://github.com/uchidalab/time_series_augmentation/blob/master/example.ipynb>. [cited by applicant]
Terryum, “Data Augmentation For Wearable Sensor Data”, 3 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://github.com/terryum/Data-Augmentation-For-Wearable-Sensor-Data>. [cited by applicant]
RainBoltz, “time series augmentation toolkit”, 2 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://github.com/RainBoltz/time-series-augmentation-toolkit/blob/master/time_series_augmentation_toolkit.py>. [cited by applicant]
KDD-OpenSource, “agots”, 4 pgs. Retrieved on Mar. 31, 2021 from the Internet URL: <https://github.com/KDD-OpenSource/agots>. [cited by applicant]
Mell, Peter, et al., “The NIST Definition of Cloud Computing”, NIST Special Publication 800-145, Sep. 2011, Gaithersburg, MD, 7 pgs. [cited by applicant]