IP Library › Granted Patent US 12,749,007
Granted Patent B2
US 12,749,007 · App. 17/398,215 · Granted Sep 29, 2026

Initialize optimized parameter in data processing system

Inventors: A Peng Zhang (Xian, CN); Lei Gao (Xian, CN); Jin Wang (Xi'an, CN); Jia Xing Tang (Xian, CN); Kai Li (Xian, CN); Geng Wu Yang (Xian, CN); Zhen Liu (Xian, CN)
Assignee: International Business Machines Corporation
G06N20/00G06F8/60G06F11/3452G06F11/3684G06F11/3688G06F11/3692
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,749,007
App. No.
17/398,215
Granted
Sep 29, 2026
Kind
B2
Abstract

An approach is provided in which the approach loads a machine learning model and a set of test case statistical data into a user system. The set of test case statistical data is based on a set of test cases corresponding to the machine learning model and includes a plurality of input parameter sets and a corresponding set of output quality measurements. The approach compares user data on the user system against the set of test case statistical data and identifies one of the plurality of input parameter sets to optimize the machine learning model based on the set of output quality measurements. The approach generates an optimized machine learning model using the machine learning model and the identified input parameter set at the user system.

Claims (114)

1 . A computer-implemented method comprising:

loading a machine learning model and a set of test case statistical data into a user system, wherein the set of test case statistical data is based on a set of test cases corresponding to the machine learning model and comprises a plurality of input parameter sets and a corresponding set of output quality measurements, and wherein the machine learning model is built by:

transforming univariate and bivariate statistics into a scaled format such that the statistics are combined through feature scaling, wherein transforming the univariate and bivariate statistics comprises:

sorting scaled statistic values and counting the values within predefined intervals to form fixed-length statistical distribution vectors, wherein the scaled statistic values are processed so that number of values falling within each predefined equal-width interval are determined in a manner that enables a uniform representation of data characteristics; and

building the machine learning model based on the transformed univariate and bivariate statistics, and the plurality of input parameter sets;

comparing user data comprising a second transformed univariate and bivariate statistics on the user system against the set of test case statistical data, wherein the comparison identifies one of the plurality of input parameter sets to optimize the machine learning model based on the set of output quality measurements;

generating, at the user system, an optimized machine learning model using the machine learning model and the identified input parameter set;

displaying one or more similar sets of test data on a user interface;

responsive to receiving through the user interface a selection choice from the one or more similar sets of test data, applying the selection choice to the optimized machine learning model and predicting an accuracy of the optimized machine learning model;

outputting the optimized machine learning model to the user through the user interface; and

responsive to receiving acceptance of the optimized machine learning model through the user interface executing the optimized machine learning model and refreshing the optimized machine learning model, by collecting information to add a new pre-trained machine learning model.

2 . The computer-implemented method of claim 1 wherein, at a developer system, the method further comprises:

collecting a set of test data corresponding to the set of test cases;

transforming the set of test data into a set of transformed descriptive statistics, wherein the transforming comprises a set of analytic computations, a set of scaling computations, and a set of sorting computations;

running the set of test cases with the plurality of input parameter sets to generate the set of output quality measurements; and

constructing the set of test case statistical data by combining the set of transformed descriptive statistics, the set of output quality measurements, and the plurality of input parameter sets.

3 . The computer-implemented method of claim 2 further comprising:

packaging the test case statistical data, the machine learning model, and an initial parameter optimizer into a deployment package; and

deploying the deployment package from the developer system to the user system.

4 . The computer-implemented method of claim 3 wherein, at the user system, the method further comprises:

collecting, by the initial parameter optimizer, a set of user data characteristics of the user data;

transforming the set of user data characteristics into a set of transformed user data statistics;

receiving a set of user parameters from a user;

predicting a model quality of the optimized machine learning model based on the set of user parameters, the set of transformed user data statistics, and the set of output quality measurements; and

displaying the predicted model quality of the optimized machine learning model to the user at the user system.

5 . The computer-implemented method of claim 3 further comprising:

collecting, by the initial parameter optimizer, a set of user data characteristics of the user data;

transforming the set of user data characteristics into a set of transformed user data statistics;

calculating a set of data similarity values between the set of transformed user data statistics and the set of transformed descriptive statistics; and

selecting a subset of the plurality of input parameter sets from the test case statistical data based on the set of data similarity values.

6 . The computer-implemented method of claim 5 further comprising:

predicting a set of model quality values of the optimized machine learning model based on the subset of input parameter sets, the set of transformed user data statistics, and the set of output quality measurements;

displaying the set of model quality values and corresponding subset of input parameter sets to the user at the user system;

receiving a selection from the user that selects one of the subsets of input parameter sets; and

creating the optimized machine learning model using the selected subset of input parameter sets.

7 . The computer-implemented method of claim 6 further comprising:

receiving, from the user, a different selection that selects a different one of the subsets of input parameter sets, wherein the selection and the different selection are received concurrently; and

creating a different optimized machine learning model using the different subset of input parameter sets, wherein the different optimized machine learning model is created concurrently with the optimized machine learning model.

8 . An information handling system comprising:

one or more processors;

a memory coupled to at least one of the processors;

a set of computer program instructions stored in the memory and executed by at least one of the processors in order to perform actions of:

loading a machine learning model and a set of test case statistical data into a user system, wherein the set of test case statistical data is based on a set of test cases corresponding to the machine learning model and comprises a plurality of input parameter sets and a corresponding set of output quality measurements, and wherein the machine learning model is built by:

transforming univariate and bivariate statistics into a scaled format such that the statistics are combined through feature scaling, wherein transforming the univariate and bivariate statistics comprises:

sorting scaled statistic values and counting the values within predefined intervals to form fixed-length statistical distribution vectors, wherein the scaled statistic values are processed so that number of values falling within each predefined equal-width interval are determined in a manner that enables a uniform representation of data characteristics; and

building the machine learning model based on the transformed univariate and bivariate statistics, and the plurality of input parameter sets;

comparing user data comprising a second transformed univariate and bivariate statistics on the user system against the set of test case statistical data, wherein the comparison identifies one of the plurality of input parameter sets to optimize the machine learning model based on the set of output quality measurements;

generating, at the user system, an optimized machine learning model using the machine learning model and the identified input parameter set;

displaying one or more similar sets of test data on a user interface;

responsive to receiving through the user interface a selection choice from the one or more similar sets of test data, applying the selection choice to the machine learning model and predicting an accuracy of the optimized machine learning model;

outputting the optimized machine learning model to the user through the user interface; and

responsive to receiving acceptance of the optimized machine learning model through the user interface executing the optimized machine learning model and refreshing the optimized machine learning model, by collecting information to add a new pre-trained machine learning model.

9 . The information handling system of claim 8 wherein the processors perform additional actions comprising:

collecting, at a developer system, a set of test data corresponding to the set of test cases;

transforming, at the developer system, the set of test data into a set of transformed descriptive statistics, wherein the transforming comprises a set of analytic computations, a set of scaling computations, and a set of sorting computations;

running, at the developer system, the set of test cases with the plurality of input parameter sets to generate the set of output quality measurements; and

constructing, at the developer system, the set of test case statistical data by combining the set of transformed descriptive statistics, the set of output quality measurements, and the plurality of input parameter sets.

10 . The information handling system of claim 9 wherein the processors perform additional actions comprising:

packaging the test case statistical data, the machine learning model, and an initial parameter optimizer into a deployment package; and

deploying the deployment package from the developer system to the user system.

11 . The information handling system of claim 10 wherein the processors perform additional actions comprising:

collecting, at the user system by the initial parameter optimizer, a set of user data characteristics of the user data;

transforming the set of user data characteristics into a set of transformed user data statistics;

receiving a set of user parameters from a user;

predicting a model quality of the optimized machine learning model based on the set of user parameters, the set of transformed user data statistics, and the set of output quality measurements; and

displaying the predicted model quality of the optimized machine learning model to the user at the user system.

12 . The information handling system of claim 10 wherein the processors perform additional actions comprising:

collecting, by the initial parameter optimizer, a set of user data characteristics of the user data;

transforming the set of user data characteristics into a set of transformed user data statistics;

calculating a set of data similarity values between the set of transformed user data statistics and the set of transformed descriptive statistics; and

selecting a subset of the plurality of input parameter sets from the test case statistical data based on the set of data similarity values.

13 . The information handling system of claim 12 wherein the processors perform additional actions comprising:

predicting a set of model quality values of the optimized machine learning model based on the subset of input parameter sets, the set of transformed user data statistics, and the set of output quality measurements;

displaying the set of model quality values and corresponding subset of input parameter sets to the user at the user system;

receiving a selection from the user that selects one of the subsets of input parameter sets; and

creating the optimized machine learning model using the selected subset of input parameter sets.

14 . The information handling system of claim 13 wherein the processors perform additional actions comprising:

receiving, from the user, a different selection that selects a different one of the subsets of input parameter sets, wherein the selection and the different selection are received concurrently; and

creating a different optimized machine learning model using the different subset of input parameter sets, wherein the different optimized machine learning model is created concurrently with the optimized machine learning model.

15 . A computer program product stored in a computer readable storage medium, comprising computer program code that, when executed by an information handling system, causes the information handling system to perform actions comprising:

loading a machine learning model and a set of test case statistical data into a user system, wherein the set of test case statistical data is based on a set of test cases corresponding to the machine learning model and comprises a plurality of input parameter sets and a corresponding set of output quality measurements, and wherein the machine learning model is built by:

transforming univariate and bivariate statistics into a scaled format such that the statistics are combined through feature scaling, wherein transforming the univariate and bivariate statistics comprises:

sorting scaled statistic values and counting the values within predefined intervals to form fixed-length statistical distribution vectors, wherein the scaled statistic values are processed so that number of values falling within each predefined equal-width interval are determined in a manner that enables a uniform representation of data characteristics; and

building the machine learning model based on the transformed univariate and bivariate statistics, and the plurality of input parameter sets;

comparing user data comprising a second transformed univariate and bivariate statistics on the user system against the set of test case statistical data, wherein the comparison identifies one of the plurality of input parameter sets to optimize the machine learning model based on the set of output quality measurements;

generating, at the user system, an optimized machine learning model using the machine learning model and the identified input parameter set;

displaying one or more similar sets of test data on a user interface;

responsive to receiving through the user interface a selection choice from the one or more similar sets of test data, applying the selection choice to the machine learning model and predicting an accuracy of the optimized machine learning model;

outputting the optimized machine learning model to the user through the user interface; and

responsive to receiving acceptance of the optimized machine learning model through the user interface executing the optimized machine learning model and refreshing the optimized machine learning model, by collecting information to add a new pre-trained machine learning model.

16 . The computer program product of claim 15 wherein, at a developer system, the information handling system performs further actions comprising:

collecting a set of test data corresponding to the set of test cases;

transforming the set of test data into a set of transformed descriptive statistics, wherein the transforming comprises a set of analytic computations, a set of scaling computations, and a set of sorting computations;

running the set of test cases with the plurality of input parameter sets to generate the set of output quality measurements; and

constructing the set of test case statistical data by combining the set of transformed descriptive statistics, the set of output quality measurements, and the plurality of input parameter sets.

17 . The computer program product of claim 16 wherein the information handling system performs further actions comprising:

packaging the test case statistical data, the machine learning model, and an initial parameter optimizer into a deployment package; and

deploying the deployment package from the developer system to the user system.

18 . The computer program product of claim 17 wherein, at the user system, the information handling system performs further actions comprising:

collecting, by the initial parameter optimizer, a set of user data characteristics of the user data;

transforming the set of user data characteristics into a set of transformed user data statistics;

receiving a set of user parameters from a user;

predicting a model quality of the optimized machine learning model based on the set of user parameters, the set of transformed user data statistics, and the set of output quality measurements; and

displaying the predicted model quality of the optimized machine learning model to the user at the user system.

19 . The computer program product of claim 17 wherein the information handling system performs further actions comprising:

collecting, by the initial parameter optimizer, a set of user data characteristics of the user data;

transforming the set of user data characteristics into a set of transformed user data statistics;

calculating a set of data similarity values between the set of transformed user data statistics and the set of transformed descriptive statistics; and

selecting a subset of the plurality of input parameter sets from the test case statistical data based on the set of data similarity values.

20 . The computer program product of claim 19 wherein the information handling system performs further actions comprising:

predicting a set of model quality values of the optimized machine learning model based on the subset of input parameter sets, the set of transformed user data statistics, and the set of output quality measurements;

displaying the set of model quality values and corresponding subset of input parameter sets to the user at the user system;

receiving a selection from the user that selects one of the subsets of input parameter sets; and

creating the optimized machine learning model using the selected subset of input parameter sets.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 10, 2021
From: ZHANG, A PENG; GAO, LEI; WANG, JIN; TANG, JIA XING; LI, KAI; YANG, GENG WU; LIU, ZHEN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 057132/0581 →
Continuity (1)
Related Publication 20230052848A1 · Feb 16, 2023
References Cited (21)
US 8862185B2 · Callard · 2014 [cited by applicant]
US 20060161403A1 · Jiang · 2006 [cited by examiner]
US 20190370684A1 · Gunes · 2019 [cited by applicant]
US 20200050717A1 · Kiribuchi · 2020 [cited by examiner]
US 20200097847A1 · Convertino · 2020 [cited by examiner]
US 20200167691A1 · Golovin · 2020 [cited by applicant]
US 20210027182A1 · Harris · 2021 [cited by applicant]
US 20210174210A1 · Hsyu · 2021 [cited by examiner]
US 20210303442A1 · Chenguttuvan · 2021 [cited by examiner]
US 20220004822A1 · Vaid · 2022 [cited by examiner]
US 20220076164A1 · Conort · 2022 [cited by examiner]
US 20220138004A1 · Nandakumar · 2022 [cited by examiner]
US 20230132064A1 · Zhang · 2023 [cited by examiner]
CN 105631518A · 2016 [cited by applicant]
CN 111260074A · 2020 [cited by applicant]
CN 110717535B · 2020 [cited by applicant]
Feurer, nitializing Bayesian Hyperparameter Optimization via Meta-Learning (Year: 2015). [cited by examiner]
Joseph, “Grid Search for model tuning,” Towards Data Science, Dec. 2018, 10 pages. [cited by applicant]
Feurer et al., “Initializing Bayesian Hyperparameter Optimization via Meta-Learning,” Association for the Advancement of Artificial Intelligence, 2015, 8 pages. [cited by applicant]
Luo, “PredicT-ML: a tool for automating machine learning model building with big clinical data,” Health Inf Sci Syst (2016) 4:5, Jun. 2016, 16 pages. [cited by applicant]
Bergstra et al., “Random Search for Hyper-Parameter Optimization,” Journal of Machine Learning Research 13, Feb. 2012, pp. 281-305. [cited by applicant]