IP Library › Granted Patent US 12,536,464
Granted Patent B2
US 12,536,464 · App. 16/262,443 · Granted Jan 27, 2026

System for constructing effective machine-learning pipelines with optimized outcomes

Inventors: Evelyn Duesterwald (Yorktown Heights, NY); Martin Hirzel (Yorktown Heights, NY); Darrell Reimer (Yorktown Heights, NY)
Assignee: International Business Machines Corporation
G06N20/00G06F16/90335
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,464
App. No.
16/262,443
Granted
Jan 27, 2026
Kind
B2
Abstract

A system, apparatus and a method for constructing pipelines, including enumerate a subspace of valid integrated pipelines, collecting a set of metrics for each of the plurality of pipelines, reducing the set of metrics to a single metric, selecting a final integrated pipeline from among the valid integrate pipelines based on reduced metric.

Claims (40)

1 . A method comprising:

enumerating a subspace of valid integrated machine learning pipelines, wherein each valid integrated machine learning pipeline integrates at least one distinct value-add via a respective machine learning pipeline transformation, wherein each valid integrated machine learning pipeline is a transformed version of same initial machine learning pipeline;

collecting a set of objective and side-effect metrics for each of the valid integrated machine learning pipelines;

reducing the set of objective and side-effect metrics to a single metric for each of the valid integrated machine learning pipelines;

ordering the valid integrated machine learning pipelines based on the single metric;

and selecting a final integrated machine learning pipeline from the valid integrated machine learning pipelines based on the final integrated machine learning pipeline having a top ranked single metric.

2 . The method according to claim 1 , wherein during the enumeration step, dynamically pruning machine learning pipelines from the subspace using preconditions, where pruning using preconditions includes removing any machine learning pipelines from the subspace having a metric that violates a metric range threshold.

3 . The method according to claim 1 , wherein during the enumerating, pruning machine learning pipelines from the subspace using a statically produced partial ordering of value-adds.

4 . The method according to claim 1 , wherein during the enumerating, pruning machine learning pipelines from the subspace based on a time threshold.

5 . The method according to claim 1 , wherein input data used to train machine learning pipelines considered for the subspace is multi-modal data from a plurality of sources, and

wherein machine learning pipelines are added to the subspace by relaxing upper and lower bounds of value-add metrics until a threshold number of machine learning pipelines is met.

6 . The method according to claim 1 , wherein the final integrated machine learning pipeline is selected based on the top ranked single metric being on a frontier of a Pareto frontier.

7 . The method according to claim 1 being cloud implemented.

8 . The method of claim 1 , wherein the subspace of valid integrated machine learning pipelines includes at least one valid integrated machine learning pipeline that is a direct transformation of the initial machine learning pipeline and at least one valid integrated machine learning pipeline that is a transformation of the direct transformation of the initial machine learning pipeline.

9 . The method of claim 1 , wherein machine learning pipeline transformations applied to valid integrated machine learning pipelines within the subspace include de-biasing, compression, and hyperparameter optimization.

10 . The method of claim 1 , wherein a first valid integrated machine learning pipeline is a first transformed version of the initial machine learning pipeline by a de-biasing transformation, wherein a second valid integrated machine learning pipeline is a second transformed version of the initial machine learning pipeline by a compression transformation, wherein a third valid integrated machine learning pipeline is a third transformed version of the initial machine learning pipeline by a compression transformation.

11 . A system comprising:

one or more processors; and

one or more computer-readable storage media collectively storing program instructions which, when executed by the one or more processors, are configured to cause the one or more processors to perform a method comprising:

enumerating a subspace of valid integrated machine learning pipelines, wherein each valid integrated machine learning pipeline integrates at least one distinct value-add via a respective machine learning pipeline transformation, wherein each valid integrated machine learning pipeline is a transformed version of same initial machine learning pipeline;

collecting a set of objective and side-effect metrics for each of the valid integrated machine learning pipelines;

reducing the set of objective and side-effect metrics to a single metric for each of the valid integrated machine learning pipelines;

ordering the valid integrated machine learning pipelines based on the single metric;

and selecting a final integrated machine learning pipeline from the valid integrated machine learning pipelines based on the final integrated machine learning pipeline having a top ranked single metric.

12 . The system according to claim 11 , wherein during the enumeration step, dynamically pruning machine learning pipelines from the subspace using preconditions, where pruning using preconditions includes removing any machine learning pipelines from the subspace having a metric that violates a metric range threshold.

13 . The system according to claim 11 , wherein during the enumerating, pruning machine learning pipelines from the subspace using a statically produced partial ordering of value-adds.

14 . The system according to claim 11 , wherein input data used to train machine learning pipelines considered for the subspace is multi-modal data from a plurality of sources, and

wherein machine learning pipelines are added to the subspace by relaxing upper and lower bounds of value-add metrics until a threshold number of machine learning pipelines is met.

15 . The system according to claim 11 , wherein the set of objective and side-effect metrics for each valid integrated machine learning pipeline are normalized yielding a tuple of metric values for each valid integrated machine learning pipeline, wherein the set of objective and side-effect metrics are reduced to the single metric by computing a geometric mean over the tuple of metric values for each valid integrated machine learning pipeline.

16 . A computer program product comprising one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions comprising instructions configured to cause one or more processors to perform a method comprising:

enumerating a subspace of valid integrated machine learning pipelines, wherein each valid integrated machine learning pipeline integrates at least one distinct value-add via a respective machine learning pipeline transformation, wherein each valid integrated machine learning pipeline is a transformed version of same initial machine learning pipeline;

collecting a set of objective and side-effect metrics for each of the valid integrated machine learning pipelines;

reducing the set of objective and side-effect metrics to a single metric for each of the valid integrated machine learning pipelines;

ordering the valid integrated machine learning pipelines based on the single metric;

and selecting a final integrated machine learning pipeline from the valid integrated machine learning pipelines based on the final integrated machine learning pipeline having a top ranked single metric.

17 . The computer program product according to claim 16 , wherein during the enumeration step, dynamically pruning machine learning pipelines from the subspace using preconditions, where pruning using preconditions includes removing any machine learning pipelines from the subspace having a metric that violates a metric range threshold.

18 . The computer program product according to claim 16 , wherein during the enumerating, pruning machine learning pipelines from the subspace using a statically produced partial ordering of value-adds.

19 . The computer program product according to claim 16 , wherein input data used to train machine learning pipelines considered for the subspace is multi-modal data from a plurality of sources, and

wherein machine learning pipelines are added to the subspace by relaxing upper and lower bounds of value-add metrics until a threshold number of machine learning pipelines is met.

20 . The computer program product according to claim 16 , wherein the set of objective and side-effect metrics for each valid integrated machine learning pipeline are normalized yielding a tuple of metric values for each valid integrated machine learning pipeline, wherein the set of objective and side-effect metrics are reduced to the single metric by computing a weighted average over the tuple of metric values for each valid integrated machine learning pipeline.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2019
From: DUESTERWALD, EVELYN; HIRZEL, MARTIN; REIMER, DARRELL
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 048233/0876 →
Continuity (1)
Related Publication 20200242510A1 · Jul 30, 2020
References Cited (22)
US 20110313953A1 · Lane · 2011 [cited by examiner]
US 20170193392A1 · Liu · 2017 [cited by examiner]
US 20180067732A1 · Seetharaman et al. · 2018 [cited by applicant]
US 20190018866A1 · Ormont · 2019 [cited by examiner]
US 20190188605A1 · Zavesky · 2019 [cited by examiner]
CN 104838355A · 2015 [cited by applicant]
CN 105978704A · 2016 [cited by applicant]
WO WO2017189725 · 2017 [cited by applicant]
Olson, Randal, Evaluation of a Tree-based Pipeline Optimization Tool, ACM ISBN 978-1-4503-4206-3/16/07. (Year: 2016). [cited by examiner]
Koehrsen Will, Hyperparameter Tuning the Random Forest in Python (Year: 2018). [cited by examiner]
Doyle, Patrick Al Qual Summary Search Method (Year: 1997). [cited by examiner]
Ian Gemp, Automated Data Cleansing through Meta-Learning (Year: 2017). [cited by examiner]
Scikit-Learn, 4.1. Pipelines and composite estimators, Version 0.20.0 (Year: 2018). [cited by examiner]
How to Build a Machine Learning Pipeline with Scikit-learn (Year: 2022). [cited by examiner]
Sklearn.metrics.f1_scored scikit-learn 0.20.0 documentation, Version 0.20. (Year: 2018). [cited by examiner]
Mel, et al. “The NIST Definition of Cloud Computing”. Recommendations of the National Institute of Standards and Technology, Nov. 16, 2015. [cited by applicant]
Feurer, M. et al., “Efficient and Robust Automated Machine Learning” Department of Computer Science University of Freiburg, Germany. 2017. [cited by applicant]
Zinkevich, M., “Rules of Machine Learning: Best Practices for ML Engineering” 2016. [cited by applicant]
Malik, S. et al., “A Visual Analytics Approach to Comparing Cohorts of Event Sequences” Doctor of Philosophy, 2016. [cited by applicant]
Grano, G. et al., “How High will it be? Using Machine Learning Models to Predict Branch Coverage in Automated Testing” University of Zurich, Department of Informatics, Switzerland, IEEE 2018. [cited by applicant]
Kalavri, V. “Performance Optimization Techniques and Tools for Distributed Graph Processing” School of Information and Communication Technology, KTH Royal Institute of Technology Stockholm, Sweden 2016 and Institute of … [cited by applicant]
Chinese Office Action, dated Apr. 20, 2023, in Chinese Application No. 202010054732.4 . [cited by applicant]