IP Library › Granted Patent US 12,579,464
Granted Patent B1
US 12,579,464 · App. 17/084,356 · Granted Mar 17, 2026

Meta-models for predicting machine learning model performance using features obtained via optimization

Inventors: Lichao Wang (Redmond, WA); Dmitry Vladimir Zhiyanov (Redmond, WA); Archiman Dutta (Shoreline, WA)
Assignee: Amazon Technologies, Inc.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,464
App. No.
17/084,356
Granted
Mar 17, 2026
Kind
B1
Abstract

Result quality metrics of a set of machine learning tasks conducted on various record groups using a plurality of machine learning models are obtained. Based on applying an algorithm to the record groups, respective sets of intermediary results corresponding to records of the groups are obtained. A meta-model for predicting result quality metrics for respective record-group-and-model combinations is trained using a training data set which includes statistical features obtained from the intermediary results. The trained meta-model is stored.

Claims (59)

1 . A system, comprising:

one or more computing devices;

wherein the one or more computing devices include instructions that upon execution on or across the one or more computing devices cause the one or more computing devices to:

obtain respective result quality metrics of a plurality of machine learning tasks, wherein individual ones of the plurality of machine learning tasks comprise providing respective record groups of a plurality of record groups as input to one or more machine learning models of a plurality of machine learning models, the one or more machine learning models to generate results to an inference problem for individual records, wherein individual ones of the record groups comprise one or more records;

determine an optimization algorithm to be used to generate one or more features representing individual record groups of the plurality of record groups, wherein an objective of the optimization algorithm is expressed via a loss function;

obtain, based at least in part on applying the optimization algorithm to individual records of the plurality of record groups, respective sets of intermediate optimization results corresponding to individual ones of the record groups and distinct from the results of the one or more machine learning models, including a first set of intermediate optimization results corresponding to a first record group of the plurality of record groups, wherein the first set of intermediate optimization results comprises alternative results to the same inference problem for which the one or more machine learning models generated results for individual records of the first record group, wherein the first set of intermediate optimization results comprises one or more non-linear learned transformations of individual records of the first record group;

generate, using statistical analysis of at least the first set of intermediate optimization results distinct from the results of the one or more machine learning models and the individual records of the first record group, one or more statistical features representing the first record group;

prepare a training data set of a meta-model for predicting respective result quality metrics ranges associated with respective record-group-and-model combinations, wherein the training data set includes, with respect to a combination of the first record group and a first machine learning model of the plurality of machine learning models, at least (a) one or more data properties of the first record group, and (b) the one or more statistical features;

train the meta-model using at least the training data set; and

in response to a query indicating a target machine learning task to be performed on a new record group which was not part of the plurality of record groups, execute a trained version of the meta-model to provide (a) an indication of a particular machine learning model of the plurality of machine learning models whose predicted result quality metrics with respect to the target machine learning task are within a particular range and (b) an explanation of the predicted result quality metrics based on alternative results generated by the optimization algorithm for the new record group.

2 . The system as recited in claim 1 , wherein the training data set further includes, with respect to the combination of the first record group and the first machine learning model, an encoding representing the first machine learning model.

3 . The system as recited in claim 1 , wherein to determine the optimization algorithm, the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:

analyze input obtained via one or more programmatic interfaces of an analytics service of a provider network.

4 . The system as recited in claim 1 , wherein to determine the optimization algorithm, the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:

examine one or more records of a knowledge base of an analytics service.

5 . The system as recited in claim 1 , wherein to train the meta-model, the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:

conduct a plurality of experiments with respective result quality prediction machine learning models of a plurality of result quality prediction machine learning models, wherein individual ones of the experiments comprise using a respective set of hyper-parameters; and

select a particular result quality prediction machine learning model of the respective result quality prediction machine learning models to be used to predict result quality metrics after the training is terminated, wherein the indication of the particular machine learning model whose predicted result quality metrics with respect to the target machine learning task are within the particular range is obtained using the particular result quality prediction machine learning model.

6 . A computer-implemented method, comprising:

obtaining respective result quality metrics of a plurality of machine learning tasks performed using a first set of one or more machine learning models on a first data set comprising a plurality of record groups, wherein a first machine learning task of the plurality of machine learning tasks comprises providing a first record group of the plurality of record groups to a first machine learning model of the first set to generate results to an inference problem for individual records of the first record group;

determining, based at least in part on applying a first optimization algorithm to individual records of the plurality of record groups, respective sets of intermediary results corresponding to individual ones of the record groups and distinct from the results of the first machine learning model, wherein a first set of intermediary results comprises alternative results to the same inference problem for which the first machine learning model generated results for the individual records of the first record group, wherein the first set of intermediate results comprises one or more non-linear learned transformations of individual records of the first record group;

preparing a training data set for a meta-model for predicting respective result quality metrics ranges associated with respective record-group-and-model combinations, wherein the training data set includes, with respect to a combination of the first record group and a first machine learning model, at least (a) one or more data quality features of the first record group, and (b) one or more statistical features obtained from the first set of intermediary results;

training the meta-model using at least the training data set; and

executing a trained version of the meta-model to provide (a) an indication of a predicted result quality metric range of a particular machine learning model of the first set of one or more machine learning models with respect to a machine learning task on a particular record group which was not part of the first data set and (b) an explanation of the predicted result quality metrics based on alternative results generated by the optimization algorithm for the particular record group.

7 . The computer-implemented method as recited in claim 6 , wherein applying the first optimization algorithm comprises executing an auxiliary machine learning model which is not in the first set of one or more machine learning models.

8 . The computer-implemented method as recited in claim 6 , wherein the first optimization algorithm comprises a plurality of stages including a first stage and a second stage, the computer-implemented method further comprising:

obtaining, in the first stage, via a first auxiliary machine learning model which is not in the first set of one or more machine learning models, respective results corresponding to individual records of one or more record groups of the plurality of record groups; and

training, using another training data set comprising labels derived from the respective results, a second auxiliary machine learning model which is not in the first set of one or more machine learning models, wherein the second stage comprises executing a trained version of the second auxiliary machine learning model, and wherein the first set of intermediary results comprises at least some results obtained from the trained version of the second auxiliary machine learning model.

9 . The computer-implemented method as recited in claim 6 , further comprising:

obtaining, via one or more programmatic interfaces of an analytics service, metadata pertaining to the plurality of machine learning tasks, wherein the metadata indicates an objective of at least the first machine learning task;

identifying, by the analytics service, based at least in part on the metadata, one or more candidate optimization algorithms for preparing at least a portion of the training data set, wherein the one or more candidate optimization algorithms include the first optimization algorithm; and

obtaining, at the analytics service via the one or more programmatic interfaces, an indication that the first optimization algorithm has been selected to prepare at least a portion of the training data set.

10 . The computer-implemented method as recited in claim 6 , further comprising:

obtaining, via a programmatic interface, an indication of the first optimization algorithm.

11 . The computer-implemented method as recited in claim 6 , further comprising:

identifying, at an analytics service, one or more statistical algorithms for analyzing the first set of intermediary results, wherein the one or more statistical features are obtained from the one or more statistical algorithms.

12 . The computer-implemented method as recited in claim 6 , wherein the training data set further includes, with respect to the combination of the first record group and the first machine learning model, an encoding representing the first machine learning model.

13 . The computer-implemented method as recited in claim 6 , wherein the training data set further includes, with respect to the combination of the first record group and the first machine learning model, an encoding representing one or more resources used to train or execute the first machine learning model.

14 . The computer-implemented method as recited in claim 6 , further comprising:

providing, based at least in part on statistical analysis of results obtained from the meta-model, a recommendation for one or more changes to record groups provided as input to one or more of (a) the meta-model or (b) individual ones of the first set of one or more machine learning models to improve one or more result quality metrics associated with the record groups.

15 . The computer-implemented method as recited in claim 6 , further comprising:

providing, based at least in part on statistical analysis of results obtained from the meta-model, a recommendation for one or more change to a particular machine learning model of the first set of one or more machine learning models to improve one or more result quality metrics of the particular machine learning model.

16 . One or more non-transitory computer-accessible storage media storing program instructions that when executed on or across one or more processors cause the one or more processors to:

obtain respective result quality metrics of a plurality of machine learning tasks performed using one or more machine learning models on a first data set comprising a plurality of record groups, wherein a first machine learning task of the plurality of machine learning tasks comprises providing a first record group of the plurality of record groups to a first machine learning model, the first machine learning model to generate results to an inference problem for individual records of the first record group;

determine, based at least in part on applying an algorithm to individual records of the plurality of record groups, respective sets of intermediary results corresponding to individual ones of the record groups and distinct from the results of the first machine learning model, wherein a first set of intermediary results comprises alternative results to the same inference problem for which the first machine learning model generated results for the individual records of the first record group, wherein the first set of intermediate results comprises one or more non-linear learned transformations of the individual records of the first record group;

train a meta-model for predicting respective result quality metrics ranges associated with respective record-group-and-model combinations, wherein a training data set used for training the meta-model includes, for a combination of the first record group and a first machine learning model, one or more statistical features obtained from the first set of intermediary results;

store a trained version of the meta-model; and

execute the trained version of the meta-model to provide (a) an indication of a predicted result quality metric range of the first machine learning model with respect to a machine learning task on a particular record group which was not part of the first data set and (b) an explanation of the predicted result quality metrics based on alternative results generated by the algorithm for the particular record group.

17 . The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein applying the algorithm comprises executing a machine learning model which was not utilized in the plurality of machine learning tasks.

18 . The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the algorithm comprises a plurality of stages including a first stage and a second stage, wherein the one or more non-transitory computer-accessible storage media store further program instructions that when executed on or across the one or more processors further cause the one or more processors to:

obtain, in the first stage, via a first auxiliary machine learning model, respective results corresponding to individual records of one or more record groups of the plurality of record groups; and

train a second auxiliary machine learning model using another training data set comprising labels derived from the respective results, wherein the second stage comprises executing a trained version of the second auxiliary machine learning model, and wherein the first set of intermediary results comprises at least some results obtained from the trained version of the second auxiliary machine learning model.

19 . The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the one or more non-transitory computer-accessible storage media store further program instructions that when executed on or across the one or more processors further cause the one or more processors to:

obtain, via one or more programmatic interfaces of an analytics service, metadata pertaining to the plurality of machine learning tasks, wherein the metadata indicates an objective of at least the first machine learning task;

identify, by the analytics service, based at least in part on the metadata, one or more candidate algorithms for preparing at least a portion of the training data set, wherein the one or more candidate algorithms include the algorithm; and

obtain, at the analytics service via the one or more programmatic interfaces, an indication that the algorithm has been selected to prepare at least a portion of the training data set.

20 . The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the one or more non-transitory computer-accessible storage media store further program instructions that when executed on or across the one or more processors further cause the one or more processors to:

obtain, from the trained version of the model, (a) a first result quality metrics range predicted for a first tuning parameter setting of the first machine learning model and (b) a second result quality metrics range predicted for a second tuning parameter setting of the first machine learning model; and

provide, based at least in part on an analysis of the first and second result quality metrics ranges, a recommendation for one or more tuning parameter settings the first machine learning model to improve one or more result quality metrics of the first machine learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2020
From: WANG, LICHAO; ZHIYANOV, DMITRY VLADIMIR; DUTTA, ARCHIMAN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 054229/0104 →
References Cited (40)
US 8977622B1 · Dutta · 2015 [cited by applicant]
US 9830344B2 · Dutta · 2017 [cited by applicant]
US 10217080B1 · Dutta · 2019 [cited by applicant]
US 10339470B1 · Dutta et al. · 2019 [cited by applicant]
US 10565385B1 · Ravi et al. · 2020 [cited by applicant]
US 10726060B1 · Dutta et al. · 2020 [cited by applicant]
US 10783167B1 · Dutta et al. · 2020 [cited by applicant]
US 11200511B1 · London · 2021 [cited by examiner]
US 11640447B1 · Stone · 2023 [cited by examiner]
US 20020072828A1 · Turner · 2002 [cited by examiner]
US 20090091802A1 · Brown · 2009 [cited by examiner]
US 20150032783A1 · Sareen · 2015 [cited by examiner]
US 20160162802A1 · Chickering · 2016 [cited by examiner]
US 20170364831A1 · Ghosh · 2017 [cited by examiner]
US 20170372000A1 · Ethington · 2017 [cited by examiner]
US 20180074797A1 · Ludwig · 2018 [cited by examiner]
US 20190095756A1 · Agrawal · 2019 [cited by examiner]
US 20190147298A1 · Rabinovich · 2019 [cited by examiner]
US 20190370607A1 · Lecue · 2019 [cited by examiner]
US 20210312323A1 · Arnold · 2021 [cited by examiner]
US 20220015643A1 · Rajbhandary · 2022 [cited by examiner]
US 20220067520A1 · Dalli · 2022 [cited by examiner]
CN 111797990A · 2020 [cited by examiner]
Neural-network-based metamodeling for financial time series forecasting Lai et. al. https://scholars.cityu.edu.hk/en/publications/neuralnetworkbased-metamodeling-for-financial-time-series-forecasting(330e6fbb-28f1-44cf-… [cited by examiner]
A Neural Network Meta-Model and its Application for Manufacturing by Lechebalier et. al. https://tsapps.nist.gov/publication/get_pdf.cfm?pub_id=919287 (Year: 2015). [cited by examiner]
Meta-Learning: A Survey Vanchoren et. al. (Year: 2018). [cited by examiner]
U.S. Appl. No. 16/808,162, filed Mar. 3, 2020, Xianshun Chen, et al. [cited by applicant]
U.S. Appl. No. 16/455,601, filed Jun. 27, 2019, Lichao Wang, et al. [cited by applicant]
U.S. Appl. No. 16/817,218, filed Mar. 12, 2020, Dmitry Vladimir Zhiyanov, et al. [cited by applicant]
U.S. Appl. No. 16/824,480, filed Mar. 19, 2020, Xianshun Chen et al. [cited by applicant]
U.S. Appl. No. 16/900,620, filed Jun. 12, 2020, Xianshun Chen et al. [cited by applicant]
U.S. Appl. No. 16/945,572, filed Jul. 31, 2020, Xianshun Chen et al. [cited by applicant]
Aaron Klein, et al., “Learning Curve Prediction With Bayesian Neural Networks”, Published as a conference paper at ICLR 2017, p. 1-16. [cited by applicant]
Joaquin Vanschoren, “Meta-Learning: A Survey”, arXiv:1810.03548v1, Oct. 8, 2018, pp. 1-29. [cited by applicant]
Patricia Maforte dos Santos, et al., “Selection of Time Series Forecasting Models based on Performance Information”, In Proceedings of the Fourth International Conference on Hybrid Intelligent Systems (HIS'04), IEEE Com… [cited by applicant]
Jan N. van Rijn, et al., “The online performance estimation framework: heterogeneous ensemble learning for data streams”, Machine Learning, 107, 2018, Springer, pp. 149-176. [cited by applicant]
Amina Adadi, et al., “Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI)”, IEEE Access, vol. 6, 2018 , pp. 52138-52160. [cited by applicant]
Sylvain Arlot, et al., “A survey of cross-validation procedures for model selection”, hal-00407906, version 1, Jul. 27, 2009, Published in Statistics Surveys 4 (2010), pp. 1-40. [cited by applicant]
Rengian Luo, et al., “Neural Architecture Optimization”, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montreal, Canada, pp. 1-12. [cited by applicant]
Bowen Baker, et al., Accelerating Neural Architecture Search Using Performance Prediction, arXiv:1705.10823v2, Nov. 8, 2017, pp. 1-14. [cited by applicant]