IP Library Granted Patent US 12,572,791
Granted Patent B2
US 12,572,791 · App. 16/950,570 · Granted Mar 10, 2026

Method, device and computer program for predicting a suitable configuration of a machine learning system for a training data set

Inventors: Arber Zela (Bad Krotzingen, DE); Frank Hutter (Freiburg, DE); Julien Siems (Freiburg, DE); Lucas Zimmer (Lörrach, DE)
Assignee: ROBERT BOSCH GMBH
G06N3/08G06F18/217G06F18/29G06N20/20G06V10/454G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,791
App. No.
16/950,570
Granted
Mar 10, 2026
Kind
B2
Abstract

A method for predicting a suitable configuration of a machine learning system for a first training data set. The method starts by training a plurality of machine learning systems on the first training data set, the machine learning systems and/or the training methods used being configured differently. This is followed by a creation of a second training data set including ascertained performances of the trained machine learning systems and the assigned configuration of the particular machine learning systems and/or training methods. This is followed by a training of a graph isomorphism network, depending on the second training data set, and a prediction in each case of the performance of a plurality of configurations not used for the training, with the aid of the GIN. A computer program and a device for carrying out the method and a machine-readable memory element, on which the computer program is stored, are also described.

Claims (43)

1 . A method for predicting a suitable configuration of a machine learning system and/or a training method, for a first training data set, comprising the following steps, which are carried out on a computer:

training a plurality of machine learning systems using the first training data set, the machine learning systems and/or training methods used for the training being configured differently;

creating a second training data set which includes ascertained respective performance capabilities of the trained machine learning systems on the first training data set and respective assigned configurations;

training a graph isomorphism network, depending on the second training data set, so that the graph isomorphism network ascertains the respective performance capabilities, depending on the respective assigned configurations;

predicting performance capabilities for a provided plurality of configurations using the graph isomorphism network; and

selecting a configuration from the provided plurality for configurations for which a best performance capability was predicted, wherein:

a further machine learning system is initialized, depending on the selected configuration, the further machine learning system being trained, and the trained, further machine learning system being used to ascertain a control variable for an actuator,

each of the respective assigned configurations includes at least one parameter which characterizes a structure of the respective machine learning system, the structure being defined using DARTS cells,

the parameters which characterize the structure of the machine learning systems and different DARTS cells are grouped into disjoint graphs for the second training data set, where further parameters of the configurations, which characterize a predefinable number of total stacked cells and/or a predefinable number of training epochs, are concatenated for each DARTS cell of the machine learning system, and

a predefinable set of values for different versions of the further parameters of the configurations are provided in each case, the machine learning systems first being trained using configurations which include the further parameters, starting with lowest values from the predefinable set of values in each case, further configurations then being selected from the predefinable set of values, depending on a predefinable computational budget, and the machine learning systems being trained using the selected further configurations, depending on the computational budget.

2 . The method as recited in claim 1 wherein the different DART cells include a normal cell and a reduction cell.

3 . The method as recited in claim 1 , wherein the further configurations are ascertained using a successive halving method, depending on the predefinable computational budget, until highest values of the predefinable set of values of the further parameters have been reached.

4 . The method as recited in claim 1 wherein, during the training, multiple different further configurations are randomly used in addition for the selected, lowest values of the predefinable set of values.

5 . A method for predicting a suitable configuration of a machine learning system and/or a training method, for a first training data set, comprising the following steps, which are carried out on a computer:

training a plurality of machine learning systems using the first training data set, the machine learning systems and/or training methods used for the training being configured differently;

creating a second training data set which includes ascertained respective performance capabilities of the trained machine learning systems on the first training data set and respective assigned configurations;

training a differentiable graph pooling network (DiffPool) or XGBoost or LGBoost, depending on the second training data set, so that the DiffPool or XGBoost or LGBoost ascertains the respective performance capabilities, depending on the respective assigned configurations;

predicting performance capabilities for a provided plurality of configurations using the DiffPool or XGBoost or LGBoost; and

selecting a configuration from the plurality for configurations for which a best performance capability was predicted, wherein;

a further machine learning system is initialized, depending on the selected configuration, the further machine learning system being trained, and the trained, further machine learning system being used to ascertain a control variable for an actuator,

each of the respective assigned configurations includes at least one parameter which characterizes a structure of the respective machine learning system, the structure being defined using DARTS cells,

the parameters which characterize the structure of the machine learning systems and different DARTS cells are grouped into disjoint graphs for the second training data set, where further parameters of the configurations, which characterize a predefinable number of total stacked cells and/or a predefinable number of training epochs, are concatenated for each DARTS cell of the machine learning system, and

a predefinable set of values for different versions of the further parameters of the configurations are provided in each case, the machine learning systems first being trained using configurations which include the further parameters, starting with lowest values from the predefinable set of values in each case, further configurations then being selected from the predefinable set of values, depending on a predefinable computational budget, and the machine learning systems being trained using the selected further configurations, depending on the computational budget.

6 . A non-transitory machine-readable memory element on which is stored a computer program for predicting a suitable configuration of a machine learning system and/or a training method, for a first training data set, the computer program, when executed by a computer, causing the computer to perform the following steps:

training a plurality of machine learning systems using the first training data set, the machine learning systems and/or training methods used for the training being configured differently;

creating a second training data set which includes ascertained respective performance capabilities of the trained machine learning systems on the first training data set and respective assigned configurations;

training a graph isomorphism network, depending on the second training data set, so that the graph isomorphism network ascertains the respective performance capabilities, depending on the respective assigned configurations;

predicting performance capabilities for a provided plurality of configurations using the graph isomorphism network; and

selecting a configuration from the provided plurality for configurations for which a best performance capability was predicted, wherein:

a further machine learning system is initialized, depending on the selected configuration, the further machine learning system being trained, and the trained, further machine learning system being used to ascertain a control variable for an actuator,

each of the respective assigned configurations includes at least one parameter which characterizes a structure of the respective machine learning system, the structure being defined using DARTS cells,

the parameters which characterize the structure of the machine learning systems and different DARTS cells are grouped into disjoint graphs for the second training data set, where further parameters of the configurations, which characterize a predefinable number of total stacked cells and/or a predefinable number of training epochs, are concatenated for each DARTS cell of the machine learning system, and

a predefinable set of values for different versions of the further parameters of the configurations are provided in each case, the machine learning systems first being trained using configurations which include the further parameters, starting with lowest values from the predefinable set of values in each case, further configurations then being selected from the predefinable set of values, depending on a predefinable computational budget, and the machine learning systems being trained using the selected further configurations, depending on the computational budget.

7 . A device configured to predict a suitable configuration of a machine learning system and/or a training method, for a first training data set, the device configured to:

train a plurality of machine learning systems using the first training data set, the machine learning systems and/or training methods used for the training being configured differently;

create a second training data set which includes ascertained respective performance capabilities of the trained machine learning systems on the first training data set and respective assigned configurations;

train a graph isomorphism network, depending on the second training data set, so that the graph isomorphism network ascertains the respective performance capabilities, depending on the respective assigned configurations;

predict performance capabilities for a provided plurality of configurations using the graph isomorphism network; and

select a configuration from the provided plurality for configurations for which a best performance capability was predicted, wherein:

a further machine learning system is initialized, depending on the selected configuration, the further machine learning system being trained, and the trained, further machine learning system being used to ascertain a control variable for an actuator,

each of the respective assigned configurations includes at least one parameter which characterizes a structure of the respective machine learning system, the structure being defined using DARTS cells,

the parameters which characterize the structure of the machine learning systems and different DARTS cells are grouped into disjoint graphs for the second training data set, where further parameters of the configurations, which characterize a predefinable number of total stacked cells and/or a predefinable number of training epochs, are concatenated for each DARTS cell of the machine learning system, and

a predefinable set of values for different versions of the further parameters of the configurations are provided in each case, the machine learning systems first being trained using configurations which include the further parameters, starting with lowest values from the predefinable set of values in each case, further configurations then being selected from the predefinable set of values, depending on a predefinable computational budget, and the machine learning systems being trained using the selected further configurations, depending on the computational budget.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2021
From: ZELA, ARBER; HUTTER, FRANK; SIEMS, JULIEN; ZIMMER, LUCAS
To: ROBERT BOSCH GMBH
Reel/Frame 055987/0542 →
Priority Claims (1)
DE 202020101012.3 · Feb 25, 2020 · national
Continuity (1)
Related Publication 20210264256A1 · Aug 26, 2021
References Cited (13)
US 20190370684A1 · Gunes · 2019 [cited by examiner]
DE 102019207911A1 · 2020 [cited by applicant]
Karpenko, Mark & Anderson, John & Sepehri, Nariman. (2006). Coordination of hydraulic manipulators by reinforcement learning. Proceedings of the American Control Conference. 2006. 6 pp . . . 10.1109/ACC.2006.1657214. (Y… [cited by examiner]
Y. Li, H Li, F. Pickard IV, B. Narayanan, F. Sen, M. Chan, S. Sankaranarayanan, B. Brooks, and B. Roux. (2017) Machine Learning Force Field Parameters from Ab Initio Data. Journal of Chemical Theory and Computation 2017… [cited by examiner]
Yiming Yang. 2018. Large-scale Machine Learning over Graphs. In Proceedings of the 2018 ACM SIGIR International Conference on Theory of Information Retrieval (ICTIR '18). Association for Computing Machinery, New York, N… [cited by examiner]
Noutahi, Emmanuel & Beaini, Dominique & Horwood, Julien & Tossou, Prudencio. (2020). Towards interpretable molecular graph representation learning. (Year: 2020). [cited by examiner]
L. Lopez, M. Guynn and M. Lu, “Predicting Computer Performance Based on Hardware Configuration Using Multiple Neural Networks,” 2018 17th IEEE International Conference on Machine Learning and Applications (ICMLA), Orlan… [cited by examiner]
Xu et al., “How Powerful Are Graph Neural Networks?,” in International Conference on Learning Representations, 2019, pp. 1-17. <https://openreview.net/forum?id=rygs6ia5km> Downloaded on Nov. 17, 2020. [cited by applicant]
Liu et al., “DARTS: Differentiable Architecture Search,” Cornell University Online Library, 2019, pp. 1-13. <https://arxiv.org/abs/1806.09055v2> Downloaded on Nov. 17, 2020. [cited by applicant]
Jamieson et al., “A Non-Stochastic Best Arm Identification and Hyperparameter Optimization,” in Proceedings of the Seventeenth International Conference on Artificial Intelligence and Statistics (AISTATS), Cornell Univer… [cited by applicant]
Ying et al., “Hierarchical Graph Representation Learning With Differentiable Pooling,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems, 2018, pp. 1-11. Downloaded on Nov. 17,… [cited by applicant]
Chen et al., “Xgboost: a Scalable Tree Boosting System,” Cornell University Online Library, 2016, pp. 1-13. <https://arxiv.org/abs/1603.02754v3>. Downloaded on Nov. 17, 2020. [cited by applicant]
Ke et al., “ Lightgbm: A Highly Efficient Gradient Boosting Decision Tree,” in Proceedings of the 31st International Conference on Neural Information Processing Systems, 2017, pp. 1-9. <https://papers.nips.cc/paper/6907… [cited by applicant]