IP Library › Granted Patent US 12,579,474
Granted Patent B2
US 12,579,474 · App. 18/152,931 · Granted Mar 17, 2026

Method and device for continual machine learning of a sequence of different tasks

Inventor: Thomas Elsken (Sindelfingen, DE)
Assignee: ROBERT BOSCH GMBH
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,474
App. No.
18/152,931
Granted
Mar 17, 2026
Kind
B2
Abstract

A method for parameterizing a function, which outputs an ideal parameterization of a machine learning system for a large number of different data sets. A first training of a machine learning system is carried out in succession on multiple training data sets, the individual optimized parameterizations of the machine learning system being stored for each of the training data sets. A second training of the machine learning system simultaneously on all data sets then follows, the optimal parameterization of the machine learning system being stored. An optimization of the parameterization of the function thereupon follows in such a way that, given an optimal parameterization of the first training, the function outputs the associated optimal parameterization of the second training.

Claims (44)

1 . A method for parameterizing a function, which outputs a parameterization of a machine learning system for a large number of different data sets, comprising the following steps:

providing a plurality of training data sets, an index (k) being associated with each of the training data sets;

repeating the following sequence of steps i. through iv. multiple times:

i. randomly drawing an index (k∈1<=k<=n),

ii. first training the machine learning system in succession on the training data sets having an associated index less than or equal to the index drawn, individual optimized parameterizations of the machine learning system being stored for each of the training data sets;

iii. second training the machine learning system on all data sets having an associated index less than or equal to the index drawn, optimal parameterization of the machine learning system being stored,

iv. associating the optimal parameterizations of the first training with the parameterization of the second training;

optimizing a parameterization of the function in such a way that, given the optimal parameterization of the first training, the function outputs the associated optimal parameterization of the second training;

wherein the machine learning system has quantified parameters for the second training or parameters of the machine learning system are quantified after the second training, the function being parameterized in such a way that the function maps the parameterizations of the first training on the associated quantified parameterization of the second training.

2 . The method as recited in claim 1 , wherein the optimization of the parameterization of the function takes place in such a way that a cost function is minimized, the cost function characterizing a difference between the parameterizations by the function and the parameterizations in the second training and/or a difference between a prediction accuracy of the machine learning system based on the parameterizations by the function and the prediction accuracy of the machine learning system base on the parameterization of the second training.

3 . The method as recited in claim 1 , wherein a machine learning system for the second training has a smaller architecture than in the first training and/or a machine learning system is compressed after the second training with respect to its architecture, the function being parameterized in such a way that the function maps the parameterizations of the first training on the parameterization of the second training.

4 . The method as recited in claim 1 , wherein the function is a linear function or neural network.

5 . A method for further training of a machine learning system, so that the machine learning system retains its previously learned properties, comprising the following steps:

parameterizing a function, which outputs a parameterization of the machine learning system for a large number of different data sets, including:

providing a plurality of training data sets, an index (k) being associated with each of the training data sets,

repeating the following sequence of steps i. through iv. multiple times:

i. randomly drawing an index (k ∈ 1<=k<=n),

ii. first training the machine learning system in succession on the training data sets having an associated index less than or equal to the index drawn, individual optimized parameterizations of the machine learning system being stored for each of the training data sets;

iii. second training the machine learning system on all data sets having an associated index less than or equal to the index drawn, optimal parameterization of the machine learning system being stored,

iv. associating the optimal parameterizations of the first training with the parameterization of the second training,

optimizing a parameterization of the function in such a way that, given the optimal parameterization of the first training, the function outputs the associated optimal parameterization of the second training;

providing a new data set;

training the machine learning system based on the new data set, to obtain a new, optimal parameterization for the new data set;

adapting the parameterization of the machine learning system using the function as a function of the new, optimal parameterization;

wherein the machine learning system has quantified parameters for the second training or parameters of the machine learning system are quantified after the second training, the function being parameterized in such a way that the function maps the parameterizations of the first training on the associated quantified parameterization of the second training.

6 . The method as recited in claim 5 , wherein the machine learning system having adapted parameters ascertains a second variable as a function of a first variable, wherein the first variable characterizes an operating state of a technical system or a state of surroundings of the technical system, and wherein the second variable characterizes an operating state of the technical system, or an activation variable for activating the technical system, or a setpoint variable for regulating the technical system.

7 . A device configured to parameterize a function, which outputs a parameterization of a machine learning system for a large number of different data sets, the device comprises one or more processors, the one or more processors is configured to:

provide a plurality of training data sets, an index (k) being associated with each of the training data sets;

repeat the following sequence of steps i. through iv. multiple times:

i. randomly drawing an index (k ∈ 1<=k<=n),

ii. first training the machine learning system in succession on the training data sets having an associated index less than or equal to the index drawn, individual optimized parameterizations of the machine learning system being stored for each of the training data sets;

iii. second training the machine learning system on all data sets having an associated index less than or equal to the index drawn, optimal parameterization of the machine learning system being stored,

iv. associating the optimal parameterizations of the first training with the parameterization of the second training;

optimize a parameterization of the function in such a way that, given the optimal parameterization of the first training, the function outputs the associated optimal parameterization of the second training;

wherein the machine learning system has quantified parameters for the second training or parameters of the machine learning system are quantified after the second training, the function being parameterized in such a way that the function maps the parameterizations of the first training on the associated quantified parameterization of the second training.

8 . A non-transitory machine-readable memory medium on which is stored a computer program parameterizing a function, which outputs a parameterization of a machine learning system for a large number of different data sets, the computer program, when executed by a computer, causing the computer to perform the following steps:

providing a plurality of training data sets, an index (k) being associated with each of the training data sets;

repeating the following sequence of steps i. through iv. multiple times:

i. Randomly drawing an index (k € 1<=k<=n),

ii. first training the machine learning system in succession on the training data sets having an associated index less than or equal to the index drawn, individual optimized parameterizations of the machine learning system being stored for each of the training data sets;

iii. second training the machine learning system on all data sets having an associated index less than or equal to the index drawn, optimal parameterization of the machine learning system being stored,

iv. associating the optimal parameterizations of the first training with the parameterization of the second training;

optimizing a parameterization of the function in such a way that, given the optimal parameterization of the first training, the function outputs the associated optimal parameterization of the second training;

wherein the machine learning system has quantified parameters for the second training or parameters of the machine learning system are quantified after the second training, the function being parameterized in such a way that the function maps the parameterizations of the first training on the associated quantified parameterization of the second training.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2023
From: ELSKEN, THOMAS
To: ROBERT BOSCH GMBH
Reel/Frame 062550/0550 →
Priority Claims (1)
DE 10 2022 200 546.5 · Jan 18, 2022 · national
Continuity (1)
Related Publication 20230229969A1 · Jul 20, 2023
References Cited (12)
US 9826149B2 · Chalom · 2017 [cited by examiner]
US 11995553B2 · Groh · 2024 [cited by examiner]
US 12194631B2 · Berkenkamp · 2025 [cited by examiner]
US 12288382B2 · Guo · 2025 [cited by examiner]
US 20210142286A1 · Hajian · 2021 [cited by examiner]
US 20220012636A1 · Lindauer · 2022 [cited by examiner]
US 20220076114A1 · Shaker · 2022 [cited by examiner]
US 20220292349A1 · Stoll · 2022 [cited by examiner]
US 20220398262A1 · Derakhshani · 2022 [cited by examiner]
Mirzadeh et al., “Linear Mode Connectivity in Multitask and Continual Learning,” Cornell University, 2020, pp. 1-21. <https://arxiv.org/pdf/2010.04495.pdf> Downloaded Dec. 30, 2022. [cited by applicant]
Ha et al., “Hypernetworks,” Cornell University, 2016, pp. 1-29. <https://arxiv.org/pdf/1609.09106.pdf> Downloaded Dec. 30, 2022. [cited by applicant]
Schmidhuber, Jürgen: “Learning to control fast-weight memories: An alternative to dynamic recurrent networks,” Neural Computation, 1992, 4(1), (1992), pp. 131-139. [cited by applicant]