IP Library › Granted Patent US 12,511,547
Granted Patent B2
US 12,511,547 · App. 17/979,052 · Granted Dec 30, 2025

Smoothed reward system transfer for actor- critic reinforcement learning models

Inventors: Christoph Kroener (Freiberg am Neckar, DE); Jared Evans (Sunnyvale, CA)
Assignee: Robert Bosch GmbH
G06N3/092
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,547
App. No.
17/979,052
Granted
Dec 30, 2025
Kind
B2
Abstract

Methods and systems for smoothening the transition of reward systems or datasets for actor-critic reinforcement learning models. A reinforcement model such as an actor-critic model is trained on a first dataset and a first reward system. The weights of the actor model and the critic model are frozen. While these weights are frozen, an affine transformation layer is attached to a final layer of the critic model, and the affine transformation layer is trained with a second dataset and a second reward system in order to adjust a weight of the final layer of the critic model. Then, the weights of the critic model are unfrozen which allows the adjusted weight of the final layer of the critic model to be implemented. The reinforcement learning model is retrained on the second dataset and second reward system, first with just the critic weights unfrozen, and then with both actor and critic weights unfrozen.

Claims (57)

1 . A method of training a reinforcement learning model comprising:

training a reinforcement learning model on a first dataset, wherein the reinforcement learning model utilizes a first reward system, an actor model, and a critic model;

freezing weights of both the actor model and the critic model;

while the actor model and critic model are frozen:

(i) attaching an affine transformation (AT) layer to a final layer of the critic model, and

(ii) training the AT layer with a second dataset and a second reward system in order to modify a weight of the final layer of the critic model;

unfreezing the weights of the critic model to allow implementation of the modified weight of the final layer;

retraining the reinforcement learning model on the second dataset and the second reward system while the weights of the critic model are unfrozen and the weights of the actor model are frozen;

unfreezing the weights of the actor model; and

retraining the reinforcement learning model on the second dataset and the second reward system while the weights of both the critic model and the actor model are unfrozen.

2 . The method of claim 1 , wherein the first dataset is the same as the second dataset.

3 . The method of claim 1 , wherein the first dataset is different than the second dataset.

4 . The method of claim 1 , wherein the first reward system is the same as the second reward system.

5 . The method of claim 1 , wherein the first reward system is different than the second reward system.

6 . The method of claim 1 , wherein the training of the AT layer includes scaling or shifting Q-values output by the critic model.

7 . The method of claim 1 , further comprising:

utilizing one or more battery state sensors to determine the first and second datasets, wherein the one or more battery state sensors detect at least one of a voltage, current, and temperature of a battery.

8 . A system of training a reinforcement learning model, the system comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the one or more processors to:

train a reinforcement learning model on a first dataset, wherein the reinforcement learning model utilizes a first reward system, an actor model, and a critic model;

freeze weights of both the actor model and the critic model;

while the actor model and critic model are frozen:

(i) attach an affine transformation (AT) layer to a final layer of the critic model, and

(ii) train the AT layer with a second dataset and a second reward system in order to modify weights of the final layer of the critic model based;

unfreeze the weights of the critic model;

retrain the reinforcement learning model on the second dataset and the second reward system while the weights of the critic model are unfrozen and the weights of the actor model are frozen;

unfreeze the weights of the actor model; and

retrain the reinforcement learning model on the second dataset and the second reward system while the weights of both the critic model and the actor model are unfrozen.

9 . The system of claim 8 , wherein the first dataset is the same as the second dataset.

10 . The system of claim 8 , wherein the first dataset is different than the second dataset.

11 . The system of claim 8 , wherein the first reward system is the same as the second reward system.

12 . The system of claim 8 , wherein the first reward system is different than the second reward system.

13 . The system of claim 8 , wherein the training of the AT layer includes scaling or shifting Q-values output by the critic model.

14 . The system of claim 8 , further comprising:

one or more battery state sensors configured to output the first and second datasets, wherein the one or more battery state sensors detect at least one of a voltage, current, and temperature of a battery.

15 . A method of training a reinforcement learning model comprising:

providing a reinforcement learning model with an actor model and a critic model;

training the reinforcement learning model with a first reward system;

freezing weights of both the actor model and the critic model;

while the actor model and critic model are frozen:

(i) attaching an affine transformation (AT) layer to the critic model, and

(ii) training the AT layer with a second reward system;

unfreezing the weights of the critic model;

while the weights of the critic model are unfrozen and the weights of the actor model are frozen, retraining the reinforcement learning model with the second reward system;

unfreezing the weights of the actor model; and

while the weights of both the critic model and the actor model are unfrozen, retraining the reinforcement learning model with the second reward system.

16 . The method of claim 15 , wherein:

the step of training the reinforcement learning model is performed on a first dataset,

the step of training the AT layer is on a second dataset,

the steps of retraining the reinforcement learning model are on the second dataset.

17 . The method of claim 16 , wherein the first dataset is different than the second dataset.

18 . The method of claim 15 , wherein the first reward system is different than the second reward system.

19 . The method of claim 15 , further comprising:

utilizing one or more battery state sensors to determine the first and second datasets, wherein the one or more battery state sensors detect at least one of a voltage, current, and temperature of a battery.

20 . The method of claim 15 , further comprising:

outputting a trained reinforcement learning model configured to optimizing charging of a vehicle battery, wherein the first and second datasets are associated with at least one of a battery voltage, battery current, or battery temperature.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2022
From: KROENER, CHRISTOPH; EVANS, JARED
To: ROBERT BOSCH GMBH
Reel/Frame 061650/0243 →
Continuity (1)
Related Publication 20240144023A1 · May 2, 2024
References Cited (29)
US 11803750B2 · Lillicrap · 2023 [cited by examiner]
US 12045272B2 · Mahapatra et al. · 2024 [cited by applicant]
US 20120105009A1 · Yao · 2012 [cited by applicant]
US 20140253039A1 · Barsukov · 2014 [cited by applicant]
US 20190229378A1 · Zhang et al. · 2019 [cited by applicant]
US 20190236455A1 · Taylor · 2019 [cited by examiner]
US 20200086483A1 · Li · 2020 [cited by examiner]
US 20200410351A1 · Lillicrap · 2020 [cited by examiner]
US 20210009226A1 · Yamamoto et al. · 2021 [cited by applicant]
US 20210326595A1 · Goldberg · 2021 [cited by examiner]
US 20220147897A1 · Liebman et al. · 2022 [cited by applicant]
US 20220209562A1 · Kessner et al. · 2022 [cited by applicant]
US 20220406046A1 · Mummadi · 2022 [cited by examiner]
US 20230130896A1 · Lee et al. · 2023 [cited by applicant]
US 20230196382A1 · Dev · 2023 [cited by examiner]
US 20230206111A1 · Alam et al. · 2023 [cited by applicant]
US 20230229957A1 · Li · 2023 [cited by examiner]
US 20230268770A1 · Howlett, III et al. · 2023 [cited by applicant]
US 20240053403A1 · Wang et al. · 2024 [cited by applicant]
US 20240059170A1 · Khamis et al. · 2024 [cited by applicant]
US 20240079900A1 · Kessner · 2024 [cited by applicant]
US 20240127788A1 · Hsieh · 2024 [cited by examiner]
US 20240303973A1 · Ramos Dos Santos · 2024 [cited by examiner]
US 20240429730A1 · Abbott et al. · 2024 [cited by applicant]
CN 115015786A · 2022 [cited by examiner]
Nicolas Heess et al., “Memory-based control with recurrent neural networks.” arXiv:1512.04455v1 [cs.LG] Dec. 14, 2015, 11 Pages. [cited by applicant]
Peter M. Attia et al. “Closed-loop optimization of fast-charging protocols for batteries with machine learning.” Nature Feb. 20, 2020, vol. 578, pp. 397-418. [cited by applicant]
Saehong Park et al., “Reinforcement Learning-based Fast Charging Control Strategy for Li-ion Batteries.” arXiv:2002.02060v2 [eess.SY] Jun. 25, 2020, 8 Pages. [cited by applicant]
Yu Sui et al., “A Multi-Agent Reinforcement Learning Framework for Lithium-ion Battery Scheduling Problems.” Energies 2020, 13(8), 1982; https://doi.org/10.3390/en13081982, 13 Pages. [cited by applicant]