IP Library › Granted Patent US 12,485,795
Granted Patent B2
US 12,485,795 · App. 17/979,047 · Granted Dec 2, 2025

Reinforcement learning for continued learning of optimal battery charging

Inventors: Reinhardt Klein (Mountain View, CA); Nikhil Ravi (Redwood City, CA); Christoph Kroener (Freiberg am Neckar, DE); Jared Evans (Sunnyvale, CA)
Assignee: Robert Bosch GmbH
B60L58/16B60L53/62G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,485,795
App. No.
17/979,047
Granted
Dec 2, 2025
Kind
B2
Abstract

Methods and systems of optimizing battery charging are disclosed. Battery state sensors are used to determine anode overpotential of a battery multiple times during multiple charge cycles. In a first phase, a reinforcement learning model (e.g., actor-critic model) is trained with rewards given throughout each charge cycle of the battery to optimize training. The reinforcement learning model can determine state-of-health characteristics of the battery over the charge cycles, and in a second phase, the reinforcement learning model is augmented accordingly. During this augmentation, the reinforcement learning model is trained with rewards given on a charge cycle-by-cycle basis, wherein rewards are given after looking at the charging optimization after the conclusion of each charge cycle. Commands are given to charge an on-field battery based on the augmented reinforcement learning model, associated state-of-health characteristics of the on-field battery are determined, and the reinforcement learning model is further augmented accordingly.

Claims (58)

1 . A method of optimizing a charging of a vehicle battery based on both simulation data and field data, the method comprising:

receiving, at a simulator, simulation battery charge data associated with charging of a simulation vehicle battery;

determining, at the simulator, an anode overpotential of the simulation vehicle battery during each charge cycle of the simulation vehicle battery;

training a reinforcement learning model based on the anode overpotential, wherein the reinforcement learning model is trained with rewards given throughout each charge cycle of the simulation vehicle battery;

determining, via the reinforcement learning model, state-of-health characteristics of the simulation vehicle battery over multiple charge cycles;

augmenting the reinforcement learning model based on the state-of-health characteristics of the simulation vehicle battery over multiple charge cycles, wherein the reinforcement learning model is trained with rewards given on a charge cycle-by-cycle basis;

commanding a charging of an on-field vehicle battery based on an output of the augmented reinforcement learning model;

based on the charging of the on-field vehicle battery, determining state-of-health characteristics of the on-field vehicle battery;

further augmenting the augmented reinforcement learning model based on the state-of-health characteristics of the on-field vehicle battery, wherein the reinforcement learning model is further trained with rewards given based on multiple charge cycles; and

based on the further augmenting, outputting a trained reinforcement learning model configured to optimize charging of on-field vehicle batteries.

2 . The method of claim 1 , wherein the reinforcement learning model utilizes a twin delayed deep deterministic (TD3) policy gradient.

3 . The method of claim 1 , wherein the step of augmenting the reinforcement learning model based on the state-of-health characteristics of the simulation battery over multiple charge cycles further includes:

training the reinforcement learning model with rewards given only after each charge cycle has completed.

4 . The method of claim 1 , wherein the step of training the reinforcement learning model includes applying random noise.

5 . The method of claim 1 , wherein the step of commanding is sent via a remote server to a battery management system of an on-field vehicle.

6 . The method of claim 1 , wherein the step of augmenting the reinforcement learning model utilizes slowly varying Ornstein-Uhlenbeck noise.

7 . The method of claim 1 , further comprising:

transferring a reward system of the reinforcement learning model from (i) a first reward system in which the rewards are given throughout each charge cycle of the simulation vehicle battery to (ii) a second reward system in which the rewards are given on a charge cycle-by-cycle basis;

wherein the transferring uses an affine layer attached to a critic model.

8 . A system of optimizing a charging of a battery based on both simulation data and field data, the system comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the one or more processors to:

receive, at a simulator, simulation battery charge data associated with charging of a simulation battery;

determine, at the simulator, an anode overpotential of the simulation battery during each charge cycle of the simulation battery;

train reinforcement learning model based on the anode overpotential, wherein the reinforcement learning model is trained with rewards given throughout each charge cycle of the simulation battery;

determine, via the reinforcement learning model, state-of-health characteristics of the simulation battery over multiple charge cycles;

augment the reinforcement learning model based on the state-of-health characteristics of the simulation battery over multiple charge cycles, wherein the reinforcement learning model is trained with rewards given on a charge cycle-by-cycle basis;

command a charging of an on-field battery based on an output of the augmented reinforcement learning model;

based on the charging of the on-field battery, determine state-of-health characteristics of the on-field battery;

further augment the augmented reinforcement learning model based on the state-of-health characteristics of the on-field battery, wherein the reinforcement learning model is further trained with rewards given based on multiple charge cycles; and

based on the further augmenting, output a trained reinforcement learning model configured to optimize charging of on-field batteries.

9 . The system of claim 8 , wherein the reinforcement learning utilizes a twin delayed deep deterministic (TD3) policy gradient.

10 . The system of claim 8 , wherein the augmenting of the reinforcement learning model based on the state-of-health characteristics of the simulation battery over multiple charge cycles further includes:

training the reinforcement learning model with rewards given only after each charge cycle has completed.

11 . The system of claim 8 , wherein the reinforcement learning model is trained by applying random noise.

12 . The system of claim 8 , wherein at least one of the at least one processors resides on a remote server, and wherein the command of the charging of an on-field battery is sent via the remote server to a battery management system associated with the on-field battery.

13 . The system of claim 8 , wherein slowly varying Ornstein-Uhlenbeck noise is utilized in perturbing a policy of the reinforcement learning model.

14 . The system of claim 8 , wherein the memory stores further instructions that, when executed by the one or more processors, cause the one or more processors to:

transfer a reward system of the reinforcement learning model from (i) a first reward system in which the rewards are given throughout each charge cycle of the simulation battery to (ii) a second reward system in which the rewards are given on a charge cycle-by-cycle basis;

wherein the transfer uses an affine layer attached to a critic model.

15 . The system of claim 8 , wherein the simulation battery is a vehicle battery, and wherein the on-field battery is an on-field vehicle battery.

16 . A non-transitory computer readable storage medium containing a plurality of program instructions, which when executed by a processor, cause the processor to perform the steps of:

receiving, at a simulator, simulation battery charge data associated with charging of a simulation vehicle battery;

determining, at the simulator, an anode overpotential of the simulation vehicle battery during each charge cycle of the simulation vehicle battery;

training reinforcement learning model based on the anode overpotential, wherein the reinforcement learning model is trained with rewards given throughout each charge cycle of the simulation vehicle battery;

determining, via the reinforcement learning model, state-of-health characteristics of the simulation vehicle battery over multiple charge cycles;

augmenting the reinforcement learning model based on the state-of-health characteristics of the simulation vehicle battery over multiple charge cycles, wherein the reinforcement learning model is trained with rewards given on a charge cycle-by-cycle basis;

commanding a charging of an on-field vehicle battery based on an output of the augmented reinforcement learning model;

based on the charging of the on-field vehicle battery, determining state-of-health characteristics of the on-field vehicle battery;

further augmenting the augmented reinforcement learning model based on the state-of-health characteristics of the on-field vehicle battery, wherein the reinforcement learning model is further trained with rewards given based on multiple charge cycles; and

based on the further augmenting, outputting a trained reinforcement learning model configured to optimize charging of on-field vehicle batteries.

17 . The non-transitory computer readable storage medium of claim 16 , wherein the reinforcement learning utilizes a twin delayed deep deterministic (TD3) policy gradient.

18 . The non-transitory computer readable storage medium of claim 16 , wherein the step of augmenting the reinforcement learning model based on the state-of-health characteristics of the simulation battery over multiple charge cycles further includes:

training the reinforcement learning model with rewards given only after each charge cycle has completed.

19 . The non-transitory computer readable storage medium of claim 16 , wherein the step of commanding is sent via a remote server to a battery management system of an on-field vehicle.

20 . The non-transitory computer readable storage medium of claim 16 , wherein the program instructions, when executed by the processor, cause the processor to further perform the step of:

transferring a reward system of the reinforcement learning model from (i) a first reward system in which the rewards are given throughout each charge cycle of the simulation vehicle battery to (ii) a second reward system in which the rewards are given on a charge cycle-by-cycle basis;

wherein the transferring uses an affine layer attached to a critic model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2022
From: KLEIN, REINHARDT; RAVI, NIKHIL; KROENER, CHRISTOPH; EVANS, JARED
To: ROBERT BOSCH GMBH
Reel/Frame 061695/0610 →
Continuity (1)
Related Publication 20240144078A1 · May 2, 2024
References Cited (29)
US 11803750B2 · Lillicrap et al. · 2023 [cited by applicant]
US 12045272B2 · Mahapatra et al. · 2024 [cited by applicant]
US 20120105009A1 · Yao · 2012 [cited by examiner]
US 20140253039A1 · Barsukov · 2014 [cited by examiner]
US 20190229378A1 · Zhang · 2019 [cited by examiner]
US 20190236455A1 · Taylor et al. · 2019 [cited by applicant]
US 20200086483A1 · Li et al. · 2020 [cited by applicant]
US 20200410351A1 · Lillicrap et al. · 2020 [cited by applicant]
US 20210009226A1 · Yamamoto et al. · 2021 [cited by applicant]
US 20210326595A1 · Goldberg et al. · 2021 [cited by applicant]
US 20220147897A1 · Liebman et al. · 2022 [cited by applicant]
US 20220209562A1 · Kessner · 2022 [cited by examiner]
US 20220406046A1 · Mummadi et al. · 2022 [cited by applicant]
US 20230130896A1 · Lee · 2023 [cited by examiner]
US 20230196382A1 · Dev et al. · 2023 [cited by applicant]
US 20230206111A1 · Alam et al. · 2023 [cited by applicant]
US 20230229957A1 · Li et al. · 2023 [cited by applicant]
US 20230268770A1 · Howlett, III · 2023 [cited by examiner]
US 20240053403A1 · Wang · 2024 [cited by examiner]
US 20240059170A1 · Khamis et al. · 2024 [cited by applicant]
US 20240079900A1 · Kessner · 2024 [cited by examiner]
US 20240127788A1 · Hsieh et al. · 2024 [cited by applicant]
US 20240303973A1 · Ramos Dos Santos et al. · 2024 [cited by applicant]
US 20240429730A1 · Abbott et al. · 2024 [cited by applicant]
CN 115015786A · 2022 [cited by applicant]
Nicolas Heess et al., “Memory-based control with recurrent neural networks.” arXiv:1512.04455v1 [cs.LG] Dec. 14, 2015, 11 Pages. [cited by applicant]
Peter M. Attia et al. “Closed-loop optimization of fast-charging protocols for batteries with machine learning.” Nature Feb. 20, 2020, vol. 578, pp. 397-418. [cited by applicant]
Saehong Park et al., “Reinforcement Learning-based Fast Charging Control Strategy for Li-ion Batteries.” arXiv:2002.02060v2 [eess.SY] Jun. 25, 2020, 8 Pages. [cited by applicant]
Yu Sui et al., “A Multi-Agent Reinforcement Learning Framework for Lithium-ion Battery Scheduling Problems.” Energies 2020, 13(8), 1982; https://doi.org/10.3390/en13081982, 13 Pages. [cited by applicant]
Cited By (1)
US 12,611,962