IP Library › Granted Patent US 12,555,023
Granted Patent B2
US 12,555,023 · App. 16/009,815 · Granted Feb 17, 2026

Reinforcement learning exploration by exploiting past experiences for critical events

Inventors: Asim Munawar (Ichikawa, JP); Giovanni De Magistris (Kawasaki, JP); Ryuki Tachibana (Yokohama, JP)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N20/00G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,023
App. No.
16/009,815
Granted
Feb 17, 2026
Kind
B2
Abstract

A computer-implemented method is provided for reinforcement learning performed by a processor. The method includes obtaining, from an environment, a given experience that includes an action, a state and a reward. The method further includes storing the given experience in an experience buffer responsive to a value of the reward included in the given experience exceeding a first threshold. The method also includes responsive to obtaining another experience having another reward that less than or equal to the first threshold, searching the experience buffer for a candidate experience with a similar state to the other experience and copying the candidate experience into an event buffer. The method additionally includes during exploration, selecting an action to be taken to the environment from the event buffer with a predetermined probability.

Claims (35)

1 . A computer-implemented method for data determination for reinforcement learning training performed by a processor, the method comprising:

obtaining, from an environment, a given experience that includes a vehicle action, a state and a reward;

configuring an experience buffer and an event buffer for cooperative model training usage such that the experience buffer and the event buffer are no longer used for training after a pre-defined number of training steps and model-free reinforcement learning is performed;

during training, storing the given experience in the experience buffer responsive to a value of the reward included in the given experience not being below an average award amount for a plurality of experiences by a first threshold amount, while excluding from the experience buffer events where an agent dies corresponding to the value of the reward included in the given experience being below the average award amount by the first threshold amount;

further during training, responsive to obtaining another experience, searching the experience buffer for a candidate experience with a similar state but with a better reward and different vehicle action to the other experience and copying the candidate experience into the event buffer storing events where an agent survives;

during exploration, selecting a vehicle action to be taken to the environment from the event buffer with a predetermined probability; and

performing the vehicle action taken from the event buffer to avoid the vehicle action becoming an event where the agent dies, the vehicle action being a controlling of a motor vehicle to perform a braking or steering action for accident avoidance.

2 . The computer-implemented method of claim 1 , further comprising stopping the selecting of the vehicle action from the event buffer after a pre-defined number of steps during a training stage of the reinforcement learning.

3 . The computer-implemented method of claim 2 , further comprising performing random exploration responsive to said stopping step.

4 . The computer-implemented method of claim 1 , further comprising storing in the experience buffer any experiences previously observed except for the experiences that resulted in a corresponding reward that fails to exceed the first threshold.

5 . The computer-implemented method of claim 1 , wherein the state represents a local state in a low-dimensional space.

6 . The computer-implemented method of claim 1 , wherein the method is applied to the plurality of experiences, wherein the first threshold is used to identify critical events, and wherein any of the plurality of experiences unrelated to the critical events are stored in the experience buffer and any of the plurality of experiences related to the critical events are stored in the event buffer.

7 . The computer-implemented method of claim 1 , further comprising plotting, on a display device, a plurality of experiences in a visualization, each of the plurality of experiences having a respective reward to form a plurality of rewards across the plurality of experiences, wherein said plotting step uses the plurality of rewards as weights for the visualization.

8 . A computer program product for data determination for reinforcement learning training performed by a processor, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer having the processor to cause the computer to perform a method comprising:

obtaining, from an environment, a given experience that includes a vehicle action, a state and a reward;

configuring an experience buffer and an event buffer for cooperative model training usage such that the experience buffer and the event buffer are no longer used for training after a pre-defined number of training steps and model-free reinforcement learning is performed;

during training, storing the given experience in the experience buffer responsive to a value of the reward included in the given experience not being below an average award amount for a plurality of experiences by a first threshold amount, while excluding from the experience buffer events where an agent dies corresponding to the value of the reward included in the given experience being below the average award amount by the first threshold amount;

further during training, responsive to obtaining another experience, searching the experience buffer for a candidate experience with a similar state but with a better reward and different vehicle action to the other experience and copying the candidate experience into the event buffer storing events where an agent survives;

during exploration, selecting a vehicle action to be taken to the environment from the event buffer with a predetermined probability; and

performing the vehicle action taken from the event buffer to avoid the vehicle action becoming an event where the agent dies, the vehicle action being a controlling of a motor vehicle to perform a braking or steering action for accident avoidance.

9 . The computer program product of claim 8 , wherein the method further comprises stopping the selecting of the vehicle action from the event buffer after a pre-defined number of steps during a training stage of the reinforcement learning.

10 . The computer program product of claim 9 , wherein the method further comprises performing random exploration responsive to said stopping step.

11 . The computer program product of claim 8 , wherein the method further comprises storing in the experience buffer any experiences previously observed except for the experiences that resulted in a corresponding reward that fails to exceed the first threshold.

12 . The computer program product of claim 8 , wherein the state represents a local state in a low-dimensional space.

13 . A computer processing system for data determination for reinforcement learning training, comprising:

a memory for storing program code; and

a processor, operatively coupled to the memory, for running the program code to

obtain, from an environment, a given experience that includes vehicle action, a state and a reward;

configuring an experience buffer and an event buffer for cooperative model training usage such that the experience buffer and the event buffer are no longer used for training after a pre-defined number of training steps and model-free reinforcement learning is performed;

during training, store the given experience in the experience buffer responsive to a value of the reward included in the given experience not being below an average award amount for a plurality of experiences by a first threshold amount, while excluding from the experience buffer events where an agent dies corresponding to the value of the reward included in the given experience being below the average award amount by the first threshold amount;

further during training, responsive to obtaining another experience, search the experience buffer for a candidate experience with a similar state but with a better reward and different vehicle action to the other experience and copying the candidate experience into the event buffer storing events where an agent survives;

during exploration, select vehicle action to be taken to the environment from the event buffer with a predetermined probability; and

perform the vehicle action taken from the event buffer to avoid the vehicle action becoming an event where the agent dies, the vehicle action being a controlling of a motor vehicle to perform a braking or steering action for accident avoidance.

14 . The computer-implemented method of claim 1 , wherein during training, the plurality of experiences and their respective frequencies of occurrence with rewards are stored and used to plot the experiences using respective ones of the rewards as weights for visualization.

15 . The computer-implemented method of claim 1 , wherein during training, a next vehicle action is determined by sampling actions for similar events in the event buffer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2018
From: MUNAWAR, ASIM; DE MAGISTRIS, GIOVANNI; TACHIBANA, RYUKI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 046101/0557 →
Continuity (1)
Related Publication 20190385091A1 · Dec 19, 2019
References Cited (35)
US 8494980B2 · Hans · 2013 [cited by examiner]
US 20070220303A1 · Kimura · 2007 [cited by examiner]
US 20090327011A1 · Petroff · 2009 [cited by examiner]
US 20100070098A1 · Sterzing · 2010 [cited by examiner]
US 20100094786A1 · Gupta · 2010 [cited by examiner]
US 20100318478A1 · Yoshiike · 2010 [cited by examiner]
US 20120084237A1 · Hasuo · 2012 [cited by examiner]
US 20150100530A1 · Mnih · 2015 [cited by examiner]
US 20160232445A1 · Srinivasan · 2016 [cited by examiner]
US 20170024346A1 · Lillicrap et al. · 2017 [cited by applicant]
US 20170032245A1 · Osband · 2017 [cited by examiner]
US 20170140269A1 · Schaul · 2017 [cited by examiner]
US 20180032082A1 · Shalev-Shwartz · 2018 [cited by examiner]
US 20190061147A1 · Luciw · 2019 [cited by examiner]
US 20190232488A1 · Levine · 2019 [cited by examiner]
Foerster et al, “Stabilising Experience Replay for Deep Multi-Agent Reinforcement Learning”, Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, PMLR 70, 2017 (Year: 2017). [cited by examiner]
Horgan et al. “Distributed prioritized experience replay.” arXiv preprint arXiv:1803.00933 (2018). (Year: 2018). [cited by examiner]
Lin, “Self-Improving Reactive Agents Based On Reinforcement Learning, Planning and Teaching”, Mach Learn 8, 293-321 (1992). https://doi.org/10.1007/BF00992699 (Year: 1992). [cited by examiner]
Gabor et al, “Multi-criteria Reinforcement Learning”, Proceedings of the Fifteenth International Conference on Machine Learning (ICML 1998), Madison, Wisconsin, USA, Jul. 24-27, 1998 (Year: 1998). [cited by examiner]
Adam et al, “Experience Replay for Real-Time Reinforcement Learning Control”, IEEE Transactions on Systems, Man, and Cybernetics—Part C: Applications and Reviews, vol. 42, No. 2, Mar. 2012 (Year: 2012). [cited by examiner]
Mnih et al. “Human-level control through deep reinforcement learning”. Nature 518, 529-533 (2015). https://doi.org/10.1038/nature14236 (Year: 2015). [cited by examiner]
Van Hasselt et al., “Deep Reinforcement Learning with Double Q-learning.” arXiv preprint arXiv:1509.06461 (2015). (Year: 2015). [cited by examiner]
Zuo et al, “Continuous reinforcement learning from human demonstrations with integrated experience replay for autonomous driving ,” 2017 IEEE International Conference on Robotics and Biomimetics (ROBIO), Macau, Macao, 2… [cited by examiner]
Hester et al. “Deep Q-learning from Demonstrations.” arXiv preprint arXiv:1704.03732 (2017). (Year: 2017). [cited by examiner]
Berkenkamp et al. “Safe Model-based Reinforcement Learning with Stability Guarantees.” arXiv e-prints (2017): arXiv-1705. (Year: 2017). [cited by examiner]
“Partially Observable Markov Decision Process” Wikipedia, modified May 24, 2018 (Year: 2018). [cited by examiner]
Wang et al. “Sample efficient actor-critic with experience replay.” arXiv preprint arXiv:1611.01224 (2017). (Year: 2017). [cited by examiner]
Hausknecht et al. “Deep Recurrent Q-Learning for Partially Observable MDPs.” arXiv preprint arXiv:1507.06527 (2015). (Year: 2017). [cited by examiner]
Isele et al. “Selective Experience Replay for Lifelong Learning.” arXiv preprint arXiv:1802.10269 (Feb. 2018). (Year: 2018). [cited by examiner]
Andrychowicz et al. “Hindsight Experience Replay.” arXiv preprint arXiv:1707.01495 (Feb. 2018). (Year: 2018). [cited by examiner]
Szita et al., “Learning to Play Using Low-Complexity Rule-Based Policies: Illustrations through Ms. Pac-Man”, Journal of Artificial Intelligence Research 30 (2007), Dec. 2007, pp. 659-684. [cited by applicant]
Śnieżyński et al., “Combining Rule Induction and Reinforcement Learning”, 2010 Ninth International Conference on Machine Learning and Applications, Dec. 2010, pp. 851-856. [cited by applicant]
Geibel et al., “Risk-Sensitive Reinforcement Learning Applied to Control under Constraints”, Journal of Artificial Intelligence Research 24 (2005), pp. 81-108, Jul. 2005. [cited by applicant]
Schaul et al., “Prioritized Experience Replay”, arXIV:1511.05952v4 [cs.LG], pp. 1-21, Feb. 2016. [cited by applicant]
Pathak et al., “Curiosity-driven Exploration by Self-supervised Prediction”, arXiv:1705.05363v1 [cs.LG], 12 pages, May 2017. [cited by applicant]