IP Library › Granted Patent US 12,688,427
Granted Patent B2
US 12,688,427 · App. 17/763,901 · Granted Jul 21, 2026

Training action selection neural networks using hindsight modelling

Inventors: Arthur Clement Guez (London, GB); Fabio Viola (Montréal, CA); Theophane Guillaume Weber (London, GB); Lars Buesing (Letchworth Garden City, GB); Nicolas Manfred Otto Heess (London, GB)
Assignee: GDM Holding LLC
G06N3/084G06N3/08G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,427
App. No.
17/763,901
Granted
Jul 21, 2026
Kind
B2
Abstract

A reinforcement learning method and system that selects actions to be performed by a reinforcement learning agent interacting with an environment. A causal model is implemented by a hindsight model neural network and trained using hindsight i.e. using future environment state trajectories. As the method and system does not have access to this future information when selecting an action, the hindsight model neural network is used to train a model neural network which is conditioned on data from current observations, which learns to predict an output of the hindsight model neural network.

Claims (34)

1 . A computer implemented method of reinforcement learning, comprising:

training an action selection neural network system to select actions to be performed by an agent in an environment for performing a task,

wherein the action selection neural network system is configured to receive input data from i) an observation characterizing a current state of the environment and ii) an output of a model neural network, and to process the input data in accordance with action selection neural network system parameters to generate an action selection output for selecting the actions to be performed by the agent; and

wherein the model neural network is configured to receive an input derived from the observation characterizing the current state of the environment and the output of the model neural network characterizes a predicted state trajectory comprising a series of k predicted future states of the environment starting from the current state;

wherein the method further comprises: training a hindsight model neural network having an output characterizing a state trajectory comprising a series of k states of the environment starting from a state of the environment at a time step t, by processing data from one or more observations characterizing the state of the environment at the time step t and at a series of k subsequent time steps the training the hindsight model comprising:

processing, using a hindsight value neural network, the output of the hindsight model neural network and data from the observation characterizing the state of the environment at the time step t, to generate an estimated hindsight value or state-action value for the state of the environment at the time step t, and backpropagating gradients of an objective function dependent upon a difference between the estimated hindsight value or state-action value for the state of the environment at the time step t and a training goal for the time step t to update parameters of the hindsight value neural network and the parameters of the hindsight model neural network; and

training the model neural network to generate an output that approximates the output of the hindsight model neural network generated from the observation characterizing the current state of the environment and the one or more observations characterizing the state of the environment at the series of k subsequent time steps after the time step t.

2 . The method of claim 1 wherein the action selection neural network system includes a state value neural network for selecting, or learning to select, the actions to be performed by the agent.

3 . The method of claim 2 wherein training the action selection neural network system comprises backpropagating gradients of an objective function dependent upon a difference between a state value or state-action value for the current state of the environment determined using the state value neural network and an estimated return or state-action value for the current state of the environment.

4 . The method of claim 2 wherein the action selection neural network system has parameters in common with the model neural network.

5 . The method of claim 1 further comprising providing, as an input to the action selection neural network system and to the hindsight value neural network, an internal state of one or more recurrent neural network (RNN) layers which receive data from the observations characterizing the state of the environment.

6 . The method of claim 1 further comprising providing, as an input to the hindsight model neural network and to the model neural network, an internal state of one or more recurrent neural network (RNN) layers which receive data from the observations characterizing the state of the environment.

7 . The method of claim 1 wherein the training goal for the time step t comprises a state value or state-action value target for the time step t.

8 . The method of claim 1 wherein the training goal for the time step t comprises an estimated return for the time step t.

9 . The method of claim 1 wherein training the output of the model neural network to approximate the output of the hindsight model neural network comprises backpropagating gradients of an objective function dependent upon a difference between the output of the hindsight model neural network characterizing the state trajectory and the output of the model neural network characterizing the predicted state trajectory.

10 . The method of claim 1 wherein the output of the hindsight model neural network characterizing the state trajectory and the output of the model neural network characterizing the predicted state trajectory each comprise features of a reduced dimensionality representation of one or more observations of the environment.

11 . The method of claim 1 wherein the output of the hindsight model neural network characterizing the state trajectory and the output of the model neural network characterizing the predicted state trajectory each have a dimensionality of less than their input, less than 20, or less than 10.

12 . The method of claim 1 wherein k is less than 20.

13 . The method of claim 1 further comprising maintaining a memory that stores data representing trajectories generated as a result of interaction of the agent with the environment, each trajectory comprising data at each of a series of time steps identifying at least an observation characterizing a state of the environment and a series of subsequent observations characterizing subsequent states of the environment for training the hindsight model neural network.

14 . A reinforcement learning neural network system, comprising: an action selection neural network system to select actions to be performed by an agent in an environment for performing a task,

wherein the action selection neural network system is configured to receive input data from i) an observation characterizing a current state of the environment and ii) an output of a model neural network, and to process the input data in accordance with action selection neural network system parameters to generate an action selection output for selecting the actions to be performed by the agent; and the model neural network,

wherein the model neural network is configured to receive an input from the observation characterizing the current state of the environment and the output of the model neural network characterizes a predicted state trajectory comprising a series of k predicted future states of the environment starting from the current state; wherein the system is configured to:

train a hindsight model neural network having an output characterizing a state trajectory comprising a series of k states of the environment starting from a state of the environment at a time step t, by processing observations characterizing the state of the environment at the time step t and at a series of k subsequent time steps training the hindsight model comprising:

processing, using a hindsight value neural network, the output of the hindsight model neural network and data from the observation characterizing the state of the environment at the time step t, to generate an estimated hindsight value or state-action value for the state of the environment at the time step t, and backpropagating gradients of an objective function dependent upon a difference between the estimated hindsight value or state-action value for the state of the environment at the time step t and a training goal for the time step t to update parameters of the hindsight value neural network and the parameters of the hindsight model neural network; and

train the model neural network to generate, an output that approximates the output of the hindsight model neural network generated from the observation characterizing the current state of the environment and the one or more observations characterizing the state of the environment at the series of k subsequent time steps after the time step t.

15 . A neural network computer system for performing a task, comprising: an action selection neural network system to select actions to be performed by an agent in an environment for performing a task,

wherein the action selection neural network system is configured to receive input data from i) an observation characterizing a current state of the environment and ii) an output of a model neural network, and to process the input data in accordance with action selection neural network system parameters to generate an action selection output for selecting the actions to be performed by the agent; and

a model neural network, wherein the model neural network is configured to receive an input from the observation characterizing the current state of the environment and the output of the model neural network characterizes a predicted state trajectory comprising a series of k predicted future states of the environment starting from the current state, wherein the model neural network has been trained to generate, an output that approximates an output of a hindsight model neural network that the hindsight model neural network would have generated from the observation characterizing the current state of the environment and one or more observations characterizing the state of the environment at the series of k subsequent time steps after a time step t, wherein the hindsight model neural network has been trained by processing the one or more observations characterizing the state of the environment at the time step t and at a series of k subsequent time steps, the training the hindsight model comprising:

processing, using a hindsight value neural network, the output of the hindsight model neural network and data from the observation characterizing the state of the environment at the time step t, to generate an estimated hindsight value or state-action value for the state of the environment at the time step t, and backpropagating gradients of an objective function dependent upon a difference between the estimated hindsight value or state-action value for the state of the environment at the time step t and a training goal for the time step t to update parameters of the hindsight value neural network and the parameters of the hindsight model neural network.

16 . The system of claim 14 wherein the action selection neural network system includes a state value neural network for selecting, or learning to select, the actions to be performed by the agent.

17 . The system of claim 16 wherein training the action selection neural network system comprises backpropagating gradients of an objective function dependent upon a difference between a state value or state-action value for the current state of the environment determined using the state value neural network and an estimated return or state-action value for the current state of the environment.

18 . The system of claim 16 wherein the action selection neural network system has parameters in common with the model neural network.

19 . The system of claim 14 further comprising providing, as an input to the action selection neural network system and to the hindsight value neural network, an internal state of one or more recurrent neural network (RNN) layers which receive data from the observations characterizing the state of the environment.

20 . The system of claim 14 further comprising providing, as an input to the hindsight model neural network and to the model neural network, an internal state of one or more recurrent neural network (RNN) layers which receive data from the observations characterizing the state of the environment.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2022
From: GUEZ, ARTHUR CLEMENT; VIOLA, FABIO; WEBER, THEOPHANE GUILLAUME; BUESING, LARS; HEESS, NICOLAS MANFRED OTTO
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 059430/0311 →
Continuity (2)
Provisional Application 62905915 · Sep 25, 2019
Related Publication 20220366245A1 · Nov 17, 2022
References Cited (56)
US 8447706B2 · Schneegaß · 2013 [cited by examiner]
US 20010044719A1 · Casey · 2001 [cited by examiner]
US 20200241878A1 · Chandak · 2020 [cited by examiner]
US 20210073563A1 · Karianakis · 2021 [cited by examiner]
CN 105637540A · 2016 [cited by applicant]
CN 108885717A · 2018 [cited by applicant]
CN 110023965A · 2019 [cited by applicant]
CN 110088774A · 2019 [cited by applicant]
CN 110114783A · 2019 [cited by applicant]
CN 110235148A · 2019 [cited by applicant]
WO WO2018224471A1 · 2018 [cited by applicant]
WO WO2021058588A1 · 2021 [cited by applicant]
Fang, et al. “Curriculum-guided hindsight experience replay” (Year: 2019). [cited by examiner]
Garcia, et al., “A comprehensive survey on safe reinforcement learning.” (Year: 2015). [cited by examiner]
Oh, et al., “Value prediction network.” (Year: 2017). [cited by examiner]
Fang, et al. “DHER: Hindsight experience replay for dynamic goals” (Year: 2018). [cited by examiner]
Lillicrap, et al. “Continuous control with deep reinforcement learning.” (Year: 2015). [cited by examiner]
Shao, et al., “Starcraft micromanagement with reinforcement learning and curriculum transfer learning.” (Year: 2018). [cited by examiner]
Rauber, Paulo, et al. “Hindsight policy gradients.” (Year: 2017). [cited by examiner]
Eysenbach, Benjamin, et al. Diversity is all you need: Learning skills without a reward function. (Year: 2018). [cited by examiner]
Fang, et al. “Curriculum-guided hindsight experience replay”, (Publication date: Sep. 6, 2019) (Year: 2019). [cited by examiner]
Zhao et al., “Curiosity-driven experience prioritization via density estimation.” (Year: 2019). [cited by examiner]
Andrychowicz et al., “Hindsight Experience Replay”, Arxiv.org, Jul. 2017. [cited by applicant]
Bellemare et al., “The arcade learning environment: An evaluation platform for general agents,” Journal of Artificial Intelligence Research, 47:253-279, Jun. 2013. [cited by applicant]
Buesing et al., “Woulda, coulda, shoulda: Counterfactually-guided policy search,” arXiv preprint arXiv:1811.06272, 2018. [cited by applicant]
Decision to Grant a Patent in Japanese Appln. No. 2022-519019, dated Jul. 18, 2023, 5 pages (with English translation). [cited by applicant]
Espholt et al., “Impala: Scalable distributed deep-RL with importance weighted actor-learner architectures,” 35th International Conference on Machine Learning, 1407-1416, Jul. 2018. [cited by applicant]
Farahmand, “Iterative value-aware model learning,” Advances in Neural Information Processing Systems, 9072-9083, Dec. 2018. [cited by applicant]
Gelada et al., “Deep MDP: Learning continuous latent space models for representation learning,” 36th International Conference on Machine Learning, 2170-2179, May 2019. [cited by applicant]
Guez et al., “An investigation of model-free planning,” 36th International Conference on Machine Learning, 2464-2473, May 2019. [cited by applicant]
Haarnoja et al., “Soft actor-critic: Off policy maximum entropy deep reinforcement learning with a stochastic actor,” 35th International Conference on Machine Learning, 1861-1870, Jul. 2018. [cited by applicant]
Harutyunyan et al., “Hindsight credit assignment,” Advances in Neural Information Processing Systems 32, 12488-12497, Dec. 2019. [cited by applicant]
Hasselt et al., “Deep reinforcement learning with double Q-learning,” 30th AAAI Conference on Artificial Intelligence, Feb. 2016. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/EP2020/076604, mailed on Sep. 12, 2020, 10 pages. [cited by applicant]
Jaderberg et al., “Reinforcement learning with unsupervised auxiliary tasks,” 5th International Conference on Learning Representations, Nov. 2016. [cited by applicant]
Kapturowski et al., “Recurrent experience replay in distributed reinforcement learning,” 7th International Conference on Learning Representations, May 2019. [cited by applicant]
Lillicrap et al., “Continuous control with deep reinforcement learning,” 4th International Conference on Learning Representations, ICLR, 2016. [cited by applicant]
Mnih et al., “Asynchronous Methods for Deep Reinforcement Learning,” International Conference on Machine Learning, Jun. 2016, 1928-1937. [cited by applicant]
Mnih et al., “Human-level control through deep reinforcement learning,” Nature, 518(7540):529-33, Feb. 2015. [cited by applicant]
Oh et al., “Value prediction network,” Advances in Neural Information Processing Systems 30, 6118-6128, 2017. [cited by applicant]
Pinto et al., “Asymmetric actor critic for image-based robot learning,” Robotics: Science and Systems (RSS), 2018. [cited by applicant]
Racanière et al., “Imagination-augmented agents for deep reinforcement learning,” Advances in neural information processing systems, 5690-5701, 2017. [cited by applicant]
Schrittwieser et al., “Mastering Atari, Go, Chess and Shogi by planning with a learned model,” arXiv preprint arXiv:1911.08265, 2019. [cited by applicant]
Schulman et al., Proximal policy optimization algorithms, arXiv preprint arXiv:1707.06347, 2017. [cited by applicant]
Silver et al., “The predictron: End-to end learning and planning,” International Conference on Machine Learning, 3191-3199, 2017. [cited by applicant]
Sutton et al., “Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction,” 10th International Conference on Autonomous Agents and Multiagent Systems—vol. (2) 761-768, 20… [cited by applicant]
“Reinforcement learning: An introduction,” 2nd ed., MIT Press, Nov. 2018. [cited by applicant]
Talvitie, “Model regularization for stable sample rollouts,” Proceedings of the Thirtieth Conference on Uncertainty in Artificial Intelligence, 780-789. AUAI Press, 2014. [cited by applicant]
Vapnik et al., “Learning using privileged information: similarity control and knowledge transfer,” Journal of machine learning research, 16(2023-2049):2, 2015. [cited by applicant]
Veeriah et al., “Discovery of Useful Questions as Auxiliary Tasks”, Advances in Neural Information Processing Systems, 9306-9317, Sep. 2019. [cited by applicant]
Weber et al., “Credit assignment techniques in stochastic computation graphs,” International Conference on Artificial Intelligence and Statistics, 2650-2660, 2019. [cited by applicant]
Zhu et al., “Reinforcement and imitation learning for diverse visuomotor skills,” CoRR, abs/1802.09564, 2018. [cited by applicant]
Office Action in European Appln. No. 20780639.9, dated Jul. 11, 2024, 5 pages. [cited by applicant]
Office Action in Chinese Appln. No. 202080066633.6, mailed on Jan. 23, 2025, 15 pages (with English translation). [cited by applicant]
Veeriah et al., “Discovery of Useful Questions as Auxiliary Tasks,” Advances in Neural Information Processing Systems 32, 2019, 12 pages. [cited by applicant]
Zhong et al., “An Intelligent Control System based on Neural Networks and Reinforcement Learning,” Journal of Southwest University, Natural Science Edition, Nov. 29, 2013, 35(11):173 (English abstract). [cited by applicant]