IP Library › Granted Patent US 12,614,067
Granted Patent B2
US 12,614,067 · App. 17/620,164 · Granted Apr 28, 2026

Robust reinforcement learning for continuous control with model misspecification

Inventors: Daniel J. Mankowitz (St. Albans, GB); Nir Levine (Sunnyvale, CA); Rae Chan Jeong (New York, CA); Abbas Abdolmaleki (London, GB); Jost Tobias Springenberg (London, GB); Todd Andrew Hester (Seattle, WA); Timothy Arthur Mann (Harpenden, GB); Martin Riedmiller (Balgheim, DE)
Assignee: GDM Holding LLC
G06N3/08G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,067
App. No.
17/620,164
Granted
Apr 28, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a policy neural network having policy parameters. One of the methods includes sampling a mini-batch comprising one or more observation-action-reward tuples generated as a result of interactions of a first agent with a first environment; determining an update to current values of the Q network parameters by minimizing a robust entropy-regularized temporal difference (TD) error that accounts for possible perturbations of the states of the first environment represented by the observations in the observation-action-reward tuples; and determining, using the Q-value neural network, an update to the policy network parameters using the sampled mini-batch of observation-action-reward tuples.

Claims (56)

1 . A method of training a policy neural network having a plurality of policy network parameters,

wherein the policy neural network is configured to receive a policy input comprising an observation characterizing a current state of an environment and to process the network input in accordance with the policy network parameters to generate a policy output that defines a probability distribution over a space of possible actions to be performed by an agent interacting with the environment,

wherein the policy neural network is trained jointly with a Q-value neural network (i) having a plurality of Q network parameters and (ii) configured to receive a Q network input comprising data identifying an action and the observation and to process the Q network input in accordance with the Q network parameters to generate a Q value for the action, and

wherein the method comprises:

sampling a mini-batch comprising one or more observation—action—reward tuples generated as a result of interactions of a first agent with a first environment;

determining an update to current values of the Q network parameters comprising:

for each tuple and for each possible perturbation of the environment in a set of a plurality of possible perturbations of the environment:

causing the agent to perform the action in the tuple when the state of the environment represented by the observation in the tuple has been perturbed by applying the possible perturbation to the state of the environment represented by the observation; and

in response, obtaining a next observation characterizing a perturbed next state; and

minimizing a robust entropy-regularized temporal difference (TD) error that accounts for possible perturbations of the states of the first environment represented by the observations in the sampled mini-batch of observation-action-reward tuples and measures, for each tuple, an error between (i) a target Q value determined based on the reward in the tuple and an infimum or average of entropy-regularized Q values for the perturbed next states for the plurality of possible perturbations and (ii) a Q value for the observation-action pair in the tuple; and

determining, using the Q-value neural network, an update to the policy network parameters using the sampled mini-batch of observation-action-reward tuples.

2 . The method of claim 1 , further comprising:

providing data specifying the trained policy neural network for use in controlling a second agent interacting with a second, different environment.

3 . The method of claim 2 , wherein the first environment is a simulation of a real-world environment and the second, different environment is the real-world environment, wherein the first agent is a real-world mechanical agent, and wherein the first agent is a simulation of the second agent.

4 . The method of claim 2 , wherein the first environment is a first real-world environment and the second, different environment is a second, different real-world environment.

5 . The method of claim 4 , wherein the first agent is a first real-world mechanical agent and the second agent is a second, different real-world mechanical agent.

6 . The method of claim 1 , wherein the robust entropy-regularized temporal difference (TD) error measures, for each tuple, an error between (i) a sum of the reward in the tuple and an infimum of entropy-regularized Q values for the perturbed next states and (ii) a Q value for the observation-action pair in the tuple.

7 . The method of claim 1 , wherein the robust entropy-regularized temporal difference (TD) error measures, for each tuple, an error between (i) a sum of the reward in the tuple and an average of entropy-regularized Q values for the perturbed next states and (ii) a Q value for the observation-action pair in the tuple.

8 . The method of claim 6 , further comprising generating a respective entropy-regularized Q value for each of the perturbed next states, comprising:

processing the next observation characterizing the perturbed next state using the policy neural network to generate a next probability distribution over possible actions;

sampling a next action from the next probability distribution;

determining a Q value for the next observation-next action pair;

determining an entropy regularization penalty based on a divergence between the next probability distribution and a reference next probability distribution; and

determining the respective entropy-regularized Q value for the perturbed next state from at least the Q value for the next observation-next action pair and the entropy regularization penalty.

9 . The method of claim 8 , wherein the reference next probability distribution is a probability distribution generated by the policy neural network in accordance with earlier values of the policy parameters.

10 . The method of claim 8 , wherein the Q value for the next observation-next action pair is generated by a target Q neural network having the same architecture as the Q-value neural network but parameter values that change more slowly during training than the Q network parameters.

11 . The method of claim 1 , wherein the possible perturbations include perturbations to one or more of: dimensions of portions of the agent or dimensions of objects in the environment, orientation of portions of the agent, or orientations of objects in the environment.

12 . The method of claim 1 , wherein the policy output comprises output parameters of a probability distribution over a continuous space of actions.

13 . The method of claim 12 , wherein the parameters are means and covariances of a multi-variate Normal distribution over the continuous space of actions.

14 . The method of claim 1 , wherein determining, using the Q network, an update to the policy network parameters using the sampled batch of observation—action—reward tuples comprises applying an actor-critic technique to the sampled batch using the Q network as a critic.

15 . The method of claim 14 , wherein the actor-critic technique is Maximum a Posteriori Policy Optimisation.

16 . The method of claim 14 , wherein the actor-critic technique is stochastic value gradients (SVG).

17 . One or more non-transitory computer-readable storage media encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations for training a policy neural network having a plurality of policy network parameters,

wherein the policy neural network is configured to receive a policy input comprising an observation characterizing a current state of an environment and to process the network input in accordance with the policy network parameters to generate a policy output that defines a probability distribution over a space of possible actions to be performed by an agent interacting with the environment,

wherein the policy neural network is trained jointly with a Q-value neural network (i) having a plurality of Q network parameters and (ii) configured to receive a Q network input comprising data identifying an action and the observation and to process the Q network input in accordance with the Q network parameters to generate a Q value for the action, and

wherein the operations comprise:

sampling a mini-batch comprising one or more observation—action—reward tuples generated as a result of interactions of a first agent with a first environment;

determining an update to current values of the Q network parameters, comprising:

for each tuple and for each possible perturbation of the environment in a set of a plurality of possible perturbations of the environment:

causing the agent to perform the action in the tuple when the state of the environment represented by the observation in the tuple has been perturbed by applying the possible perturbation to the state of the environment represented by the observation; and

in response, obtaining a next observation characterizing a perturbed next state; and

minimizing a robust entropy-regularized temporal difference (TD) error that accounts for possible perturbations of the states of the first environment represented by the observations in the sampled mini-batch of observation—action—reward tuples and measures, for each tuple, an error between (i) a target Q value determined based on the reward in the tuple and an infimum or average of entropy-regularized Q values for the perturbed next states for the plurality of possible perturbations and (ii) a Q value for the observation-action pair in the tuple; and

determining, using the Q-value neural network, an update to the policy network parameters using the sampled mini-batch of observation-action-reward tuples.

18 . A system comprising one or more computers and one or more storage devices storing instruction that when executed by one or more computers cause the one or more computers to perform operations fort raining a policy neural network having a plurality of policy network parameters,

wherein the policy neural network is configured to receive a policy input comprising an observation characterizing a current state of an environment and to process the network input in accordance with the policy network parameters to generate a policy output that defines a probability distribution over a space of possible actions to be performed by an agent interacting with the environment,

wherein the policy neural network is trained jointly with a Q-value neural network (i) having a plurality of Q network parameters and (ii) configured to receive a Q network input comprising data identifying an action and the observation and to process the Q network input in accordance with the Q network parameters to generate a Q value for the action, and

wherein the operations comprise:

sampling a mini-batch comprising one or more observation—action—reward tuples generated as a result of interactions of a first agent with a first environment;

determining an update to current values of the Q network parameters, comprising:

for each tuple and for each possible perturbation of the environment in a set of a plurality of possible perturbations of the environment:

causing the agent to perform the action in the tuple when the state of the environment represented by the observation in the tuple has been perturbed by applying the possible perturbation to the state of the environment represented by the observation; and

in response, obtaining a next observation characterizing a perturbed next state; and

minimizing a robust entropy-regularized temporal difference (TD) error that accounts for possible perturbations of the states of the first environment represented by the observations in the sampled mini-batch of observation—action—reward tuples and measures, for each tuple, an error between (i) a target Q value determined based on the reward in the tuple and an infimum or average of entropy-regularized Q values for the perturbed next states for the plurality of possible perturbations and (ii) a Q value for the observation-action pair in the tuple; and

determining, using the Q-value neural network, an update to the policy network parameters using the sampled mini-batch of observation—action—reward tuples.

19 . The system of claim 18 , wherein the robust entropy-regularized temporal difference (TD) error measures, for each tuple, an error between (i) a sum of the reward in the tuple and an infimum of entropy-regularized Q values for the perturbed next states and (ii) a Q value for the observation-action pair in the tuple.

20 . The system of claim 18 , wherein the robust entropy-regularized temporal difference (TD) error measures, for each tuple, an error between (i) a sum of the reward in the tuple and an average of entropy-regularized Q values for the perturbed next states and (ii) a Q value for the observation-action pair in the tuple.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2021
From: MANKOWITZ, DANIEL J.; LEVINE, NIR; JEONG, RAE CHAN; ABDOLMALEKI, ABBAS; SPRINGENBERG, JOST TOBIAS; HESTER, TODD ANDREW; MANN, TIMOTHY ARTHUR; RIEDMILLER, MARTIN
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 058452/0150 →
Continuity (2)
Provisional Application 62862616 · Jun 17, 2019
Related Publication 20220343157A1 · Oct 27, 2022
References Cited (48)
US 10960539B1 · Kalakrishnan · 2021 [cited by examiner]
US 11453121B2 · Sermanet · 2022 [cited by examiner]
US 20190155969A1 · Haaland · 2019 [cited by examiner]
EP 3525136A1 · 2019 [cited by examiner]
WO WO2017139507A1 · 2017 [cited by examiner]
Schulman et al., “Equivalence between policy gradients and soft q-learning,” CoRR, Apr. 2017, arXiv:1704.06440, 15 pages (Year: 2017). [cited by examiner]
Dai et al., “From Credit Assignment to Entropy Regularization: Two New Algorithms for Neural Sequence Prediction,” CoRR, Apr. 2018, arxiv.org/abs/1804.10974, 19 pages (Year: 2018). [cited by examiner]
Abdolmaleki et al., “Maximum a posteriori policy optimisation, ” CoRR, Jun. 2018, arXiv:1806.06920, 23 pages. [cited by applicant]
Abdolmaleki et al., “Relative entropy regularized policy iteration,” CoRR, Dec. 2018, arxiv.org/abs/1812.02256. 23 pages. [cited by applicant]
Andrychowicz et al., “Learning dexterons in-hand manipulation,” The International Journal of Robotics Research, Nov. 2019, 18 pages. [cited by applicant]
Braun et al., “Path integral control and bounded rationality,” 2011 IEEE Symposium on Adaptive Dynamic Programming and Reinforcement Learning, Apr. 2011, pp. 202-209. [cited by applicant]
Christiano et al., “Transfer from simulation to real world through learning deep inverse dynamics model,” CoRR, Oct. 2016, arxiv.org/abs/1610.03518, 8 pages. [cited by applicant]
Dai et al., “From Credit Assignment to Entropy Regularization: Two New Algorithms for Neural Sequence Prediction,” CoRR, Apr. 2018, arxiv.org/abs/1804.10974, 19 pages. [cited by applicant]
Derman et al., “Soft-robust actor-critic policy-gradient,” CoRR, Mar. 2018, arXiv:1803.04848, 17 pages. [cited by applicant]
Devin et al., “Learning modular neural network policies for multi-task and multi-robot transfer,” 2017 IEEE International Conference on Robotics and Automation, Jun. 2017, pp. 2169-2176. [cited by applicant]
Di Castro et al., “Policy gradients with variance related risk criteria,” CoRR, Jun. 2012, arXiv:1206.6404, 8 pages. [cited by applicant]
Duan et al., “Benchmarking deep reinforcement learning for continuous control, ” Proceedings of The 33rd International Conference on Machine Learning, 2016, 48:1329-1338. [cited by applicant]
Dulac-Arnold et al., “Challenges of real-world reinforcement learning,” CoRR, Apr. 2019, arxiv.org/abs/1904.12901, 13 pages. [cited by applicant]
Fox et al., “G-learning: Taming the noise in reinforcement learning via soft updates,” CoRR, Dec. 2015, arxiv.org/abs/1512.08562, 11 pages. [cited by applicant]
Gao, “Machine learning applications for data center optimization,” Google, 2014, 13 pages. [cited by applicant]
Gran-Moya et al., “Planning with informationprocessing constraints and model uncertainty in markov decision processes,” CoRR, Sep. 2016, arxiv.org/abs/1604.02080, pp. 475-491. [cited by applicant]
Haarnoja et al., “Soft actor-critic algorithms and applications,” CoRR, Dec. 2018, arxiv.org/abs/1812.05905, 17 pages. [cited by applicant]
Heess et al., “Learning continuous control policies by stochastic value gradients,” Advances in Neural Information Processing Systems 28, 2015, 9 pages. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/EP2020/066755, dated Dec. 30, 2021, 9 pages. [cited by applicant]
International Serach Report and Written Opinion in International Appln. No. PCT/EP2020/066755, dated Sep. 23, 2020, 14 pages. [cited by applicant]
Iyengar, “Robust dynamic programming,” Mathematics of Operations Research, May 2005, 30(2):257-280. [cited by applicant]
Kalashnikov et al., “Scalable deep reinforcement learning for vision-based robotic manipulation,” Proceedings of The 2nd Conference on Robot Learning, 87:651-673. [cited by applicant]
Kingma et al., “Anto-encoding variational bayes,” CoRR, Dec. 2013, arXiv: 1312.6114, 14 pages. [cited by applicant]
Mankowitz et al., “A bayesian approach to robust reinforcement learning,” Proceedings of The 35th Uncertainty in Artificial Intelligence Conference, 2020, 115:648-658. [cited by applicant]
Mankowitz et al., “Learning robust options,” Thirty-Second AAAI Conference on Artificial Intelligence, 2018, 32(1):6409-6416. [cited by applicant]
Mankowitz et al., “Unicorn: Continual learning with a universal, off-policy agent,” CoRR, Feb. 2018, arxiv.org/abs/ 1802.08294, 17 pages. [cited by applicant]
Mnih et al., “Human-level control through deep reinforcement learning,” Nature, Feb. 2015, 518(7540):529-533. [cited by applicant]
Nachum et al., “Bridging the gap between value and policy based reinforcement learning,” Advances: in Neural Information Processing Systems, 2017, pp. 2775-2785. [cited by applicant]
Nilim et al., “Robust control of markov decision processes with uncertain transition matrices,” Operations Research, Oct. 2005, 53(5):780-798. [cited by applicant]
Peng et al., “Sim-to-real transfer of robotic control with dynamics randomization,” 2018 IEEE International Conference on Robotics and Automation, May 2018, pp. 1-8. [cited by applicant]
Rastogi et al., “Sample-efficient reinforcement learning via difference models,” Third Machine Learning in Planning and Control of Robot Motion Workshop at ICRA, 2018, 6 pages. [cited by applicant]
Rezende et al., “Stochastic backpropagation and approximate inference in deep generative models,” Proceedings of the 31st International Conference on Machine Learning, 2014, 32(2):1278-1286. [cited by applicant]
Riedmiller et al., “Learning by playingsolving sparse reward tasks from scratch,” Proceedings of the 35th International Conference on Machine Learning, 2018, 80:4344-4353. [cited by applicant]
Rubin et al., “Trading value and information in MDPs,” Decision Making with Imperfect Decision Makers, 2012, 16 pages. [cited by applicant]
Rusu et al., “Policy distillation,” CoRR, Nov. 2015, arxiv.org/abs/1511.06295, 13 pages. [cited by applicant]
Schulman et al., “Equivalence between policy gradients and soft q-learning,” CoRR, Apr. 2017, arXiv:1704.06440, 15 pages. [cited by applicant]
shadowrobot.com [online], “Shadow Dexterous Hand Technical Specification,” Sep. 2020, retrieved on Mar. 16, 2022, retrieved from URL<https://www.shadowrobot.com/wp-content/uploads/shadow dexterous hand technical specifi… [cited by applicant]
Tamar et al., “Scaling up robust mdps using function approximation,” International Conference on Machine Learning, 2014, pp. 181-189. [cited by applicant]
Tassa et al., “Deepmind control suite,” CoRR, Jan. 2018, arxiv.org/abs/1801.00690, 24 pages. [cited by applicant]
Teh et al., “Distral: Robust multitask reinforcement learning,” Advances in Neural Information Processing Systems 30, 2017, 11 pages. [cited by applicant]
Tessler et al., “A deep hierarchical approach to lifelong learning in minecraft,” Thirty-First AAAI Conference on Artificial Intelligence, 2017, 31(1):1553-1561. [cited by applicant]
Wiesemann et al., “Robust Markov Decision Processes,” Math. Oper. Res., Nov. 2012, 38(1):153-183. [cited by applicant]
Office Action in European Appln. No. 20733750.2, dated Nov. 7, 2023, 9 pages. [cited by applicant]