IP Library Granted Patent US 12,205,032
Granted Patent B2
US 12,205,032 · App. 18/542,476 · Granted Jan 21, 2025

Distributional reinforcement learning using quantile function neural networks

Inventors: Georg Ostrovski (London, GB); William Clinton Dabney (London, GB)
Assignee: DeepMind Technologies Limited
G06N3/08G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,032
App. No.
18/542,476
Granted
Jan 21, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for selecting an action to be performed by a reinforcement learning agent interacting with an environment. In one aspect, a method comprises: receiving a current observation; for each action of a plurality of actions: randomly sampling one or more probability values; for each probability value: processing the action, the current observation, and the probability value using a quantile function network to generate an estimated quantile value for the probability value with respect to a probability distribution over possible returns that would result from the agent performing the action in response to the current observation; determining a measure of central tendency of the one or more estimated quantile values; and selecting an action to be performed by the agent in response to the current observation using the measures of central tendency for the actions.

Claims (46)

1. A method performed by one or more computers, the method comprising:

selecting an action to be performed by a reinforcement learning agent interacting with an environment, comprising:

receiving a current observation characterizing a current state of the environment;

for each action of a plurality of actions that can be performed by the agent to interact with the environment:

processing a network input that comprises the action, the current observation, a risk-sensitive value that increases a level of risk aversion in selecting the action to be performed by the agent, using an action selection neural network, in accordance with values of a set of action selection neural network parameters, to generate an action selection neural network output,

wherein the action selection neural network output characterizes a return distribution that defines a probability distribution over possible returns that would result from the agent performing the action in response to the observation; and

selecting an action from the plurality of actions to be performed by the agent in response to the current observation using the action selection neural network outputs.

2. The method of claim 1 , wherein for each action of the plurality of actions, the risk-sensitive value is generated as an output of a risk measure function.

3. The method of claim 2 , wherein for each action of the plurality of actions, the risk-sensitive value is generated as an output of a distortion risk measure function.

4. The method of claim 2 , wherein for each action of the plurality of actions, generating the risk-sensitive value as the output of the risk measure function comprises:

sampling a random value; and

processing the random value using the risk measure function to generate the risk-sensitive value.

5. The method of claim 4 , wherein the random value is sampled from a uniform distribution over an interval [0,1].

6. The method of claim 4 , wherein the risk measure function is a non-decreasing function mapping a source domain to a target range.

7. The method of claim 6 , wherein the risk measure function maps a point 0 in the source domain to a point 0 in the target range.

8. The method of claim 6 , wherein the risk measure function maps a point 1 in the source domain to a point 1 in the target range.

9. The method of claim 2 , wherein the risk measure function is configured to generate risk-sensitive values that increase a level of risk aversion in selecting the action to be performed by the agent.

10. A system comprising:

one or more computers; and

one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

selecting an action to be performed by a reinforcement learning agent interacting with an environment, comprising:

receiving a current observation characterizing a current state of the environment;

for each action of a plurality of actions that can be performed by the agent to interact with the environment:

processing a network input that comprises the action, the current observation, a risk-sensitive value that increases a level of risk aversion in selecting the action to be performed by the agent, using an action selection neural network, in accordance with values of a set of action selection neural network parameters, to generate an action selection neural network output,

wherein the action selection neural network output characterizes a return distribution that defines a probability distribution over possible returns that would result from the agent performing the action in response to the observation; and

selecting an action from the plurality of actions to be performed by the agent in response to the current observation using the action selection neural network outputs.

11. The system of claim 10 , wherein for each action of the plurality of actions, the risk-sensitive value is generated as an output of a risk measure function.

12. The system of claim 11 , wherein for each action of the plurality of actions, the risk-sensitive value is generated as an output of a distortion risk measure function.

13. The system of claim 11 , wherein for each action of the plurality of actions, generating the risk-sensitive value as the output of the risk measure function comprises:

sampling a random value; and

processing the random value using the risk measure function to generate the risk-sensitive value.

14. The system of claim 13 , wherein the random value is sampled from a uniform distribution over an interval [0,1].

15. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

selecting an action to be performed by a reinforcement learning agent interacting with an environment, comprising:

receiving a current observation characterizing a current state of the environment;

for each action of a plurality of actions that can be performed by the agent to interact with the environment:

processing a network input that comprises the action, the current observation, a risk-sensitive value that increases a level of risk aversion in selecting the action to be performed by the agent, using an action selection neural network, in accordance with values of a set of action selection neural network parameters, to generate an action selection neural network output,

wherein the action selection neural network output characterizes a return distribution that defines a probability distribution over possible returns that would result from the agent performing the action in response to the observation; and

selecting an action from the plurality of actions to be performed by the agent in response to the current observation using the action selection neural network outputs.

16. The non-transitory computer storage media of claim 15 , wherein for each action of the plurality of actions, the risk-sensitive value is generated as an output of a risk measure function.

17. The non-transitory computer storage media of claim 16 , wherein for each action of the plurality of actions, the risk-sensitive value is generated as an output of a distortion risk measure function.

18. The non-transitory computer storage media of claim 16 , wherein for each action of the plurality of actions, generating the risk-sensitive value as the output of the risk measure function comprises:

sampling a random value; and

processing the random value using the risk measure function to generate the risk-sensitive value.

19. The non-transitory computer storage media of claim 18 , wherein the random value is sampled from a uniform distribution over an interval [0,1].

20. The non-transitory computer storage media of claim 18 , wherein the risk measure function is a non-decreasing function mapping a source domain to a target range.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071109/0414 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2024
From: OSTROVSKI, GEORG; DABNEY, WILLIAM CLINTON
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 066012/0216 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2023
From: OSTROVSKI, GEORG; DABNEY, WILLIAM CLINTON
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 065936/0819 →
Continuity (5)
Continuation 18169803 · Feb 15, 2023
Continuation 16767046
Provisional Application 62646154 · Mar 21, 2018
Provisional Application 62628875 · Feb 9, 2018
Related Publication 20240135182A1 · Apr 25, 2024
References Cited (56)
US 20170076201A1 · van Hasselt · 2017 [cited by examiner]
US 20190332923A1 · Gendron-Bellemare · 2019 [cited by examiner]
CN 106094516 · 2016 [cited by applicant]
WO WO2017091629 · 2017 [cited by examiner]
[No Author Listed], “Autoregressive Quantile Networks for Generative Modeling,” International Conference on Machine Learning, 2018, 16 pages. [cited by applicant]
Allais, “Allais paradox,” Utility and Probability, 1990, pp. 3-9. [cited by applicant]
Arjovsky et al., “Wasserstein GAN,” CoRR, Dec. 2017, arXiv;1701.07875v3, 32 pages. [cited by applicant]
Azar et al., “On the sample complexity of reinforcement learning with a generative model,” CoRR, Jun. 2012, https://arxiv.org/ftp/arxiv/papers/1206/1206.6461, 8 pages. [cited by applicant]
Barth-Maron et al., “Distributed Distributional Deterministic Policy Gradients,” CoRR, Apr. 2018, https://arxiv.org/abs/1804.08617, 16 pages. [cited by applicant]
Bellemare et al., “A distributional perspective on reinforcement learning,” CoRR, Jul. 2017, https://arxiv.org/abs/1707.06887, 19 pages. [cited by applicant]
Bellemare et al., “The Arcade Learning Environment; an evaluation platform for general agents,” Journal of Artificial Intelligence Research, Jun. 2013, 47:253-279. [cited by applicant]
Bellman, “Dynamic Programming,” Science, Jul. 1966, 153(3731):34-37. [cited by applicant]
Bousquet et al., “From optimal transport to generative modeling: the vegan cookbook,” CoRR, May 2017, https://arxiv.org/abs/1705.07642, 15 pages. [cited by applicant]
Chow et al., “Algorithms for CVaR optimization in MDPs,” Advances in Neural Information Processing Systems 27 (NIPS 2014), 2014, 9 pages. [cited by applicant]
Dabney et al., “Distributional reinforcement learning with quantile regression,” CoRR, Oct. 2017, https://arxiv.org/abs/1710.10044, 14 pages. [cited by applicant]
Dabney et al., “Implicit Quantile Networks for Distributional Reinforcement Learning,” CoRR, Jun. 2018, arXiv:1806.06923v1, 14 pages. [cited by applicant]
Dhaene et al., “Remarks on quantiles and distortion risk measures.,” European Actuarial Journal, Nov. 2012, 2(2):319-328. [cited by applicant]
Fortunato et al., “Noisy networks for exploration,” CoRR, Jun. 2017, https://arxiv.org/abs/1706.10295, 21 pages. [cited by applicant]
Geist et al., “Kalman temporal differences,” Journal of Artificial Intelligence Research, Oct. 2010, 39:483-532. [cited by applicant]
Gonzalez et al., “On the shape of the probability weighting function,” Cognitive Psychology, Feb. 1999, 38(1):129-166. [cited by applicant]
Gruslys et al., “The Reactor: a fast and sample efficient actor-critic agent for reinforcement learning,” CoRR, Apr. 2017, https://arxiv.org/abs/1704.04651, 18 pages. [cited by applicant]
Hasselt et al., “Deep reinforcement learning with double Q-learning,” CoRR, Sep. 2015, https://arxiv.org/abs/1509.06461, 13 pages. [cited by applicant]
Hessel et al., “Rainbow: combining improvements in deep reinforcement learning,” CoRR, Oct. 2017, https://arxiv.org/abs/1710.02298, 14 pages. [cited by applicant]
Howard et al., “Risk-sensitive markov decision processes,” Management Science, Mar. 1972, 18(7):356-369. [cited by applicant]
Huber et al., “Robust estimation of a location parameter,” Breakthrough in Statistics, Jun. 1963, pp. 492-518. [cited by applicant]
Jaquette, “Markov decision processes with a new optimality criterion: discrete time,” The Annals of Statistics, May 1973, 1(3):496-505. [cited by applicant]
Lattimore et al., “PAC bounds for discounted MDPs,” International Conference on Algorithmic Learning Theory, 2012, pp. 320-334. [cited by applicant]
Maddison et al., “Particle value functions,” CoRR, Mar. 2017, https://arxiv.org/abs/1703.05820, 12 pages. [cited by applicant]
Majumdar et al., “How should a robot assess risk? Towards an axiomatic theory of risk in robotics,” CoRR, Oct. 2017, https://arxiv.org/abs/1710.11040, 16 pages. [cited by applicant]
Marcus et al., “Risk sensitive markov decision processes,” Systems and Control in the Twenty-First Century, 1997, pp. 263-279. [cited by applicant]
Mnih et al., “Human-level control through deep reinforcement learning,” Nature, Feb. 2015, 518(7540):529-533. [cited by applicant]
Moerland et al., “Efficient exploration with double uncertain value networks,” CoRR, Nov. 2017, https://arxiv.org/abs/1711.10789, 17 pages. [cited by applicant]
Morimura et al., “Nonparametric return distribution approximation for reinforcement learning,” Proceedings of the 27th International Conference on Machine Learning (ICML), 2010, 8 pages. [cited by applicant]
Morimura et al., “Parametric return density estimation for reinforcement learning,” CoRR, Mar. 2012, https://arxiv.org/ftp/arxiv/papers/1203/1203.3497, 8 pages. [cited by applicant]
Muller, “Integral probability metrics and their generating classes of functions,” Advances in Applied Probability, Jun. 1997, 29(2):429-443. [cited by applicant]
Nair et al., “Massively parallel methods for deep reinforcement learning,” CoRR, Jul. 2015, https://arxiv.org/abs/1507.04296, 14 pages. [cited by applicant]
Office Action in European Appln. No. 19704796.2, dated Aug. 25, 2023, 16 pages. [cited by applicant]
Office Action in European Appln. No. 19704796.2, dated Jan. 27, 2023, 13 pages. [cited by applicant]
Osband et al., “(More) efficient reinforcement learning via posterior sampling,” Advances in Neural Information Processing Systems 26 (NIPS 2013), 2013, 10 pages. [cited by applicant]
PCT International Preliminary Report on Patentability in International Appln. No. PCT/EP2019053315, dated Aug. 11, 2020, 14 pages. [cited by applicant]
PCT International Search Report and Written Opinion in International Appln. No. PCT/EP2019053315, dated May 8, 2019, 22 pages. [cited by applicant]
Rowland et al., “An analysis of categorical distributional reinforcement learning,” CoRR, Feb. 2018, https://arxiv.org/abs/1802.08163, 19 pages. [cited by applicant]
Schaul et al., “Prioritized experience replay,” CoRR, Nov. 2015, https://arxiv.org/abs/1511.05952, 21 pages. [cited by applicant]
Schaul et al., “Universal value function approximators,” In International Conference on Machine Learning, 2015, 37:1312-1320. [cited by applicant]
Sobel, “The variance of discounted markov decision processes,” Journal of Applied Probability, Dec. 1982, 19(4):794-802. [cited by applicant]
Sutton, “Learning to predict by the methods of temporal differences,” Machine Learning, 1988, 3(1):9-44. [cited by applicant]
Tolstikhin et al., “Wasserstein auto-encoders,” CoRR, Nov. 2017, https://arxiv.org/abs/1711.01558, 20 pages. [cited by applicant]
Tversky et al., “Advances in prospect theory: cumulative representation of uncertainty,” Journal of Risk and Uncertainty, Oct. 1992, 5(4):297-323. [cited by applicant]
Wang et al., “A class of distortion operators for pricing financial and insurance risks,” Journal of Risk and Insurance, Mar. 2000, 67(1):15-36. [cited by applicant]
Wang et al., “Dueling network architectures for deep reinforcement learning,” Proceedings of The 33rd. International Conference on Machine Learning, PMLR, 2016, 48: 9 pages. [cited by applicant]
Wang et al., “Premium calculation by transforming the layer premium density,” ASTIN Bulletin: The Journal of the IAA, May 1996, 26(1):71-92. [cited by applicant]
Watkins, “Learning from delayed rewards,” Thesis for the degree of Doctor, King's College, May 1989, 241 pages. [cited by applicant]
White, “Mean, variance, and probabilistic criteria in finite markov decision processes: a review,” Journal of Optimization Theory and Applications, 1988, 56(1):1-29. [cited by applicant]
Wu et al., “Curvature of the probability weighting function,” Management Science, Dec. 1996, 42(12):1676-1690. [cited by applicant]
Yaari, “The dual theory of choice under risk,” Econometrica: Journal of the Econometric Society, Jan. 1987, 55(1):95-115. [cited by applicant]
Yu et al., “More than a million ways to be pushed. a high-fidelity experimental dataset of planar pushing,” 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Oct. 9-14, 2016, 8 pages. [cited by applicant]