IP Library › Granted Patent US 12,444,182
Granted Patent B2
US 12,444,182 · App. 18/016,746 · Granted Oct 14, 2025

Training action selection neural networks using auxiliary tasks of controlling observation embeddings

Inventors: Markus Wulfmeier (Balgheim, DE); Tim Hertweck (Lauchringen, DE); Martin Riedmiller (Balgheim, DE)
Assignee: GDM Holding LLC
G06V10/82G06V10/87
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,182
App. No.
18/016,746
Granted
Oct 14, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for selecting actions to be performed by an agent interacting with an environment to accomplish a goal. In one aspect, a method comprises: obtaining an observation characterizing a state of the environment, processing the observation using an embedding model to generate a lower-dimensional embedding of the observation, determining an auxiliary task reward based on a value of a particular dimension of the embedding, determining an overall reward based at least in part on the auxiliary task reward, and determining an update to values of multiple parameters of an action selection neural network based on the overall reward using a reinforcement learning technique.

Claims (49)

1. A method for training an action selection neural network having a plurality of parameters that is used to select actions to be performed by an agent interacting with an environment, wherein the action selection neural network is configured to process an input comprising an observation characterizing a state of the environment to generate an action selection output that comprises a respective action score for each action in a set of possible actions that can be performed by the agent, and select the action to be performed by the agent from the set of possible actions based on the action scores, the method comprising:

obtaining an observation characterizing a state of the environment at a time step;

processing the observation using an embedding model to generate a lower-dimensional embedding of the observation, wherein the lower-dimensional embedding of the observation has a plurality of dimensions;

determining an auxiliary task reward for the time step based on a value of a particular dimension of the embedding, wherein the auxiliary task reward corresponds to an auxiliary task of controlling the value of the particular dimension of the embedding;

determining an overall reward for the time step based at least in part on the auxiliary task reward for the time step; and

determining an update to values of the plurality of parameters of the action selection neural network based on the overall reward for the time step using a reinforcement learning technique.

2. The method of claim 1 , wherein the auxiliary task of controlling the value of the particular dimension of the embedding comprises maximizing or minimizing the value of the particular dimension of the embedding.

3. The method of claim 2 , wherein determining the auxiliary task reward for the time step comprises:

determining a maximum value of the particular dimension of embeddings of respective observations characterizing the state of the environment at each of a plurality of time steps;

determining a minimum value of the particular dimension of embeddings of respective observations characterizing the state of the environment at each of the plurality of time steps; and

determining the auxiliary task reward for the time step based on: (i) the value of the particular dimension of the embedding at the time step, (ii) the maximum value corresponding to the particular dimension of the embedding, and (iii) the minimum value corresponding to the particular dimension of the embedding.

4. The method of claim 3 , wherein determining the auxiliary task reward for the time step comprises:

determining a ratio of: (i) a difference between the maximum value corresponding to the particular dimension of the embedding and the value of the particular dimension of the embedding at the time step, and (ii) a difference between the maximum value and the minimum value corresponding to the particular dimension of the embedding.

5. The method of claim 3 , wherein determining the auxiliary task reward for the time step comprises:

determining a ratio of: (i) a difference between the value of the particular dimension of the embedding at the time step and the minimum value corresponding to the particular dimension of the embedding, and (ii) a difference between the maximum value and the minimum value corresponding to the particular dimension of the embedding.

6. The method of claim 1 , further comprising selecting the auxiliary task of controlling the value of the particular dimension of the embedding from a set of possible auxiliary tasks in accordance with a task selection policy, wherein each possible auxiliary task corresponds to controlling a value of a respective dimension of the embedding.

7. The method of claim 1 , wherein the reinforcement learning technique is an off-policy reinforcement learning technique.

8. The method of claim 1 , wherein the embedding model comprises a random matrix, and processing the observation using the embedding model comprises:

applying the random matrix to a vector representation of the observation to generate a projection of the observation; and

applying a non-linear activation function to the projection of the observation.

9. The method of claim 8 , further comprising generating the vector representation of the observation by flattening the observation into a vector.

10. The method of claim 1 , wherein the embedding model comprises an embedding neural network.

11. The method of claim 10 , wherein the embedding neural network comprises an encoder neural network of an auto-encoder neural network.

12. The method of claim 11 , wherein the auto-encoder neural network is a variational auto-encoder (VAE) neural network.

13. The method of claim 12 , wherein the variational auto-encoder neural network is a β-variational auto-encoder (β-VAE) neural network.

14. The method of claim 12 , wherein processing the observation using the embedding model to generate the lower-dimensional embedding of the observation comprises:

processing the observation using the encoder neural network to generate parameters defining a probability distribution over a latent space; and

determining the lower-dimensional embedding of the observation based on a mean of the probability distribution over the latent space.

15. The method of claim 1 , wherein the observation comprises an image and the lower-dimensional embedding of the observation comprises respective coordinates for each of a plurality of key points in the image.

16. The method of claim 1 , wherein the observation comprises an image and the lower-dimensional embedding of the observation comprises a set of statistics characterizing a spatial color distribution in the image.

17. The method of claim 1 , further comprising:

determining a main task reward for the time step that corresponds to a main task being performed by the agent in the environment; and

determining the overall reward for the time step based on the auxiliary task reward for the time step and the main task reward for the time step.

18. The method of claim 17 , wherein the agent is a mechanical agent interacting with a real-world environment, and the main task being performed by the agent comprises physically manipulating objects in the environment.

19. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for training an action selection neural network having a plurality of parameters that is used to select actions to be performed by an agent interacting with an environment, wherein the action selection neural network is configured to process an input comprising an observation characterizing a state of the environment to generate an action selection output that comprises a respective action score for each action in a set of possible actions that can be performed by the agent, and select the action to be performed by the agent from the set of possible actions based on the action scores, the operations comprising:

obtaining an observation characterizing a state of the environment at a time step;

processing the observation using an embedding model to generate a lower-dimensional embedding of the observation, wherein the lower-dimensional embedding of the observation has a plurality of dimensions;

determining an auxiliary task reward for the time step based on a value of a particular dimension of the embedding, wherein the auxiliary task reward corresponds to an auxiliary task of controlling the value of the particular dimension of the embedding;

determining an overall reward for the time step based at least in part on the auxiliary task reward for the time step; and

determining an update to values of the plurality of parameters of the action selection neural network based on the overall reward for the time step using a reinforcement learning technique.

20. A system comprising:

one or more computers; and

one or more storage devices communicatively coupled to the one or more computers,

wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for training an action selection neural network having a plurality of parameters that is used to select actions to be performed by an agent interacting with an environment, wherein the action selection neural network is configured to process an input comprising an observation characterizing a state of the environment to generate an action selection output that comprises a respective action score for each action in a set of possible actions that can be performed by the agent, and select the action to be performed by the agent from the set of possible actions based on the action scores, the operations comprising:

obtaining an observation characterizing a state of the environment at a time step;

processing the observation using an embedding model to generate a lower-dimensional embedding of the observation, wherein the lower-dimensional embedding of the observation has a plurality of dimensions;

determining an auxiliary task reward for the time step based on a value of a particular dimension of the embedding, wherein the auxiliary task reward corresponds to an auxiliary task of controlling the value of the particular dimension of the embedding;

determining an overall reward for the time step based at least in part on the auxiliary task reward for the time step; and

determining an update to values of the plurality of parameters of the action selection neural network based on the overall reward for the time step using a reinforcement learning technique.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2023
From: WULFMEIER, MARKUS; HERTWECK, TIM; RIEDMILLER, MARTIN
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 063932/0514 →
Continuity (2)
Provisional Application 63057795 · Jul 28, 2020
Related Publication 20230290133A1 · Sep 14, 2023
References Cited (69)
JP 2018537880 · 2018 [cited by applicant]
JP 2019534517 · 2019 [cited by applicant]
WO WO2018224471 · 2018 [cited by applicant]
Gu et al., “Deep Reinforcement Learning for Robotic Manipulation” arXiv: 1610.00633, 2016, 9 pages (Year: 2016). [cited by examiner]
Minh et al., “Asynchronous Methods for Deep Reinforcement Learning” arXiv: 1602.01783v2, Jun. 16, 2016, 19 pages (Year: 2016). [cited by examiner]
Notice of Allowance in Japanese Appln. No. 2023-506026, dated Jun. 17, 2024, 5 pages (with English translation). [cited by applicant]
Abdolmaleki et al., “Relative entropy regularized policy iteration,” CoRR, submitted on Dec. 5, 2018, arXiv:1812.02256v1, 23 pages. [cited by applicant]
Andrychowicz et al., “Hindsight experience replay,” Advances in neural information processing systems, 2017, 11 pages. [cited by applicant]
Aytar et al., “Playing hard exploration games by watching youtube,” Advances in neural information processing systems, 2018, 12 pages. [cited by applicant]
Bengio et al., “Curriculum learning,” in Proceedings of the 26th annual international conference on machine learning, Jun. 14, 2009, p. 41-48. [cited by applicant]
Bingham et al., “Random projection in dimensionality reduction: applications to image and text data,” in Proceedings of the seventh ACM SIGKDD international conference on Knowledge discovery and data mining, Aug. 26, 20… [cited by applicant]
Burgess et al., “Monet:Unsupervised scene decomposition and representation,” CoRR, submitted on Jan. 22, 2019, arXiv:1901.11390v1, 22 pages. [cited by applicant]
Byravan et al., “Imagined value gradients: Model-based policy optimization with tranferable latent dynamics models,” in Conference on Robot Learning, May 12, 2020, 24 pages. [cited by applicant]
Cabi et al., “The intentional unintentional agent: Learning to solve many continuous control tasks simultaneously,” in Conference on Robot Learning, Oct. 18, 2017, 10 pages. [cited by applicant]
Chen et al., “A simple framework for contrastive learning of visual representations,” in International conference on machine learning, Nov. 21, 2020, p. 1597-1607. [cited by applicant]
Donahue et al., “Large scale adversarial representation learning,” Advances in neural information processing systems, 2019, 11 pages. [cited by applicant]
Dosovitskiy et al., “Learning to act by predicting the future,” CoRR, submitted on Feb. 14, 2017, arXiv:1611.01779v2, 14 pages. [cited by applicant]
Eysenbach et al., “Diversity is all you need: Learning skills without a reward function,” CoRR, submitted on Oct. 9, 2018, arXiv:1802.06070v6, 22 pages. [cited by applicant]
Florensa et al., “Reverse curriculum generation for reinforcement learning,” in Conference on robot learning, Oct. 18, 2017, p. 482-495. [cited by applicant]
Forestier et al., “Intrinsically motivated goal exploration processes with automatic curriculum learning,” The Journal of Machine Learning Research, Jan. 1, 2022, 23(1):6818-58. [cited by applicant]
Graves et al., “Automated curriculum learning for neural networks,” in international conference on machine learning, Jul. 17, 2017, p. 1311-1320. [cited by applicant]
Gregor et al., “Towards conceptual compression,” in NeurIPS, 2016, p. 3549-3557. [cited by applicant]
Gregor et al., “Variational intrinsic control,” CoRR, submitted on Nov. 22, 2016, arXiv:1611.07507v1, 15 pages. [cited by applicant]
Grill et al., “Bootstrap your own latent—a new approach to self-supervised learning,” Advances in neural information processing systems, 2020, 33:21271-84. [cited by applicant]
Grimm et al., “Disentangled cumulants help successor representations transfer to new tasks,” CoRR, submitted on Nov. 25, 2019, arXiv:1911.10866v1, 15 pages. [cited by applicant]
Hafner et al., “Dream to control: Learning behaviors by latent imagination,” CoRR, submitted on Mar. 17, 2020, arXiv:1912.01603v3, 20 pages. [cited by applicant]
Heess et al., “Emergence oflocomotion behaviours in rich environments,” CoRR, submitted on Jul. 10, 2017, arXiv:1707.02286v2, 14 pages. [cited by applicant]
Hertweck et al., “Simple Sensor Intentions for Exploration,” CoRR, submitted on May 15, 2020, arXiv:2005.07541v1, 13 pages. [cited by applicant]
Higgins et al., “beta-vae: Learning basic visual concepts with a constrained variational framework,” in International conference on learning representations, Nov. 4, 2016, 13 pages. [cited by applicant]
Higgins et al., “Darla: Improving zero-shot transfer in reinforcement learning,” in International Conference on Machine Learning, Jul. 17, 2017, p. 1480-1490. [cited by applicant]
Hinton et al., “Reducing the dimensionality of data with neural networks,” Science, Jul. 28, 2006, 313(5786):504-7. [cited by applicant]
International Preliminary Report on Patentability in Appln. No. PCT/EP2021/071078, mailed on Jan. 31, 2023, 10 pages. [cited by applicant]
International Search Report and Written Opinion in Appln. No. PCT/EP2021/071078, mailed on Nov. 19, 2021, 15 pages. [cited by applicant]
Jaderberg et al., “Reinforcement learning with unsupervised auxiliary tasks,” CoRR, submitted on Nov. 16, 2016, arXiv:1611.05397v1, 14 pages. [cited by applicant]
Jeong et al., “Self-supervised sim-to-real adaptation for visual robotic manipulation,” CoRR, submitted on Oct. 21, 2019, arXiv:1910.09470v1, 7 pages arXiv:1910.09470, 2019. [cited by applicant]
Kaiser et al., “Model-based reinforcement learning for atari,” CoRR, submitted on Feb. 19, 2020, arXiv: 1903.00374v4, 28 pages. [cited by applicant]
Kalashnikov et al., “Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation,” CoRR, submitted on Nov. 28, 2018, arXiv:1806.10293v3, 23 pages. [cited by applicant]
Kingma et al., “Auto-encoding variational bayes,” CoRR, submitted on Dec. 10, 2022, arXiv:1312.6114v11, 14 pages. [cited by applicant]
Klyubin et al., “All else being equal be empowered,” in European Conference on Artificial Life, 2005, pp. 744-753. [cited by applicant]
Kulkarni et al., “Unsupervised learning of object keypoints for perception and control,” in NeurIPS, 2019, p. 10723-10733. [cited by applicant]
Lange et al., “Autonomous reinforcement learning on raw visual input data in a real world application,” in the 2012 international joint conference on neural networks (IJCNN), Jun. 10, 2012, 8 pages. [cited by applicant]
Laskin et al., “Curl: Contrastive unsupervised representations for reinforcement learning,” in International Conference on Machine Learning, Nov. 21, 2020, p. 5639-5650. [cited by applicant]
Laversanne-Finot et al., “Curiosity driven exploration of learned disentangled goal spaces,” in Conference on Robot Learning, Oct. 23, 2018, p. 487-504. [cited by applicant]
Lee et al., “Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model,” Advances in Neural Information Processing Systems, 2020, 33:74-52. [cited by applicant]
Lillicrap et al., “Continuous control with deep reinforcement learning,” CoRR, submitted on Jul. 5, 2019, arXiv:1509.02971v6, 14 pages. [cited by applicant]
Locatello et al., “Challenging common assumptions in the unsupervised learning of disentangled representations,” in international conference on machine learning, May 24, 2019, p. 4114-4124. [cited by applicant]
Mirowski et al., “Learning to navigate in complex environments,” CoRR, submitted on Jan. 13, 2017, arXiv:1611.03673v3, 16 pages. [cited by applicant]
Mnih et al., “Playing atari with deep reinforcement learning,” CoRR, submitted on Dec. 19, 2013, arXiv:1312.5602v1, 9 pages. [cited by applicant]
Nachum et al., “Data-efficient hierarchical reinforcement learning,” Advances in neural information processing systems, 2018, 11 pages. [cited by applicant]
Nair et al., “Visual reinforcement learning with imagined goals,” in NeurIPS, 2018, p. 9191-9200. [cited by applicant]
Office Action in Japanese Appln. No. 2023506026, mailed on Jan. 9, 2024, 8 pages (with English translation). [cited by applicant]
Oord et al., “Representation learning with contrastive predictive coding,” CoRR, submitted on Jan. 22, 2019, arXiv:1807.03748v2, 13 pages. [cited by applicant]
Pathak et al., “Curiosity-driven exploration by self-supervised prediction,” in International conference on machine learning, Jul. 17, 2017, 10 pages. [cited by applicant]
Rezende et al., “Stochastic backpropagation and approximate inference in deep generative models,” in International conference on machine learning, Jun. 18, 2014, p. 1278-1286. [cited by applicant]
Riedmiller et al., “Learning by Playing—Solving Sparse Reward Tasks from Scratch,” CoRR, submitted on Feb. 28, 2018, arXiv:1802.10567v1, 18 pages. [cited by applicant]
Schmidhuber, “Powerplay: Training an increasingly general problem solver by continually searching for the simplest still unsolvable problem,” Frontiers in psychology, Jun. 7, 2013, 4:313, 14 pages. [cited by applicant]
Schwarzer et al., “Data-efficient reinforcement learning with momentum predictive representations,” CoRR, submitted on Jul. 12, 2020, arXiv:2007.05929v1, 14 pages. [cited by applicant]
Sharma et al., “Dynamics-aware unsupervised discovery of skills,” CoRR, submitted on Feb. 14, 2020, arXiv:1907.01657v2, 21 pages. [cited by applicant]
Smith et al., “Learning skills diverse via value-relevant features,” in Conference on Lifelong Learning Agents, Nov. 28, 2022, p. 1174-1194. [cited by applicant]
Tassa et al., “Deepmind control suite,” CoRR, submitted on Jan. 2, 2018, arXiv:1801.00690v1, 24 pages. [cited by applicant]
Teh et al., “Distral: Robust multitask reinforcement learning,” Advances in neural information processing systems, 2017, 11 pages. [cited by applicant]
Thomas et al., “Independently controllable features,” CoRR, submitted on Aug. 3, 2017, arXiv:1708.01289v1, 12 pages. [cited by applicant]
Todorov et al., “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ international conference on intelligent robots and systems, Oct. 7, 2012, p. 5026-5033. [cited by applicant]
Wang et al., “Paired open-ended trailblazer (poet): Endlessly generating increasingly complex and diverse learning environments and their solutions,” CoRR, submitted on Feb. 21, 2019, arXiv:1901.01753v3, 28 pages. [cited by applicant]
Watters et al., “Spatial broadcast decoder: A simple architecture for learning disentangled representations in vaes,” CoRR, submitted on Aug. 14, 2019, arXiv:1901.07017v2, 35 pages. [cited by applicant]
Wulfmeier et al., “Regularized hierarchical policies for compositional transfer in robotics,” CoRR, submitted on Jun. 27, 2019, 32 pages. [cited by applicant]
Yu et al., “Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,” in Conference on robot learning, May 12, 2020, 17 pages. [cited by applicant]
Li, “Deep Reinforcement Learning,” cs.LG, Oct. 15, 2018, arXiv:1810.06339v1, 150 pages. [cited by applicant]
Office Action in European Appln. No. 21751552.7, mailed on Jul. 29, 2025, 10 pages. [cited by applicant]