IP Library › Granted Patent US 12,748,440
Granted Patent B2
US 12,748,440 · App. 18/850,857 · Granted Sep 29, 2026

Controlling robots using latent action vector conditioned controller neural networks

Inventors: Steven Bohez (London, GB); Saran Tunyasuvunakool (London, GB)
Assignee: GDM Holding LLC
G05D1/648G06N3/044G06N3/0499G06N3/092G05D2101/15G05D2109/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,440
App. No.
18/850,857
Filed
Sep 25, 2024
Granted
Sep 29, 2026
Kind
B2
Art Unit
3665
USPC
701/23
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for controlling agents. In particular, an agent can be controlled using a hierarchical controller that includes a task policy neural network and a low-level controller neural network.

Claims (54)

1 . A method for controlling an agent interacting with an environment to perform a task, the method comprising, at each of a plurality of time steps:

receiving an observation comprising data characterizing a state of the environment at the time step, wherein the data characterizing the state of the environment comprises sensor data generated from sensor readings of sensors of the agent at the time step;

processing the observation using a task policy neural network for the task to generate a task output that defines a latent action vector from a latent action space;

processing a low-level input comprising (i) the sensor data and (ii) the latent action vector defined by the task output using a low-level controller neural network to generate a policy output that defines a control input for controlling the agent in response to the observation, wherein the low-level controller neural network is configured to:

process the sensor data through a first neural network branch comprising a plurality of first neural network layers to generate a first branch output;

process a second branch input comprising the first branch output and the latent action vector defined by the task output through a second neural network branch comprising a plurality of second neural network layers to generate a second branch output; and

generate the policy output from the first branch output and the second branch output; and

controlling the agent using the control input defined by the policy output.

2 . The method of claim 1 , wherein the agent is a robot and wherein the environment is a real-world environment.

3 . The method of claim 2 , wherein the task policy neural network has been trained through reinforcement learning to control a simulated agent to perform the task in a computer simulation of the real-world environment.

4 . The method of claim 3 , wherein the low-level controller neural network is pre-trained prior to training the task policy neural network through reinforcement learning and is held fixed during the training of the task policy neural network through reinforcement learning.

5 . The method of claim 3 , wherein the task policy neural network has been trained jointly with a value neural network through an actor-critic reinforcement learning technique, and wherein the value neural network is configured to:

receive a value input that includes additional information characterizing an input state of the computer simulation of the real-world environment that is not provided to the task policy neural network or the low-level controller neural network, and

process the value input to generate a value output that estimates a value of the input state of the environment to performing the task.

6 . The method of claim 5 , wherein the additional information comprises one or more of:

(i) data characterizing one or more future states of the computer simulation of the environment, or

(ii) ground truth state data obtained from the computer simulation of the environment.

7 . The method of claim 1 , wherein the first neural network branch comprises one or more recurrent neural network layers and the second neural network branch comprises only feedforward neural network layers.

8 . The method of claim 1 , wherein processing the sensor data through a first neural network branch comprising a plurality of first neural network layers to generate a first branch output comprises:

applying a normalization to the sensor data to generate normalized sensor data; and

processing the normalized sensor data using the first neural network branch to generate the first branch output.

9 . The method of claim 8 , wherein the second branch input further comprises the normalized sensor data.

10 . The method of claim 1 , wherein generating the policy output from the first branch output and the second branch output comprises:

computing a linear combination of the first branch output and the second branch output.

11 . The method of claim 1 , wherein the method further comprises:

generating, from the task output, parameters of a probability distribution over the latent action space; and

selecting, as the latent action defined by the task output, a latent action from the latent action space using the probability distribution.

12 . The method of claim 11 , wherein the task output includes (i) a mean of a multi-variate Gaussian distribution over the latent action space and (ii) a covariance matrix of the multi-variate Gaussian distribution over the latent action space.

13 . The method of claim 12 , wherein the task output includes (iii) a filtering value, and wherein generating the parameters of the probability distribution comprises:

applying the filtering value to the mean in the task output to generate a mean of the probability distribution.

14 . The method of claim 13 , wherein applying the filtering value to the mean in the task output to generate a mean of the probability distribution comprises:

computing a product between the filtering value and the mean.

15 . The method of claim 13 , wherein the low-level controller neural network is pre-trained prior to training the task policy neural network through reinforcement learning and is held fixed during the training of the task policy neural network through reinforcement learning, and wherein applying the filtering value to the mean in the task output to generate a mean of the probability distribution comprises:

clipping the mean included in the task output based on a range of latent actions provided as input to the low-level controller neural network during the pre-training of the low-level controller neural network; and

computing a product between the filtering value and the clipped mean.

16 . The method of claim 13 , wherein:

an objective for the training of the task policy neural network includes a regularization term that penalizes the task policy neural network for generating task outputs that specify multi-variate Gaussian distributions that diverge from an AR(1) prior distribution over the latent action space having a scaling factor.

17 . The method of claim 13 , wherein:

for the training of the task policy neural network, the task policy neural network is initialized to generate filtering values that equal the scaling factor.

18 . The method of claim 1 , wherein the observation further comprises task data characterizing the task.

19 . The method of claim 18 , wherein the task data comprises one or more of:

data characterizing a target state of the agent for completing the task,

data characterizing a target position of one or more objects in the environment for completing the task; or

data characterizing one or more target locations in the environment to be reached for completing the task.

20 . A system comprising:

one or more computers; and

one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for controlling an agent interacting with an environment to perform a task, the operations comprising, at each of a plurality of time steps:

receiving an observation comprising data characterizing a state of the environment at the time step, wherein the data characterizing the state of the environment comprises sensor data generated from sensor readings of sensors of the agent at the time step;

processing the observation using a task policy neural network for the task to generate a task output that defines a latent action vector from a latent action space;

processing a low-level input comprising (i) the sensor data and (ii) the latent action vector defined by the task output using a low-level controller neural network to generate a policy output that defines a control input for controlling the agent in response to the observation, wherein the low-level controller neural network is configured to:

process the sensor data through a first neural network branch comprising a plurality of first neural network layers to generate a first branch output;

process a second branch input comprising the first branch output and the latent action vector defined by the task output through a second neural network branch comprising a plurality of second neural network layers to generate a second branch output; and

generate the policy output from the first branch output and the second branch output; and

controlling the agent using the control input defined by the policy output.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2025
From: BOHEZ, STEVEN; TUNYASUVUNAKOOL, SARAN
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 071184/0355 →
Continuity (2)
Provisional Application 63323992 · Mar 25, 2022
Related Publication 20250224737A1 · Jul 10, 2025
References Cited (65)
US 12585941B2 · Anthony · 2026 [cited by examiner]
US 20200104685A1 · Hasenclever · 2020 [cited by examiner]
US 20220076099A1 · Sermanet · 2022 [cited by examiner]
US 20230244906A1 · Becnel · 2023 [cited by examiner]
Abdolmaleki et al., A distributional view on multi-objective policy optimization, in Proceedings of the 37th International Conference on Machine Learning (ICML), Nov. 21, 2020, 12 pages. [cited by applicant]
Alemi et al.,“Deep variational information bottleneck,” CoRR, submitted on Dec. 1, 2016, arXiv:1612.00410v7, 19 pages. [cited by applicant]
Andrychowicz et al., “Learning dexterous in-hand manipulation,” The International Journal of Robotics Research, Jan. 2020, 39(1):3-20. [cited by applicant]
Apgar et al., “Fast online trajectory optimization for the bipedal robot cassie,” In Robotics: Science and Systems, Jun. 26, 2018, p. 14. [cited by applicant]
Bellicoso et al., “Dynamic locomotion through online nonlinear motion optimization for quadrupedal robots,” IEEE Robotics and Automation Letters, Jan. 17, 2018, 3(3):2261-2268. [cited by applicant]
Bloesch et al., “Towards real robot learning in the wild: a case study in bipedal locomotion,” In 5th Annual Conference on Robot Learning, Jan. 11, 2022, p. 1502-1511. [cited by applicant]
Bohez et al., “Value constrained model-free continuous control,” CoRR, submitted on Feb. 12, 2019, arXiv:1902.04623v1, 12 pages. [cited by applicant]
Brakel et al., “Learning coordinated terrain-adaptive locomotion by imitating a centroidal dynamics planner,” CoRR, submitted on Oct. 23, 2022, arXiv:2111.00262v1, 10 pages. [cited by applicant]
Carius et al., “Trajectory optimization with implicit hard contacts,” IEEE Robotics and Automation Letters, Jul. 4, 2018, 3(4):3316-3323. [cited by applicant]
Chentanez et al., “Physics-based motion capture imitation with deep reinforcement learning,” In Proceedings of the 11th annual international conference on motion, interaction, and games, Nov. 8, 2018, p. 1-10. [cited by applicant]
Galashov et al., “Information asymmetry in KL-regularized RL,” CoRR, submitted on May 3, 2019, arXiv:1905.01240v1, 25 pages. [cited by applicant]
Gangapurwala et al., “Real-time trajectory adaptation for quadrupedal locomotion using deep reinforcement learning,” 2021 IEEE International Conference on Robotics and Automation (ICRA), May 30, 2021, p. 5973-5979. [cited by applicant]
Gleicher, “Retargetting motion to new characters,” In Proceedings of the 25th Annual Conference on Computer Graphics and Interactive Techniques, Apr. 27, 1998, 10 pages. [cited by applicant]
Goyal et al., “Transfer and exploration via the information bottleneck,” CoRR, submitted on Jan. 30, 2019, arXiv:1901.10902v5, 20 pages. [cited by applicant]
Haarnoja et al., “Learning to walk via deep reinforcement learning,” CoRR, submitted on Dec. 26, 2018, arXiv:1812.11103v3, 10 pages. [cited by applicant]
Hafner et al., “Towards general and autonomous learning of core skills: a case study in locomotion,” In Conference on Robot Learning, Oct. 4, 2021, p. 1084-1099. [cited by applicant]
Hasenclever et al.,“Comic: Complementary task learning & mimicry for reusable skills,” In International Conference on Machine Learning, Nov. 21, 2020, p. 4105-4115. [cited by applicant]
Heess et al., “Emergence of locomotion behaviours in rich environments,” CoRR, submitted on Jul. 7, 2017, arXiv:1707.02286v2, 14 pages. [cited by applicant]
Hochreiter et al., “Long Short-Term Memory,” Neural Computation, 1997, 9(8):1735-1780. [cited by applicant]
Holden et al., “Phase-functioned neural networks for character control,” ACM Transactions on Graphics (TOG), Jul. 20, 2017, 36(4):1-13. [cited by applicant]
Hutter et al., “Anymal-a highly mobile and dynamic quadrupedal robot,” In 2016 IEEE/RSJ international conference on intelligent robots and systems (IROS), Oct. 9, 2016, p. 38-44. [cited by applicant]
Hwangbo et al., “Learning agile and dynamic motor skills for legged robots,” CoRR, submitted on Jan. 24, 2019, arXiv:1901.08652v1, 20 pages. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/EP2023/057855, mailed on Oct. 10, 2024, 18 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/EP2023/057855, mailed on Aug. 11, 2023, 25 pages. [cited by applicant]
Invitation to Pay Additional Fees in International Appln. No. PCT/EP2023/057855, mailed on Jun. 21, 2023, 18 pages. [cited by applicant]
Kalakrishnan et al., “Fast, robust quadruped locomotion over challenging terrain,” In 2010 IEEE International Conference on Robotics and Automation, May 3, 2010, p. 2665-2670. [cited by applicant]
Lee et al., “Learning quadrupedal locomotion over challenging terrain,” CoRR, submitted on Oct. 21, 2020, arXiv:2010.11251v1, 22 pages. [cited by applicant]
Li et al., “Reinforcement learning for robust parameterized locomotion control of bipedal robots,” CoRR, submitted on Mar. 26, 2021, arXiv:2103.14295v1, 7 pages. [cited by applicant]
Liu et al., “From motor control to team play in simulated humanoid football,” CoRR, submitted on May 25, 2021, arXiv:2105.12196v1, 57 pages. [cited by applicant]
Merel et al., “Catch & carry: reusable neural controllers for vision-guided whole-body tasks,” ACM Transactions on Graphics (TOG), Jul. 8, 2020, 39(4):39-1. [cited by applicant]
Merel et al., “Learning human behaviors from motion capture by adversarial imitation,” CoRR, submitted on Jul. 7, 2017, arXiv:1707.02201v2, 12 pages. [cited by applicant]
Merel et al., “Neural probabilistic motor primitives for humanoid control,” CoRR, submitted on Jan. 15, 2019, arXiv:1811.11711v2, 14 pages. [cited by applicant]
Miki et al., “Learning robust perceptive locomotion for quadrupedal robots in the wild,” Science Robotics, 7(62):eabk2822. [cited by applicant]
Mocap.cs.cmu.edu [online], “CMU Graphics Lab Motion Capture Database,” available on or before Mar. 21, 2003, via Internet Archive: Wayback Machine URL <https://web.archive.org/web/20250000000000*/http://mocap.cs.cmu.edu… [cited by applicant]
Neunert et al., “Trajectory optimization through contacts and automatic gait discovery for quadrupeds,” CoRR, submitted on Jul. 15, 2016, arXiv:1607.04537v1, 11 pages. [cited by applicant]
Oord et al., “Wavenet: a generative model for raw audio,” CoRR, submitted on Sep. 12, 2016, arXiv:1609.03499v2, 15 pages. [cited by applicant]
Peng et al., “Deepmimic: Example-guided deep reinforcement learning of physics-based character skills,” ACM Transactions on Graphics, Jul. 30, 2018, 37(4):1-4. [cited by applicant]
Peng et al., “Mcp: Learning composable hierarchical control with multiplicative compositional policies,” CoRR, submitted on May 23, 2019, arXiv:1905.09808v1, 18 pages. [cited by applicant]
Peng et al., “Simto-real transfer of robotic control with dynamics randomization,” CoRR, submitted on Oct. 18, 2017, arXiv:1710.06537v3, 8 pages. [cited by applicant]
Peng et al.,“Learning agile robotic locomotion skills by imitating animals,” CoRR, submitted on Apr. 2, 2020, arXiv:2004.00784v3, 14 pages. [cited by applicant]
Peng et al., “Amp: Adversarial motion priors for stylized physics-based character control,” ACM Transactions on Graphics (ToG), Jul. 19, 2021, 40(4):1-20. [cited by applicant]
Perlin et al., “An image synthesizer,” In Proceedings of the 12th Annual Conference on Computer Graphics and Interactive Techniques, Jul. 1, 1985, p. 287-296. [cited by applicant]
Pollard et al., “Adapting human motion for the control of a humanoid robot,” In Proceedings 2002 IEEE international conference on robotics and automation, May 11, 2002, 1390-1397. [cited by applicant]
Robotis.com [online], “Open Platform Humanoid Project,” available on or before Feb. 11, 2018, via Internet Archive: Wayback Machine URL<https://web.archive.org/web/20250000000000*/https://emanual.robotis.com/docs/en/pla… [cited by applicant]
Sadeghi et al., “CAD2RL: Real single-image flight without a single real image,” CoRR, submitted on Nov. 13, 2016, arXiv:1611.04201v4, 12 pages. [cited by applicant]
Safonova et al., “Construction and optimal search of interpolated motion graphs,” In ACM SIGGRAPH 2007 papers, Jul. 29, 2007, p. 106-es. [cited by applicant]
Siekmann et al., “Blind bipedal stair traversal via sim-to-real reinforcement learning,” CoRR, submitted on May 18, 2021, arXiv:2105.08328v1, 9 pages. [cited by applicant]
Song et al., “V-mpo: on-policy maximum a posteriori policy optimization for discrete and continuous control,” CoRR, submitted on Sep. 26, 2019, arXiv:1909.12238v1, 19 pages. [cited by applicant]
Tan et al., “Sim-to-real: Learning agile locomotion for quadruped robots,” CoRR, submitted on Apr. 27, 2018, arXiv:1804.10332v2, 11 pages. [cited by applicant]
Tassa et al., “dm control: Software and tasks for continuous control,” Software Impacts, Nov. 1, 2020, 6:100022. [cited by applicant]
Tirumala et al., “Behavior priors for efficient reinforcement learning,” Journal of Machine Learning Research, 2022, 23(221):1-68. [cited by applicant]
Tishby et al., “The information bottleneck method,” CoRR, submitted on Apr. 24, 2000, arXiv:physics/0004057v1, 16 pages. [cited by applicant]
Tobin et al., “Domain randomization for transferring deep neural networks from simulation to the real world,” CoRR, submitted on Mar. 20, 2017, arXiv:1703.06907v1, 8 pages. [cited by applicant]
Todorov et al., “Mujoco: a physics engine for model-based control,” In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, p. 5026-5033. [cited by applicant]
Wang et al., “Robust imitation of diverse behaviors,” CoRR, submitted on Jul. 10, 2017, arXiv:1707.02747v2, 12 pages. [cited by applicant]
Xie et al., “Dynamics randomization revisited: a case study for quadrupedal locomotion,” CoRR, submitted on Nov. 4, 2020, arXiv:2011.02404v3, 7 pages. [cited by applicant]
Xie et al., “Learning locomotion skills for cassie: Iterative design and sim-to-real,” In Conference on Robot Learning, May 12, 2020, p. 317-329. [cited by applicant]
Yang et al., “Multi-expert learning of adaptive legged locomotion,” Science Robotics, Dec. 9, 2020, 5(49): eabb2174. [cited by applicant]
Yu et al., “Sim-to-real transfer for biped locomotion,” CoRR, submitted on Mar. 4, 2019, arXiv:1903.01390v2, 8 pages. [cited by applicant]
Zhang et al., “Mode-adaptive neural networks for quadruped motion control,” ACM Transaction on Graphics, Jul. 30, 2018, 37(4):1-4. [cited by applicant]
Office Action in European Appln. No. 23714736.8, Apr. 1, 2026, 8 pages. [cited by applicant]