Controlling robots using latent action vector conditioned controller neural networks
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for controlling agents. In particular, an agent can be controlled using a hierarchical controller that includes a task policy neural network and a low-level controller neural network.
1 . A method for controlling an agent interacting with an environment to perform a task, the method comprising, at each of a plurality of time steps:
receiving an observation comprising data characterizing a state of the environment at the time step, wherein the data characterizing the state of the environment comprises sensor data generated from sensor readings of sensors of the agent at the time step;
processing the observation using a task policy neural network for the task to generate a task output that defines a latent action vector from a latent action space;
processing a low-level input comprising (i) the sensor data and (ii) the latent action vector defined by the task output using a low-level controller neural network to generate a policy output that defines a control input for controlling the agent in response to the observation, wherein the low-level controller neural network is configured to:
process the sensor data through a first neural network branch comprising a plurality of first neural network layers to generate a first branch output;
process a second branch input comprising the first branch output and the latent action vector defined by the task output through a second neural network branch comprising a plurality of second neural network layers to generate a second branch output; and
generate the policy output from the first branch output and the second branch output; and
controlling the agent using the control input defined by the policy output.
2 . The method of claim 1 , wherein the agent is a robot and wherein the environment is a real-world environment.
3 . The method of claim 2 , wherein the task policy neural network has been trained through reinforcement learning to control a simulated agent to perform the task in a computer simulation of the real-world environment.
4 . The method of claim 3 , wherein the low-level controller neural network is pre-trained prior to training the task policy neural network through reinforcement learning and is held fixed during the training of the task policy neural network through reinforcement learning.
5 . The method of claim 3 , wherein the task policy neural network has been trained jointly with a value neural network through an actor-critic reinforcement learning technique, and wherein the value neural network is configured to:
receive a value input that includes additional information characterizing an input state of the computer simulation of the real-world environment that is not provided to the task policy neural network or the low-level controller neural network, and
process the value input to generate a value output that estimates a value of the input state of the environment to performing the task.
6 . The method of claim 5 , wherein the additional information comprises one or more of:
(i) data characterizing one or more future states of the computer simulation of the environment, or
(ii) ground truth state data obtained from the computer simulation of the environment.
7 . The method of claim 1 , wherein the first neural network branch comprises one or more recurrent neural network layers and the second neural network branch comprises only feedforward neural network layers.
8 . The method of claim 1 , wherein processing the sensor data through a first neural network branch comprising a plurality of first neural network layers to generate a first branch output comprises:
applying a normalization to the sensor data to generate normalized sensor data; and
processing the normalized sensor data using the first neural network branch to generate the first branch output.
9 . The method of claim 8 , wherein the second branch input further comprises the normalized sensor data.
10 . The method of claim 1 , wherein generating the policy output from the first branch output and the second branch output comprises:
computing a linear combination of the first branch output and the second branch output.
11 . The method of claim 1 , wherein the method further comprises:
generating, from the task output, parameters of a probability distribution over the latent action space; and
selecting, as the latent action defined by the task output, a latent action from the latent action space using the probability distribution.
12 . The method of claim 11 , wherein the task output includes (i) a mean of a multi-variate Gaussian distribution over the latent action space and (ii) a covariance matrix of the multi-variate Gaussian distribution over the latent action space.
13 . The method of claim 12 , wherein the task output includes (iii) a filtering value, and wherein generating the parameters of the probability distribution comprises:
applying the filtering value to the mean in the task output to generate a mean of the probability distribution.
14 . The method of claim 13 , wherein applying the filtering value to the mean in the task output to generate a mean of the probability distribution comprises:
computing a product between the filtering value and the mean.
15 . The method of claim 13 , wherein the low-level controller neural network is pre-trained prior to training the task policy neural network through reinforcement learning and is held fixed during the training of the task policy neural network through reinforcement learning, and wherein applying the filtering value to the mean in the task output to generate a mean of the probability distribution comprises:
clipping the mean included in the task output based on a range of latent actions provided as input to the low-level controller neural network during the pre-training of the low-level controller neural network; and
computing a product between the filtering value and the clipped mean.
16 . The method of claim 13 , wherein:
an objective for the training of the task policy neural network includes a regularization term that penalizes the task policy neural network for generating task outputs that specify multi-variate Gaussian distributions that diverge from an AR(1) prior distribution over the latent action space having a scaling factor.
17 . The method of claim 13 , wherein:
for the training of the task policy neural network, the task policy neural network is initialized to generate filtering values that equal the scaling factor.
18 . The method of claim 1 , wherein the observation further comprises task data characterizing the task.
19 . The method of claim 18 , wherein the task data comprises one or more of:
data characterizing a target state of the agent for completing the task,
data characterizing a target position of one or more objects in the environment for completing the task; or
data characterizing one or more target locations in the environment to be reached for completing the task.
20 . A system comprising:
one or more computers; and
one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for controlling an agent interacting with an environment to perform a task, the operations comprising, at each of a plurality of time steps:
receiving an observation comprising data characterizing a state of the environment at the time step, wherein the data characterizing the state of the environment comprises sensor data generated from sensor readings of sensors of the agent at the time step;
processing the observation using a task policy neural network for the task to generate a task output that defines a latent action vector from a latent action space;
processing a low-level input comprising (i) the sensor data and (ii) the latent action vector defined by the task output using a low-level controller neural network to generate a policy output that defines a control input for controlling the agent in response to the observation, wherein the low-level controller neural network is configured to:
process the sensor data through a first neural network branch comprising a plurality of first neural network layers to generate a first branch output;
process a second branch input comprising the first branch output and the latent action vector defined by the task output through a second neural network branch comprising a plurality of second neural network layers to generate a second branch output; and
generate the policy output from the first branch output and the second branch output; and
controlling the agent using the control input defined by the policy output.