Learning robust legged robot locomotion with implicit terrain imagination via deep reinforcement learning
Disclosed is technology for controlling deep reinforcement learning-based legged robot locomotion by inferring implicit terrain information. A legged robot control method may include inferring an action of a quadrupedal robot from proprioception through a deep reinforcement learning-legged robot model, and a locomotion policy that implicitly infers properties of terrains through which the quadrupedal robot moves may be learned in the legged robot model.
1 . A legged robot control method performed by a computer device,
wherein the computer device comprises at least one processor configured to execute computer-readable instructions included in a memory,
the legged robot control method comprises inferring, by the at least one processor, an action of a quadrupedal robot from proprioception through a deep reinforcement learning-legged robot model, and
a locomotion policy that implicitly infers properties of terrains through which the quadrupedal robot moves is learned in the legged robot model, wherein (a) a context-aided estimator that estimates surrounding environmental information during a learning process of the locomotion policy is jointly learned in the legged robot model, (b) the legged robot model is a neural network that infers the action when a proprioceptive observation, a body velocity, and a latent state are given as a policy network configured as an actor network in an asymmetric actor-critic network, and (c) the context-aided estimator is optimized using a hybrid loss function that includes body velocity estimation loss and variational auto-encoder (VAE) loss.
2 . The legged robot control method of claim 1 , wherein a locomotion policy that enables a blind locomotion of the quadrupedal robot using an asymmetric actor-critic architecture is learned in the legged robot model.
3 . The legged robot control method of claim 1 , wherein the policy network is trained with an interplay with a value network configured as a critic network in the asymmetric actor-critic network, and the value network is trained using a disturbance force randomly applied to a robot's body and height information of the robot's surrounding environment.
4 . The legged robot control method of claim 1 , wherein the proprioceptive observation is measured using a joint encoder and an inertial measurement unit (IMU), and the body velocity and the latent state are estimated using the context-aided estimator.
5 . The legged robot control method of claim 1 , wherein the proprioceptive observation includes at least one of a body angular velocity, a gravity vector in a body frame, a body velocity command, a joint angle, a joint angular velocity, and a previous action.
6 . The legged robot control method of claim 1 , wherein the policy network is trained to infer a joint angle around a robot's stand still pose.
7 . The legged robot control method of claim 1 , wherein the context-aided estimator includes a body velocity estimation model and an auto-encoder model that shares a unified encoder.
8 . The legged robot control method of claim 1 , wherein the context-aided estimator includes a single encoder and a multi-head decoder and encodes the proprioceptive observation into the body velocity and the latent state through the encoder.
9 . The legged robot control method of claim 1 , wherein a power distribution reward for a motor used on the robot is included in a reward function to train the policy network.
10 . A legged robot control method performed by a computer device, wherein the computer device comprises at least one processor configured to execute computer-readable instructions included in a memory, the legged robot control method comprises inferring, by the at least one processor, an action of a quadrupedal robot from proprioception through a deep reinforcement learning-legged robot model, and a locomotion policy that implicitly infers properties of terrains through which the quadrupedal robot moves is learned in the legged robot model, wherein (a) a context-aided estimator that estimates surrounding environmental information during a learning process of the locomotion policy is jointly learned in the legged robot model, (b) the legged robot model is a neural network that infers the action when a proprioceptive observation, a body velocity, and a latent state are given as a policy network configured as an actor network in an asymmetric actor-critic network, and (c) the context-aided estimator is optimized using a hybrid loss function that includes body velocity estimation loss and variational auto-encoder (VAE) loss, and (d) adaptive bootstrapping for adaptively tuning a bootstrapping probability is performed according to a reward coefficient of variation by the context-aided estimator during training of the policy network.
11 . A non-transitory computer-readable recording medium storing instructions that, when executed by a processor, cause the processor to perform a legged robot control method comprising inferring an action of a quadrupedal robot from proprioception through a deep reinforcement learning-legged robot model, wherein a locomotion policy that implicitly infers properties of terrains through which the quadrupedal robot moves is learned in the legged robot model, wherein (a) a context-aided estimator that estimates surrounding environmental information during a learning process of the locomotion policy is jointly learned in the legged robot model, (b) the legged robot model is a neural network that infers the action when a proprioceptive observation, a body velocity, and a latent state are given as a policy network configured as an actor network in an asymmetric actor-critic network, and (c) the context-aided estimator is optimized using a hybrid loss function that includes body velocity estimation loss and variational auto-encoder (VAE) loss.
12 . A computer-implemented legged robot control system comprising:
at least one processor configured to execute computer-readable instructions included in a memory,
wherein the at least one processor is configured to process a process of inferring an action of a quadrupedal robot from proprioception through a deep reinforcement learning-legged robot model, and
a locomotion policy that implicitly infers properties of terrains through which the quadrupedal robot moves is learned in the legged robot model, wherein (a) a context-aided estimator that estimates surrounding environmental information during a learning process of the locomotion policy is jointly learned in the legged robot model, (b) the legged robot model is a neural network that infers the action when a proprioceptive observation, a body velocity, and a latent state are given as a policy network configured as an actor network in an asymmetric actor-critic network, and (c) the context-aided estimator is optimized using a hybrid loss function that includes body velocity estimation loss and variational auto-encoder (VAE) loss.
13 . The legged robot control system of claim 12 , wherein:
the policy network is trained with an interplay with a value network configured as a critic network in the asymmetric actor-critic network, and
the value network is trained using a disturbance force randomly applied to a robot's body and height information of the robot's surrounding environment.
14 . The legged robot control system of claim 13 , wherein:
the proprioceptive observation is measured using a joint encoder and an inertial measurement unit (IMU),
the body velocity and the latent state are estimated using the context-aided estimator, and
the context-aided estimator includes a single encoder and a multi-head decoder and encodes the proprioceptive observation into the body velocity and the latent state through the encoder.
15 . The legged robot control system of claim 13 , wherein a power distribution reward for a motor used on the robot is included in a reward function to train the policy network.
16 . The legged robot control system of claim 13 , wherein adaptive bootstrapping for adaptively tuning a bootstrapping probability is performed according to a reward coefficient of variation by the context-aided estimator during training of the policy network.