IP Library › Granted Patent US 11,904,467
Granted Patent B2
US 11,904,467 · App. 17/056,104 · Granted Feb 20, 2024

System and methods for pixel based model predictive control

Inventor: Danijar Hafner (Toronto, CA)
Assignee: GOOGLE LLC
B25J9/161B25J9/163B25J9/1661G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,904,467
App. No.
17/056,104
Granted
Feb 20, 2024
Kind
B2
Abstract

Techniques are disclosed that enable model predictive control of a robot based on a latent dynamics model and a reward function. In many implementations, the latent space can be divided into a deterministic portion and stochastic portion, allowing the model to be utilized in generating more likely robot trajectories. Additional or alternative implementations include many reward functions, where each reward function corresponds to a different robot task.

Claims (44)

1. A method implemented by one or more processors, comprising:

training a latent robot dynamics model using unsupervised robot trajectories, wherein each of the unsupervised robot trajectories includes a corresponding sequence of:

partial robotic observations, each of the partial robotic observations being for a corresponding time step of the sequence, and

robotic actions, each of the robotic actions being for a corresponding time step of the sequence;

identifying supervised robot task trajectories for a robot task, wherein each of the supervised robot task trajectories includes a corresponding sequence of:

partial task robotic observations during a corresponding performance of the robot task,

task robotic actions during the corresponding performance of the robot task, and

labeled task rewards for the corresponding performance of the robot task;

training a reward function for the robot task using the supervised robot task trajectories;

controlling a robot to perform the robot task, wherein controlling the robot to perform the robot task comprises:

determining a sequence of actions for the robot using both the trained robot latent dynamics model and the trained reward function for the robot task; and

controlling the robot by implementing the sequence of actions.

2. The method of claim 1 , wherein the partial robotic observations of the unsupervised robot trajectories, and the partial task robotic observations of the supervised robot task trajectories, are each a corresponding image that captures a corresponding robot.

3. The method of claim 1 , wherein the latent robot dynamics model is a deterministic belief state model (DBSM).

4. The method of claim 3 , wherein the DBSM comprises an encoder network, a transition function, a posterior function, and a decoder network.

5. The method of claim 4 , wherein training the DBSM comprises training the transition function using latent overshooting.

6. The method of claim 5 , wherein the latent overshooting comprises performing a fixed number of open-loop predictions from a corresponding posterior at every time step.

7. The method of claim 6 , wherein the latent overshooting further comprises determining a Kullback-Leibler divergence between the open-loop predictions and the corresponding posterior.

8. The method of claim 3 , wherein training the DBSM comprises training the encoder network to deterministically update a deterministic activation vector at every time step.

9. The method of claim 8 , wherein determining a sequence of actions for the robot using both the trained robot latent dynamics model and the trained reward function for the robot task comprises using model predictive control in view of the trained robot latent dynamics model and the trained reward function.

10. The method of claim 9 , wherein training the reward function comprises training the reward function based on a first quantity of the supervised robot task trajectories, wherein the first quantity is less than a second quantity of the unsupervised robot trajectories on which the latent robot dynamics model is trained.

11. The method of claim 10 , wherein the first quantity is less than one percent of the second quantity.

12. The method of claim 10 , wherein the first quantity is less than fifty.

13. The method of claim 12 , wherein the first quantity is less than twenty-five.

14. The method of claim 13 , further comprising:

identifying second supervised robot task trajectories for a second robot task, wherein each of the second supervised robot task trajectories includes a corresponding sequence of:

second partial task robotic observations during a corresponding performance of the second robot task,

second task robotic actions during the corresponding performance of the second robot task, and

second labeled task rewards for the corresponding performance of the second robot task;

training a second reward function for the second robot task using the supervised robot task trajectories;

controlling a robot to perform the second robot task, wherein controlling the robot to perform the second robot task comprises:

determining a second sequence of actions for the robot using both the trained robot latent dynamics model and the trained second reward function for the second robot task; and

controlling the robot by implementing the second sequence of actions.

15. A method implemented by one or more processors, comprising:

identifying unsupervised robot trajectories, wherein each of the unsupervised robot trajectories includes:

observation images, each of the observation images capturing a corresponding robot and being for a corresponding time step of the sequence, and

robotic actions, each of the robotic actions being for a corresponding time step of the sequence;

training a latent robot dynamics model using the unsupervised robot trajectories;

using the trained latent robot dynamics model in latent planning for one or more robotic control tasks through generation, at each of a plurality of time steps and using the trained robot dynamics model, of a corresponding deterministic activation vector.

16. The method of claim 15 , wherein the latent robot dynamic model comprises an encoder network, a transition function, a posterior function, and a decoder network.

17. The method of claim 16 , wherein training the latent robot dynamics model comprises training the transition function using latent overshooting.

18. The method of claim 17 , wherein latent overshooting comprises performing a fixed number of open-loop predictions from a corresponding posterior at every time step.

19. The method of claim 18 , wherein latent overshooting further comprises determining a Kullback-Leibler divergence between the open-loop predictions and the corresponding posterior.

20. The method of claim 16 , wherein training the DBSM comprises training the encoder network to deterministically update the deterministic activation vector at every time step.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2021
From: HAFNER, DANIJAR
To: GOOGLE LLC
Reel/Frame 055425/0051 →
Continuity (2)
Provisional Application 62673744 · May 18, 2018
Related Publication 20210205984A1 · Jul 8, 2021
Cited By (1)
US 12,715,114