IP Library › Granted Patent US 12,569,984
Granted Patent B2
US 12,569,984 · App. 18/436,684 · Granted Mar 10, 2026

System and methods for pixel based model predictive control

Inventor: Danijar Hafner (Toronto, CA)
Assignee: GOOGLE LLC
B25J9/161B25J9/163B25J9/1661G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,569,984
App. No.
18/436,684
Granted
Mar 10, 2026
Kind
B2
Abstract

Techniques are disclosed that enable model predictive control of a robot based on a latent dynamics model and a reward function. In many implementations, the latent space can be divided into a deterministic portion and stochastic portion, allowing the model to be utilized in generating more likely robot trajectories. Additional or alternative implementations include many reward functions, where each reward function corresponds to a different robot task.

Claims (53)

1 . A robot comprising:

one or more actuators;

memory storing a trained latent robot dynamics model and a trained reward function for a robot task;

one or more processors configured to:

predict multiple candidate sequences of actions using a latent state observation of the robot and the latent robot dynamics model;

select, from the multiple candidate sequences of actions, using the trained reward function, a sequence of actions for performing the robot task; and

control one or more actuators based on the selected sequence of actions to cause performance of the robot task.

2 . The robot of claim 1 , wherein the latent robot dynamics model is a deterministic belief state model (DBSM).

3 . The robot of claim 2 , wherein the DBSM includes an encoder network and a decoder network.

4 . The robot of claim 3 , wherein the encoder network is used in deterministically updating a deterministic activation vector based on processing the latent state observation of the robot, and one or more historical latent state observations of the robot.

5 . The robot of claim 4 , wherein the DBSM further includes a transition function and wherein the transition function is used to generate one or more future states of the robot based on the latent state observation and the sequence of actions for performing the robot task.

6 . The robot of claim 5 , wherein the decoder network is used to generate the sequence of actions for performing the robot task based on the latent state observation of the robot, the one or more historical latent state observations of the robot, and the one or more future states of the robot.

7 . A robot comprising:

one or more actuators;

memory storing a trained latent robot dynamics model and a trained reward function for a robot task;

one or more processors configured to:

use a latent state observation of the robot, the latent robot dynamics model, and the trained reward function, to determine a sequence of actions for performing the robot task; and

control one or more actuators based on the determined sequence of actions to cause performance of the robot task;

wherein, subsequent to controlling the one or more actuators based on the determined sequence of actions to cause performance of the robot task, one or more of the processors are further configured to:

identify an updated latent state observation of the robot;

use the updated latent state observation of the robot, the latent robot dynamics model, and the trained reward function, to determine an updated sequence of actions for performing the robot task; and

control the one or more actuators based on the determined updated sequence of actions to cause performance of the robot task.

8 . A robot comprising:

one or more actuators;

memory storing a trained latent robot dynamics model and a trained reward function for a robot task;

one or more processors configured to:

use a latent state observation of the robot, the latent robot dynamics model, and the trained reward function, to determine a sequence of actions for performing the robot task; and

control one or more actuators based on the determined sequence of actions to cause performance of the robot task;

wherein, subsequent to controlling the one or more actuators based on the determined sequence of actions to cause performance of the robot task, one or more of the processors are further configured to:

use an additional latent state observation of the robot, the latent robot dynamics model, and the trained reward function, to determine an additional sequence of actions for performing an additional robot task; and

control the one or more actuators based on the additional sequence of actions to cause performance of the additional robot task.

9 . A method implemented by one or more processors, the method comprising:

predicting multiple candidate sequences of actions using a latent state observation of a robot and a latent robot dynamics model;

selecting, from the multiple candidate sequences of actions, using a trained reward function a sequence of actions for performing a robot task, wherein the trained reward function is for the robot task; and

controlling one or more actuators based on the selected sequence of actions to cause performance of the robot task.

10 . The method of claim 9 , wherein the latent robot dynamics model is a deterministic belief state model (DBSM).

11 . The method of claim 10 , wherein the DBSM includes an encoder network and a decoder network.

12 . The method of claim 11 , wherein the encoder network is used in deterministically updating a deterministic activation vector based on processing the latent state observation of the robot, and one or more historical latent state observations of the robot.

13 . The method of claim 12 , wherein the DBSM further includes a transition function and wherein the transition function is used to generate one or more future states of the robot based on the latent state observation and the sequence of actions for performing the robot task.

14 . The method of claim 13 , wherein the decoder network is used to generate the sequence of actions for performing the robot task based on the latent state observation of the robot, the one or more historical latent state observations of the robot, and the one or more future states of the robot.

15 . A method implemented by one or more processors, the method comprising:

using a latent state observation of a robot, a latent robot dynamics model, and a trained reward function, to determine a sequence of actions for performing a robot task, wherein the trained reward function is for the robot task; and

controlling one or more actuators based on the determined sequence of actions to cause performance of the robot task;

wherein, subsequent to controlling the one or more actuators based on the determined sequence of actions to cause performance of the robot task, the method further comprises:

identifying an updated latent state observation of the robot;

using the updated latent state observation of the robot, the latent robot dynamics model, and the trained reward function, to determine an updated sequence of actions for performing the robot task; and

controlling the one or more actuators based on the determined updated sequence of actions to cause performance of the robot task.

16 . A method implemented by one or more processors, the method comprising:

using a latent state observation of a robot, a latent robot dynamics model, and a trained reward function, to determine a sequence of actions for performing a robot task, wherein the trained reward function is for the robot task; and

controlling one or more actuators based on the determined sequence of actions to cause performance of the robot task;

wherein, subsequent to controlling the one or more actuators based on the determined sequence of actions to cause performance of the robot task, the method further comprises:

using an additional latent state observation of the robot, the latent robot dynamics model, and the trained reward function, to determine an additional sequence of actions for performing an additional robot task; and

controlling the one or more actuators based on the additional sequence of actions to cause performance of the additional robot task.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2024
From: HAFNER, DANIJAR
To: GOOGLE LLC
Reel/Frame 066776/0171 →
Continuity (3)
Continuation 17056104
Provisional Application 62673744 · May 18, 2018
Related Publication 20240173854A1 · May 30, 2024
References Cited (24)
US 9070083B2 · Suh · 2015 [cited by applicant]
US 10603797B2 · Ozaki · 2020 [cited by examiner]
US 10766136B1 · Porter · 2020 [cited by applicant]
US 10800040B1 · Beckman et al. · 2020 [cited by applicant]
US 20150100530A1 · Mnih · 2015 [cited by examiner]
US 20170334066A1 · Levine et al. · 2017 [cited by applicant]
US 20180089553A1 · Liu · 2018 [cited by applicant]
US 20180099408A1 · Shibata · 2018 [cited by examiner]
US 20190143541A1 · Nemallan · 2019 [cited by applicant]
US 20210078168A1 · Mehnert · 2021 [cited by applicant]
US 20210205984A1 · Hafner · 2021 [cited by applicant]
CN 103810500 · 2014 [cited by applicant]
CN 106448670 · 2017 [cited by applicant]
WO 2017201023 · 2017 [cited by applicant]
WO 2017223192 · 2017 [cited by applicant]
WO 2019222597 · 2019 [cited by applicant]
China National Intellectual Property Administration; Notice of Allowance issued in Application No. 201980033351.3; 5 pages; dated Nov. 1, 2023. [cited by applicant]
European Patent Office; Examination Report issued in Application No. 19728281.7; 6 pages; dated Jan. 4, 2023. [cited by applicant]
China National Intellectual Property Administration; First Office Action issued in Application No. 201980033351.3; 36 pages; dated Feb. 9, 2023. [cited by applicant]
European Patent Office; International Search Report and Written Opinion of PCT Ser. No. PCT/US2019/032823; 22 pages; dated Sep. 3, 2019. [cited by applicant]
Wahlstrom, N. et al. “From Pixels to Torques: Policy Leaming with Deep Dynamical Models”; arxiv.org; retrieved from internet: URL:https://arxiv.org/pdf/1502.02251.pdf [retrieved Aug. 26, 2019]; 9 pages; Jun. 18, 2015. [cited by applicant]
Hafner, D. et al. “Leaming Latent Dynamics for Planning for Pixels”; retrieved from internet: URL:https://arxiv.org/pdf/1811.04551v4.pdf [retrieved on Aug. 26, 2019]; 20 pages; May 15, 2019. [cited by applicant]
Argall, B. et al. “A Survey of Robot Leaming from Demonstration”; Robotics and Autonomous Systems, Elsevier BV, vol. 57, No. 5, pp. 469-483; May 31, 2009. [cited by applicant]
European Patent Office; Summons issued in Application No. 19728281.7; 11 pages; dated Dec. 19, 2025. [cited by applicant]