IP Library Patent Application 19173425
Patent Application
App. No. 19/173,425

TRANSFORMER DIFFUSION FOR ROBOTIC TASK LEARNING

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/173,425
Abstract

Implementations are provided for learning dexterous tasks. In various implementations, a plurality of images may be retrieved that capture an environment in which a robot operates from multiple different perspectives. Data indicative of the plurality of images and a proprioceptive state of the robot may be processed using a diffusion model that includes a transformer-encoder and a transformer-decoder. The transformer-encoder may be used to generate latent embeddings representing the plurality of images and proprioceptive state of the robot. The transformer-decoder may be used to process the latent embeddings and data indicative of a diffusion timestep to generate robot control data. The robot control data may include a series of actions to be performed by the robot over a time interval. The robot may be operated in accordance with the robot control data.

Claims (43)

1 . A method implemented using one or more processors and comprising:

retrieving a plurality of images that capture an environment in which a robot operates from multiple different perspectives;

processing data indicative of the plurality of images and a proprioceptive state of the robot using a transformer-encoder to generate latent embeddings representing the plurality of images and proprioceptive state of the robot;

processing the latent embeddings and data indicative of a diffusion timestep using a transformer-decoder to generate robot control data, wherein the robot control data comprises a series of actions to be performed by the robot over a time interval;

causing the robot to be operated in accordance with the robot control data.

2 . The method of claim 1 , wherein the series of actions comprise a series of absolute joint positions of a plurality of joints of the robot.

3 . The method of claim 2 , wherein the series of actions further comprise a series of gripper positions for two or more grippers.

4 . The method of claim 3 , wherein the series of gripper positions are continuous.

5 . The method of claim 1 , wherein the series of actions comprise:

joint commands and/or torque commands;

Cartesian commands for an end effector of the robot;

a target robot pose; or

code specifying reward functions for motion controller optimization; or selected predefined robot primitives.

6 . The method of claim 1 , wherein the transformer-encoder and transformer-decoder form a diffusion policy.

7 . The method of claim 1 , further comprising processing each of the plurality of images using a respective convolutional neural network to generate feature maps.

8 . The method of claim 7 , further comprising flattening the feature maps into a sequence of tokens that comprise the data indicative of the plurality of images that is processed using the transformer encoder.

9 . The method of claim 1 , wherein the transformer-decoder comprises a diffusion denoiser.

10 . The method of claim 1 , wherein the diffusion timestep is represented as a one-hot vector.

11 . The method of claim 1 , wherein the robot is a simulated robot or a real robot.

12 . The method of claim 1 , wherein one or both of the transformer-encoder and transformer decoder are trained using training data collected using imitation learning.

13 . The method of claim 12 , wherein the imitation learning comprises teleoperation of one or more robots using a puppeteering interface.

14 . The method of claim 13 , wherein the puppeteering interface comprises two leader arms of a first size that are synchronized with two follower arms of a second size that is greater than the first size.

15 . The method of claim 13 , wherein the imitation learning comprises one or more of the following tasks:

folding a shirt;

hanging a shirt on a hanger;

shoelace tying;

robot finger placement;

gear insertion; or

stacking random collections of dishware.

16 . The method of claim 1 , wherein at least the transformer-decoder is trained with a diffusion loss.

17 . The method of claim 16 , wherein both the transformer-encoder and transformer-decoder are trained with diffusion loss.

18 . A method implemented using one or more processors and comprising:

retrieving a plurality of images that capture, from multiple different perspectives, an environment in which a robot was operated to perform a sequence of actions;

processing data indicative of the plurality of images and a proprioceptive state of the robot using a transformer-encoder to generate latent embeddings representing the plurality of images and proprioceptive state of the robot;

adding noise to the sequence of actions performed by the robot to generate a plurality of noisy actions;

processing the latent embeddings and the plurality of noisy actions using a diffusion-based transformer decoder to predict noise values;

based on the predicted noise values, training the diffusion-based transformer decoder.

19 . The method of claim 18 , wherein predicted actions are determined using the predicted noise values, and the diffusion-based transformer-decoder is trained based on a comparison of the predicted actions and the sequence of actions performed by the robot.

20 . A system comprising one or more processors and memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to:

retrieve a plurality of images that capture an environment in which a robot operates from multiple different perspectives;

process data indicative of the plurality of images and a proprioceptive state of the robot using a transformer-encoder to generate latent embeddings representing the plurality of images and proprioceptive state of the robot;

process the latent embeddings and data indicative of a diffusion timestep using a transformer-decoder to generate robot control data, wherein the robot control data comprises a series of actions to be performed by the robot over a time interval;

cause the robot to be operated in accordance with the robot control data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2025
From: ZHAO, ZIHAO; TOMPSON, JONATHAN; DRIESS, DANNY MICHAEL; FLORENCE, PETER RAYMOND; FINN, CHELSEA; WAHID, AYZAAN
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 071495/0948 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071550/0092 →