IP Library Patent Application 17191264
Patent Application
App. No. 17/191,264

Transformer-Based Meta-Imitation Learning Of Robots

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/191,264
Abstract

A training system for a robot includes: a model having a transformer architecture and configured to determine how to actuate at least one of arms and an end effector of the robot; a training dataset including sets of demonstrations for the robot to perform training tasks, respectively; and a training module configured to: meta-train a policy of the model using first ones of the sets of demonstrations for first ones of the training tasks, respectively; and optimize the policy of the model using second ones of the sets of demonstrations for second ones of the training tasks, respectively, where the sets of demonstrations for the training tasks each include more than one demonstration and less than a first predetermined number of demonstrations.

Claims (46)

1 . A training system for a robot, comprising:

a model having a transformer architecture and configured to determine how to actuate at least one of arms and an end effector of the robot;

a training dataset including sets of demonstrations for the robot to perform training tasks, respectively; and

a training module configured to:

meta-train a policy of the model using first ones of the sets of demonstrations for first ones of the training tasks, respectively; and

optimize the policy of the model using second ones of the sets of demonstrations for second ones of the training tasks, respectively,

wherein the sets of demonstrations for the training tasks each include more than one demonstration and less than a first predetermined number of demonstrations.

2 . The training system of claim 1 wherein the training module is configured to meta-train the policy using reinforcement learning.

3 . The training system of claim 1 wherein the training module is configured to meta-train the policy using one of the Reptile algorithm and the model-agnostic meta-learning (MAML) algorithm.

4 . The training system of claim 1 wherein the training module is configured to meta-train the policy of the model before optimizing the policy.

5 . The training system of claim 1 wherein the model is configured determine how to actuate at the least one of the arms and the end effector of the robot to advance toward or to completion of a task.

6 . The training system of claim 5 wherein the task is different than the training tasks.

7 . The training system of claim 5 wherein, after the meta-training and the optimization, the model is configured to perform the task using less than or equal to a second predetermined number of user input demonstrations for performing the task,

wherein the second predetermined number is an integer greater than zero.

8 . The training system of claim 7 wherein the second predetermined number is 5.

9 . The training system of claim 7 wherein the user input demonstrations include: (a) positions of joints of the robot; and (b) a pose of the end effector of the robot.

10 . The training system of claim 9 wherein the pose of the end effector includes a position of the end effector and an orientation of the end effector.

11 . The training system of claim 9 wherein the user input demonstrations also include a position of an object to be interacted with by the robot during performance of the task.

12 . The training system of claim 11 wherein the user input demonstrations also include a position of a second object in an environment of the robot.

13 . The training system of claim 1 wherein the first predetermined number is an integer less than or equal to ten.

14 . A training system, comprising:

a model having a transformer architecture and configured to determine an action;

a training dataset including sets of demonstrations for training tasks, respectively; and

a training module configured to:

meta-train a policy of the model using first ones of the sets of demonstrations for first ones of the training tasks, respectively; and

optimize the policy of the model using second ones of the sets of demonstrations for second ones of the training tasks, respectively,

wherein the sets of demonstrations for the training tasks each include more than one demonstration and less than a first predetermined number of demonstrations.

15 . A training method for a robot, comprising:

storing a model having a transformer architecture and configured to determine how to actuate at least one of arms and an end effector of the robot;

storing a training dataset including sets of demonstrations for the robot to perform training tasks, respectively;

meta-training a policy of the model using first ones of the sets of demonstrations for first ones of the training tasks, respectively; and

optimizing the policy of the model using second ones of the sets of demonstrations for second ones of the training tasks, respectively,

wherein the sets of demonstrations for the training tasks each include more than one demonstration and less than a first predetermined number of demonstrations.

16 . The training method of claim 15 wherein the meta-training includes meta-training the policy using reinforcement learning.

17 . The training method of claim 15 wherein the meta-training includes meta-training the policy using one of the Reptile algorithm and the model-agnostic meta-learning (MAML) algorithm.

18 . The training method of claim 15 wherein the meta-training includes meta-training the policy of the model before optimizing the policy.

19 . The training method of claim 15 wherein the model is configured determine how to actuate at the least one of the arms and the end effector of the robot to advance toward or to completion of a task.

20 . The training method of claim 19 wherein the task is different than the training tasks.

21 . The training method of claim 19 wherein, after the meta-training and the optimization, the model is configured to perform the task using less than or equal to a second predetermined number of user input demonstrations for performing the task,

wherein the second predetermined number is an integer greater than zero.

22 . The training method of claim 21 wherein the second predetermined number is 5.

23 . The training method of claim 21 wherein the user input demonstrations include: (a) positions of joints of the robot; and (b) a pose of the end effector of the robot.

24 . The training method of claim 23 wherein the pose of the end effector includes a position of the end effector and an orientation of the end effector.

25 . The training method of claim 23 wherein the user input demonstrations also include a position of an object to be interacted with by the robot during performance of the task.

26 . The training method of claim 25 wherein the user input demonstrations also include a position of a second object in an environment of the robot.

27 . The training method of claim 15 wherein the first predetermined number is an integer less than or equal to ten.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2024
From: NAVER LABS CORPORATION
To: NAVER CORPORATION
Reel/Frame 068820/0495 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2021
From: PEREZ, JULIEN; KIM, SEUNGSU; CACHET, THÉO
To: NAVER CORPORATION; NAVER LABS CORPORATION
Reel/Frame 055482/0839 →