IP Library Granted Patent US 12,153,414
Granted Patent B2
US 12,153,414 · App. 17/652,607 · Granted Nov 26, 2024

Imitation learning in a manufacturing environment

Inventors: Matthew C. Putman (Brooklyn, NY); Andrew Sundstrom (Brooklyn, NY); Damas Limoge (Brooklyn, NY); Vadim Pinskiy (Wayne, NJ); Aswin Raghav Nirmaleswaran (Brooklyn, NY); Eun-Sol Kim (Cliffside Park, NJ)
Assignee: Nanotronics Imaging, Inc.
G05B19/423
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,153,414
App. No.
17/652,607
Granted
Nov 26, 2024
Kind
B2
Abstract

A computing system identifies a trajectory example generated by a human operator. The trajectory example includes trajectory information of the human operator while performing a task to be learned by a control system of the computing system. Based on the trajectory example, the computing system trains the control system to perform the task exemplified in the trajectory example. Training the control system includes generating an output trajectory of a robot performing the task. The computing system identifies an updated trajectory example generated by the human operator based on the trajectory example and the output trajectory of the robot performing the task. Based on the updated trajectory example, the computing system continues to train the control system to perform the task exemplified in the updated trajectory example.

Claims (58)

1. A method for training a control system, comprising:

receiving, by a computing system, an initial teacher policy based on a trajectory example generated by a human operator in a first action space, the trajectory example captured using one or more sensors monitoring movements of the human operator, the trajectory example comprising trajectory information of the human operator while performing a task to be learned by a control system of the computing system;

based on the initial teacher policy, generating, by the computing system, an initial student policy by training the control system to perform the task exemplified in the trajectory example, wherein the control system exists is a second action space that is lower dimension from the first action space, wherein movements of the control system in the second action space are limited compared to movements of the human operator in the first action space, wherein training the control system comprises:

causing the control system to mimic the movements of the human operator while performing the task, and

monitoring the movements of the control system using sensors, and

generating an output trajectory of the control system performing the task based on the monitored movements;

providing, by the computing system, the output trajectory of the control system to the human operator for determining a reproducibility of the trajectory example based on the output trajectory generated by the control system;

receiving, by the computing system, an updated teacher policy based on an updated trajectory example generated by the human operator responsive to the determined reproducibility of the trajectory example; and

based on the updated teacher policy, generating, by the computing system, an updated student policy by training the control system to perform the task exemplified in the updated trajectory example.

2. The method of claim 1 , wherein generating, by the computing system, the updated student policy by training the control system to perform the task exemplified in the updated trajectory example comprises:

outputting an updated output trajectory of the control system performing the task.

3. The method of claim 2 , further comprising:

receiving, by the computing system, a further updated teacher policy comprising a further updated trajectory example generated by the human operator based on the trajectory example, the output trajectory, the updated trajectory example, and the updated output trajectory of the control system performing the task; and

based on the further updated teacher policy, generating, by the computing system, a further updated student policy by training the control system to perform the task exemplified in the further updated trajectory example.

4. The method of claim 1 , wherein the trajectory example is projected from a first environment in which the human operator performs the task into a second environment in which the control system performs the task, wherein the first environment is a higher dimensional environment than the second environment.

5. The method of claim 4 , wherein the updated trajectory example is projected from the first environment in which the human operator performs the task into the second environment in which the control system performs the task.

6. The method of claim 1 , further comprising:

minimizing a distance between the task as performed by the human operator and the task as performed by the control system.

7. The method of claim 1 , wherein the task is a manufacturing task.

8. A system comprising:

a processor; and

a memory having programming instructions stored thereon, which, when executed by the processor, causes the system to perform operations, comprising:

receiving an initial teacher policy based on a trajectory example generated by a human operator in a first action space, the trajectory example captured using one or more sensors monitoring movements of the human operator, the trajectory example comprising trajectory information of the human operator while performing a task to be learned by a control system of the system;

based on the initial teacher policy, generating an initial student policy by training the control system to perform the task exemplified in the trajectory example, wherein the control system exists is a second action space that is lower dimension from the first action space, wherein movements of the control system in the second action space are limited compared to movements of the human operator in the first action space, wherein training the control system comprises:

causing the control system to mimic the movements of the human operator while performing the task, and

monitoring the movements of the control system using sensors, and

generating an output trajectory of the control system performing the task based on the monitored movements;

providing the output trajectory of the control system to the human operator for determining a reproducibility of the trajectory example based on the output trajectory generated by the control system;

receiving an updated teacher policy based on an updated trajectory example generated by the human operator responsive to the determined reproducibility of the trajectory example; and

based on the updated teacher policy, generating an updated student policy by training the control system to perform the task exemplified in the updated trajectory example.

9. The system of claim 8 , wherein generating the updated student policy by training the control system to perform the task exemplified in the updated trajectory example comprises:

outputting an updated output trajectory of the robot control system performing the task.

10. The system of claim 9 , wherein the operations further comprise:

receiving a further updated teacher policy based on a further updated trajectory example generated by the human operator based on the trajectory example, the output trajectory, the updated trajectory example, and the updated output trajectory of the control system performing the task; and

based on the further updated teacher policy, generating a further updated student policy by training the control system to perform the task exemplified in the further updated trajectory example.

11. The system of claim 8 , wherein the trajectory example is projected from a first environment in which the human operator performs the task into a second environment in which the control system performs the task, wherein the first environment is a higher dimensional environment than the second environment.

12. The system of claim 11 , wherein the updated trajectory example is projected from the first environment in which the human operator performs the task into the second environment in which the control system performs the task.

13. The system of claim 8 , wherein the operations further comprise:

minimizing a distance between the task as performed by the human operator and the task as performed by the control system.

14. The system of claim 8 , wherein the task is a manufacturing task.

15. A non-transitory computer readable medium comprising one or more sequences of instructions, which, when executed by a processor, causes a computing system to perform operations comprising:

receiving, by a computing system, an initial teacher policy based on a trajectory example generated by a human operator in a first action space, the trajectory example captured using one or more sensors monitoring movements of the human operator, the trajectory example comprising trajectory information of the human operator while performing a task to be learned by a control system of the computing system;

based on the initial teacher policy, generating, by the computing system, an initial student policy by training the control system to perform the task exemplified in the trajectory example, wherein the control system exists is a second action space that is lower dimension from the first action space, wherein movements of the control system in the second action space are limited compared to movements of the human operator in the first action space, wherein training the control system comprises:

causing the control system to mimic the movements of the human operator while performing the task, and

monitoring the movements of the control system using sensors, and

generating an output trajectory of the control system performing the task based on the monitored movements;

providing, by the computing system, the output trajectory of the control system to the human operator for determining a reproducibility of the trajectory example based on the output trajectory generated by the control system;

receiving, by the computing system, an updated teacher policy based on an updated trajectory example generated by the human operator responsive to the determined reproducibility of the trajectory example; and

based on the updated teacher policy, generating, by the computing system, an updated student policy by training the control system to perform the task exemplified in the updated trajectory example.

16. The non-transitory computer readable medium of claim 15 , wherein generating, by the computing system, the updated student policy by training the control system to perform the task exemplified in the updated trajectory example comprises:

outputting an updated output trajectory of the control system performing the task.

17. The non-transitory computer readable medium of claim 16 , further comprising:

receiving, by the computing system, a further updated teacher policy comprising a further updated trajectory example generated by the human operator based on the trajectory example, the output trajectory, the updated trajectory example, and the updated output trajectory of the control system performing the task; and

based on the further updated trajectory example, generating, by the computing system, a further updated student policy by training the control system to perform the task exemplified in the further updated trajectory example.

18. The non-transitory computer readable medium of claim 15 , wherein the trajectory example is projected from a first environment in which the human operator performs the task into a second environment in which the control system performs the task, wherein the first environment is a higher dimensional environment than the second environment.

19. The non-transitory computer readable medium of claim 18 , wherein the updated trajectory example is projected from the first environment in which the human operator performs the task into the second environment in which the control system performs the task.

20. The non-transitory computer readable medium of claim 15 , further comprising:

minimizing a distance between the task as performed by the human operator and the task as performed by the control system.

Assignments (2)
SECURITY INTEREST Recorded Nov 30, 2023
From: NANOTRONICS IMAGING, INC.; NANOTRONICS HEALTH LLC; CUBEFABS INC.
To: ORBIMED ROYALTY & CREDIT OPPORTUNITIES IV, LP
Reel/Frame 065726/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2022
From: PUTMAN, MATTHEW C.; SUNDSTROM, ANDREW; LIMOGE, DAMAS; PINSKIY, VADIM; NIRMALESWARAN, ASWIN RAGHAV; KIM, EUN-SOL
To: NANOTRONICS IMAGING, INC.
Reel/Frame 059182/0632 →
Continuity (2)
Provisional Application 63153811 · Feb 25, 2021
Related Publication 20220269254A1 · Aug 25, 2022