IP Library › Granted Patent US 12,583,111
Granted Patent B2
US 12,583,111 · App. 18/472,767 · Granted Mar 24, 2026

Eye-on-hand reinforcement learner for dynamic grasping with active pose estimation

Inventors: Siddarth Jain (Cambridge, MA); Baichuan Huang (Piscataway, NJ); Jingjin Yu (Edison, NJ)
Assignee: Mitsubishi Electric Research Laboratories, Inc.
B25J9/1664B25J9/161B25J9/163B25J9/1671B25J9/1697B25J19/023
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,583,111
App. No.
18/472,767
Filed
Sep 22, 2023
Granted
Mar 24, 2026
Kind
B2
Art Unit
3656
USPC
700/250
Abstract

A controller is provided for performing dynamic grasping of a target object using visual sensory inputs. The controller includes a robotic interface connected to a robotic arm including links connected by joints having actuators and encoders, and a gripper of the end-effector of the robotic arm configured to grasp the target object in response to robot control signals, and a vision sensor configured to continuously provide visual observations for tracking poses of the target object in a workspace and compute grasp poses, wherein the vision sensor is mounted on a distal end of the robotic arm adjacent to the gripper. The controller trains the Eye-on-Hand reinforcement learner policy, tracks the poses of the target object, and generates robot control signals to follow the target object while keeping it in the field of view of the vision sensor and grasp the target object in the workspace.

Claims (44)

1 . A controller for performing dynamic grasping of a target object using a robotic arm and visual sensory inputs, comprising:

a data input/output interface configured to receive state measurements of the robotic arm and the target object from sensors arranged on the robotic arm, wherein the robotic arm includes links connected by joints having actuators and encoders, and a gripper of an end-effector of the robotic arm configured to grasp the target object in response to robot control signals, wherein the sensors include a vision sensor configured to continuously provide visual observations for tracking poses of the target object in a workspace and compute grasp poses;

a memory configured to store an Eye-on-Hand (EoH) reinforcement leaner (EARL) policy, a physics-based simulator, an arm motion generation program; and

a processor, in connection with the memory, configured to perform steps of:

training the Eye-on-Hand reinforcement learner (EARL) policy, wherein

for the training of the EARL policy, the processor is configured to train a neural network to learn a control policy for the dynamic grasping, and

the training of the neural network to learn the control policy is based on a curriculum design that gradually increases task difficulty and adapts reward function design;

tracking the poses of the target object moving in the workspace based on the state measurements;

computing a set of grasp poses on the target object and dynamically selecting a desired grasp pose on the target object moving in the workspace;

computing robotic arm motion commands using the trained Eye-on-Hand reinforcement learner policy;

generating the robot control signals based on the computed robotic arm motion commands; and

transmitting, via the data input/output interface, the robot control signals to the actuators of the joints and the gripper to follow the target object while keeping the target object in a field of view of the vision sensor and grasp the target object in the workspace.

2 . The controller of claim 1 , wherein the vision sensor is mounted on a distal end of the robotic arm adjacent to the gripper.

3 . The controller of claim 1 , wherein the vision sensor is an onboard camera sensor, wherein the visual observations provided by the onboard camera sensor constitutes image data comprising information of a depth channel, a first color channel, a second color channel, and a third color channel.

4 . The controller of claim 3 , wherein the visual observations are encoded as a high-level representation or an encoding of the target object.

5 . The controller of claim 3 , wherein the vision sensor provides the visual observations to the processor, wherein the processer computes a spatial location of the target object as a six-dimensional (6D) pose of the target object in the workspace indicative of a position and orientation of the target object relative to a camera frame of the vision sensor, or the gripper, or the base frame of the robotic arm.

6 . The controller of claim 1 , wherein the EARL policy generalizes to work on target objects which are not used for training the EARL policy, and a model or identity of the target object is not known, and the motion of the target object is not explicitly modeled or known a priori for the dynamic grasping.

7 . The controller of claim 1 , wherein the dynamic grasping is performed in six degrees of freedom (DoF).

8 . The controller of claim 1 , wherein the target object moves in a linear, circular, or random pattern including random motions in the workspace.

9 . The controller of claim 1 , wherein the processor computes the set of grasp poses on the target object relative to a first pose of the target object derived from the visual observations, and dynamically selects a location for the desired grasp pose on the target object based on the first pose of the target object.

10 . The controller of claim 1 , wherein the control policy maps the desired grasp pose and joint states to at least one of a desired joint velocity of the robotic arm or a position of the robotic arm for tracking the target object and gripper actions including opening and/or closing the gripper to perform a desired grasp on the target object.

11 . The controller of claim 10 , wherein the joint states include one or more of the joint positions, velocity, and torque values.

12 . The controller of claim 10 , wherein parameters of the neural network for the control policy to perform the dynamic grasping are learned with reinforcement learning.

13 . The controller of claim 10 , wherein the control policy is learned in simulations, wherein parallel training of the control policy is performed in simultaneous simulations with one of more independent Eye-on-Hand systems in a physics-based simulator.

14 . The controller of claim 10 , wherein the control policy for the dynamic grasping is learned using one or more real robotic arms.

15 . The controller of claim 10 , wherein the reward function design to train the control policy has multiple components with diverse guidance.

16 . The controller of claim 1 , wherein the learned control policy runs in real-time for the continuous degree of freedom of the Eye-on-Hand system.

17 . The controller of claim 1 , wherein the control policy trained in a simulator is applicable to real-world Eye-on-Hand systems with parameter tuning.

18 . A system for performing dynamic grasping of a target object using visual sensory inputs, comprising:

a robotic arm configured to include links connected by joints having actuators and encoders, and a gripper of an end-effector of the robotic arm configured to grasp the target object in response to robot control signals;

sensors arranged on the robotic arm, wherein the sensors are configured to measure state measurements of the robotic arm and the target object, wherein the sensors include a vision sensor configured to continuously provide visual observations for tracking poses of the target object in a workspace and compute grasp poses; and

a controller comprising:

a data input/output interface configured to receive the state measurements of the robotic arm and the target object from the sensors;

a memory configured to store an Eye-on-Hand (EoH) reinforcement leaner (EARL) policy, a physics-based simulator, an arm motion generation program; and

a processor, in connection with the memory, configured to perform steps of:

training the Eye-on-Hand reinforcement learner (EARL) policy, wherein

for the training of the EARL policy, the processor is configured to train a neural network to learn a control policy for the dynamic grasping, and

the training of the neural network to learn the control policy is based on a curriculum design that gradually increases task difficulty and adapts reward function design;

tracking the poses of the target object moving in the workspace;

computing a set of grasp poses on the target object and dynamically selecting a desired grasp pose on the target object moving in the workspace;

computing robotic arm motion commands using the trained Eye-on-Hand reinforcement learner policy;

generating the robot control signals based on the computed robotic arm motion commands; and

transmitting, via the data input/output interface, the robot control signals to the actuators of the joints and the gripper to follow the target object while keeping the target object in a field of view of the vision sensor and grasp the target object in the workspace.

19 . The system of claim 18 , wherein the control policy maps the desired grasp pose and joint states to at least one of a desired joint velocity of the robotic arm or a position of the robotic arm for tracking the target object and gripper actions including opening and/or closing the gripper to perform a desired grasp on the target object.

Continuity (2)
Provisional Application 63580744 · Sep 6, 2023
Related Publication 20250100141A1 · Mar 27, 2025
References Cited (10)
US 9924213B2 · Kelsen et al. · 2018 [cited by applicant]
US 10800040B1 · Beckman · 2020 [cited by examiner]
US 20200061811A1 · Iqbal · 2020 [cited by examiner]
US 20210118166A1 · Tremblay · 2021 [cited by examiner]
US 20240316763A1 · Hoffman · 2024 [cited by examiner]
US 20250214249A1 · Hosomi · 2025 [cited by examiner]
JP 2021084210A · 2021 [cited by examiner]
Plappert, Asymmetric Self-Play for Automatic Goal Discovery in Robotic Manipulation (Year: 2021). [cited by examiner]
Akinola, Iretiayo, et al. “Dynamic grasping with reachability and motion awareness.” 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2021. [cited by applicant]
Morrison, Douglas, Peter Corke, and Jürgen Leitner. “Closing the loop for robotic grasping: A real-time, generative grasp synthesis approach.” arXiv preprint arXiv:1804.05172 (2018). [cited by applicant]