IP Library Granted Patent US 12,552,021
Granted Patent B2
US 12,552,021 · App. 18/489,789 · Granted Feb 17, 2026

Techniques for training and implementing reinforcement learning policies for robot control

Inventors: Bingjie Tang (Los Angeles, CA); Yashraj Shyam Narang (Seattle, WA); Dieter Fox (Seattle, WA); Fabio Tozeto Ramos (Seattle, WA)
Assignee: NVIDIA CORPORATION
B25J9/163B25J9/1605B25J9/1653B25J9/1664B25J9/1671G05B2219/40499
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,552,021
App. No.
18/489,789
Granted
Feb 17, 2026
Kind
B2
Abstract

One embodiment of a method for training a machine learning model to control a robot includes causing a model of the robot to move within a simulation based on one or more outputs of the machine learning model, computing an error within the simulation, computing at least one of a reward or an observation based on the error, and updating one or more parameters of the machine learning model based on the at least one of a reward or an observation.

Claims (50)

1 . A computer-implemented method for training a machine learning model to control a robot, the method comprising:

causing a model of the robot to move within a simulation based on one or more outputs of the machine learning model;

computing an error within the simulation based on simulated sensor data of a physically articulated manipulator of the robot and movement of the model of the robot;

computing at least one of a reward or an observation based on the error; and

training the machine learning model by updating one or more parameters of the machine learning model based on the at least one of the reward or the observation.

2 . The computer-implemented method of claim 1 , further comprising computing a distance between one or more points on a model of an object being grasped by the model of the robot during the simulation and a signed distance field (SDF) associated with a target pose of the model of the object, wherein the at least one of the reward or the observation is further computed based on the distance.

3 . The computer-implemented method of claim 2 , wherein the distance comprises a root-mean-square SDF distance.

4 . The computer-implemented method of claim 1 , further comprising:

performing one or more operations to determine a success rate of the machine learning model when controlling the model of the robot to perform a task involving an object in one or more simulations; and

responsive to determining that the success rate is greater than a predefined threshold, increasing a starting distance between the model of the robot and a model of the object in one or more subsequent simulations.

5 . The computer-implemented method of claim 1 , wherein computing the reward comprises, responsive to determining that the error is greater than an error threshold:

computing a weight value based on the error, and

computing the reward based on the weight value.

6 . The computer-implemented method of claim 1 , wherein the error is associated with at least one of an interpenetration between two objects during the simulation, a solver residual, a deviation from a ground truth, a deviation of the simulation from a slower simulation, a deviation from a reference value, or a deviation from an analytical solution.

7 . The computer-implemented method of claim 1 , wherein the steps of causing the model of the robot to move within the simulation, computing the error, computing the at least of one of the reward or the observation, and updating the one or more parameters are repeated for each time step included in a plurality of time steps.

8 . The computer-implemented method of claim 1 , further comprising performing one or more operations to control the robot based on one or more additional outputs of the machine learning model.

9 . The computer-implemented method of claim 1 , further comprising:

generating one or more control signals based on one or more additional outputs of the machine learning model; and

causing the robot to move within a real-world environment based on the one or more control signals.

10 . The computer-implemented method of claim 9 , further comprising processing one or more sensor signals using the machine learning model to generate the one or more additional outputs.

11 . One or more non-transitory computer-readable media storing instructions that, when executed by at least one processor, cause the at least one processor to perform the steps of:

causing a model of a robot to move within a simulation based on one or more outputs of a machine learning model;

computing an error within the simulation based on simulated sensor data of a physically articulated manipulator of the robot and movement of the model of the robot;

computing at least one of a reward or an observation based on the error; and

training the machine learning model by updating one or more parameters of the machine learning model based on the at least one of the reward or the observation.

12 . The one or more non-transitory computer-readable media of claim 11 , wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of computing a distance between one or more points on a model of an object being grasped by the model of the robot during the simulation and a signed distance field (SDF) associated with a target pose of the model of the object, wherein the at least one of the reward or the observation is further computed based on the distance.

13 . The one or more non-transitory computer-readable media of claim 12 , wherein the distance comprises a root-mean-square SDF distance.

14 . The one or more non-transitory computer-readable media of claim 11 , wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the steps of:

performing one or more operations to determine a success rate of the machine learning model when controlling the model of the robot to perform a task involving an object in one or more simulations; and

responsive to determining that the success rate is greater than a predefined threshold, increasing a starting distance between the model of the robot and a model of the object in one or more subsequent simulations.

15 . The one or more non-transitory computer-readable media of claim 11 , wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the steps of, responsive to determining that the error is greater than an error threshold:

computing a weight value based on the error, and

computing the reward based on the weight value.

16 . The one or more non-transitory computer-readable media of claim 11 , wherein computing the error comprises:

performing one or more operations to determine one or more intermediate errors between one or more points on the model of the robot and one or more points on a model of an object within the simulation; and

computing the error based on the one or more intermediate errors.

17 . The one or more non-transitory computer-readable media of claim 16 , wherein computing the error based on the one or more intermediate errors comprises determining a maximum of the one or more intermediate errors.

18 . The one or more non-transitory computer-readable media of claim 11 , wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of performing one or more operations to control the robot based on one or more additional outputs of the machine learning model.

19 . The one or more non-transitory computer-readable media of claim 11 , wherein the simulation includes at least one of picking a first object, placing the first object, or inserting the first object into a second object.

20 . A system, comprising:

one or more memories storing instructions; and

one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to:

cause a model of a robot to move within a simulation based on one or more outputs of a machine learning model,

perform one or more operations to determine an error within the simulation based on simulated sensor data of a physically articulated manipulator of the robot and movement of the model of the robot,

compute at least one of a reward or an observation based on the error, and

train the machine learning model by updating one or more parameters of the machine learning model based on the at least one of the reward or the observation.

21 . The system of claim 20 , further comprising the robot, wherein the one or more processors, when executing the instructions, are further configured to:

process one or more sensor signals associated with the robot using the machine learning model to generate one or more additional outputs; and

perform one or more operations to control the robot based on the one or more additional outputs.

22 . The system of claim 21 , wherein the robot is controlled to perform a task that comprises at least one of picking an object, placing an object, or inserting an object into another object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2023
From: TANG, BINGJIE; NARANG, YASHRAJ SHYAM; FOX, DIETER; TOZETO RAMOS, FABIO
To: NVIDIA CORPORATION
Reel/Frame 065285/0573 →
Continuity (2)
Provisional Application 63488667 · Mar 6, 2023
Related Publication 20240300099A1 · Sep 12, 2024
References Cited (112)
US 10556336B1 · Bai · 2020 [cited by examiner]
US 10800040B1 · Beckman · 2020 [cited by examiner]
US 11213946B1 · Bai · 2022 [cited by examiner]
US 11244145B2 · Eshima · 2022 [cited by examiner]
US 11461589B1 · Bai · 2022 [cited by examiner]
US 11494632B1 · Bai · 2022 [cited by examiner]
US 11584008B1 · Beckman et al. · 2023 [cited by applicant]
US 11685045B1 · Herzog et al. · 2023 [cited by applicant]
US 11707838B1 · Vogelsong et al. · 2023 [cited by applicant]
US 11735309B2 · Purdie · 2023 [cited by examiner]
US 12157226B2 · Kranski · 2024 [cited by examiner]
US 12168296B1 · Bennice · 2024 [cited by examiner]
US 20040267404A1 · Danko · 2004 [cited by examiner]
US 20160364880A1 · Barratt · 2016 [cited by examiner]
US 20200159648A1 · Ghare · 2020 [cited by examiner]
US 20200265302A1 · Sanyal · 2020 [cited by examiner]
US 20210187733A1 · Lee et al. · 2021 [cited by applicant]
US 20210213973A1 · Carillo Peña · 2021 [cited by examiner]
US 20210358595A1 · Tamersoy · 2021 [cited by examiner]
US 20210390452A1 · Kohata · 2021 [cited by examiner]
US 20220105626A1 · Luo · 2022 [cited by examiner]
US 20220111517A1 · Bennice · 2022 [cited by examiner]
US 20220172107A1 · Butterfoss · 2022 [cited by examiner]
US 20220262100A1 · Chandler · 2022 [cited by examiner]
US 20220382246A1 · Heiden · 2022 [cited by examiner]
US 20230109398A1 · Kranski · 2023 [cited by examiner]
US 20230226696A1 · Mandlekar et al. · 2023 [cited by applicant]
US 20230298263A1 · Yang · 2023 [cited by examiner]
US 20230410404A1 · Barsan · 2023 [cited by examiner]
US 20240253215A1 · Bennice et al. · 2024 [cited by applicant]
US 20240300096A1 · Cherian et al. · 2024 [cited by applicant]
US 20250128419A1 · Lai et al. · 2025 [cited by applicant]
B. Lee, C. Zhang, Z. Huang and D. D. Lee, “Online Continuous Mapping using Gaussian Process Implicit Surfaces,” 2019, IEEE, International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada, 2019, pp. 688… [cited by examiner]
Oliff et al; Reinforcement learning for facilitating human-robot-interaction in manufacturing; 2020; Elsevier; Journal of Manufacturing Systems, 56(2020) 326-340; https://doi.org/10.1016/j.jmsy.2020.06.018; (Year: 2020). [cited by examiner]
Abeyruwan et al., “i-Sim2Real: Reinforcement Learning of Robotic Policies in Tight Human-Robot Interaction Loops”, 6th Conference on Robot Learning, arXiv:2207.06572, Nov. 22, 2022, pp. 1-32. [cited by applicant]
Akkaya et al., “Solving Rubik's Cube with a Robot Hand”, arXiv:1910.07113, Oct. 16, 2019, pp. 1-51. [cited by applicant]
Allshire et al., “Transferring Dexterous Manipulation from GPU Simulation to a Remote Real-World TriFinger”, arXiv:2108.09779, Oct. 20, 2022, 8 pages. [cited by applicant]
Andrews et al., “Contact and Friction Simulation for Computer Graphics”, SIGGRAPH Courses, https://doi.org/10.1145/3532720.3535640, Aug. 7-11, 2022, pp. 1-172. [cited by applicant]
Andrychowicz et al., “Learning Dexterous In-Hand Manipulation”, The International Journal of Robotics Research, DOI: 10.1177/0278364919887447, vol. 39, 2020, pp. 3-20. [cited by applicant]
Apolinarska et al., “Robotic Assembly of Timber Joints Using Reinforcement Learning”, Automation in Construction, https://doi.org/10.1016/j.autcon.2021.103569, vol. 125, No. 103569, Feb. 27, 2021, pp. 1-8. [cited by applicant]
Beltran-Hernandez et al., “Variable Compliance Control for Robotic Peg-in-Hole Assembly: A Deep-Reinforcement-Learning Approach”, Applied Sciences, doi:10.3390/app10196923, vol. 10, No. 6923, Oct. 2, 2020, pp. 1-17. [cited by applicant]
Bengio et al., “Curriculum Learning”, Proceedings of the 26th International Conference on Machine Learning, 2009, pp. 41-48. [cited by applicant]
Chen et al., “Visual Dexterity: In-hand Dexterous Manipulation from Depth”, arXiv:2211.11744, Nov. 21, 2022, pp. 1-61. [cited by applicant]
Chen et al., “Midas: A Multi-Joint Robotics Simulator with Intersection-Free Frictional Contact”, arXiv:2210.00130, Sep. 30, 2022, 6 pages. [cited by applicant]
Davchev et al., “Residual Learning from Demonstration: Adapting DMPs for Contact-rich Manipulation”, arXiv:2008.07682, Sep. 14, 2021, 8 pages. [cited by applicant]
Drake, Samuel Hunt, “Using Compliance in Lieu of Sensory Feedback for Automatic Assembly”, 1978, 173 pages. [cited by applicant]
Fan et al., “A Learning Framework for High Precision Industrial Assembly”, International Conference on Robotics and Automation (ICRA), IEEE, May 20-24, 2019, pp. 811-817. [cited by applicant]
Ferguson et al., “Intersection-free Rigid Body Dynamics”, ACM Transactions on Graphics, https://doi.org/10.1145/3450626.3459802, vol. 40, No. 4, Article 183, Aug. 2021, pp. 183:1-183:16. [cited by applicant]
Fu et al., “Safely Learning Visuo-Tactile Feedback Policies in Real for Industrial Insertion”, arXiv:2210.01340, Oct. 4, 2022, 7 pages. [cited by applicant]
Gaz et al., “Dynamic Identification of the Franka Emika Panda Robot With Retrieval of Feasible Parameters Using Penalty-Based Optimization”, IEEE Robotics and Automation Letters, DOI 10.1109/LRA.2019.2931248, vol. 4, No… [cited by applicant]
Handa et al., “DeXtreme: Transfer of Agile In-hand Manipulation from Simulation to Reality”, arXiv:2210.13702, Oct. 25, 2022, pp. 1-28. [cited by applicant]
He et al., “Mask R-CNN”, IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2961-2969. [cited by applicant]
Hebecker et al., “Towards Real-World Force-Sensitive Robotic Assembly through Deep Reinforcement Learning in Simulations”, IEEE/ASME International Conference on Advanced Intelligent Mechatronics, 2021, pp. 1045-1051. [cited by applicant]
Hou et al., “Data-efficient Hierarchical Reinforcement Learning for Robotic Assembly Control Applications”, IEEE Transactions on Industrial Electronics, DOI 10.1109/TIE.2020.3038072, 2020, 11 pages. [cited by applicant]
Huang et al., “Fast Peg-and-Hole Alignment Using Visual Compliance”, International Conference on Intelligent Robots and Systems (IROS), Nov. 3-7, 2013, pp. 286-292. [cited by applicant]
Inoue et al., “Deep Reinforcement Learning for High Precision Assembly Tasks”, International Conference on Intelligent Robots and Systems (IROS), Sep. 24-28, 2017, pp. 819-825. [cited by applicant]
Johannink et al., “Residual Reinforcement Learning for Robot Control”, International Conference on Robotics and Automation (ICRA), May 20-24, 2019, pp. 6023-6029. [cited by applicant]
Kimble et al., “Benchmarking Protocols for Evaluating Small Parts Robotic Assembly Systems”, IEEE Robotics and Automation Letters, vol. 5, No. 2, Apr. 2020, pp. 883-889. [cited by applicant]
Kimble et al., “Performance Measures to Benchmark the Grasping, Manipulation, and Assembly of Deformable Objects Typical to Manufacturing Applications”, Frontiers in Robotics and AI, DOI 10.3389/frobt.2022.999348, Nov. … [cited by applicant]
Lambeta et al., “DIGIT: A Novel Design for a Low-Cost Compact High-Resolution Tactile Sensor With Application to In-Hand Manipulation”, IEEE Robotics and Automation Letters, vol. 5, No. 3, Jul. 2020, pp. 3838-3845. [cited by applicant]
Lan et al., “Affine Body Dynamics: Fast, Stable and Intersection-free Simulation of Stiff Materials”, arXiv:2201.10022, Jan. 31, 2022, pp. 1-14. [cited by applicant]
Lee et al., “Learning Quadrupedal Locomotion Over Challenging Terrain”, Science Robotics, DOI: 10.1126/scirobotics.abc5986, vol. 5, Oct. 21, 2020, pp. 1-13. [cited by applicant]
Lee et al., “Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning”, IEEE International Conference on Robotics and Automation (ICRA), May 31-Aug… [cited by applicant]
Lin et al., “Microsoft COCO: Common Objects in Context”, arXiv: 1405.0312, May 1, 2014, pp. 1-14. [cited by applicant]
Lozano-Perez et al., “Automatic Synthesis of Fine-Motion Strategies for Robots”, The International Journal of Robotics Research, vol. 3, No. 1, 1984, pp. 3-24. [cited by applicant]
Luo et al., “Reinforcement Learning on Variable Impedance Controller for High-Precision Robotic Assembly”, International Conference on Robotics and Automation (ICRA), May 20-24, 2019, pp. 3080-3087. [cited by applicant]
Luo et al., “Robust Multi-Modal Policies for Industrial Assembly via Reinforcement Learning and Demonstrations: A Large-Scale Study”, arXiv:2103.11512, Jul. 31, 2021, 10 pages. [cited by applicant]
Luo et al., “Dynamic Experience Replay”, 3rd Conference on Robot Learning, 2019, pp. 1-10. [cited by applicant]
Luo et al., “A Learning Approach to Robot-Agnostic Force-Guided High Precision Assembly”, arXiv:2010.08052, Aug. 2, 2021, 7 pages. [cited by applicant]
Macklin, Miles, “Nvidia Warp: A High-performance Python Framework for GPU Simulation and Graphics”, Retrieved from https://github.com/nvidia/warp, Mar. 2022, 9 pages. [cited by applicant]
Macklin et al., “Local Optimization for Robust Signed Distance Field Collision”, Proceedings of the ACM on Computer Graphics and Interactive Techniques, https://doi.org/10.1145/3384538, vol. 3, No. 1, Article 8, May 202… [cited by applicant]
Makoviichuk et al., “RL Games: A High performance RL Library”, RL Implementations, https://github.com/Denys88/rl_games, May 2022, 18 pages. [cited by applicant]
Makoviychuk et al., “Isaac Gym: High Performance GPU-Based Physics Simulation for Robot Learning”, arXiv:2108.10470, Aug. 25, 2021, pp. 1-32. [cited by applicant]
Margolis et al., “Rapid Locomotion via Reinforcement Learning”, arXiv:2205.02824, May 5, 2022, 12 pages. [cited by applicant]
Morgan et al., “Vision-driven Compliant Manipulation for Reliable, High-Precision Assembly Tasks” arXiv:2106.14070, Jun. 26, 2021, 13 pages. [cited by applicant]
Muratore et al., “Assessing Transferability from Simulation to Reality for Reinforcement Learning”, IEEE Transactions on Pattern Analysis and Machine Intelligence, DOI 10.1109/TPAMI.2019.2952353, Oct. 2, 2019, pp. 1-12. [cited by applicant]
Narang et al., “Factory: Fast Contact for Robotic Assembly”, arXiv:2205.03532, May 7, 2022, 21 pages. [cited by applicant]
Peng et al., “Sim-to-Real Transfer of Robotic Control with Dynamics Randomization”, IEEE International Conference on Robotics and Automation (ICRA), May 21-25, 2018, pp. 3803-3810. [cited by applicant]
Pinto et al., “Asymmetric Actor Critic for Image-Based Robot Learning”, arXiv:1710.06542, Oct. 18, 2017, 8 pages. [cited by applicant]
Rudin et al., “Learning toWalk in Minutes Using Massively Parallel Deep Reinforcement Learning”, 5th Conference on Robot Learning, 2021, pp. 1-10. [cited by applicant]
Schoettler et al., “Meta-Reinforcement Learning for Robotic Industrial Insertion Tasks”, IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), DOI: 10.1109/IROS45743.2020.9340848, Oct. 25-29, 2020,… [cited by applicant]
Schulman et al., “Proximal Policy Optimization Algorithms”, arXiv:1707.06347, Aug. 28, 2017, pp. 1-12. [cited by applicant]
Shao et al., “Learning to Scaffold the Development of Robotic Manipulation Skills”, IEEE International Conference on Robotics and Automation (ICRA), May 31-Aug. 31, 2020, pp. 5671-5677. [cited by applicant]
Si et al., “Taxim: An Example-based Simulation Model for GelSight Tactile Sensors”, arXiv:2109.04027, Dec. 14, 2021, 8 pages. [cited by applicant]
Son et al., “Sim-to-Real Transfer of Bolting Tasks with Tight Tolerance”, IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), DOI: 10.1109/IROS45743.2020.9341644, Oct. 25-29, 2020, pp. 9056-9063. [cited by applicant]
Spector et al., “InsertionNet—A Scalable Solution for Insertion”, IEEE Robotics and Automation Letters, 2021, pp. 2-9. [cited by applicant]
Spector et al., “Deep Reinforcement Learning for Contact-Rich Skills Using Compliant Movement Primitives”, arXiv:2008.13223, Oct. 25, 2020, pp. 1-27. [cited by applicant]
Spector et al., “InsertionNet 2.0: Minimal Contact Multi-Step Insertion Using Multimodal Multiview Sensory Input”, arXiv:2203.01153, Mar. 2, 2022, 7 pages. [cited by applicant]
Tang et al., “Autonomous Alignment of Peg and Hole by Force/Torque Measurement for Robotic Assembly”, IEEE International Conference on Automation Science and Engineering (CASE), Aug. 21-24, 2016, pp. 162-167. [cited by applicant]
Thomas et al., “Learning Robotic Assembly from CAD”, IEEE International Conference on Robotics and Automation (ICRA), May 21-25, 2018, pp. 3524-3531. [cited by applicant]
Tian et al., “Assemble Them All: Physics-Based Planning for Generalizable Assembly by Disassembly”, ACM Transactions on Graphics, https://doi.org/10.1145/3550454.3555525, vol. 41, No. 6, Article 278, Dec. 2022, pp. 278:… [cited by applicant]
Tsai et al., “A New Technique for Fully Autonomous and Efficient 3D Robotics Hand/Eye Calibration”, IEEE Transactions on Robotics and Automation, vol. 5, No. 3, Jun. 1989, pp. 345-358. [cited by applicant]
Vecerik et al., “Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards”, arXiv:1707.08817, Oct. 8, 2018, pp. 1-10. [cited by applicant]
Vecerik et al., “A Practical Approach to Insertion with Variable Socket Position Using Deep Reinforcement Learning”, International Conference on Robotics and Automation (ICRA), May 20-24, 2019, pp. 754-760. [cited by applicant]
Drigalski et al., “Robots Assembling Machines: Learning from the World Robot Summit 2018 Assembly Challenge”, Advanced Robotics, arXiv:1911.05884, vol. 1, No. 1, Jan. 2020, pp. 1-18. [cited by applicant]
Vuong et al., “Learning Sequences of Manipulation Primitives for Robotic Assembly”, arXiv:2011.00778, Mar. 26, 2021, 7 pages. [cited by applicant]
Wang et al., “TACTO: A Fast, Flexible, and Open-source Simulator for High-Resolution Vision-based Tactile Sensors”, IEEE Robotics and Automation Letters, DOI: 10.1109/LRA.2022.3146945, Feb. 10, 2022, pp. 1-8. [cited by applicant]
Wen et al., “You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration”, arXiv:2201.12716, May 6, 2022, 22 pages. [cited by applicant]
Whitney, D. E., “Quasi-Static Assembly of Compliantly Supported Rigid Parts”, Journal of Dynamic Systems, Measurement, and Control, vol. 104, Mar. 1982, pp. 65-77. [cited by applicant]
Wu et al., “Learning Dense Rewards for Contact-Rich Manipulation Tasks”, arXiv:2011.08458, Nov. 17, 2020, 8 pages. [cited by applicant]
Xia et al., “Dynamic Analysis for Peg-In-Hole Assembly With Contact Deformation”, The International Journal of Advanced Manufacturing Technology, vol. 30, DOI 10.1007/s00170-005-0047-4, 2006, pp. 118-128. [cited by applicant]
Xu et al., “Efficient Tactile Simulation with Differentiability for Robotic Manipulation”, 6th Conference on Robot Learning, 2022, pp. 1-11. [cited by applicant]
Xu et al., “Compare Contact Model-based Control and Contact Model-free Learning: A Survey of Robotic Peg-in-hole Assembly Strategies”, IEEE, arXiv:1904.05240, Mar. 2019, pp. 1-15. [cited by applicant]
Yoon et al., “Fast and Accurate Data-Driven Simulation Framework for Contact-Intensive Tight-Tolerance Robotic Assembly Tasks”, arXiv:2202.13098, Feb. 26, 2022, pp. 1-16. [cited by applicant]
Yuan et al., “GelSight: High-Resolution Robot Tactile Sensors for Estimating Geometry and Force”, Sensors, vol. 17, No. 2762, doi:10.3390/s17122762, Nov. 29, 2017, pp. 1-21. [cited by applicant]
Zhang et al., “A Modular Robotic Arm Control Stack for Research: Franka-Interface and FrankaPy” arXiv:2011.02398, Nov. 4, 2020, pp. 1-7. [cited by applicant]
Zhang et al., “Learning Insertion Primitives with Discrete-Continuous Hybrid Action Space for Robotic Assembly Tasks”, arXiv:2110.12618, Oct. 25, 2021, 7 pages. [cited by applicant]
Zhao et al., “Offline Meta-Reinforcement Learning for Industrial Insertion”, arXiv:2110.04276, Sep. 1, 2022, 8 pages. [cited by applicant]
Lee et al., “Making sense of vision and touch: Learning multimodal representations for contact-rich tasks”, arXiv:2305.17110v1, Jul. 28, 2019, 14 pages. [cited by applicant]
Non Final Office Action received for U.S. Appl. No. 18/490,630, dated May 22, 2025, 24 pages. [cited by applicant]
Final Office Action received for U.S. Appl. No. 18/490,630, dated Aug. 27, 2025, 18 pages. [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 18/490,630, dated Nov. 21, 2025, 7 pages. [cited by applicant]