IP Library › Granted Patent US 12,629,839
Granted Patent B2
US 12,629,839 · App. 18/825,339 · Granted May 19, 2026

System and methods for robotic teleoperation intention estimation

Inventors: Mingyu Cai (Tustin, CA); Karankumar Ashokbhai Patel (San Jose, CA); Soshi Iba (Mountain View, CA); Songpo Li (San Jose, CA)
Assignee: Honda Motor Co., Ltd.
B25J9/1689G05B13/027G05B2219/40174
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,629,839
App. No.
18/825,339
Granted
May 19, 2026
Kind
B2
Abstract

A system and method for robotic teleoperation enable a teleoperated robotic element to perform a sequence of actions based on intention estimation, eliminating the need for continuous human control. The system includes a robotic teleoperation input that receives motion inputs and gaze data from a human operator performing a robotic teleoperation task. A robotic teleoperation feature extractor analyzes and processes the motion inputs and the gaze data into sequential input data. A multi-window model assigns hierarchical prediction windows to the input data, generating windowed sequential data. A hierarchical neural network processes the windowed sequential data to determine low-level action intentions and high-level task intentions, and generates an intention estimation based on the low-level action intentions and the high-level task intentions. A hierarchical dependency model incorporates hierarchical dependent loss to refine the intention estimation.

Claims (54)

1 . A system for robotic teleoperation, comprising:

a robotic teleoperation input configured to receive inputs of a human operator performing a robotic teleoperation task;

a robotic teleoperation feature extractor configured to:

analyze the inputs,

extract features from the inputs, and

process the extracted features into sequential input data;

a multi-window model configured to generate windowed sequential data based on the sequential input data;

a hierarchical neural network configured to determine intentions of the human operator based on the windowed sequential data, and generate an intention estimation based on the intentions; and

a teleoperated robotic element configured to perform a sequence of actions associated with the intention estimation.

2 . The system of claim 1 , wherein the hierarchical neural network is further configured to train a recurrent neural network based on the intention estimation.

3 . The system of claim 1 , wherein the hierarchical neural network includes:

a root neural network that processes the windowed sequential data to generate a root latent space representation;

a task encoder that generates task-level representations from the root latent space representation; and

an action encoder that generates action-level representations from the root latent space representation, conditioned on the task-level representations.

4 . The system of claim 1 , wherein the multi-window model implements a sequential mask vector to discard unnecessary horizons of the sequential input data.

5 . The system of claim 1 , further including an assistive control module configured to automatically switch to an autonomous control mode in a presence of communication latency to execute the sequence of actions associated with the intention estimation based on a confidence calculation associated with the intention estimation.

6 . The system of claim 1 , wherein the hierarchical neural network is trained by a dataset including labeled teleoperation sequences, including various assembly tasks and corresponding action sequences.

7 . The system of claim 1 , further including a hierarchical dependency model configured to incorporate hierarchical dependent loss to refine the intention estimation.

8 . The system of claim 1 , wherein the multi-window model is configured to assign hierarchical prediction windows having varying lengths to the sequential input data in generating the windowed sequential data.

9 . The system of claim 1 , wherein the hierarchical neural network includes a Slow-Fast model for processing the inputs, which includes:

a Slow Pathway operating at a predetermined low frame rate to capture spatial semantics as spatial information;

a Fast Pathway operating at a predetermined high frame rate to capture motion as temporal information; and

a multi-perception layer configured to integrate the spatial information and the temporal information to generate task and action embeddings.

10 . A method for robotic teleoperation, comprising:

receiving inputs of a human operator performing a robotic teleoperation task;

analyzing the inputs;

extracting features from the inputs;

processing the extracted features into sequential input data;

generating windowed sequential data based on the sequential input data;

determining, with a hierarchical neural network, intentions of the human operator based on the windowed sequential data, and generating an intention estimation based on the intentions; and

performing, with a teleoperated robotic element, a sequence of actions associated with the intention estimation.

11 . The method of claim 10 , further including training a recurrent neural network based on the intention estimation.

12 . The method of claim 10 , wherein the hierarchical neural network includes:

a root neural network that processes the windowed sequential data to generate a root latent space representation;

a task encoder that generates task-level representations from the root latent space representation; and

an action encoder that generates action-level representations from the root latent space representation, conditioned on the task-level representations.

13 . The method of claim 10 , further including implementing a sequential mask vector to discard unnecessary horizons of the sequential input data.

14 . The method of claim 10 , further including automatically switching to an autonomous control mode of the teleoperated robotic element in a presence of communication latency to execute the sequence of actions associated with the intention estimation based on a confidence calculation associated with the intention estimation.

15 . The method of claim 10 , wherein the hierarchical neural network is trained by a dataset including labeled teleoperation sequences, including various assembly tasks and corresponding action sequences.

16 . The method of claim 10 , further including incorporating hierarchical dependent loss to refine the intention estimation.

17 . The method of claim 16 , wherein the hierarchical dependent loss is combined with a classification entropy loss to form a total loss function for training the hierarchical neural network, and the total loss function includes a weighted sum of weights of the classification entropy loss and the hierarchical dependent loss.

18 . The method of claim 10 , further including assigning hierarchical prediction windows having varying lengths to the sequential input data in generating the windowed sequential data.

19 . The method of claim 10 , wherein the hierarchical neural network includes a Slow-Fast model for processing the inputs, which includes:

a Slow Pathway operating at a predetermined low frame rate to capture spatial semantics as semantic information;

a Fast Pathway operating at a predetermined high frame rate to capture motion as temporal information; and

a multi-perception layer configured to integrate the spatial information and the temporal information to generate task and action embeddings.

20 . A non-transitory computer readable storage medium storing instructions that when executed by a computer having a processor to perform a method for robotic teleoperation, the method comprising:

receiving inputs of a human operator performing a robotic teleoperation task;

analyzing the inputs;

extracting features from the inputs;

processing the extracted features into sequential input data;

generating windowed sequential data based on the sequential input data;

determining, with a hierarchical neural network, intentions of the human operator based on the windowed sequential data, and generating an intention estimation based on the intentions; and

performing, with a teleoperated robotic element, a sequence of actions associated with the intention estimation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2024
From: CAI, MINGYU; PATEL, KARANKUMAR ASHOKBHAI; IBA, SOSHI; LI, SONGPO
To: HONDA MOTOR CO., LTD.
Reel/Frame 068496/0938 →
Continuity (2)
Provisional Application 63582337 · Sep 13, 2023
Related Publication 20250083325A1 · Mar 13, 2025
References Cited (53)
US 11351671B2 · Thackston · 2022 [cited by examiner]
US 12444411B1 · Derezhenets · 2025 [cited by examiner]
US 12448001B2 · Bagnell · 2025 [cited by examiner]
US 12499879B1 · Soliman · 2025 [cited by examiner]
US 20220314449A1 · Hosomi · 2022 [cited by examiner]
US 20230226698A1 · Chaki · 2023 [cited by examiner]
US 20230234231A1 · Dong · 2023 [cited by examiner]
US 20230286159A1 · Kondapally · 2023 [cited by examiner]
US 20240051143A1 · Dong · 2024 [cited by examiner]
US 20240149458A1 · Watabe · 2024 [cited by examiner]
US 20250001605A1 · Felip · 2025 [cited by examiner]
US 20250091214A1 · Kuang · 2025 [cited by examiner]
US 20250108512A1 · Lima · 2025 [cited by examiner]
US 20250205896A1 · Naramura · 2025 [cited by examiner]
US 20250217114A1 · Roper, Jr. · 2025 [cited by examiner]
US 20250375891A1 · Wen · 2025 [cited by examiner]
Darvish, Kourosh, et al. “A hierarchaical architecture for human-robot cooperation processes.” IEEE Transactions on Robotics 37.2 (2020): 567-586. [cited by applicant]
A. Graves and A. Graves, “Long short-term memory,” Supervised sequence labelling with recurrent neural networks, pp. 37-45, 2012. [cited by applicant]
D. Aarno and D. Kragic, “Motion intention recognition in robot assisted applications,” Robotics and Autonomous Systems, vol. 56, No. 8, pp. 692-705, 2008. [cited by applicant]
F. Abi-Farraj, T. Osa, N. P. J. Peters, G. Neumann, and P. R. Giordano, “A learning-based shared control architecture for interactive task execution,” in 2017 IEEE international conference on robotics and automation (IC… [cited by applicant]
K. I. Alevizos, C. P. Bechlioulis, and K. J. Kyriakopoulos, “Physical human-robot cooperation based on robust motion intention estimation,” Robotica, vol. 38, No. 10, pp. 1842-1866, 2020. [cited by applicant]
A. Belardinelli, A. R. Kondapally, D. Ruiken, D. Tanneberg, and T. Watabe, “Intention estimation from gaze and motion features for human-robot shared-control object manipulation,” in 2022 IEEE/RSJ International Conferen… [cited by applicant]
E. De Momi, L. Kranendonk, M. Valenti, N. Enayati, and G. Ferrigno, “A neural network-based approach for trajectory planning in robot-human handover tasks,” Frontiers in Robotics and AI, vol. 3, p. 34, 2016. [cited by applicant]
A. D. Dragan and S. S. Srinivasa, “A policy-blending formalism for shared control,” The International Journal of Robotics Research, vol. 32, No. 7, pp. 790-805, 2013. [cited by applicant]
C. Fang, L. Peternel, A. Seth, M. Sartori, K. Mombaur, and E. Yoshida, “Human modeling in physical human-robot interaction: A brief survey,” IEEE Robotics and Automation Letters, 2023. [cited by applicant]
J. T. Feddema and O. R. Mitchell, “Vision-guided servoing with feature-based trajectory generation (for robots),” IEEE Transactions on Robotics and Automation, vol. 5, No. 5, pp. 691-700, 1989. [cited by applicant]
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 6202-6211. [cited by applicant]
D. Gao, W. Yang, H. Zhou, Y. Wei, Y. Hu, and H. Wang, “Deep hierarchical classification for category prediction in e-commerce system,” arXiv preprint arXiv:2005.06692, 2020. [cited by applicant]
J.-H. Han, S.-H. Choi, and J.-H. Kim, “Interactive human intention reading by learning hierarchical behavior knowledge networks for human-robot interaction,” ETRI Journal, vol. 38, No. 6, pp. 1229-1239, 2016. [cited by applicant]
K. Hauser, “Recognition, prediction, and planning for assisted teleoperation of freeform tasks,” Autonomous Robots, vol. 35, pp. 241-254, 2013. [cited by applicant]
B. Hayes and B. Scassellati, “Autonomously constructing hierarchical task networks for planning and human-robot collaboration,” in 2016 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2016, pp. 54… [cited by applicant]
S. Holtzen, Y. Zhao, T. Gao, J. B. Tenenbaum, and S.-C. Zhu, “Inferring human intent from video by sampling hierarchical plans,” in 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, … [cited by applicant]
Z. Huang, Y.-J. Mun, X. Li, Y. Xie, N. Zhong, W. Liang, J. Geng, T. Chen, and K. Driggs-Campbell, “Hierarchical intention tracking for robust human-robot collaboration in industrial assembly tasks,” in 2023 IEEE Interna… [cited by applicant]
G. Li, Z. Li, and Z. Kan, “Assimilation control of a robotic exoskeleton for physical human-robot interaction,” IEEE Robotics and Automation Letters, vol. 7, No. 2, pp. 2977-2984, 2022. [cited by applicant]
Y. Li and S. S. Ge, “Human-robot collaboration based on motion intention estimation,” IEEE/ASME Transactions on Mechatronics, vol. 19, No. 3, pp. 1007-1014, 2013. [cited by applicant]
W. Lu, Z. Hu, and J. Pan, “Human-robot collaboration using variable admittance control and human intention prediction,” in 2020 IEEE 16th International Conference on Automation Science and Engineering (CASE). IEEE, 2020… [cited by applicant]
S. Manschitz and D. Ruiken, “Shared autonomy for intuitive teleoperation,” ICRA Workshop: Shared Autonomy in Physical Human-Robot Interaction: Adaptability and Trust, May 2022. [cited by applicant]
T. Marcucci, M. Petersen, D. von Wrangel, and R. Tedrake, “Motion planning around obstacles with convex optimization,” arXiv preprint arXiv:2205.04422, 2022. [cited by applicant]
L. R. Medsker and L. Jain, “Recurrent neural networks,” Design and Applications, vol. 5, No. 64-67, p. 2, 2001. [cited by applicant]
D. Nicolis, A. M. Zanchettin, and P. Rocco, “Human intention estimation based on neural networks for enhanced collaboration with robots,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS… [cited by applicant]
L. R. Rabiner, “A tutorial on hidden markov models and selected applications in speech recognition,” Proceedings of the IEEE, vol. 77, No. 2, pp. 257-286, 1989. [cited by applicant]
L. Rabiner and B. Juang, “An introduction to hidden markov models,” ieee assp magazine, vol. 3, No. 1, pp. 4-16, 1986. [cited by applicant]
L. Rozo, S. Calinon, D. G. Caldwell, P. Jimenez, and C. Torras, “Learning physical collaborative robot behaviors from human demonstrations,” IEEE Transactions on Robotics, vol. 32, No. 3, pp. 513-527, 2016. [cited by applicant]
P. Schydlo, M. Rakovic, L. Jamone, and J. Santos-Victor, “Anticipation in human-robot cooperation: A recurrent neural network approach for multiple action sequences prediction,” in 2018 IEEE International Conference on … [cited by applicant]
T. B. Sheridan, “Teleoperation, telerobotics and telepresence: A progress report,” Control Engineering Practice, vol. 3, No. 2, pp. 205-214, 1995. [cited by applicant]
L. X. Shi, A. Sharma, T. Z. Zhao, and C. Finn, “Waypointbased imitation learning for robotic manipulation,” arXiv preprint arXiv:2307.14326, 2023. [cited by applicant]
Sohn, Sungryull, Junhyuk Oh, and Honglak Lee. “Hierarchical reinforcement learning for zero-shot generalization with subtask dependencies.” Advances in neural information processing systems 31 (2018). [cited by applicant]
A. K. Tanwani and S. Calinon, “A generative model for intention recognition and manipulation assistance in teleoperation,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2017, … [cited by applicant]
A. K. Tanwani and S. Calinon, “Learning robot manipulation tasks with task-parameterized semitied hidden semi-markov model,” IEEE Robotics and Automation Letters, vol. 1, No. 1, pp. 235-242, 2016. [cited by applicant]
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017. [cited by applicant]
R. Wilcox, S. Nikolaidis, and J. Shah, “Optimization of temporal dynamics for adaptive human-robot interaction in assembly manufacturing,” Robotics, vol. 8, No. 441, pp. 10-15, 2013. [cited by applicant]
W. Yu, R. Alqasemi, R. Dubey, and N. Pernalete, “Telemanipulation assistance based on motion intention recognition,” in Proceedings of the 2005 IEEE international conference on robotics and automation. IEEE, 2005, pp. 1… [cited by applicant]
C. Zhu, Q. Cheng, and W. Sheng, “Human intention recognition in smart assisted living systems using a hierarchical hidden markov model,” in 2008 IEEE International Conference on Automation Science and Engineering. IEEE,… [cited by applicant]