IP Library › Granted Patent US 12,678,960
Granted Patent B2
US 12,678,960 · App. 18/806,621 · Granted Jul 14, 2026

Systems and methods for robotic teleoperation intention estimation

Inventors: Lijing Kuang (La Jolla, CA); Songpo Li (San Jose, CA); Soshi Iba (Mountain View, CA)
Assignee: Honda Motor Co., Ltd.
B25J9/1689
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,678,960
App. No.
18/806,621
Filed
Aug 15, 2024
Granted
Jul 14, 2026
Kind
B2
Art Unit
3656
USPC
700/264
Abstract

Systems and methods for robotic teleoperation are provided. The system receives motion input and gaze data of a human operator performing a robotic teleoperation task, and a plurality of features are extracted from the received data. A recurrent neural network determines latent actions from the extracted features and generates temporally associated action inferences. A model determines inferred intention from a sequenced pair of action inferences, an intention hierarchy, and a probabilistic uncertainty function. A teleoperated robotic element performs a task associated with the inferred intention without a corresponding control signal from the human operator. The recurrent neural network is trained using extracted features and contrastive learning. The intention hierarchy is generated from action inferences. The probabilistic uncertainty function is generated using Bayesian hierarchical modeling. The system and method improves robotic teleoperation by estimating the intention of the human operator.

Claims (73)

1 . A system for robotic teleoperation, comprising:

a robotic teleoperation input configured to:

receive a plurality of motion inputs of a human operator while performing a robotic teleoperation task, and

receive a plurality of gaze data of the human operator while performing the robotic teleoperation task;

a robotic teleoperation feature extractor configured to:

analyze the motion inputs and the gaze data associatively, and

extract a plurality of features from the motion inputs and the gaze data;

a robotic teleoperation neural network configured to:

determine a plurality of latent actions, each associated with one or more extracted features, and

generate a plurality of temporally associated action inferences, each action inference being associated with a determined latent action and associated extracted features;

a robotic teleoperation transition model configured to:

determine elements of a transition matrix associated with a sequenced pair of action inferences;

a robotic teleoperation dynamic transition model configured to:

determine an inferred intention associated with the determined elements of the transition matrix,

wherein the dynamic transition model includes an intention hierarchy associated with a plurality of inferred intentions, and

wherein the inferred intention is based at least in part on an estimated probabilistic uncertainty applied to the intention hierarchy; and

a teleoperated robotic element that performs a task associated with the inferred intention without a corresponding control signal from the human operator.

2 . The system of claim 1 , wherein the robotic teleoperation neural network is further configured to train a recurrent neural network with the extracted features.

3 . The system of claim 1 , wherein the robotic teleoperation transition model generates one or more transition matrices associated with one or more sequenced pairs of action inferences.

4 . The system of claim 3 , wherein the robotic teleoperation transition model is refined using contrastive learning when a sequenced pair of action inferences is not associated with an intended teleoperation of a robotic element.

5 . The system of claim 3 , wherein the robotic teleoperation transition model performs a contrastive loss function to generate weights associated with one or more transition matrices and applies a conditional rule to a transition matrix to redetermine at least one action inference associated with the transition matrix.

6 . The system of claim 1 , wherein the robotic teleoperation dynamic transition model is further configured to:

perform hierarchy discovery using a hierarchical clustering algorithm to generate an intention hierarchy from action inferences;

estimate probabilistic uncertainty of at least a part of the intention hierarchy uncertainty using Bayesian hierarchical modeling; and

generate a likelihood estimation of an inferred intention based on the estimated probabilistic uncertainty applied to the intention hierarchy.

7 . The system of claim 6 , wherein the hierarchical clustering algorithm includes a Ward's linkage algorithm, and wherein Bayesian hierarchical modeling includes one or more of a Markov Chain Monte Carlo and a Variational Inference algorithm.

8 . A method for robotic teleoperation, comprising:

receiving a plurality of gaze data and a plurality of motion inputs associated with a human operator performing a robotic teleoperation task;

extracting a plurality of features from the motion inputs and the gaze data;

determining, by a neural network, a plurality of latent actions, each associated with one or more extracted features;

generating, by the neural network, a plurality of temporally associated action inferences, each action inference being associated with a determined latent action and associated extracted features;

determining elements of a transition matrix associated with a sequenced pair of action inferences;

determining, using a dynamic transition model, an inferred intention associated with the determined elements of the transition matrix,

wherein the dynamic transition model includes an intention hierarchy associated with a plurality of inferred intentions, and

wherein the inferred intention is based at least in part on an estimated probabilistic uncertainty applied to the intention hierarchy; and

performing, by a teleoperated robotic element, a task associated with the inferred intention without a corresponding control signal from the human operator.

9 . The method of claim 8 , wherein neural network is a recurrent neural network and further comprising:

training the recurrent neural network with the extracted features.

10 . The method of claim 8 , further comprising:

generating one or more transition matrices associated with one or more sequenced pairs of action inferences.

11 . The method of claim 10 , further comprising:

refining the robotic teleoperation transition model using contrastive learning when a sequenced pair of action inferences is not associated with an intended teleoperation of a robotic element.

12 . The method of claim 10 , further comprising:

performing a contrastive loss function to generate weights associated with one or more transition matrices; and

applying a conditional rule to a transition matrix to redetermine at least one action inference associated with the transition matrix.

13 . The method of claim 8 , further comprising:

performing hierarchy discovery using a hierarchical clustering algorithm to generate an intention hierarchy from action inferences;

estimating probabilistic uncertainty of at least a part of the intention hierarchy uncertainty using Bayesian hierarchical modeling; and

generating a likelihood estimation of an inferred intention based on the estimated probabilistic uncertainty applied to the intention hierarchy.

14 . The method of claim 13 , wherein the hierarchical clustering algorithm includes a Ward's linkage algorithm, and wherein Bayesian hierarchical modeling includes one or more of a Markov Chain Monte Carlo and a Variational Inference algorithm.

15 . A non-transitory computer readable storage medium storing instructions that when executed by a computer having a processor to perform a method for robotic teleoperation, the method comprising:

receiving a plurality of gaze data and a plurality of motion inputs associated with a human operator performing a robotic teleoperation task;

extracting a plurality of features from the motion inputs and the gaze data;

determining, by a neural network, a plurality of latent actions, each associated with one or more extracted features;

generating, by the neural network, a plurality of temporally associated action inferences, each action inference being associated with a determined latent action and associated extracted features;

determining elements of a transition matrix associated with a sequenced pair of action inferences;

determining, using a dynamic transition model, an inferred intention associated with the determined elements of the transition matrix,

wherein the dynamic transition model includes an intention hierarchy associated with a plurality of inferred intentions, and

wherein the inferred intention is based at least in part on an estimated probabilistic uncertainty applied to the intention hierarchy; and

performing, by a teleoperated robotic element, a task associated with the inferred intention without a corresponding control signal from the human operator.

16 . The non-transitory computer readable storage medium of claim 15 , wherein neural network is a recurrent neural network and further comprising:

training the recurrent neural network with the extracted features.

17 . The non-transitory computer readable storage medium of claim 15 , further comprising:

generating one or more transition matrices associated with one or more sequenced pairs of action inferences.

18 . The non-transitory computer readable storage medium of claim 17 , further comprising:

refining the robotic teleoperation transition model using contrastive learning when a sequenced pair of action inferences is not associated with an intended teleoperation of a robotic element.

19 . The non-transitory computer readable storage medium of claim 17 , further comprising:

performing a contrastive loss function to generate weights associated with one or more transition matrices; and

applying a conditional rule to a transition matrix to redetermine at least one action inference associated with the transition matrix.

20 . The non-transitory computer readable storage medium of claim 15 , further comprising:

performing hierarchy discovery using a hierarchical clustering algorithm to generate an intention hierarchy from action inferences;

estimating probabilistic uncertainty of at least a part of the intention hierarchy uncertainty using Bayesian hierarchical modeling; and

generating a likelihood estimation of an inferred intention based on the estimated probabilistic uncertainty applied to the intention hierarchy.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE 18/133837 PREVIOUSLY RECORDED ON REEL 68304 FRAME 107. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 20, 2024
From: KUANG, LIJING; LI, SONGPO; IBA, SOSHI
To: HONDA MOTOR CO., LTD.
Reel/Frame 068686/0784 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2024
From: KUANG, LIJING; LI, SONGPO; IBA, SOSHI
To: HONDA MOTOR CO., LTD.
Reel/Frame 068304/0107 →
Continuity (2)
Provisional Application 63584070 · Sep 20, 2023
Related Publication 20250091214A1 · Mar 20, 2025
References Cited (53)
US 12444411B1 · Derezhenets · 2025 [cited by examiner]
US 12499879B1 · Soliman · 2025 [cited by examiner]
US 20090088634A1 · Zhao · 2009 [cited by examiner]
US 20180064499A1 · Itkowitz · 2018 [cited by examiner]
US 20190201047A1 · Yates · 2019 [cited by examiner]
US 20210146531A1 · Tremblay · 2021 [cited by examiner]
US 20210170590A1 · Laftchiev · 2021 [cited by examiner]
US 20220258351A1 · Thackston · 2022 [cited by examiner]
US 20220331973A1 · Taniguchi · 2022 [cited by examiner]
US 20230148120A1 · Kranski · 2023 [cited by examiner]
US 20250001605A1 · Felip · 2025 [cited by examiner]
US 20250083325A1 · Cai · 2025 [cited by examiner]
US 20250091214A1 · Kuang · 2025 [cited by examiner]
US 20250108512A1 · Lima · 2025 [cited by examiner]
US 20250217114A1 · Roper, Jr. · 2025 [cited by examiner]
US 20250375891A1 · Wen · 2025 [cited by examiner]
US 20260028045A1 · Bagnell · 2026 [cited by examiner]
CN 107097227A · 2017 [cited by examiner]
CN 108527370A · 2019 [cited by examiner]
CN 110605714A · 2019 [cited by examiner]
CN-107097227-A translation (Year: 2017). [cited by examiner]
CN-108527370-A translation (Year: 2018). [cited by examiner]
CN-110605714-A translation (Year: 2019). [cited by examiner]
Duarte_Action_Alignment_from_Gaze_Cues_in_Human-Human_and_Human-Robot_Interaction_ECCVW_2018_paper (Year: 2018). [cited by examiner]
Abi-Farraj, F., Pacchierotti, C., Arenz, O., Neumann, G., and Giordano, P. R. (2019). A haptic shared-control architecture for guided multi-target robotic grasping. IEEE transactions on haptics, 13(2):270-285. [cited by applicant]
Abi-Farraj, F., Pedemonte, N., and Giordano, P. R. (2016). A visual-based shared control architecture for remote telemanipulation. In 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. … [cited by applicant]
Amor, H. B., Neumann, G., Kamthe, S., Kroemer, O., and Peters, J. (2014). Interaction primitives for human-robot cooperation tasks. In 2014 IEEE international conference on robotics and automation (ICRA), pp. 2831-2837.… [cited by applicant]
Balachandran, R., Mishra, H., Cappelli, M., Weber, B., Secchi, C., Ott, C., and Albu-Schaeffer, A. (2020). Adaptive authority allocation in shared control of robots using bayesian filters. In 2020 IEEE International Con… [cited by applicant]
Campbell, J., Stepputtis, S., and Amor, H. B. (2019). Probabilistic multimodal modeling for human-robot interaction tasks. arXiv preprint arXiv:1908.04955. [cited by applicant]
Carmichael, M. G., Aldini, S., Khonasty, R., Tran, A., Reeks, C., Liu, D., Waldron, K. J., and Dissanayake, G. (2019). The anbot: An intelligent robotic co-worker for industrial abrasive blasting. In 2019 IEEE/RSJ Inter… [cited by applicant]
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. (2020). A simple framework for contrastive learning of visual representations. In International conference on machine learning, pp. 1597-1607. PMLR. [cited by applicant]
K. Darvish, L. Penco, J. Ramos, R. Cisneros, J. Pratt, E. Yoshida, S. Ivaldi, and D. Pucci, “Teleoperation of humanoid robots: A survey,” IEEE Transactions on Robotics, 2023. [cited by applicant]
T. Gao, X. Yao, and D. Chen, “Simcse: Simple contrastive learning of sentence embeddings,” arXiv preprint arXiv:2104.08821, 2021. [cited by applicant]
Hamaya, M., Matsubara, T., Noda, T., Teramae, T., and Morimoto, J. (2017). Learning assistive strategies for exoskeleton robots from user-robot physical interaction. Pattern Recognition Letters, 99:67-76. [cited by applicant]
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9729-9738, 20… [cited by applicant]
Huang, Z., Mun, Y.- J., Li, X., Xie, Y., Zhong, N., Liang, W., Geng, J., Chen, T., and Driggs-Campbell, K. (2023). Hierarchical intention tracking for robust human-robot collaboration in industrial assembly tasks. In 20… [cited by applicant]
S. Jain and B. Argall, “Probabilistic human intent recognition for shared autonomy in assistive robotics,” ACM Transactions on Human-Robot Interaction (THRI), vol. 9, No. 1, pp. 1-23, 2019. [cited by applicant]
Khosla, P., Teterwak, P., Wang, C., Sarna, A., Tian, Y., Isola, P., Maschinot, A., Liu, C., and Krishnan, D. (2020). Supervised contrastive learning. Advances in neural information processing systems, 33:18661-18673. [cited by applicant]
Koert, D., Pajarinen, J., Schotschneider, A., Trick, S., Rothkopf, C., and Peters, J. (2019). Learning intention aware online adaptation of movement primitives. IEEE Robotics and Automation Letters, 4(4):3719-3726. [cited by applicant]
Lai, Y., Paul, G., Cui, Y., and Matsubara, T. (2022). User intent estimation during robot learning using physical human robot interaction primitives. Autonomous Robots, 46(2):421-436. [cited by applicant]
S. Li, M. Bowman, H. Nobarani, and X. Zhang, “Inference of manipulation intent in teleoperation for robotic assistance,” Journal of Intelligent & Robotic Systems, vol. 99, pp. 29-43, 2020. [cited by applicant]
D. P. Losey, C. G. McDonald, E. Battaglia, and M. K. O'Malley, “A review of intent detection, arbitration, and communication aspects of shared control for physical human-robot interaction,” Applied Mechanics Reviews, vo… [cited by applicant]
Murtagh, F. and Contreras, P. (2012). Algorithms for hierarchical clustering: an overview. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 2(1):86-97. [cited by applicant]
Rumelhart, D. E., Hinton, G. E., Williams, R. J., et al. (1985). Learning internal representations by error propagation. [cited by applicant]
Selvaggio, M., Moccia, R., Ficuciello, F., Siciliano, B., et al. (2019). Haptic-guided shared control for needle grasping optimization in minimally invasive robotic surgery. In 2019 IEEE/RSJ International Conference on … [cited by applicant]
Shen, Y., Mo, X., Krisciunas, V., Hanson, D., and Shi, B. E. (2023). Intention estimation via gaze for robot guidance in hierarchical tasks. In Annual Conference on Neural Information Processing Systems, pp. 140-164. PM… [cited by applicant]
Singh, S., Arora, C., and Jawahar, C. (2016). First person action recognition using deep learned descriptors. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2620-2628. [cited by applicant]
A. Stooke, K. Lee, P. Abbeel, and M. Laskin, “Decoupling representation learning from reinforcement learning,” in International Conference on Machine Learning, pp. 9870-9879, PMLR, 2021. [cited by applicant]
Tanwani, A. K. and Calinon, S. (2017). A generative model for intention recognition and manipulation assistance in teleoperation. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 4… [cited by applicant]
A. Yan, Z. He, X. Lu, J. Du, E. Chang, A. Gentili, J. McAuley, and C.-N. Hsu, “Weakly supervised contrastive learning for chest x-ray report generation,” arXiv preprint arXiv:2109.12242, 2021. [cited by applicant]
A. Yan, Z. He, J. Li, T. Zhang, and J. McAuley, “Personalized showcases: Generating multi-modal explanations for recommendations,” in Proceedings of the 46th International Acm Sigir Conference on Research and Developmen… [cited by applicant]
J. Yang, J. Duan, S. Tran, Y. Xu, S. Chanda, L. Chen, B. Zeng, T. Chilimbi, and J. Huang, “Vision-language pre- training with triple contrastive learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision an… [cited by applicant]
Zhu, Y., Fusano, K., Aoyama, T., and Hasegawa, Y. (2023). Intention-reflected predictive display for operability improvement of time-delayed teleoperation system. ROBOMECH Journal, 10(1):1-11. [cited by applicant]