IP Library Granted Patent US 12,370,678
Granted Patent B2
US 12,370,678 · App. 17/716,520 · Granted Jul 29, 2025

Demonstration-conditioned reinforcement learning for few-shot imitation

Inventors: Theo Cachet (Grenoble, FR); Christopher Dance (Grenoble, FR); Julien Perez (Grenoble, FR)
Assignee: NAVER CORPORATION
B25J9/163G05B13/0265G05D1/0088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,370,678
App. No.
17/716,520
Granted
Jul 29, 2025
Kind
B2
Abstract

A computer-implemented method for performing few-shot imitation is disclosed. The method comprises obtaining at least one set of training data, wherein each set of training data is associated with a task and comprises (i) one of samples of rewards and a reward function, (ii) one of samples of state transitions and a transition distribution, and (iii) a set of first demonstrations, training a policy network embodied in an agent using reinforcement learning by inputting at least one set of first demonstrations of the at least one set of training data into the policy network, and by maximizing a risk measure or an average return over the at least one set of first demonstrations of the at least one set of training data based on respective one or more reward functions or respective samples of rewards, obtaining a set of second demonstrations associated with a new task, and inputting the set of second demonstrations and an observation of a state into the trained policy network for performing the new task.

Claims (66)

1. A method, performed by a processor and memory, embodied in an agent comprising a trained policy network for performing at least one task, the method performed using the agent comprising:

obtaining a set of first demonstrations and an observation as input for the trained policy network;

inputting the set of first demonstrations and the observation into the trained policy network for performing a task associated with the first demonstrations; and

determining, by the trained policy network, at least one action to be performed via at least one of an actuator of the agent and a motor of the agent based on the inputted set of first demonstrations and the inputted observation,

wherein the trained policy network is trained using reinforcement learning.

2. The method of claim 1 further comprising determining at least one action comprising at least one of controlling a robot or a part of the robot, controlling a machine, controlling a vehicle, and manipulating a state of an environment.

3. The method of claim 1 , wherein the trained policy network has the transformer architecture with axial attention.

4. The method of claim 1 , wherein the trained policy network includes at least one of a first self-attention module for processing the set of first demonstrations, a second self-attention module for processing the observation, and a cross-attention module for processing the set of first demonstrations and the observation.

5. The method of claim 1 wherein a demonstration of the set of first demonstrations includes a sequence of observations, wherein each of the observations includes at least one of a state-action pair, a state, a position, an image, and a sensor measurement.

6. The method of claim 1 wherein:

the task includes a manipulation task for manipulating an object by a robot;

the observation includes information on one or more parts of the robot;

the set of first demonstrations includes a sequence including at least one of positions and orientations of the one or more parts of the robot; and

the method further includes controlling at least one actuator of the robot based on the determined action to be performed.

7. The method of claim 1 wherein:

the task includes a navigation task for navigating a robot;

the observation includes information on one or more parts of the robot;

the set of first demonstrations includes a sequence of positions of the robot; and

the method further includes controlling at least one actuator of the robot based on the determined action to be performed.

8. A computer-implemented method for performing few-shot imitation, the computer implemented method comprising:

obtaining at least one set of training data, wherein each set of training data is associated with a task and includes (i) at least one of samples of rewards and a reward function, (ii) at least one of samples of state transitions and a transition distribution, and (iii) a set of first demonstrations;

training a policy network embodied in an agent using reinforcement learning by:

inputting at least one set of first demonstrations of the at least one set of training data into the policy network; and

maximizing a risk measure or an average return over the at least one set of first demonstrations of the at least one set of training data based on respective one or more reward functions or respective samples of rewards; and

after the training:

obtaining a set of second demonstrations associated with a new task not included in the training data; and

inputting the set of second demonstrations and an observation of a state into the trained policy network for performing the new task.

9. The computer-implemented method of claim 8 , wherein the policy network includes one of:

the transformer architecture with axial attention; and

at least one of a first self-attention module configured to process the set of second demonstrations, a second self-attention module configured to process the observation of the state, and a cross-attention module configured to process the set of second demonstrations and the observation of the state.

10. The computer-implemented method of claim 8 wherein the policy network includes the transformer architecture with axial attention and the computer-implemented method further comprises at least one of:

encoding the inputted at least one set of first demonstrations as a first multidimensional tensor and applying attention by a first transformer of the policy network is along a single axis of the first multidimensional tensor; and

encoding the inputted set of second demonstrations as a second multidimensional tensor and applying attention of a second transformer of the policy network along a single axis of the second multidimensional tensor.

11. The computer-implemented method of claim 8 , wherein at least one of:

the obtaining a set of second demonstrations associated with a new task and the inputting the set of second demonstrations and an observation of a state into the trained policy network for performing the new task are performed at inference time; and

the obtaining at least one set of training data and training a policy network using reinforcement learning are performed during training time.

12. The computer-implemented method claim 8 , wherein the inputting the at least one set of first demonstrations of the at least one set of training data into the policy network to train the policy network includes inputting at least one of a state of the agent, a state-action pair, and an observation-action history into the policy network for training the policy network.

13. The computer-implemented method of claim 8 , wherein the at least one set of first demonstrations of the at least one set of training data include demonstrations of at least two tasks, and wherein maximizing the average return over the at least one set of first demonstrations of the at least one set of training data includes maximizing an average cumulative reward over the at least two tasks.

14. A system comprising:

a control module including a policy network that:

is trained based on a set of training tasks; and

includes the transformer architecture with axial attention on a single axis of a multi-dimensional tensor generated by the transformer architecture; and

a training module configured to:

input to the policy network a set of demonstrations for a task that is different than the training tasks; and

train weight parameters of encoder modules of the transformer architecture based on the single axis of the multi-dimensional tensor generated based on the input set of demonstrations.

15. The system of claim 14 wherein the training module is configured to train the weight parameters of the encoder modules based on maximizing an average return of the policy network.

16. The system of claim 14 wherein the control module is configured to selectively actuate an actuator based on an output of the policy network.

17. The system of claim 14 wherein the transformer architecture includes an encoding module configured to generate the multi-dimensional tensor based on the set of demonstrations.

18. The system of claim 14 wherein each demonstration of the set of demonstrations includes a time series of observations.

19. The system of claim 18 wherein the time series of observations are of random lengths.

20. The system of claim 18 wherein each observation includes at least one of:

a state-action pair;

a state;

a position;

an image; and

a measurement.

21. The system of claim 14 wherein the task is manipulating an object, and wherein the set of demonstrations includes a sequence of positions and orientations of a robot.

22. The system of claim 14 wherein the task includes navigating toward a target position, and wherein the set of demonstrations includes a sequence of positions of a navigating robot.

23. The system of claim 14 wherein the policy network includes L encoder layers connected in series, wherein L is an integer greater than one.

24. The system of claim 23 wherein the policy network further includes L decoder layers configured to determine an action based on an output of the L encoder layers.

25. The system of claim 14 further comprising a processor executing instructions stored in a memory, wherein the instructions stored in the memory further comprise instructions for the control module, including the policy network, and the training module.

26. The system of claim 25 wherein the instructions further comprise instructions for training the policy network with the training tasks using reinforcement learning.

27. The system of claim 26 wherein the instructions further comprise instructions for an agent, including the policy network, that is configured to determine at least one action to be performed based the set of demonstrations for the task that is different than the training tasks.

28. The system of claim 27 wherein the at least one action is a navigation action.

29. The system of claim 28 wherein the agent determines the navigation action for a robot.

30. The system of claim 26 wherein the instructions using reinforcement learning use proximal policy optimization.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2024
From: NAVER LABS CORPORATION
To: NAVER CORPORATION
Reel/Frame 068820/0495 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2022
From: CACHET, THÉO; DANCE, CHRISTOPHER; PEREZ, JULIEN
To: NAVER CORPORATION; NAVER LABS CORPORATION
Reel/Frame 059546/0644 →
Priority Claims (1)
EP 21305799 · Jun 10, 2021 · regional
Continuity (1)
Related Publication 20220395975A1 · Dec 15, 2022
References Cited (97)
US 10452978B2 · Shazeer et al. · 2019 [cited by applicant]
US 20190102676A1 · Nazari · 2019 [cited by examiner]
US 20210205988A1 · James · 2021 [cited by examiner]
US 20210308863A1 · Liu · 2021 [cited by examiner]
US 20220105624A1 · Kalakrishnan · 2022 [cited by examiner]
US 20220297303A1 · Khansari Zadeh · 2022 [cited by examiner]
US 20230256593A1 · Zolna · 2023 [cited by examiner]
Abbeel, P. and Ng, A. Y. Apprenticeship learning via inverse reinforcement learning. In Proceedings of the 21st International Conference on Machine Learning (ICML), pp. 1. Association for Computing Machinery, 2004. [cited by applicant]
Allan Zhou, Eric Jang, Daniel Kappler, Alexander Herzog, Mohi Khansari, Paul Wohlhart, Yunfei Bai, Mrinal Kalakrishnan, Sergey Levine, and Chelsea Finn. Watch, try, learn: Meta-learning from demonstrations and rewards. … [cited by applicant]
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Man'e, D. Concrete problems in AI safety. arXiv preprint arXiv:1606.06565, 2016. [cited by applicant]
Anirban Santara, Abhishek Naik, Balaraman Ravindran, Dipankar Das, Dhee-vatsa Mudigere, Sasikanth Avancha, and Bharat Kaul. RAIL: Risk-averse imitation learning. arXiv preprint arXiv:1707.06658, 2017. [cited by applicant]
Anirudha Majumdar and Marco Pavone. How should a robot assess risk? towards an axiomatic theory of risk in robotics. In Robotics Research, pp. 75{84. Springer, 2020. [cited by applicant]
Arora, S., Du, S. S., Kakade, S., Luo, Y., and Saunshi, N. Provable representation learning for imitation learning via bi-level optimization. In International Conference on Machine Learning, ICML, arXiv preprint arXiv:2… [cited by applicant]
Ashvin Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine. Visual reinforcement learning with imagined goals. In Advances in Neural Information Processing Systems, NeurIPS, pp. 9209-9220, 201… [cited by applicant]
Avi Singh, Eric Jang, Alexander Irpan, Daniel Kappler, Murtaza Dalal, Sergey Levine, Mohi Khansari, and Chelsea Finn. Scalable multi-task imitation learning with autonomous improvement. In IEEE International Conference … [cited by applicant]
Alessandro Bonardi, Stephen James, and Andrew J. Davison. Learning one-shot imitation from humans without humans. IEEE Robotics and Automation. Letters, 5(2):3533-3539, 2020. [cited by applicant]
Boyd, R., Richerson, P. J., and Henrich, J. The cultural niche: Why social learning is essential for human adaptation. Proceedings of the National Academy of Sciences, 108:10918-10925, 2011. [cited by applicant]
Brown, D. S., Goo, W., Nagarajan, P., and Niekum, S. Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations. arXiv preprint arXiv:1904.06387, 2019. [cited by applicant]
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. … [cited by applicant]
Cachet Théo et al: “Transformer-based Meta-Imitation Learning for Robotic Manipulation”, 3rd Workshop on Robot Learning, Thirty-fourth Conference on Neural Information Processing Systems (NeurIPS), virtual only conferen… [cited by applicant]
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. End-to-end object detection with transformers. In European Conference on Computer Vision, ECCV, vol. 12346 of Lecture Notes in Computer S… [cited by applicant]
Charpentier, C. J., Iigaya, K., and O'Doherty, J. P. A neurocomputational account of arbitration between choice imitation and goal emulation during human observational learning. Neuron, 106(4):687-699, 2020. [cited by applicant]
Cheng, Z., Liu, L., Liu, A., Sun, H., Fang, M., and Tao, D. On the guaranteed almost equivalence between imitation learning from observation and demonstration. arXiv preprint arXiv:2010.08353, 2020. [cited by applicant]
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y. Safe exploration in continuous action spaces. arXiv, abs/1801.08757, 2018. [cited by applicant]
Daniel S. Brown, Wonjoon Goo, and Scott Niekum. Better-than-demonstrator imitation learning via automatically-ranked demonstrations. In Conference on Robot Learning, CoRL, vol. 100 of Proceedings of Machine Learning Res… [cited by applicant]
Dasari, S. and Gupta, A. Transformers for one-shot visual imitation. In Conference on Robot Learning, CoRL, 2020. [cited by applicant]
De-An Huang, Suraj Nair, Danfei Xu, Yuke Zhu, Animesh Garg, Li Fei-Fei, Silvio Savarese, and Juan Carlos Niebles. Neural task graphs: Generalizing to unseen tasks from a single video demonstration. In IEEE Conference on… [cited by applicant]
Diana Borsa, Bilal Piot, Remi Munos, and Olivier Pietquin. Observational learning by reinforcement learning. arXiv preprint arXiv:1706.06617, 2017. [cited by applicant]
Dibya Ghosh, Abhishek Gupta, and Sergey Levine. Learning actionable representations with goal conditioned policies. In International Conference on Learning Representations, ICLR, 2019. [cited by applicant]
Ding, Y., Florensa, C., Abbeel, P., and Phielipp, M. Goal conditioned imitation learning. In Advances in Neural Information Processing Systems, NeurIPS, pp. 15298-15309, 2019. [cited by applicant]
Dmytro Korenkevych, A. Rupam Mahmood, Gautham Vasan, and James Bergstra. Autoregressive policies for continuous control deep reinforcement learning. In International Joint Conference on Articial Intelligence, IJCAI, pp.… [cited by applicant]
Duan, Y., Andrychowicz, M., Stadie, B., Ho, J., Schneider, J., Sutskever, I., Abbeel, P., and Zaremba, W. One-shot imitation learning. In Advances in Neural Information Processing Systems, NIPS, pp. 1087-1098, 2017. [cited by applicant]
Dulac-Arnold, G., Mankowitz, D., and Hester, T. Challenges of real-world reinforcement learning. arXiv preprint arXiv:1904.12901, 2019. [cited by applicant]
Finn, C., Abbeel, P., and Levine, S. Model-agnostic metalearning for fast adaptation of deep networks. In International Conference on Machine Learning, pp. 1126-1135. PMLR, 2017. [cited by applicant]
Finn, C., Levine, S., and Abbeel, P. Guided cost learning: Deep inverse optimal control via policy optimization. In International Conference on Machine Learning (ICML), pp. 49-58, 2016. [cited by applicant]
Finn, C., Yu, T., Zhang, T., Abbeel, P., and Levine, S.: One-shot visual imitation learning via meta-learning, arXiv preprint arXiv:1709.04905, 2017. [cited by applicant]
Garcia, J. and Fernandez, F. Safe exploration of state and action spaces in reinforcement learning. Journal of Artificial Intelligence Research, 45:515-564, 2012. [cited by applicant]
Ghosh, D., Gupta, A., and Levine, S. Learning actionable representations with goal conditioned policies. In International Conference on Learning Representations, ICLR, 2019. [cited by applicant]
Glorot, X. and Bengio, Y. Understanding the difficulty of training deep feedforward neural networks. In International Conference on Artificial Intelligence and Statistics, AISTATS, 2010. [cited by applicant]
Goo, W. and Niekum, S. One-shot learning of multi-step tasks from observation via activity localization in auxiliary video. In International Conference on Robotics and Automation, ICRA, pp. 7755-7761, 2019. [cited by applicant]
Gu, Albert, Goael, Karan, Ré, Christopher. “Efficiently Modeling Long Sequences with Structured State Spaces” Conference paper at ICLR 2022. openreview.net/forum?id=uYLFoz1vIAC. [cited by applicant]
Guo, X., Chang, S., Yu, M., Tesauro, G., and Campbell, M. Hybrid reinforcement learning with expert state sequences. In AAAI Conference on Artificial Intelligence, AAAI, pp. 3739-3746, 2019. [cited by applicant]
He, K., Zhang, X., Ren, S., and Sun, J. Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pp. 1026-1034, 20… [cited by applicant]
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., Dulac-Arnold, G., Agapiou, J. P., Leibo, J. Z., and Gruslys, A. Deep Q-learning from demonstrat… [cited by applicant]
Ho, J. and Ermon, S. Generative adversarial imitation learning. In Advances in Neural Information Processing Systems, NIPS, pp. 4565-4573, 2016. [cited by applicant]
Ho, J., Kalchbrenner, N., Weissenborn, D., and Salimans, T. Axial attention in multidimensional transformers. arXiv preprint arXiv:1912.12180, Dec. 20, 2019. [cited by applicant]
Hoehl, S., Keupp, S., Schleihauf, H., McGuigan, N., Buttelmann, D., and Whiten, A. ‘Over-imitation’: A review and appraisal of a decade of research. Developmental Review, 51:90-108, 2019. [cited by applicant]
Huang, D., Nair, S., Xu, D., Zhu, Y., Garg, A., Fei-Fei, L., Savarese, S., and Niebles, J. C. Neural task graphs: Generalizing to unseen tasks from a single video demonstration. In IEEE Conference on Computer Vision and… [cited by applicant]
James, S., Bloesch, M., and Davison, A. J. Task-embedded control networks for few-shot imitation learning. In Conference on Robot Learning, CoRL, pp. 783-795, 2018. [cited by applicant]
Jinyoung Choi, Christopher Dance, Jung-eun Kim, Seulbin Hwang, and Kyung-sik Park. Risk-conditioned distributional soft actor-critic for risk-sensitive navigation. In Submitted to International Conference on Robotics an… [cited by applicant]
Jonathan Lacotte, Mohammad Ghavamzadeh, Yinlam Chow, and Marco Pavone. Risk-sensitive generative adversarial imitation learning. In International Conference on Artificial Intelligence and Statistics, AISTATS, pp. 2154-2… [cited by applicant]
Karl Cobbe, Chris Hesse, Jacob Hilton, and John Schulman. Leveraging procedural generation to benchmark reinforcement learning. In International Conference on Machine Learning, ICML, pp. 2048-2056, 2020. [cited by applicant]
Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Jaskowski, W. ViZDoom: A Doom-based AI research platform for visual reinforcement learning. In [cited by applicant]
Kim, Y. Design of low inertia manipulator with high stiffness and strength using tension amplifying mechanisms. IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS, pp. 5850-5856, 2015. [cited by applicant]
Laskey, M., Lee, J., Fox, R., Dragan, A., and Goldberg, K. Dart: Noise injection for robust imitation learning. arXiv preprint arXiv:1703.09327, 2017. [cited by applicant]
Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P. A simple neural attentive meta-learner. In International Conference on Learning Representations, ICLR, 2018. [cited by applicant]
Nasiriany, S., Pong, V., Lin, S., and Levine, S. Planning with goal-conditioned policies. In Advances in Neural Information Processing Systems, NeurIPS, pp. 14814-14825, 2019. [cited by applicant]
Neu, G. and Szepesvari, C. Apprenticeship learning using inverse reinforcement learning and gradient methods. arXiv preprint arXiv:1206.5264, 2012. [cited by applicant]
Ng, A. Y., Harada, D., and Russell, S. Policy invariance under reward transformations: Theory and application to reward shaping. In Proceedings of the 16th International Conference on Machine Learning (ICML), pp. 278-28… [cited by applicant]
Nichol, A., Achiam, J., and Schulman, J. On first-order meta-learning algorithms. arXiv preprint arXiv:1803.02999, 2018. [cited by applicant]
Ozan Sener and Vladlen Koltun. Multi-task learning as multi-objective optimization. In Advances in Neural Information Processing Systems, NeurIPS, pp. 527-538, 2018. [cited by applicant]
Paine, T. L., Colmenarejo, S. G., Wang, Z., Reed, S., Aytar, Y., Pfaff, T., Hoffman, M. W., Barth-Maron, G., Cabi, S., Budden, D., et al. One-shot high-fidelity imitation: Training large-scale deep nets with RL. arXiv p… [cited by applicant]
Pomerleau, D. Efficient training of artificial neural networks for autonomous navigation. Neural Computation, 3(1): 88-97, 1991. [cited by applicant]
Rajaraman, N., Yang, L., Jiao, J., and Ramchandran, K. Toward the fundamental limits of imitation learning. In Advances in Neural Information Processing Systems, NeurIPS, 2020. [cited by applicant]
Ross, S. and Bagnell, D. Efficient reductions for imitation learning. In International Conference on Artificial Intelligence and Statistics, AISTATS, pp. 661-668, 2010. [cited by applicant]
Ross, S., Gordon, G. J., and Bagnell, J. A. No-regret reductions for imitation learning and structured prediction. In AISTATS. Citeseer, 2011. [cited by applicant]
Rothkopf, C., Ballard, D., and Hayhoe, M. Task and context determine where you look. Journal of Vision, 7(14), 2007. [cited by applicant]
Sanjeev Arora, Simon S. Du, Sham Kakade, Yuping Luo, and Nikunj Saunshi. Provable representation learning for imitation learning via bi-level optimization. In International Conference on Machine Learning, ICML, 2020. [cited by applicant]
Santara, A., Naik, A., Ravindran, B., Das, D., Mudigere, D., Avancha, S., and Kaul, B. Rail: Risk-averse imitation learning. arXiv preprint arXiv:1707.06658, 2017. [cited by applicant]
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017. [cited by applicant]
Sener, O. and Koltun, V. Multi-task learning as multiobjective optimization. In Advances in Neural Information Processing Systems, NeurIPS, pp. 527-538, 2018. [cited by applicant]
Seungsu Kim and Stephane Doncieux. Learning highly diverse robot throwing movements through quality-diversity search. In Genetic and Evolutionary Computation Conference, pp. 1177-1178, 2017. [cited by applicant]
Sieb, M., Xian, Z., Huang, A., Kroemer, O., and Fragkiadaki, K. Graph-structured visual imitation. In Conference on Robot Learning (CoRL), pp. 979-989. PMLR, 2020. [cited by applicant]
Singh, A., Jang, E., Irpan, A., Kappler, D., Dalal, M., Levine, S., Khansari, M., and Finn, C. Scalable multitask imitation learning with autonomous improvement. In IEEE International Conference on Robotics and Automati… [cited by applicant]
Soroush Nasiriany, Vitchyr Pong, Steven Lin, and Sergey Levine. Planning with goal-conditioned policies. In Advances in Neural Information Processing Systems, NeurIPS, pp. 14814-14825, 2019. [cited by applicant]
Sudeep Dasari et al: “Transformers for One-Shot Visual Imitation”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov. 11, 2020 (Nov. 11, 2020), XP081811772. [cited by applicant]
Tay, Y., Dehghani, M., Bahri, D., and Metzler, D. Efficient transformers: A survey. arXiv preprint arXiv:2009.06732, 2020. [cited by applicant]
Tianhe Yu, Chelsea Finn, Sudeep Dasari, Annie Xie, Tianhao Zhang, Pieter Abbeel, and Sergey Levine. One-shot imitation from observing humans via domain-adaptive meta-learning. In Robotics: Science and Systems, RSS, 2018. [cited by applicant]
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, Gabriel Dulac-Arnold, John P. Agapiou, Joel Z. Leibo, and Audrunas Gruslys. Deep … [cited by applicant]
Todorov, E., Erez, T., and Tassa, Y. MuJoCo: A physics engine for model-based control. In IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS, pp. 5026-5033, 2012. [cited by applicant]
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. arXiv preprint arXiv:1801.01290, 2018. [cited by applicant]
Uddeshya Upadhyay, Nikunj Shah, Sucheta Ravikanti, and Mayanka Medhe. Transformer based reinforcement learning for games. arXiv preprint arXiv:1912.03918, 2019. [cited by applicant]
Van der Maaten, L. and Hinton, G. Visualizing data using t-SNE. Journal of Machine Learning Research, 9:2579-2605, 2008. [cited by applicant]
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In Advances in Neural Information Processing Systems, NIPS, pp. 5998-6008, 2017. [cited by applicant]
Venuto, D., Boussioux, L., Wang, J., Dali, R., Chakravorty, J., Bengio, Y., and Precup, D. Avoidance learning using observational reinforcement learning. arXiv preprint arXiv:1909.11228, 2019. [cited by applicant]
Wang, Z. and Liu, J.-C. Translating math formula images to latex sequences using deep neural networks with sequence-level training. International Journal on Document Analysis and Recognition (IJDAR), pp. 1-13, 2020. [cited by applicant]
Will Dabney, Georg Ostrovski, David Silver, and Rémi Munos. Implicit quantile networks for distributional reinforcement learning. In International Conference on Machine Learning, ICML, pp. 1104-1113, 2018. [cited by applicant]
Wonjoon Goo and Scott Niekum. One-shot learning of multi-step tasks from observation via activity localization in auxiliary video. In International Conference on Robotics and Automation, ICRA, pp. 7755-7761, 2019. [cited by applicant]
Wu, Y.-H., Charoenphakdee, N., Bao, H., Tangkaratt, V., and Sugiyama, M. Imitation learning from imperfect demonstration. arXiv preprint arXiv:1901.09387, 2019. [cited by applicant]
Xiao, H., Herman, M., Wagner, J., Ziesche, S., Etesami, J., and Linh, T. H. Wasserstein adversarial imitation learning. arXiv preprint arXiv:1906.08113, 2019. [cited by applicant]
Xiaoxiao Guo, Shiyu Chang, Mo Yu, Gerald Tesauro, and Murray Campbell. Hybrid reinforcement learning with expert state sequences. In AAAI Conference on Artificial Intelligence, AAAI, pp. 3739-3746, 2019. [cited by applicant]
Yan Duan, Marcin Andrychowicz, Bradly Stadie, Jonathan Ho, Jonas Schneider, Ilya Sutskever, Pieter Abbeel, and Wojciech Zaremba. One-shot imitation learning. In Advances in Neural Information Processing Systems, NIPS, p… [cited by applicant]
Yiming Ding, Carlos Florensa, Pieter Abbeel, and Mariano Phielipp. Goal-conditioned imitation learning. In Advances in Neural Information Processing Systems, NeurIPS, pp. 15298-15309, 2019. [cited by applicant]
Yu, L., Yu, T., Finn, C., and Ermon, S. Meta-inverse reinforcement learning with probabilistic context variables. In Advances in Neural Information Processing Systems, NeurIPS, pp. 11749-11760, 2019. [cited by applicant]
Yu, T., Finn, C., Dasari, S., Xie, A., Zhang, T., Abbeel, P., and Levine, S. One-shot imitation from observing humans via domain-adaptive meta-learning. In Robotics: Science and Systems, RSS, 2018. [cited by applicant]
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., Levine, S., Pong, V., et al. Meta-World source code. https://github.com/rlworkgroup/metaworld, 2019. [cited by applicant]
European Search Report for European Application No. EP21305799.5 dated Jan. 4, 2022. [cited by applicant]