IP Library › Granted Patent US 12,488,278
Granted Patent B2
US 12,488,278 · App. 17/067,284 · Granted Dec 2, 2025

Interactive agent and control using reinforcement learning

Inventors: Katja Hofmann (Cambridge, GB); Luisa Maria Zintgraf (Oxford, GB); Sam Michael Devlin (Trumpington, GB); Kamil Andrzej Ciosek (Cambridge, GB)
Assignee: Microsoft Technology Licensing, LLC.
G06N20/00G06N3/006G06N3/092G06N5/04G06N3/02G06Q10/101
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,278
App. No.
17/067,284
Granted
Dec 2, 2025
Kind
B2
Abstract

In various examples there is a method performed by a computer-implemented agent in an environment. The method comprises storing a reinforcement learning policy for controlling the computer-implemented agent. The method also comprises storing a distribution as a latent representation of a belief of the computer-implemented agent about at least one other agent in the environment. The method involves executing the computer-implemented agent according to the policy conditioned on parameters characterizing the distribution.

Claims (29)

1 . A method performed by a computer-implemented agent in an environment, the method comprising:

storing a reinforcement learning policy for use by the computer-implemented agent to act in the environment;

storing a distribution as a latent representation of a belief of the computer-implemented agent about another agent in the environment, wherein the latent representation is computed from an input of an observation made by the computer-implemented agent of a behavior of the other agent in the environment, wherein the latent representation comprises a permanent component denoting permanent behaviors of the other agent and a temporal component denoting temporal behaviors of the other agent, and wherein the latent representation includes a hierarchical structure wherein the temporal component of the other agent depends on the permanent component of the other agent, the temporal component being generated from a first layer of an encoder, the permanent component being generated from a second layer of the encoder, wherein the first layer is deeper in the hierarchical structure than the second layer; and

executing the computer-implemented agent according to the policy conditioned on parameters characterizing the distribution.

2 . The method of claim 1 comprising storing the distribution as the latent representation at an encoder having been trained using historical information of the observations of the computer-implemented agent of behavior of the other agent in the environment.

3 . The method of claim 2 comprising inputting to the encoder a current trajectory of behavior of the other agent.

4 . The method of claim 1 comprising observing a state of the computer-implemented agent and inputting the observed state into the policy together with the parameters in order to execute the computer-implemented agent according to the policy conditioned on parameters characterizing the distribution.

5 . The method of claim 1 wherein the distribution is a Gaussian distribution and the parameters comprise one or more pairs each comprising a mean and a standard deviation.

6 . The method of claim 1 wherein at least one of the permanent component is a Gaussian distribution or the temporal component is a Gaussian distribution.

7 . The method of claim 1 wherein the other agent is a first other agent, and where the latent representation is of belief of the computer-implemented agent about the first other agent and at least second and third other agents in the environment.

8 . The method of claim 1 wherein the other agent is a first other agent, and wherein the computer-implemented agent is collaborating or competing with at least one of the first other agent or a second other agent.

9 . The method of claim 1 comprising jointly learning the policy and the distribution over the latent representation using a mechanism trainable by gradient descent.

10 . The method of claim 9 wherein the mechanism comprises an encoder and a predictive component, and wherein the encoder, predictive component and policy are trained using a same loss function.

11 . The method of claim 10 wherein the loss function computes a measure of a difference between an observed future trajectory of the computer-implemented agent and a prediction of the future trajectory computed using the predictive component.

12 . The method of claim 9 wherein the mechanism is a variational autoencoder.

13 . The method of claim 1 wherein the computer-implemented agent is controlled without predicting future actions of the other agent.

14 . The method of claim 1 , wherein the temporal component of the other agent comprises a temporal state and the permanent component of the other agent comprises a permanent type.

15 . The method of claim 1 wherein the observation of the behavior of the other agent in the environment made by the computer-implemented agent is made by a sensor at the computer-implemented agent.

16 . A computer-implemented agent comprising:

a processor;

a memory storing instructions, that, when executed by the processor, cause the computer-implemented agent to perform, in an environment, a method comprising:

computing a latent representation of an inference of the computer-implemented agent about another agent in the environment, the latent representation being computed from an input of an observation made by the computer-implemented agent of a behavior of the other agent in the environment, wherein computing the latent representation comprises factorizing the latent representation into a permanent output and a temporal output, the permanent output comprising a first distribution denoting permanent behaviors of the other agent, the temporal output comprising a second distribution denoting temporal behaviors of the other agent, and wherein the temporal output is generated from a first layer of an encoder, the permanent output is generated from a second layer of the encoder, and the first layer is deeper in a hierarchical structure of the latent representation than the second layer; and

executing the computer-implemented agent according to a policy conditioned on parameters characterizing the first and second distributions.

17 . The computer-implemented agent of claim 16 deployed as any of: a self-driving vehicle, a physical robot, a computer-implemented game player, a virtual assistant.

18 . A computer-implemented method for training a computer-implemented agent, the method comprising:

storing a reinforcement learning policy for use by the computer-implemented agent to act in an environment; and

using a mechanism trainable by gradient descent, jointly learning the policy and a distribution as a latent representation of a belief of the computer-implemented agent about at least one other agent in the environment, wherein the latent representation is computed from inputs of observations made by the computer-implemented agent of behaviors of all of the other agents in the environment, wherein the latent representation comprises a permanent component denoting permanent behaviors of at least one of the other agents and a temporal component denoting temporal behaviors of at least one of the other agents, and wherein the latent representation includes a hierarchical structure wherein the temporal component of the at least one other agent depends on the permanent component of the at least one other agent, the temporal component being generated from a first layer of an encoder, the permanent component being generated from a second layer of the encoder, wherein the first layer is deeper in the hierarchical structure than the second layer.

19 . The method of claim 18 wherein the mechanism comprises an encoder, a predictive component and a loss function.

20 . The method of claim 18 , wherein the temporal component of the at least one other agent comprises a temporal state and the permanent component of the at least one other agent comprises a permanent type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2021
From: HOFMANN, KATJA; ZINTGRAF, LUISA MARIA; DEVLIN, SAM MICHAEL; CIOSEK, KAMIL ANDRZEJ
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 055919/0593 →
Continuity (1)
Related Publication 20220114474A1 · Apr 14, 2022
References Cited (58)
US 7921066B2 · Van et al. · 2011 [cited by applicant]
US 10809735B2 · Halder · 2020 [cited by examiner]
US 20200076857A1 · Van Seijen et al. · 2020 [cited by applicant]
WO 2018212918A1 · 2018 [cited by applicant]
Papoudakis, Georgios, and Stefano V. Albrecht. “Variational autoencoders for opponent modeling in multi-agent systems.” arXiv preprint arXiv:2001.10829 (2020). (Year: 2020). [cited by examiner]
Feng, Xidong, et al. “Vehicle trajectory prediction using intention-based conditional variational autoencoder.” 2019 IEEE Intelligent Transportation Systems Conference (ITSC). IEEE, 2019. (Year: 2019). [cited by examiner]
Rosello, Pol, and Mykel J. Kochenderfer. “Multi-agent reinforcement learning for multi-object tracking.” Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems. 2018. (Year: 2018). [cited by examiner]
Zhao, Shengjia, Jiaming Song, and Stefano Ermon. “Learning hierarchical features from deep generative models.” International Conference on Machine Learning. PMLR, 2017. (Year: 2017). [cited by examiner]
Krupnik, et al., “Multi Agent Reinforcement Learning with Multi-Step Generative Models”, In Repository of arXiv:1901.10251v1, Jan. 29, 2019, 13 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US21/043807”, Mailed Date: Nov. 12, 2021, 11 Pages. [cited by applicant]
Zintgraf, et al., “Exploration in Approximate Hyper-State Space for Meta Reinforcement Learning”, In Repository of arXiv:2010.01062v1, Oct. 2, 2020, 16 Pages. [cited by applicant]
Albrecht, et al., “Autonomous agents modelling other agents: A comprehensive survey and open problems”, In Journal of Artificial Intelligence, May 1, 2018, 46 Pages. [cited by applicant]
Albrecht, et al., “Belief and truth in hypothesised behaviours”, In Journal of Artificial Intelligence, vol. 235, Jun. 2016, pp. 63-94. [cited by applicant]
Albrecht, et al., “Reasoning about Hypothetical Agent Behaviours and their Parameters”, In Repository of arXiv:1906.11064v1, Jun. 26, 2019, 9 Pages. [cited by applicant]
Baker, et al., “Modeling human plan recognition using Bayesian theory of mind”, In Plan, activity, and intent recognition: Theory and practice, 2014, pp. 177-204. [cited by applicant]
Baker, et al., “Rational quantitative attribution of beliefs, desires and percepts in human mentalizing”, In Journal of Nature Human Behaviour, vol. 1, Issue 4, Mar. 13, 2017, pp. 1-10. [cited by applicant]
Bard, et al., “Particle filtering for dynamic agent modelling in simplified poker”, In Proceedings of the National Conference on Artificial Intelligence, Jul. 22, 2007, pp. 515-521. [cited by applicant]
Barrett, et al., “Teamwork with Limited Knowledge of Teammates”, In Proceedings of Twenty-Seventh AAAI Conference on Artificial Intelligence, Jun. 30, 2013, 7 Pages. [cited by applicant]
Bergstrom, et al., “On the evolution of behavioral heterogeneity in individuals and populations”, In Biology and Philosophy, vol. 13, Issue 2, Apr. 1, 1998, pp. 205-231. [cited by applicant]
Canaan, et al., “Evaluating the Rainbow DQN Agent in Hanabi with Unseen Partners”, In Repository of arXiv:2004.13291v1, Apr. 28, 2020, 8 Pages. [cited by applicant]
Carmel, et al., “Exploration strategies for model-based learning in multi-agent systems”, In Journal of Autonomous Agents and Multi-agent Systems, vol. 2, 1990, pp. 1-38. [cited by applicant]
Carroll, et al., “On the utility of learning about humans for human-ai coordination”, In Advances in Neural Information Processing Systems, Dec. 8, 2019, pp. 1-12. [cited by applicant]
Chalkiadakis, Georgios, “A Bayesian approach to multiagent reinforcement learning and coalition formation under uncertainty”, In Thesis of University of Toronto, Jan. 1, 2007, 258 Pages. [cited by applicant]
Chalkiadakis, et al., “Cooperative Games with Overlapping Coalitions”, In Journal of Artificial Intelligence Research, vol. 39, Sep. 2010, pp. 179-216. [cited by applicant]
Chalkiadakis, et al., “Coordination in multiagent reinforcement learning: A bayesian approach”, In Proceedings of the second international joint conference on Autonomous agents and multiagent systems, Jul. 14, 2003, pp.… [cited by applicant]
Zhao, “Learning Hierarchical Features from Deep Generative Models”, In International Conference on Machine earning, Jul. 17, 2017, 9 Pages. [cited by applicant]
Chung, et al., “A recurrent latent variable model for sequential data”, In Advances in neural information processing systems, 2015, pp. 1-9. [cited by applicant]
Duan, et al., “RL2 : Fast Reinforcement Learning via Slow Reinforcement Learning”, In Repository of arXiv:1611.02779v2, Nov. 10, 2016, pp. 1-14. [cited by applicant]
Duff, Michael O'Gordon, “Optimal Learning: Computational procedures for Bayes-adaptive Markov decision processes”, In Thesis of University of Massachusetts Amherst, Feb. 2002, 255 Pages. [cited by applicant]
Foerster, et al., “Bayesian action decoder for deep multi-agent rein415 forcement learning”, In Repository of arXiv:1811.01458v1, Nov. 4, 2018, 12 pages. [cited by applicant]
Haroush,et al. “Neuronal prediction of opponent's behavior during cooperative social interchange in primates”, In Journal of Cell, vol. 160, Issue 6, Mar. 2015, pp. 1233-1245. [cited by applicant]
He, et al., “Opponent modeling in deep reinforcement learning”, In International conference on machine learning, Jun. 11, 2016, 10 Pages. [cited by applicant]
Heap, et al., “Game theory: a critical text”, In Publication of Psychology Press, 2004. [cited by applicant]
Hernandez-Leal, et al., “A Survey of Learning in Multiagent Environments: Dealing with Non-Stationarity”, In Repository of arXiv:1707.09183v1, Jul. 28, 2017, pp. 1-64. [cited by applicant]
Hoang, et al., “A General Framework for Interacting Bayes-Optimally with Self-Interested Agents using Arbitrary Parametric Model and Model Prior”, In Repository of arXiv:1304.2024v1, Apr. 7, 2013, 10 Pages. [cited by applicant]
Joshen, Yedid, “VAIN: Attentional Multi-agent Predictive Modeling”, In Proceedings of 31st Conference on Neural Information Processing Systems, Dec. 4, 2017, 11 Pages. [cited by applicant]
Hu, et al., “Simplified action decoder for deep multi-agent reinforcement learning”, In Repository of arXiv:1912.02288v1, Dec. 4, 2019, pp. 1-14. [cited by applicant]
Humplik, et al., “Meta reinforcement learning as task inference”, In Proceedings of arXiv:1905.06424, May 15, 2019, 20 Pages. [cited by applicant]
Iqbal, et al., “Actor-Attention-Critic for Multi-Agent Reinforcement Learning”, In Repository of arXiv:1810.02912v1, Oct. 5, 2018, pp. 1-12. [cited by applicant]
Ivanovic, et al., “Generative Modeling of Multimodal Multi-Human Behavior”, In Proceedings of IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Oct. 1, 2018, pp. 3088-3095. [cited by applicant]
Ivanovic, et al., “The Trajectron: Probabilistic Multi-Agent Trajectory Modeling With Dynamic Spatiotemporal Graphs”, In Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 2375-2384. [cited by applicant]
Kingma, et al., “Auto-encoding variational bayes”, In Repository of arXiv:1312.6114v1, Dec. 20, 2013, pp. 1-9. [cited by applicant]
Kopf, et al., “Partner Approximating Learners (PAL): Simulation-Accelerated Learning with Explicit Partner Modeling in Multi-Agent Domains”, In Proceedings of 6th International Conference on Control, Automation and Robo… [cited by applicant]
Nachbar, John H., “Beliefs in repeated games”, In Journal of Econometrica, vol. 73, No. 2, Mar. 2005, pp. 459-480. [cited by applicant]
Nghia, Hoang Trong, “New Advances on Bayesian and Decision-Theoretic Approaches for Interactive Machine Learning”, In Thesis of National University of Singapore, Oct. 2014, 223 Pages. [cited by applicant]
Ortega, et al., “Meta-learning of Sequential Strategies”, In Repository of arXiv:1905.03030v2, Jul. 18, 2019, pp. 1-15. [cited by applicant]
Papoudakis, et al., “Opponent Modelling with Local Information Variational Autoencoders”, In Repository of arXiv:2006.09447v1, Jun. 16, 2020, 12 pages. [cited by applicant]
Papoudakis, et al., “Variational Autoencoders for Opponent Modeling in Multi-Agent Systems”, In Repository of arXiv:2001.10829v1, Jan. 29, 2020, 8 Pages. [cited by applicant]
Zintgraf, et al., “VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning”, In International Conference on Learning Representations, Apr. 26, 2020, pp. 1-20. [cited by applicant]
Rabinowitz, et al., “Machine theory of mind”, In Proceedings of the 35th International Conference on Machine Learning, Jul. 10, 2018, 21 Pages. [cited by applicant]
Raileanu, et al., “Modeling Others using Oneself in Multi-Agent Reinforcement Learning”, In Repository of arXiv:1802.09640v3, Mar. 23, 2018, 11 Pages. [cited by applicant]
Sadigh, et al., “Information gathering actions over human internal state”, In Proceedings of IEEE/RSJ International Conference on Intelligent Robots and Systems, Oct. 9, 2016, 8 Pages. [cited by applicant]
Serrino, et al., “Finding friend and foe in multi-agent games”, In Proceedings of 33rd Conference on Neural Information Processing Systems, Dec. 8, 2019, pp. 1-11. [cited by applicant]
Shapley, L. S., “Stochastic games”, In Proceedings of the national academy of sciences, Oct. 1, 1953, pp. 1095-1100. [cited by applicant]
Southey, et al., “Bayes' bluff: Opponent modelling in poker”, In Repository of arXiv preprint arXiv:1207.1411, Jul. 4, 2012, 9 Pages. [cited by applicant]
Stone, et al., “Ad Hoc Autonomous Agent Teams: Collaboration without Pre-Coordination”, In Proceedings of the Twenty-Fourth AAAI Conference on Artificial Intelligence, Jul. 11, 2010, pp. 1504-1509. [cited by applicant]
Tan, Ming, “Multi-agent reinforcement learning: independent vs. cooperative agents”, In Publications of Readings in agents, Oct. 1997, 8 Pages. [cited by applicant]
Wang, et al., “Learning to Reinforcement Learn”, In Proceedings of the 39th Annual Meeting of the Cognitive Science Society, Nov. 17, 2016, pp. 1-17. [cited by applicant]