IP Library Granted Patent US 12,354,027
Granted Patent B2
US 12,354,027 · App. 15/943,947 · Granted Jul 8, 2025

Method and system for an intelligent artificial agent

Inventors: Mark Bishop Ring (Anaheim, CA); Satinder Baveja (Ann Arbor, MI); Peter Stone (Austin, TX); James MacGlashan (Riverside, RI); Samuel Barrett (Somerville, MA); Roberto Capobianco (Itri, IT); Varun Kompella (Aachen, DE); Kaushik Subramanian (Richmond, CA); Peter Wurman (Acton, MA)
Assignee: SONY GROUP CORPORATION
G06N5/043G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,027
App. No.
15/943,947
Granted
Jul 8, 2025
Kind
B2
Abstract

A method and system for teaching an artificial intelligent agent where the agent can be placed in a state that it would like it to learn how to achieve. By giving the agent several examples, it can learn to identify what is important about these example states. Once the agent has the ability to recognize a goal configuration, it can use that information to then learn how to achieve the goal states on its own. An agent may be provided with positive and negative examples to demonstrate a goal configuration. Once the agent has learned certain goal configurations, the agent can learn policies and skills that achieve the learned goal configuration. The agent may create a collection of these policies and skills from which to select based on a particular command or state.

Claims (44)

1. A method for training an artificial intelligent agent to recognize a goal configuration, comprising:

placing the agent in the goal configuration and identifying a resulting state as a positive example;

providing negative examples to the agent that demonstrate the agent in a state failing to achieve the goal configuration;

extracting key state features when the agent is in the goal configuration, the key state features including at least one of a room feature, object positioning, ambient lighting, and ambient sounds;

determining what feature categories are important in the goal configuration during receipt of positive examples to the agent;

learning and recognizing, by the agent, the goal configuration based on the extracted key state features and the determined important feature categories;

creating policies, by the agent, based on the learned goal configuration;

converting state features into a distance function to determine how far the agent is from the goal configuration;

using goal detection as a final reward; and

using a goal distance as an intermediate reward.

2. The method of claim 1 , wherein an interface is used to indicate an example as being either the positive example or the negative example, the interface includes at least one of a spoken word received by the agent, an electronic signal received from a computing device, and a physical button on the agent.

3. The method of claim 1 , wherein the step of extracting key state features includes looking for similarity in state features in each of the positive and negative examples.

4. The method of claim 3 , further comprising increasing a confidence of the agent as the positive and negative examples are received by the agent.

5. The method of claim 4 , wherein the agent takes an action upon reaching a predetermined level of confidence.

6. The method of claim 1 , wherein the key state features are weighted according to a predetermined weight value.

7. The method of claim 1 , further comprising asking, by the agent, for human feedback regarding whether the agent is in a goal state.

8. A system comprising a processor and a computer-usable medium embodying a computer program code, the computer program code comprising instructions executable by the processor and configured to provide a method of learning to recognize a goal configuration of an artificial agent, the method comprising:

placing the agent in the goal configuration and identifying a resulting state as a positive example;

providing negative examples to the agent that demonstrate the agent in a state failing to achieve the goal configuration;

extracting key state features when the agent is in the goal configuration, the key state features including at least one of a room feature, object positioning, ambient lighting, and ambient sounds;

determining what feature categories are important in the goal configuration during receipt of positive examples to the agent;

learning and recognizing, by the agent, the goal configuration based on the extracted key state features and the determined important feature categories;

creating policies, by the agent, based on the learned goal configuration;

converting state features into a distance function to determine how far the agent is from the goal configuration;

using the distance function as an intermediate reward for the agent; and

using goal detection as a final reward.

9. The system of claim 8 , wherein the method further comprises recognizing whether the agent is in an initialization state.

10. The system of claim 8 , wherein the method further comprises self-practice by the agent.

11. The system of claim 10 , wherein a selected goal configuration for self-practice is selected based on at least one of a random determination, which goal configuration needs the most improvement, which goal configuration is most likely to improve, which goal configuration has been used least recently, and which goal configuration is most used.

12. The system of claim 10 , wherein the method further comprises biasing action choices based on which actions are important to achieving the goal configuration.

13. The system of claim 8 , wherein the method further comprises updating a policy for achieving a goal configuration based on performance of the agent.

14. A non-transitory computer-readable storage medium with an executable program stored thereon, wherein the program instructs one or more processors to perform the following steps to cause an agent to recognize and learn a goal configuration:

placing the agent in the goal configuration and identifying a resulting state as a positive example;

providing negative examples to the agent that demonstrate the agent in a state failing to achieve the goal configuration;

extracting key state features when the agent is in the goal configuration, the key state features including at least one of a room feature, object positioning, ambient lighting, and ambient sounds;

determining what feature categories are important in the goal configuration during receipt of positive examples to the agent;

learning and recognizing, by the agent, the goal configuration based on the extracted key state features and the determined important feature categories;

creating policies, by the agent, based on the learned goal configuration;

converting state features into a distance function to determine how far the agent is from the goal configuration;

using the distance function as an intermediate reward for the agent; and using goal detection as a final reward.

15. The non-transitory computer-readable storage medium of claim 14 , wherein the step of extracting key state features includes looking for similarity of the key state features in each of the positive and negative examples.

16. The non-transitory computer-readable storage medium of claim 14 , wherein the program instructs one or more processors to perform the following steps:

increasing a confidence of the agent as additional ones of the positive and negative examples are received by the agent; and

taking an action by the agent upon reaching a predetermined level of confidence.

Assignments (4)
CHANGE OF NAME Recorded May 17, 2023
From: SONY CORPORATION
To: SONY GROUP CORPORATION
Reel/Frame 063672/0079 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2020
From: COGITAI, INC.
To: SONY CORPORATION OF AMERICA; SONY CORPORATION
Reel/Frame 051588/0478 →
SECURITY INTEREST Recorded May 24, 2019
From: COGITAI, INC.
To: SONY CORPORATION OF AMERICA
Reel/Frame 049278/0735 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2018
From: RING, MARK BISHOP; BAVEJA, SATINDER; STONE, PETER; MACGLASHAN, JAMES; BARRETT, SAMUEL; CAPOBIANCO, ROBERTO; KOMPELLA, VARUN; SUBRAMANIAN, KAUSHIK; WURMAN, PETER
To: COGITAI, INC.
Reel/Frame 045821/0694 →
Continuity (1)
Related Publication 20190303776A1 · Oct 3, 2019
References Cited (42)
US 6366896B1 · Hutchison · 2002 [cited by examiner]
US 6594524B2 · Esteller et al. · 2003 [cited by applicant]
US 9997039B1 · Heaton et al. · 2018 [cited by applicant]
US 10289910B1 · Chen · 2019 [cited by examiner]
US 20080059274A1 · Holliday · 2008 [cited by applicant]
US 20140095412A1 · Agashe et al. · 2014 [cited by applicant]
US 20150290798A1 · Iwatake · 2015 [cited by examiner]
US 20190012371A1 · Campbell et al. · 2019 [cited by applicant]
US 20190130312A1 · Xiong et al. · 2019 [cited by applicant]
US 20190261566A1 · Robertson et al. · 2019 [cited by applicant]
US 20190347621A1 · White · 2019 [cited by applicant]
US 20200090042A1 · Wayne et al. · 2020 [cited by applicant]
US 20200211106A1 · Pan et al. · 2020 [cited by applicant]
US 20210187733A1 · Lee et al. · 2021 [cited by applicant]
Masakazu Hirkoawa, Coaching Robots: Online Behavior Learning from Human Subjective Feedback, I. Jordanov and L.C. Jain (Eds.): Innovations in Intelligent Machines—3, SCI 442, pp. 37-51. (Year: 2013). [cited by examiner]
Patrick Grüneberg, An Approach to Subjective Computing: A Robot That Learns From Interaction With Humans, IEEE Transactions on Autonomous Mental Development, vol. 6, No. 1, Mar. 2014 (Year: 2014). [cited by examiner]
Jens Kober, Reinforcement learning in robotics: A survey, The International Journal of Robotics Research, 32(11) 1238-1274, 2013 (Year: 2013). [cited by examiner]
Anna Gruebler, Coaching robot behavior using continuous physiological affective feedback, 2011 11th IEEE-RAS International Conference on Humanoid Robots, Bled, Slovenia, Oct. 26-28, 2011 (Year: 2011). [cited by examiner]
Ve ̌cerík (“Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards”) arXiv:1707.08817v1 [cs.AI] Jul. 27, 2017 (Year: 2017). [cited by examiner]
Katyal (“Leveraging Deep Reinforcement Learning for Reaching Robotic Tasks”) Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2017, pp. 18-19 (Year: 2017). [cited by examiner]
Grollman (“Robot Learning from Failed Demonstrations”) Int J Soc Robot (2012) 4:331-342, Jun. 30, 2012 © Springer Science & Business Media BV 2012 (Year: 2012). [cited by examiner]
Nicolescu (“Natural Methods for Robot Task Learning: Instructive Demonstrations, Generalization and Practice”) AAMAS'03, Jul. 14-18, 2003, Melbourne, Australia. (Year: 2003). [cited by examiner]
Thomaz_2008_Teachable robots: Understanding human teaching behavior to build more effective robot learners Artificial Intelligence 172 (2008) 716-737 (Year: 2008). [cited by examiner]
Baranes_2012_Active learning of inverse models with intrinsically motivated goal exploration in robots Robotics and Autonomous Systems 61 (2013) 49-73 (Year: 2012). [cited by examiner]
Rai (“Learning from failed demonstrations in unreliable systems”) 2013 13th IEEE-RAS International Conference on Humanoid Robots (Humanoids). Oct. 15-17, 2013. Atlanta, GA (Year: 2013). [cited by examiner]
Zeng (“Object Manipulation Learning by Imitation”) arXiv:1603.00964v3 [cs.RO] Nov. 19, 2017 (Year: 2017). [cited by examiner]
Shiarlis (“Inverse Reinforcement Learning from Failure”) Proceedings of the 15th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2016) (Year: 2016). [cited by examiner]
Kim (“Socially Adaptive Path Planning in Human Environments Using Inverse Reinforcement Learning”) Int J of Soc Robotics (2016) 8:51-66 DOI 10.1007/s12369-015-0310-2 (Year: 2016). [cited by examiner]
Hilleli (“Deep Learning of Robotic Tasks without a Simulator using Strong and Weak Human Supervision”) arXiv:1612.01086v3 [cs.AI] Mar. 26, 2017 (Year: 2017). [cited by examiner]
Luo (“Learning to select relevant perspective in a dynamic environment”) 2008 IEEE International Joint Conference on Neural Networks (IEEE World Congress on Computational Intelligence) (Year: 2008). [cited by examiner]
Willems, et al., “The Context-Tree Weighting Method: Basic Properties”, IEEE Transactions on Information Theory, vol. 41, No. 3, May 1995, pp. 653-664. [cited by applicant]
Begleiter et. al., “On Prediction Using Variable Order Markov Models”, 2004, Journal of Artificial Intelligence Research, 22 (2004), pp. 385-421 (Year: 2004). [cited by applicant]
Thorhallsson et. al., Visualizing the Bias Variance Tradeoff', 2017, University of British Columbia, 2017, pp. 1-9 (Year: 2017). [cited by applicant]
Wang et. al., “genCNN: A Convolutional Architecture for Word Sequence Prediction”, 2015, arXiv, 2015, pp. 1-13 (Year: 2015). [cited by applicant]
Wu et. al., “A Novel Sensory Mapping Design for Bipedal Walking on a Sloped Surface”, 2012, International Journal of Advanced Robotic Systems, 9 (2012), pp. 1-9 (Year: 2012). [cited by applicant]
Botvinick, Matthew Michael. “Hierarchical reinforcement learning and decision making.” Current opinion in neurobiology 22.6 (2012) : 956-962. [cited by applicant]
Florensa, Carlos, et al. “Reverse curriculum generation for reinforcement learning.” Conference on robot learning. PMLR, 2017. [cited by applicant]
Bellemare et. al., “Skip Context Tree Switching”, 2014, Proceedings of the 31st International Conference on Machine Learning, vol. 32(2), pp. 1458-1466 (Year: 2014). [cited by applicant]
Zhong et. al., “Toward a self-organizing pre-symbolic neural model representing sensorimotor primitives”, 2014, Frontiers in Behavioral Neuroscience, vol. 8, pp. 1-11 (Year: 2014). [cited by applicant]
Brandes et al., “ASAP: A Machine Learning Framework for Local Protein Properties”, 2016, Database, vol. 2016, pp. 1-10 (Year: 2016). [cited by applicant]
Tjalkens et al., “Context Tree Weighting: Multi-Alphabet Sources”, 1993, Proceedings of the 14th Symposium on Information Theory in the Benelux, vol. 14(1993), pp. 128-135 (Year: 1993). [cited by applicant]
Zaheer et al., “Latent LSTM Allocation: Joint Clustering and Non-Linear Dynamic Modeling of Sequence Data”, 2017, Proceedings of the 34th International Conference on Machine Learning, vol. 34 (2017), pp. 3967-3976 (Year… [cited by applicant]
Cited By (1)
US 12,650,895