IP Library Granted Patent US 11,443,229
Granted Patent B2
US 11,443,229 · App. 16/120,111 · Granted Sep 13, 2022

Method and system for continual learning in an intelligent artificial agent

Inventors: Mark Bishop Ring (Anaheim, CA); Satinder Baveja (Ann Arbor, MI); Roberto Capobianco (Itri, IT); Varun Kompella (Aachen, DE); Kaushik Subramanian (Richmond, CA); James MacGlashan (Riverside, RI)
Assignees: Sony Group Corporation; Sony Corporation of America
G06N20/00G06N3/088G06N5/043
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,443,229
App. No.
16/120,111
Granted
Sep 13, 2022
Kind
B2
Abstract

A method and system for teaching an artificial intelligent agent includes giving the agent several examples where it can learn to identify what is important about these example states. Once the agent has the ability to recognize a goal configuration, it can use that information to then learn how to achieve the goal states on its own. An agent may be provided with positive and negative examples to demonstrate a goal configuration. Once the agent has learned certain goal configurations, the agent can learn an option to achieve the goal configuration and a distance function that predicts at least one of a distance and a duration to the goal configuration under the learned option. This distance function prediction may be incorporated as a state feature of the agent.

Claims (33)

1. A method for training an artificial intelligent agent, comprising: defining, within the agent, a first continual learning block to include a first skill to achieve a first goal configuration for the agent and a first knowledge feature providing a first prediction of at least one of a distance and duration to achieve the first goal configuration;

using the first skill to move the agent in the first goal configuration;

defining, within the agent, a second continual learning block, including a second goal configuration, distinct from the first goal configuration, and a second knowledge feature providing a second prediction of at least one of a distance and duration to achieve the second goal configuration, wherein the second continual learning block builds upon the first continual learning block, and

using the first prediction by the second continual learning block to move the agent to the second goal configuration.

2. The method of claim 1 , further comprising:

using features of the first goal configuration for achievement of the second goal configuration.

3. The method of claim 1 , wherein the first knowledge feature is a value function based on the first goal configuration as a termination condition.

4. The method of claim 1 , further comprising:

providing positive examples via an interface to the agent when the agent is in the first goal configuration;

providing negative examples via the interface to the agent when the agent is not in the first goal configuration; and

extracting key state features to determine what features are important during receipt of positive examples to the agent.

5. The method of claim 1 , further comprising incorporating the first prediction as a state feature of the agent.

6. The method of claim 1 , wherein the first knowledge feature is selected from the group consisting of a distance function, a time to completion, a time to initiation of something else, and a prediction of a value of a feature at the time of completion.

7. The method of claim 1 , wherein the first knowledge feature is learned, either before, in conjunction with, interleaved with, or after a policy.

8. A method of learning to achieve a goal configuration of an artificial agent, comprising:

defining, within the agent, the goal configuration for the agent as part of a continual learning block;

determining a knowledge feature as a prediction of at least one of a distance and duration required to achieve the goal configuration;

relying on a previous learned continual learning block, having a previously learned distinct goal configuration, to move the agent in the goal configuration;

determining a first knowledge feature as a first prediction of a number of steps required to achieve the goal configuration; and

relying on a previous knowledge feature to achieve the goal configuration, the previous knowledge feature being a previous prediction of at least one of a distance and duration required to achieve the previous learned goal configuration.

9. The method of claim 8 , wherein a previous knowledge feature is used to achieve the goal configuration, the previous knowledge feature being a previous prediction of at least one of a distance and duration required to achieve the previous learned goal configuration.

10. The method of claim 8 , wherein the previous learned goal configuration is an element of a previous continual learning block.

11. The method of claim 10 , wherein the previous continual learning block includes a plurality of previous continual learning blocks, each having a respective previous learned goal configuration and a respective previous knowledge feature.

12. The method of claim 11 , further comprising planning ahead, by the agent, to determine how to most efficiently achieve the respective previous learned goal configurations in order to achieve the goal configuration.

13. The method of claim 8 , wherein the first knowledge feature is selected from the group consisting of a distance function, a time to completion, a time to initiation of something else, and a prediction of a value of a feature at the time of completion.

14. A method of learning to achieve a goal configuration of an artificial agent, comprising:

defining, within the agent, the goal configuration for the agent as part of a continual learning block;

determining, within the agent, a knowledge feature as a prediction of at least one of a duration and a distance required to achieve the goal configuration, the knowledge feature being a component of the continual learning block; and

relying, by the agent, on a previous learned distinct goal configuration, of a previously learned continual learning block, to move the agent in the goal configuration;

wherein a previous knowledge feature is used to achieve the goal configuration, wherein the previous knowledge feature is a previous prediction of at least one of a duration and a distance required to achieve the previous learned goal configuration, and wherein the previous knowledge feature, along with the previous goal configuration, are components of a previous continual learning block.

15. The method of claim 14 , wherein the previous continual learning block includes a plurality of previous continual learning blocks, each having a respective previous learned goal configuration and a respective previous knowledge feature.

16. The method of claim 15 , wherein each of the plurality of the previous continual learning blocks are relied upon to achieve the goal configuration.

17. The method of claim 16 , further comprising planning ahead, by the agent, to determine how to most efficiently achieve the respective previous learned goal configurations in order to achieve the goal configuration.

Assignments (4)
CHANGE OF NAME Recorded Apr 24, 2026
From: SONY CORPORATION
To: SONY GROUP CORPORATION
Reel/Frame 075474/0301 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2020
From: COGITAI, INC.
To: SONY CORPORATION OF AMERICA; SONY CORPORATION
Reel/Frame 051588/0522 →
SECURITY INTEREST Recorded May 24, 2019
From: COGITAI, INC.
To: SONY CORPORATION OF AMERICA
Reel/Frame 049278/0735 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2018
From: RING, MARK BISHOP; BAVEJA, SATINDER; CAPOBIANCO, ROBERTO; KOMPELLA, VARUN; SUBRAMANIAN, KAUSHIK; MACGLASHAN, JAMES
To: COGITAI, INC.
Reel/Frame 047197/0208 →
Continuity (1)
Related Publication 20200074349A1 · Mar 5, 2020