IP Library Granted Patent US 10,055,687
Granted Patent B2
US 10,055,687 · App. 14/689,052 · Granted Aug 21, 2018

Method for creating predictive knowledge structures from experience in an artificial agent

Inventors: Mark Ring (Anaheim, CA); Tom Schaul (London, GB)
Assignee: Mark B. Ring
G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,055,687
App. No.
14/689,052
Granted
Aug 21, 2018
Kind
B2
Abstract

Building a forecast for an autonomous agent at least comprises assigning a selected parameter of the autonomous agent to a scalar variable, adding a new policy to a set of policies where the new policy maps internal states of the autonomous agent to actions of the autonomous agent in which the mapping may optimize the scalar variable, and adding a new forecast to a set of forecasts where the forecast at least comprises a prediction regarding future values of the scalar variable following execution of the new policy, regardless whether the agent ever actually chooses to take actions in accordance with said new policy. A state of the autonomous agent may be evaluated following completion of each of the agent's actions by comparing the agent's state information with the predicted values of one or more forecasts. Whether to build an additional forecast may be determined based on the evaluation.

Claims (50)

1. A method comprising the steps of:

building a forecast for an autonomous agent, said building at least comprising:

selecting a policy from a set of policies, said policy mapping states of said autonomous agent to actions of said autonomous agent; and

automatically choosing and adding a new forecast to a set of forecasts, said new forecast at least comprising a prediction regarding future states of said autonomous agent during execution and termination of said policy based on a closed-loop sequence of actions, where the policy is considered with conditions for termination of the policy;

evaluating a state of said autonomous agent following termination of said policy, said evaluation at least comprising comparing said state with said prediction;

creating a new policy that optimizes a function over observable signals and forecasts;

building a further new forecast, said further new forecast at least comprising a further prediction regarding future states of said autonomous agent during execution and termination of said new policy based on a closed-loop sequence of actions, where the new policy is considered with conditions for termination of the new policy;

evaluating a state of said autonomous agent following termination of said new policy, said evaluation at least comprising comparing said state with said further prediction; and

determining whether to build an additional forecast, said determining optionally at least in part based on said evaluation.

2. The method as recited in claim 1 , further comprising the steps of:

determining if said forecast is ineffective; and pruning said forecast from said set of forecasts upon said determination.

3. The method as recited in claim 1 , in which said building the forecast further comprises:

building a new policy, said building at least comprising:

selecting a state value of said autonomous agent; and

adding a new policy to the set of policies, said new policy mapping states of said autonomous agent to actions of said autonomous agent, said actions optimizing said state value.

4. The method as recited in claim 1 , in which said set of forecasts comprises a hierarchical structure.

5. The method as recited in claim 3 , in which said new policy further comprises starting and stopping criteria.

6. The method as recited in claim 1 , in which a state, set of states or state value predicted by any forecast in said set of forecasts is associated with at least one of said policies in said set of policies.

7. The method as recited in claim 3 , in which said selected state value comprises at least one of an observation signal, a forecast of interest, a function of a combination of observation signals, and a function of forecast values in said set of forecasts.

8. The method as recited in claim 1 , in which said step of determining whether to terminate the policy is further based on a threshold value.

9. The method as recited in claim 1 , in which said set of forecasts and said set of policies comprise a hierarchical structure.

10. A method comprising:

steps for building a forecast for an autonomous agent, said steps for building at least comprising selecting a policy from a set of policies, said policy mapping states of said autonomous agent to actions of said autonomous agent, and automatically choosing and adding a new forecast to a set of forecasts, said new forecast at least comprising a prediction regarding future states of said autonomous agent during execution and termination of said policy based on a closed-loop sequence of actions, where the policy is considered with conditions for termination of the policy;

steps for evaluating a state of said autonomous agent following termination of said policy;

steps for creating a new policy that optimizes a function over observable signals and forecasts;

steps for building a further new forecast, said further new forecast at least comprising a further prediction regarding future states of said autonomous agent during execution and termination of said new policy based on a closed-loop sequence of actions, where the new policy is considered with conditions for termination of the new policy;

steps for evaluating a state of said autonomous agent following termination of said new policy, said evaluation at least comprising comparing said state with said further prediction; and

steps for determining whether to build an additional forecast.

11. The method as recited in claim 10 , further comprising:

steps for determining if said forecast is ineffective; steps for pruning said forecast from said set of forecasts upon said determination.

12. A non-transitory computer-readable storage medium with an executable program stored thereon, wherein the program instructs one or more processors to perform the following steps:

building a forecast for an autonomous agent, said building at least comprising:

selecting a policy from a set of policies, said policy mapping states of said autonomous agent to actions of said autonomous agent; and

automatically choosing and adding a new forecast to a set of forecasts, said new forecast at least comprising a prediction regarding future states of said autonomous agent during execution and termination of said policy based on a closed-loop sequence of actions, where the policy is considered with conditions for termination of the policy;

evaluating a state of said autonomous agent following termination of said policy, said evaluation at least comprising comparing said state with said prediction;

creating a new policy that optimizes a function over observable signals and forecasts;

building a further new forecast, said further new forecast at least comprising a further prediction regarding future states of said autonomous agent during execution and termination of said new policy based on a closed-loop sequence of actions, where the new policy is considered with conditions for termination of the new policy;

evaluating a state of said autonomous agent following termination of said new policy, said evaluation at least comprising comparing said state with said further prediction; and

determining whether to build an additional forecast, said determining optionally based at least in part on said evaluation.

13. The program instructing the one or more processors as recited in claim 12 , further comprising the steps of: determining if said forecast is ineffective; and pruning said forecast from said set of forecasts upon said determination.

14. The program instructing the one or more processors as recited in claim 12 , in which said building the forecast further comprises:

building a new policy, said building at least comprising:

selecting a state value of said autonomous agent; and

adding a new policy to the set of policies, said new policy mapping states of said autonomous agent to actions of said autonomous agent, said actions optimizing said state value.

15. The program instructing the one or more processors as recited in claim 12 , in which said set of forecasts comprises a hierarchical structure.

16. The program instructing the one or more processors as recited in claim 14 , in which said new policy further comprises starting and stopping criteria.

17. The program instructing the one or more processors as recited in claim 12 , in which a state, set of states or state value predicted by any forecast in said set of forecasts is associated with at least one of said policies in said set of policies.

18. The program instructing the one or more processors as recited in claim 14 , in which said selected state value comprises at least one of an observation signal, a forecast of interest, a function of a combination of observation signals, and a function of forecast values in said set of forecasts.

19. The program instructing the one or more processors as recited in claim 14 , in which said step of determining whether to terminate the policy is further based on a threshold value.

20. The program instructing the one or more processors as recited in claim 12 , in which said set of forecasts and said set of policies comprise a hierarchical structure.

Assignments (6)
CHANGE OF NAME Recorded Apr 24, 2026
From: SONY CORPORATION
To: SONY GROUP CORPORATION
Reel/Frame 075474/0301 →
CORRECTIVE ASSIGNMENT TO CORRECT THE PROVISIONAL APPLICATION NUMBER 61891006 TO 61981006. PREVIOUSLY RECORDED AT REEL: 051588 FRAME: 0419. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jan 24, 2020
From: COGITAI, INC.
To: SONY CORPORATION OF AMERICA; SONY CORPORATION
Reel/Frame 051692/0807 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2020
From: COGITAI, INC.
To: SONY CORPORATION OF AMERICA; SONY CORPORATION
Reel/Frame 051588/0419 →
SECURITY INTEREST Recorded May 24, 2019
From: COGITAI, INC.
To: SONY CORPORATION OF AMERICA
Reel/Frame 049278/0735 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 22, 2019
From: RING, MARK
To: COGITAI, INC.
Reel/Frame 049258/0366 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2016
From: SCHAUL, TOM, PHD
To: RING, MARK B, PHD
Reel/Frame 038924/0847 →
Continuity (2)
Provisional Application 61981006 · Apr 17, 2014
Related Publication 20160012338A1 · Jan 14, 2016