IP Library › Granted Patent US 11,708,089
Granted Patent B2
US 11,708,089 · App. 17/021,457 · Granted Jul 25, 2023

Systems and methods for curiosity development in agents

Inventors: Ran Tian (Dublin, CA); Haiming Gang (San Jose, CA); David Francis Isele (Sunnyvale, CA)
Assignee: HONDA MOTOR CO., LTD.
B60W60/0011B60W60/0015G05D1/0221G05D2201/0213
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,708,089
App. No.
17/021,457
Granted
Jul 25, 2023
Kind
B2
Abstract

Systems and methods for curiosity development in an agent located in an uncertain environment are provided. In one embodiment, the system includes a goal state module, a curiosity module, and a planning module. The goal module is configured to calculate a goal state of a goal associated with the environment. The curiosity module is configured to determine an uncertainty value for the environment and calculate a curiosity reward based on the uncertainty value. The planning module is configured to update a motion plan based on the goal state and the curiosity reward.

Claims (30)

1. A system for agent exploration of an environment, the system comprising:

a goal state module, implemented via a processor, configured to calculate a goal state of an agent associated with the environment;

a curiosity module, implemented via a processor, configured to:

determine an uncertainty value for the environment; and

calculate a curiosity reward based on the uncertainty value; and

a planning module, implemented via a processor, configured to update a motion plan based on the goal state and the curiosity reward, wherein the updated motion plan causes the agent to collect a new observation when the curiosity reward outweighs the reward associated with the goal state.

2. The system of claim 1 , wherein the uncertainty value is based sensor accuracy of a sensor of the agent.

3. The system of claim 1 , wherein the uncertainty value is based on an object in the environment.

4. The system of claim 1 , wherein the curiosity reward is a accumulative belief state entropy over a planning horizon, wherein the planning horizon is based on an amount of time from an initial action.

5. The system of claim 1 , wherein the agent is a host vehicle, wherein the environment is a roadway, and wherein the goal state is based on a planned maneuver of the host vehicle.

6. The system of claim 5 , wherein the uncertainty value for the roadway is based on a proximate vehicle on the roadway.

7. A method for agent exploration of an agent in an environment, the method comprising:

calculating a goal state of a goal of the agent associated with the environment;

determining an uncertainty value for the environment;

calculating a curiosity reward based on the uncertainty value; and

updating a motion plan based on the goal state and the curiosity reward, wherein the updated motion plan causes the agent to collect a new observation when the curiosity reward outweighs the reward associated with the goal state.

8. The method of claim 7 , wherein the uncertainty value is based sensor accuracy of a sensor of the agent.

9. The method of claim 7 , wherein the uncertainty value is based on an object in the environment.

10. The method of claim 7 , wherein the curiosity reward is a accumulative belief state entropy over a planning horizon, wherein the planning horizon is based on an amount of time from an initial action.

11. The method of claim 7 , wherein the agent is a host vehicle, wherein the environment is a roadway, and wherein the goal state is based on a planned maneuver of the host vehicle.

12. The method of claim 11 , wherein the uncertainty value for the roadway is based on a proximate vehicle on the roadway.

13. A non-transitory computer readable storage medium storing instructions that when executed by a computer having a processor to perform a method for agent exploration of an agent in an environment, the method comprising:

calculating a goal state of a goal of the agent associated with the environment;

determining an uncertainty value for the environment;

calculating a curiosity reward based on the uncertainty value; and

updating a motion plan based on the goal state and the curiosity reward, wherein the updated motion plan causes the agent to collect a new observation when the curiosity reward outweighs the reward associated with the goal state.

14. The non-transitory computer readable storage medium of claim 13 , wherein the uncertainty value is based sensor accuracy of a sensor of the agent.

15. The non-transitory computer readable storage medium of claim 13 , wherein the uncertainty value is based on an object in the environment.

16. The non-transitory computer readable storage medium of claim 13 , the method wherein the curiosity reward is a accumulative belief state entropy over a planning horizon, wherein the planning horizon is based on an amount of time from an initial action.

17. The non-transitory computer readable storage medium of claim 13 , wherein the agent is a host vehicle, wherein the environment is a roadway, wherein the goal state is based on a planned maneuver of the host vehicle, wherein the uncertainty value for the roadway is based on a proximate vehicle on the roadway.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2020
From: TIAN, RAN; GANG, HAIMING; ISELE, DAVID FRANCIS
To: HONDA MOTOR CO., LTD.
Reel/Frame 053776/0301 →
Continuity (2)
Provisional Application 62983345 · Feb 28, 2020
Related Publication 20210269060A1 · Sep 2, 2021