IP Library › Granted Patent US 8,290,883
Granted Patent B2
US 8,290,883 · App. 12/556,872 · Granted Oct 16, 2012

Learning system and learning method comprising an event list database

Assignee: Honda Motor Co., Ltd.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,290,883
App. No.
12/556,872
Granted
Oct 16, 2012
Kind
B2
Abstract

A learning system according to the present invention includes an event list database for storing a plurality of event lists, each of the event lists being a set including a series of state-action pairs which reaches a state-action pair immediately before earning a reward, an event list managing section for classifying state-action pairs into the plurality of event lists for storing, and a learning control section for updating expectation of reward of a state-action pair which is an element of each of the event lists.

Claims (11)

1. A learning system comprising:

an event list database for storing a plurality of event lists, each of the event lists being a set including a series of state-action pairs which reaches a state-action pair immediately before earning a reward;

an event list managing section for classifying state-action pairs into the plurality of event lists for storing; and

a learning control section for updating expectation of reward of a state-action pair which is an element of each of the plurality of event lists.

2. A learning system according to claim 1 further comprising a temporary list storing section wherein every time an action is selected the event list managing section has the state-action pair stored in the temporary list storing section and every time a reward is earned the event list managing section has a state-action pair in a set of state-action pairs stored in the temporary list storing section, which has not been stored in the event list database, stored as an element of the event list of the state-action pair immediately before earning the reward in the event list database.

3. A learning system according to claim 1 wherein every time a reward is earned the learning control section updates, using a value of the reward, expectation of reward of a state-action pair which is an element of the event list of the state-action pair immediately before earning the reward and updates, using 0 as a value of reward, expectation of reward of a state-action pair which is an element of the event lists except the event list of the state-action pair immediately before earning the reward.

4. A learning method in a learning system having an event list database for storing a plurality of event lists, each of the event lists being a set including a series of state-action pairs which reaches a state-action pair immediately before earning a reward, an event list managing section and a learning control section, the method comprising the steps of:

classifying, by the event list managing section, state-action pairs into the plurality of event lists for storing; and

updating, by the learning control section, expectation of reward of a state-action pair which is an element of each of the plurality of event lists.

5. A learning method according to claim 4 wherein every time an action is selected the event list managing section has the state-action pair temporarily stored and every time a reward is earned the event list managing section has a state-action pair in a set of state-action pairs temporarily stored, which has not been stored in the event list database, stored as an element of the event list of the state-action pair immediately before earning the reward in the event list database.

6. A learning method according to claim 4 wherein every time a reward is earned the learning control section updates, using a value of the reward, expectation of reward of a state-action pair which is an element of the event list of the state-action pair immediately before earning the reward and updates, using 0 as a value of reward, expectation of reward of a state-action pair which is an element of the event lists except the event list of the state-action pair immediately before earning the reward.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 25, 2009
From: TAKEUCHI, JOHANE; TSUJINO, HIROSHI
To: HONDA MOTOR CO., LTD.
Reel/Frame 023576/0501 →
Priority Claims (1)
JP 2009-187526 · Aug 12, 2009 · national
Continuity (2)
Provisional Application 61136610 · Sep 18, 2008
Related Publication 20100070439A1 · Mar 18, 2010