IP Library › Granted Patent US 11,619,915
Granted Patent B2
US 11,619,915 · App. 16/797,573 · Granted Apr 4, 2023

Reinforcement learning method and reinforcement learning system

Inventors: Hidenao Iwane (Kawasaki, JP); Junichi Shigezumi (Kawasaki, JP); Yoshihiro Okawa (Yokohama, JP); Tomotake Sasaki (Kawasaki, JP); Hitoshi Yanami (Kawasaki, JP)
Assignee: FUJITSU LIMITED
G05B13/0265B25J9/163G06N20/00H02J3/381H02J2300/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,619,915
App. No.
16/797,573
Granted
Apr 4, 2023
Kind
B2
Abstract

A computer-implemented reinforcement learning method includes determining, based on a target probability of satisfaction of a constraint condition related to a state of a control object and a specific time within which a controller causes the state of the control object not satisfying the constraint condition to be the state of the control object satisfying the constraint condition, a parameter of a reinforcement learner that causes, in a specific probability, the state of the control object to satisfy the constraint condition at a first timing following a second timing at which the state of control object satisfies the constraint condition; and determining a control input to the control object by either the reinforcement learner or the controller, based on whether the state of the control object satisfies the constraint condition at a specific timing.

Claims (25)

1. A computer-implemented reinforcement learning method comprising:

determining, based on a target probability of satisfaction of a constraint condition related to a state of a control object and a specific time within which a controller causes the state of the control object not satisfying the constraint condition to be the state of the control object satisfying the constraint condition, a parameter of a reinforcement learner that causes, in a specific probability, the state of the control object to satisfy the constraint condition at a first timing following a second timing at which the state of control object satisfies the constraint condition, wherein the specific probability is set to a probability that is higher than the target probability and calculated based on the specific time and the target probability; and

determining a control input to the control object by either the reinforcement learner or the controller, based on whether the state of the control object satisfies the constraint condition at a specific timing, wherein

the specific time is defined by the number of steps of determining the control input, and

the determining of the parameter includes determining the specific probability as a power root of the number of steps corresponding to the target probability.

2. The reinforcement learning method according to claim 1 , wherein the reinforcement learner is configured to automatically adjust a search range of the control input so that the constraint condition is satisfied with the specific probability.

3. The reinforcement learning method according to claim 1 , wherein the control object is a wind power generation facility, and

the reinforcement learner uses a generator torque of the wind power generation facility as the control input, at least one of a power generation amount of the power generation facility, a rotation amount of a turbine of the power generation facility, a rotation speed of the turbine of the power generation facility, a wind direction for the power generation facility, and a wind speed for the power generation facility as the state, and the power generation amount of the power generation facility as a reward so as to perform reinforcement learning for learning a policy for controlling the control object.

4. The reinforcement learning method according to claim 1 , wherein the determining of the parameter includes determining the specific probability as a square root of the target probability.

5. A computer-readable medium storing therein a reinforcement learning program executable by one or more computing devices, the reinforcement learning program comprising:

one or more instructions for determining, based on a target probability of satisfaction of a constraint condition related to a state of a control object and a specific time within which a controller causes the state of the control object not satisfying the constraint condition to be the state of the control object satisfying the constraint condition, a parameter of a reinforcement learner that causes, in a specific probability, the state of the control object to satisfy the constraint condition at a first timing following a second timing at which the state of control object satisfies the constraint condition, wherein the specific probability is set to a probability that is higher than the target probability and calculated based on the specific time and the target probability; and

one or more instructions for determining a control input to the control object by either the reinforcement learner or the controller, based on whether the state of the control object satisfies the constraint condition at a specific timing, wherein

the specific time is defined by the number of steps of determining the control input, and

the determining of the parameter includes determining the specific probability as a power root of the number of steps corresponding to the target probability.

6. The computer-readable medium according to claim 5 , wherein

the reinforcement learner is configured to automatically adjust a search range of the control input so that the constraint condition is satisfied with the specific probability.

7. The computer-readable medium according to claim 5 , wherein

the control object is a wind power generation facility, and

the reinforcement learner uses a generator torque of the wind power generation facility as the control input, at least one of a power generation amount of the power generation facility, a rotation amount of a turbine of the power generation facility, a rotation speed of the turbine of the power generation facility, a wind direction for the power generation facility, and a wind speed for the power generation facility as the state, and the power generation amount of the power generation facility as a reward so as to perform reinforcement learning for learning a policy for controlling the control object.

8. The computer-readable medium according to claim 5 , wherein the determining of the parameter includes determining the specific probability as a square root of the target probability.

9. A reinforcement learning system comprising:

one or more memories; and

one or more processors coupled to the one or more memories and the one or more processors configured to:

determine, based on a target probability of satisfaction of a constraint condition related to a state of a control object and a specific time within which a controller causes the state of the control object not satisfying the constraint condition to be the state of the control object satisfying the constraint condition, a parameter of a reinforcement learner that causes, in a specific probability, the state of the control object to satisfy the constraint condition at a first timing following a second timing at which the state of control object satisfies the constraint condition, and

cause any one of the reinforcement learner and the controller to determine a control input to the control object, based on whether the state of the control object at a specific timing satisfies the constraint condition.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2020
From: IWANE, HIDENAO; SHIGEZUMI, JUNICHI; OKAWA, YOSHIHIRO; SASAKI, TOMOTAKE; YANAMI, HITOSHI
To: FUJITSU LIMITED
Reel/Frame 051888/0786 →
Priority Claims (1)
JP JP2019-039031 · Mar 4, 2019 · national
Continuity (1)
Related Publication 20200285204A1 · Sep 10, 2020