IP Library › Granted Patent US 11,119,451
Granted Patent B2
US 11,119,451 · App. 16/537,526 · Granted Sep 14, 2021

Apparatus, method, program, and recording medium

Inventors: Takamitsu Matsubara (Nara, JP); Yunduan Cui (Nara, JP); Lingwei Zhu (Nara, JP); Hiroaki Kanokogi (Tokyo, JP); Morihiro Fujisaki (Tokyo, JP); Go Takami (Tokyo, JP); Yota Furukawa (Tokyo, JP)
Assignees: Yokogawa Electric Corporation; NATIONAL UNIVERSITY CORPORATION NARA INSTITUTE OF SCIENCE AND TECHNOLOGY
G05B13/02F01K23/101G06N5/003G06N20/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,119,451
App. No.
16/537,526
Granted
Sep 14, 2021
Kind
B2
Abstract

Provided is an apparatus including a plurality of agents that each set some devices among a plurality of devices provided in a facility to be target devices, wherein each of the plurality of agents includes a state acquiring section that acquires state data indicating a state of the facility; a control condition acquiring section that acquires control condition data indicating a control condition of each target device; and a learning processing section that uses learning data including the state data and the control condition data to perform learning processing of a model that outputs recommended control condition data indicating a control condition recommended for each target device in response to input of the state data.

Claims (37)

1. An apparatus comprising a plurality of agents that each set one or more devices among a plurality of devices provided in a facility to be target devices, wherein each of the plurality of agents includes:

a control section that controls the target device according to a control condition indicated by recommended control condition data;

a state acquiring section that acquires state data indicating a state of the facility after the target device has been controlled by the control section;

a control condition acquiring section that acquires control condition data indicating a control condition of each target device;

a learning processing section that uses kernel dynamic policy programming to generate learning data including the state data and the control condition data to perform learning processing of a model that outputs recommended control condition data indicating a control condition recommended for each target device in response to input of the state data and distributes the learning processing for the target device among the plurality of agents.

2. The apparatus according to claim 1 , wherein

each learning processing section performs the learning processing of the model using the learning data and a reward value determined according to a preset reward function, and

each model outputs the recommended control condition data indicating the control condition of each target device recommended for increasing the reward value beyond a reference reward value, in response to input of the state data.

3. The apparatus according to claim 1 , wherein

each of the plurality of agents further includes a recommended control condition output section that outputs the recommended control condition data obtained by supplying the model with the state data.

4. The apparatus according to claim 3 , wherein

each recommended control condition output section uses the model to output the recommended control condition data indicating a closest control condition included in a control condition series that is most highly recommended, among a plurality of control condition series obtained by selecting any one control condition from among a plurality of control conditions of the target device at each of a plurality of future timings.

5. The apparatus according to claim 1 , wherein

the state acquiring sections in at least two agents among the plurality of agents acquire the state data that is common therebetween.

6. The apparatus according to claim 1 , wherein

the state acquiring section of at least one agent among the plurality of agents acquires the state data that further includes control condition data indicating a control condition of a device that is not a target device of the at least one agent among the plurality of devices.

7. The apparatus according to claim 1 , wherein

the learning processing section of each agent performs learning processing using kernel dynamic policy programming for the target device.

8. The apparatus according to claim 1 , wherein

a collection of target devices of each of the plurality of agents includes only target devices that are not in a collection of target devices of each other agent differing from this agent among the plurality of agents.

9. The apparatus according to claim 8 , wherein

each of the plurality of agents sets a single device to be the target device.

10. The apparatus of claim 1 further comprising:

where the number of devices among the plurality of devices is decreased in response to outputting recommended control condition data indicating a control condition.

11. A method in which a plurality of agents, which each set one or more devices among a plurality of devices provided in a facility to be target devices, comprising:

acquiring control condition data indicating a control condition of each target device;

controlling the target device according to the control condition indicated by recommended control condition data;

acquiring state data indicating a state of the facility in response to the target device being controlled according to the control condition; and

using kernel dynamic policy programming to generate learning data including the state data and the control condition data to perform learning processing of a model that outputs recommended control condition data indicating a control condition recommended for each target device in response to input of the state data.

12. A non-transitory recording medium storing thereon a program that causes one or more computers to function as a plurality of agents that each set one or more devices among a plurality of devices provided in a facility to be target devices, wherein each of the plurality of agents includes:

a state acquiring section that acquires state data indicating a state of the facility;

a control condition acquiring section that acquires control condition data indicating a control condition of each target device;

a learning processing section that uses kernel dynamic policy programming to generate learning data including the state data and the control condition data to perform learning processing of a mod& that outputs recommended control condition data indicating a control condition recommended for each target device in response to input of the state data;

a control section that controls the target device according to the control condition indicated by the recommended control condition data; each state acquiring section acquires the state data in response to the target device being controlled by the control section;

the learning processing for the target device is distributed among the plurality of agents.

13. The non-transitory recording medium of claim 12 further comprising:

where the number of devices among the plurality of devices is decreased in response to outputting recommended control condition data indicating a control condition.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2019
From: KANOKOGI, HIROAKI; FUJISAKI, MORIHIRO; TAKAMI, GO; FURUKAWA, YOTA
To: YOKOGAWA ELECTRIC CORPORATION
Reel/Frame 050069/0446 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2019
From: MATSUBARA, TAKAMITSU; CUI, YUNDUAN; ZHU, LINGWEI
To: NATIONAL UNIVERSITY CORPORATION NARA INSTITUTE OF SCIENCE AND TECHNOLOGY
Reel/Frame 050069/0452 →
Priority Claims (1)
JP JP2018-153340 · Aug 17, 2018 · national
Continuity (1)
Related Publication 20200057416A1 · Feb 20, 2020