IP Library Granted Patent US 12,699,929
Granted Patent B2
US 12,699,929 · App. 18/214,528 · Granted Aug 4, 2026

Computer-readable recording medium storing multi-agent reinforcement learning program, information processing apparatus, and multi-agent reinforcement learning method

Inventors: Yoshihiro Okawa (Yokohama, JP); Hayato Dan (Yokohama, JP); Natsuki Ishikawa (Yamato, JP); Masatoshi Ogawa (Zama, JP)
Assignee: Fujitsu Limited
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,699,929
App. No.
18/214,528
Granted
Aug 4, 2026
Kind
B2
Abstract

A process includes obtaining, according to a predetermined update order of a policy parameter of each agent, a degree of influence on a constraint specific to a first agent and a degree of influence on a system-wide constraint in which a degree of influence by an updated policy parameter of a second agent previous to the first agent in the update order is shared, and in a case where an update width of the policy parameter of the first agent exists in both ranges respectively determined depending on the degree of influence on the constraint specific to the first agent and the degree of influence on the system-wide constraint, updating the policy parameter of the first agent and causing the degree of influence on the system-wide constraints by the updated policy parameter to be shared with a third agent next to the first agent in the update order.

Claims (17)

1 . A non-transitory computer-readable recording medium storing a multi-agent reinforcement learning program for causing a computer to execute a process for a constrained control problem in which a plurality of agents are involved, the process comprising:

obtaining, according to a predetermined update order of a policy parameter of each agent of the plurality of agents, a degree of influence on a constraint specific to a first agent of the plurality of agents and a degree of influence on a system-wide constraint in which a degree of influence by an updated policy parameter of a second agent previous to the first agent in the update order is shared, the second agent being included in the plurality of agents;

in a case where an update width of the policy parameter of the first agent exists in both ranges respectively determined depending on the degree of influence on the constraint specific to the first agent and the degree of influence on the system-wide constraint, updating the policy parameter of the first agent and causing the degree of influence on the system-wide constraints by the updated policy parameter to be shared with a third agent next to the first agent in the update order, the third agent being included in the plurality of agents;

in a case where the update width of the policy parameter of the first agent does not exist in any of the ranges respectively determined depending on the degree of influence on the constraint specific to the first agent and the degree of influence on the system-wide constraint, aborting update of the policy parameter of the first agent and the policy parameter of the third agent; and

randomly determining the update order of the policy parameter of each agent of the plurality of agents to update policy parameter for the purpose of one of maximizing and minimizing a reward common to the agents, and wherein the process is configured to obtain the degree of influence on the constraint specific to the first agent and the degree of influence on the system-wide constraint according to the determined update order.

2 . An information processing apparatus comprising:

a memory; and

a processor coupled to the memory and configured to:

obtain, according to a predetermined update order of a policy parameter of each agent of a plurality of agents, a degree of influence on a constraint specific to a first agent of the plurality of agents and a degree of influence on a system-wide constraint in which a degree of influence by an updated policy parameter of a second agent previous to the first agent in the update order is shared, the second agent being included in the plurality of agents,

in a case where an update width of the policy parameter of the first agent exists in both ranges respectively determined depending on the degree of influence on the constraint specific to the first agent and the degree of influence on the system-wide constraint, update the policy parameter of the first agent and cause the degree of influence on the system-wide constraints by the updated policy parameter to be shared with a third agent next to the first agent in the update order, the third agent being included in the plurality of agents;

in a case where the update width of the policy parameter of the first agent does not exist in any of the ranges respectively determined depending on the degree of influence on the constraint specific to the first agent and the degree of influence on the system-wide constraint, abort update of the policy parameter of the first agent and the policy parameter of the third agent; and

randomly determine the update order of the policy parameter of each agent of the plurality of agents to update policy parameter for the purpose of one of maximizing and minimizing a reward common to the agents, and wherein the processor is configured to obtain the degree of influence on the constraint specific to the first agent and the degree of influence on the system-wide constraint according to the determined update order.

3 . A multi-agent reinforcement learning method for causing a computer to execute a process for a constrained control problem in which a plurality of agents are involved, the process comprising:

obtaining, according to a predetermined update order of a policy parameter of each agent of the plurality of agents, a degree of influence on a constraint specific to a first agent of the plurality of agents and a degree of influence on a system-wide constraint in which a degree of influence by an updated policy parameter of a second agent previous to the first agent in the update order is shared, the second agent being included in the plurality of agents;

in a case where an update width of the policy parameter of the first agent exists in both ranges respectively determined depending on the degree of influence on the constraint specific to the first agent and the degree of influence on the system-wide constraint, updating the policy parameter of the first agent and causing the degree of influence on the system-wide constraints by the updated policy parameter to be shared with a third agent next to the first agent in the update order, the third agent being included in the plurality of agents;

in a case where the update width of the policy parameter of the first agent does not exist in any of the ranges respectively determined depending on the degree of influence on the constraint specific to the first agent and the degree of influence on the system-wide constraint, aborting update of the policy parameter of the first agent and the policy parameter of the third agent; and

randomly determining the update order of the policy parameter of each agent of the plurality of agents to update policy parameter for the purpose of one of maximizing and minimizing a reward common to the agents, and wherein the process is configured to obtain the degree of influence on the constraint specific to the first agent and the degree of influence on the system-wide constraint according to the determined update order.