IP Library Patent Application 18814671
Patent Application
App. No. 18/814,671

POLICY TRAINING DEVICE, POLICY TRAINING METHOD, AND COMMUNICATION SYSTEM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/814,671
Abstract

A policy training device that trains, through first reinforcement learning, a first agent configured to output a first action of a control object according to an input of a first state of the control object, includes a memory, and processor circuitry coupled to the memory and configured to change a first parameter regarding a constraint condition in the first reinforcement learning for every predetermined number of times of a training operation in the first reinforcement learning, and train the first agent by using the first parameter as at least a part of the first state and by ensuring that the constraint condition is satisfied.

Claims (38)

1 . A policy training device that trains, through first reinforcement learning, a first agent configured to output a first action of a control object according to an input of a first state of the control object, the policy training device comprising:

a memory; and

processor circuitry coupled to the memory and configured to:

change a first parameter regarding a constraint condition in the first reinforcement learning for every predetermined number of times of a training operation in the first reinforcement learning; and

train the first agent by using the first parameter as at least a part of the first state and by ensuring that the constraint condition is satisfied.

2 . The policy training device according to claim 1 , wherein the processor circuitry is configured to randomly change the first parameter within a predetermined change range.

3 . The policy training device according to claim 1 ,

wherein the processor circuitry is further configured to;

acquire a cost from the first state of the control object,

wherein the constraint condition is the constraint condition related to the cost, and

the first parameter is a threshold value of the cost.

4 . The policy training device according to claim 1 ,

wherein the processor circuitry is further configured to:

acquire a cost from the first state of the control object,

wherein the constraint condition is the constraint condition related to the cost, and

the first parameter is used to acquire the cost.

5 . The policy training device according to claim 1 ,

wherein the processor circuitry is further configured to:

determine a change range of the first parameter, according to a cost when the control object performs a second action output from a second agent trained through second reinforcement learning.

6 . The policy training device according to claim 1 ,

wherein the processor circuitry is further configured to acquire a third action of the control object output from the first agent in response to an input of a second state of the control object; and

output the acquired third action,

wherein the processor circuitry is configured to input a second parameter regarding the constraint condition to the first agent as at least the part of the second state.

7 . A policy training method of a policy training device that trains, through first reinforcement learning, a first agent configured to output a first action of a control object according to an input of a first state of the control object, the policy training method for causing a computer to execute a process, the process comprising:

changing a first parameter regarding a constraint condition in the first reinforcement learning for every predetermined number of times of a training operation in the first reinforcement learning; and

training the first agent by using the first parameter as at least a part of the first state and by ensuring that the constraint condition is satisfied.

8 . The policy training method according to claim 7 , the process further comprising:

determining a change range of the first parameter, according to a cost when the control object performs a second action output from a second agent trained through second reinforcement learning.

9 . A communication system comprising:

a base station device; and

a policy training device configured to train, through first reinforcement learning, a first agent configured to output a first action of a control object according to an input of a first state of the control object, the policy training device including:

a memory, and

processor circuitry coupled to the memory and configured to:

change a first parameter regarding a constraint condition in the first reinforcement learning for every predetermined number of times of a training operation in the first reinforcement learning, and

train the first agent by using the first parameter as at least a part of the first state and by ensuring that the constraint condition is satisfied.

10 . The communication system according to claim 9 ,

wherein the processor circuitry is further configured to:

determine a change range of the first parameter, according to a cost when the control object performs a second action output from a second agent trained through second reinforcement learning.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER 19345640 TO CHANGE IT TO 19346640 AND TO CORRECT THE APPLICATION 19317078 TO CHANGE IT TO 19319078. PREVIOUSLY RECORDED ON REEL 73477 FRAME 351. ASSIGNOR(S) HEREBY CONFIRMS THE THE ASSIGNMENT. Recorded Jan 20, 2026
From: FUJITSU LIMITED
To: 1FINITY INC.
Reel/Frame 075265/0471 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2025
From: FUJITSU LIMITED
To: 1FINITY INC.
Reel/Frame 073477/0351 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2024
From: TERANISHI, YUTA; ABE, FUMIKA; OGAWA, MASATOSHI; ISHIKAWA, NATSUKI; OKAWA, YOSHIHIRO
To: FUJITSU LIMITED
Reel/Frame 068393/0765 →