IP Library Granted Patent US 12,412,075
Granted Patent B2
US 12,412,075 · App. 17/624,552 · Granted Sep 9, 2025

Behavior learning system, behavior learning method and program

Inventors: Yuichiro Dan (Tokyo, JP); Keita Hasegawa (Tokyo, JP); Takafumi Harada (Tokyo, JP); Tomoaki Washio (Tokyo, JP); Yoshihito Oshima (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G06N3/045G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,075
App. No.
17/624,552
Granted
Sep 9, 2025
Kind
B2
Abstract

An action learning system includes a memory, and a processor configured to train, based on first data indicating a property of an environment in which data is collected from multiple devices and to which an action determined by a first neural network according to a state of the environment is applied and second data indicating a property of the environment to which the action is not applied, a second neural network that calculates a similarity degree between distributions of the first data and the second data, and train, after the second neural network is trained, the first neural network that determines an action according to the state of the environment, by reinforcement learning including, in a reward, a value that changes based on a relationship between a similarity degree and a parameter set by a user, the similarity degree being calculated by the second neural network based.

Claims (21)

1. An action learning system comprising:

a memory; and

a processor configured to train, based on first data indicating a property of an environment in which data is collected from a plurality of devices and to which an action determined by a first neural network according to a state of the environment is applied and second data indicating a property of the environment to which the action is not applied, a second neural network,

train, after the second neural network is trained, the first neural network that determines an action according to the state of the environment, by reinforcement learning using a reward including a value that changes based on a relationship between a similarity degree and a parameter set by a user, the similarity degree being calculated by the second neural network based on third data indicating a property of the environment to which an action determined by the first neural network is applied, wherein the reward is a value that increases as the similarity degree is closer to the parameter set by the user; and

generating, by the trained first neural network, an example network attack action that causes a cyberattack to a network to suppress execution of the cyberattack according to the example network attack action, wherein the example network attack action, if executed, causes a damage to the network while maintaining abnormality in the network caused by the cyberattack undetected based on the similarity degree.

2. The action learning system according to claim 1 , wherein

the reward is a value that increases as the similarity degree is closer to the parameter set by the user.

3. The action learning system according to claim 2 , wherein

the processor trains the first neural network by reinforcement learning using the reward including a value of a Lorenz function that uses the parameter as a location parameter and the similarity degree as a variable.

4. The action learning system according to claim 1 , wherein

the processor trains the second neural network according to a GAN (generative adversarial network) algorithm.

5. An action learning method for execution by a computer, the action learning method comprising:

training, based on first data indicating a property of an environment in which data is collected from a plurality of devices and to which an action determined by a first neural network according to a state of the environment is applied and second data indicating a property of the environment to which the action is not applied, a second neural network; and

training, after the second neural network is trained, the first neural network that determines an action according to the state of the environment, by reinforcement learning using a reward including a value that changes based on a relationship between a similarity degree is to a parameter set by a user, the similarity degree being calculated by the second neural network based on third data indicating a property of the environment to which an action determined by the first neural network is applied,

wherein the reward is a value that increases as the similarity degree is closer to the parameter set by the user; and

generating, by the trained first neural network, an example network attack action that causes a cyberattack to a network to suppress execution of the cyberattack according to the example network attack action, wherein the example network attack action, if executed, causes a damage to the network while maintaining abnormality in the network caused by the cyberattack undetected based on the similarity degree.

6. A non-transitory computer-readable recording medium having stored therein a program for causing a computer to execute a process comprising:

training, based on first data indicating a property of an environment in which data is collected from a plurality of devices and to which an action determined by a first neural network according to a state of the environment is applied and second data indicating a property of the environment to which the action is not applied, a second neural network; and

training, after the second neural network is trained, the first neural network that determines an action according to the state of the environment, by reinforcement learning using a reward including a value that changes based on a relationship between a similarity degree is to a parameter set by a user, the similarity degree being calculated by the second neural network based on third data indicating a property of the environment to which an action determined by the first neural network is applied,

wherein the reward is a value that increases as the similarity degree is closer to the parameter set by the user; and

generating, by the trained first neural network, an example network attack action that causes a cyberattack to a network to suppress execution of the cyberattack according to the example network attack action, wherein the example network attack action, if executed, causes a damage to the network while maintaining abnormality in the network caused by the cyberattack undetected based on the similarity degree.

Assignments (2)
CHANGE OF NAME Recorded Oct 22, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 073184/0647 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2022
From: DAN, YUICHIRO; HASEGAWA, KEITA; HARADA, TAKAFUMI; WASHIO, TOMOAKI; OSHIMA, YOSHIHITO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 058531/0830 →
Continuity (1)
Related Publication 20220253677A1 · Aug 11, 2022
References Cited (5)
US 20160188981A1 · Doerring · 2016 [cited by examiner]
US 20190318244A1 · Alvarez · 2019 [cited by examiner]
US 20200019871A1 · Balakrishnan · 2020 [cited by examiner]
Ying Chen et al. (2018) “Evaluation of Reinforcement Learning Based False Data Injection Attack to Automatic Voltage Control”, IEEE Transactions on Smart Grid. [cited by applicant]
Martin Arjovsky et al. (2017) “Wasserstein GAN”, Courant Institute of Mathematical Sciences Facebook AI Research. [cited by applicant]