IP Library Granted Patent US 12,386,692
Granted Patent B2
US 12,386,692 · App. 18/252,223 · Granted Aug 12, 2025

Fault recovery support apparatus, fault recovery support method and program

Inventors: Hiroki Ikeuchi (Tokyo, JP); Yosuke Takahashi (Tokyo, JP); Kotaro Matsuda (Tokyo, JP); Tsuyoshi Toyono (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G06F11/0793
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,386,692
App. No.
18/252,223
Granted
Aug 12, 2025
Kind
B2
Abstract

A fault recovery support apparatus according to an embodiment includes: a fault insertion unit configured to insert a fault into a target system; a behavior execution unit configured to execute a behavior related to recovery from the fault for the target system; a first construction unit configured to construct an automaton representing a recovery process from the fault by using observation data acquired from the target system as a result of the behavior by the behavior execution unit; and a second construction unit configured to construct a workflow representing a behavior for separating each fault included in a plurality of faults and a recovery process of each fault by using a plurality of the automatons related to the plurality of faults.

Claims (40)

1. A fault recovery support apparatus comprising:

a memory; and

a processor configured to execute:

inserting, by a fault inserter, a plurality of faults into a target system;

executing, by a behavior executer, a behavior related to recovery from each of the plurality of the faults for the target system;

constructing, by a first constructor, a plurality of automatons, each representing a recovery process from each of the plurality of the faults, by clustering observation data acquired from the target system as a result of the behavior by the behavior executer; and

constructing, by a second constructor, a workflow representing a behavior for separating each fault included in a plurality of faults and a recovery process of each fault by using a plurality of the automatons related to the plurality of faults,

wherein the constructing by the first constructor includes

clustering the observation data to allocate discrete numbers to a third behavior sequence formed by combining a first behavior sequence for the target system before the observation data is obtained and a second behavior sequence for the target system when the observation data is obtained, and

constructing an automaton based on an observation table in which values of elements having the first behavior sequence as a row index and the second behavior sequence as a column index are set as the discrete numbers by an L* algorithm,

wherein the constructing by the second constructor includes constructing the workflow in which a type of fault and a state that is able to be a state of the target system are indicated by observation data acquired when a behavior is executed on the target system in a form of a directed graph by using a plurality of the automatons of which initial states are not distinguishable,

wherein the workflow is a directed graph comprising a plurality of nodes and a plurality of directed edges,

wherein each node is defined as a tuple of a label of the observation data and a vector indicating current states of the plurality of automatons corresponding to the plurality of faults;

wherein each directed edge from a first node to a second node represents a behavior executable on the target system and the label of the observation data obtained as a result of executing the behavior based on a current state,

wherein the executing, by the behavior executer, includes transmitting, to the target system, a command corresponding to the behavior represented by the each directed edge in the workflow, and acquire observation data from the target system in response to the command execution,

wherein the executing, by the behavior executer, includes:

assigning the label of the observation data to the observation data using a result of the clustering,

determining a transition to a next node in the workflow based on the assigned label of the observation data, and updating the vector indicating the current states of the plurality of automatons corresponding to the plurality of faults, and

executing, based on the determined transition and the updated vector, a recovery control for the target system so as to recover the target system from a fault state toward a normal state.

2. The fault recovery support apparatus according to claim 1 , wherein

the constructing by the first constructor includes

giving the discrete number to the third behavior sequence by a classifier that has been learned in advance.

3. A fault recovery support method of causing a computer to perform:

inserting a plurality of faults into a target system;

executing a behavior related to recovery from each of the plurality of the faults for the target system;

constructing a plurality of automatons, each representing a recovery process from each of the plurality of the faults, by clustering observation data acquired from the target system as a result of the behavior in the behavior executer; and

constructing a workflow representing a behavior for separating each fault included in a plurality of faults and a recovery process of each fault by using a plurality of the automatons related to the plurality of faults,

wherein the constructing the plurality of automatons includes

clustering the observation data to allocate discrete numbers to a third behavior sequence formed by combining a first behavior sequence for the target system before the observation data is obtained and a second behavior sequence for the target system when the observation data is obtained, and

constructing an automaton based on an observation table in which values of elements having the first behavior sequence as a row index and the second behavior sequence as a column index are set as the discrete numbers by an L* algorithm,

wherein the constructing the workflow includes constructing the workflow in which a type of fault and a state that is able to be a state of the target system are indicated by observation data acquired when a behavior is executed on the target system in a form of a directed graph by using a plurality of the automatons of which initial states are not distinguishable,

wherein the workflow is a directed graph comprising a plurality of nodes and a plurality of directed edges,

wherein each node is defined as a tuple of a label of the observation data and a vector indicating current states of the plurality of automatons corresponding to the plurality of faults;

wherein each directed edge from a first node to a second node represents a behavior executable on the target system and an observation label obtained as a result of executing the behavior based on a current state, and

wherein the executing includes transmitting, to the target system, a command corresponding to the behavior represented by the each directed edge in the workflow, and acquiring observation data from the target system in response to the command execution,

wherein the executing, by the behavior executer, includes:

assigning the label of the observation data to the observation data using a result of the clustering,

determining a transition to a next node in the workflow based on the assigned label of the observation data, and updating the vector indicating the current states of the plurality of automatons corresponding to the plurality of faults, and

executing, based on the determined transition and the updated vector, a recovery control for the target system so as to recover the target system from a fault state toward a normal state.

4. A non-transitory computer-readable recording medium having computer-readable instructions stored thereon, which when executed, cause a computer to function as the fault recovery support apparatus according to claim 1 .

Assignments (2)
CHANGE OF NAME Recorded Aug 15, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 072473/0885 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2023
From: IKEUCHI, HIROKI; TAKAHASHI, YOSUKE; MATSUDA, KOTARO; TOYONO, TSUYOSHI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 063579/0124 →
Continuity (1)
Related Publication 20230409425A1 · Dec 21, 2023
References Cited (7)
CN 107025344A · 2017 [cited by examiner]
H. Ikeuchi et al., “A framework for automatic failure recovery in ICT systems by deep reinforcement learning,” in 2020 IEEE 40th International Conference on Distributed Computing Systems (ICDCS). IEEE, 2020. [cited by applicant]
K. R. Joshi et al., “Probabilistic model-driven recovery in distributed systems,” IEEE Transactions on Dependable and Secure Computing, vol. 8, No. 6, pp. 913-928, 2010. [cited by applicant]
A. Watanabe et al., “Workflow extraction for service operation using multiple unstructured trouble tickets,” IEICE Transactions on Information and Systems, vol. 101, No. 4, pp. 1030-1041, 2018. [cited by applicant]
D. Angluin, “Learning regular sets from queries and counterexamples,” Information and computation, vol. 75, No. 2, pp. 87 to 106, 1987. [cited by applicant]
M. Ester et al., “A density-based algorism for discovering clusters in large spatial databases with noise.” in Kdd, vol. 96, No. 34, 1996, pp. 226-231. [cited by applicant]
L. Mclnnes et al., “Umap: Uniform manifold approximation and projection for dimension reduction,” arXiv preprint arXiv:1802.03426, Sep. 18, 2020. [cited by applicant]