IP Library Granted Patent US 12,650,907
Granted Patent B2
US 12,650,907 · App. 18/828,495 · Granted Jun 9, 2026

Storage system and failure handling method in storage system

Inventors: Kenichi Betsuno (Tokyo, JP); Norio Shimozono (Tokyo, JP)
Assignee: HITACHI VANTARA, LTD.
G06F11/2092G06F11/1471G06F11/1662
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,650,907
App. No.
18/828,495
Granted
Jun 9, 2026
Kind
B2
Abstract

In a storage system including a plurality of storage nodes and a plurality of storage devices that provide storage areas to the storage nodes, a cluster control unit that has detected occurrence of a failure in another storage node requests an external control device to execute a detach process of detaching a storage device allocated to a failure storage node from the failure storage node and an attach process of allocating the storage device to an alternative storage node. Then, the alternative cluster control unit restores, based on a log stored in a storage device allocated to a storage node in which a most recent failure has occurred among the storage nodes belonging to the same redundancy group as a failure storage control unit of the failure storage node, storage content of a memory of the alternative storage node.

Claims (54)

1 . A storage system comprising:

a plurality of storage nodes; and

a plurality of storage devices that provides storage areas to the plurality of storage nodes,

wherein each of the storage nodes includes:

a memory that stores cache data related to data read from and written to the storage areas by the storage node,

a storage control unit that reads and writes the data from and to the storage areas in response to a request from a host device, updates the cache data related to the data in the memory, creates a log related to the cache data, and stores the log in a storage device allocated to the storage node, and

a cluster control unit that manages a plurality of storage control units in a redundancy group, dispersedly arranges and manages the plurality of storage control units belonging to the same redundancy group in the plurality of storage nodes, and monitors occurrence of a failure in another storage node,

the storage control units belonging to the same redundancy group synchronize the cache data stored in the memory,

a failure detection cluster control unit that is a cluster control unit that has detected the occurrence of the failure in the other storage node requests an external control device to create an alternative storage node that is a storage node that substitutes for a failure storage node that is the storage node in which the failure has occurred,

the failure detection cluster control unit executes a detach process of separating a storage device allocated to the failure storage node from the failure storage node, then requests the control device to execute an attach process of allocating the storage device separated from the failure storage node to the alternative storage node, and

an alternative cluster control unit that is a cluster control unit included in the alternative storage node selects a specific storage node in which a most recent failure has occurred among the storage nodes including the storage control units belonging to the same redundancy group as a failure storage control unit included in the failure storage node, selects and executes, based on a log stored in a storage device allocated to the specific storage node, a first recovery method of restoring storage content in a memory included in the alternative storage node, and

activates an alternative storage control unit that is a storage control unit substituting for the failure storage control unit in the alternative storage node.

2 . The storage system according to claim 1 , wherein

the memory stores control information related to the storage control unit,

the log includes a log related to the control information, and

the storage control unit

updates the control information in the memory, creates the log related to the control information, and stores the log related to the cache data and the log related to the control information in the storage device allocated to the storage node.

3 . The storage system according to claim 1 , wherein

the failure detection cluster control unit:

determines whether a failure has occurred in all the storage control units belonging to the same redundancy group as the failure storage control unit, and

when it is determined that a failure has occurred in all the storage control units belonging to the same redundancy group as the failure storage control unit, the alternative cluster control unit selects and executes the first recovery method, and

when it is determined that any of the storage control units belonging to the same redundancy group as the failure storage control unit is normally operating, the alternative cluster control unit selects and executes, instead of the first recovery method, based on storage content stored in a memory of a storage node including a storage control unit that is normally operating and belonging to the same redundancy group as the failure storage control unit, a second recovery method of restoring the storage content in the memory included in the alternative storage node.

4 . The storage system according to claim 1 , further comprising a disk array device that includes the plurality of storage devices,

wherein the plurality of storage nodes and the disk array device are connected to each other via a predetermined network, and

the detach process and the attach process include

deletion, addition, or a configuration change of the storage devices in the disk array device based on control from a management device of the storage system via the predetermined network.

5 . The storage system according to claim 1 , wherein

a cluster control unit included in a target storage node to which a stop instruction has been input

creates a list of storage control units to be made redundant operating in the target storage node,

secures a redundancy destination storage node that makes the target storage node redundant, and

makes the storage control units to be made redundant in the redundancy destination storage node, and

executes a detach process of separating a storage device allocated to the target storage node from the target storage node, and then requests the control device to execute an attach process of allocating the storage device separated from the target storage node to the redundancy destination storage node.

6 . The storage system according to claim 1 , wherein

the alternative cluster control unit

activates all alternative storage control units to be in an active mode out of an active mode and a standby mode in the redundancy group so that all the alternative storage control units of each of the plurality of storage nodes are distributedly disposed in all the alternative storage nodes of each of the plurality of storage nodes.

7 . The storage system according to claim 1 , wherein

the alternative cluster control unit

activates all alternative storage control units of each of the plurality of storage nodes to be in an active mode out of an active mode and a standby mode in the redundancy group so that all the alternative storage control units of each of the plurality of storage nodes are centrally disposed in some of the alternative storage nodes of the plurality of storage nodes.

8 . The storage system according to claim 7 , wherein

the alternative cluster control unit

activates all alternative storage control units to be in the active mode so that all the alternative storage control units are centrally disposed in some of the alternative storage nodes, and

then redisposes all the alternative storage control units to be in the active mode so that all the alternative storage control units are distributed to all the alternative storage nodes.

9 . A failure handling method in a storage system executed by a storage system including a plurality of storage nodes and a plurality of storage devices that provide storage areas to the plurality of storage nodes,

each of the storage nodes including

a memory that stores cache data related to data read from and written to the storage areas by the storage node,

a storage control unit that reads and writes the data from and to the storage areas in response to a request from a host device, updates the cache data related to the data in the memory, creates a log related to the cache data, and stores the log in a storage device allocated to the storage node, and

a cluster control unit that manages a plurality of storage control units in a redundancy group, dispersedly arranges and manages the plurality of storage control units belonging to the same redundancy group in the plurality of storage nodes, and monitors occurrence of a failure in another storage node,

the failure handling method comprising:

the storage control units belonging to the same redundancy group synchronizing the cache data stored in the memory,

a failure detection cluster control unit that is a cluster control unit that has detected the occurrence of the failure in the other storage node

requesting an external control device to create an alternative storage node that is a storage node that substitutes for a failure storage node that is the storage node in which the failure has occurred;

an alternative cluster control unit that is a cluster control unit included in the alternative storage node

executing a detach process of separating a storage device allocated to the failure storage node from the failure storage node, then requesting the control device to execute an attach process of allocating the storage device separated from the failure storage node to the alternative storage node, selecting a specific storage node in which a most recent failure has occurred among the storage nodes including the storage control units belonging to the same redundancy group as a failure storage control unit included in the failure storage node, selecting and executing, based on a log stored in the storage device allocated to the specific storage node, a first recovery method of restoring storage content in a memory included in the alternative storage node; and

activating an alternative storage control unit that is a storage control unit substituting for the failure storage control unit in the alternative storage node.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2024
From: BETSUNO, KENICHI; SHIMOZONO, NORIO
To: HITACHI VANTARA, LTD.
Reel/Frame 068529/0019 →
Priority Claims (1)
JP 2024-008734 · Jan 24, 2024 · national
Continuity (1)
Related Publication 20250238335A1 · Jul 24, 2025
References Cited (10)
US 6553509B1 · Hanson · 2003 [cited by examiner]
US 10083100B1 · Agetsuma et al. · 2018 [cited by applicant]
US 20180285219A1 · Donlan · 2018 [cited by examiner]
US 20190163593A1 · Agetsuma · 2019 [cited by examiner]
US 20190278676A1 · Zou · 2019 [cited by examiner]
US 20200097204A1 · Sato · 2020 [cited by examiner]
US 20200409583A1 · Kusters · 2020 [cited by examiner]
US 20220261321A1 · Zakharkin · 2022 [cited by examiner]
US 20220404977A1 · Ito · 2022 [cited by examiner]
JP 2019101703A · 2019 [cited by applicant]