IP Library › Granted Patent US 12,645,664
Granted Patent B2
US 12,645,664 · App. 18/489,597 · Granted Jun 2, 2026

Automated failover for a paired set of consistency groups while storage expansion occurs within a cross-site storage system

Inventors: V Ramakrishna Rao Yadala (Karnataka, IN); Sohan Shetty (Bangalore, IN); Akhil Kaushik (San Jose, CA)
Assignee: NetApp, Inc.
G06F16/2365G06F16/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,664
App. No.
18/489,597
Granted
Jun 2, 2026
Kind
B2
Abstract

A computer implemented method includes maintaining information indicative of whether a data replication relationship between a dataset associated with the local CG (CG1) and a mirror copy of the dataset stored on a remote CG (CG2) of a remote distributed storage system is in an in-synchronization (InSync) state or an out-of-synchronization (OOS) state, determining whether a CG storage expansion process is in progress, and in response to determining the OOS state and whether the primary storage site has a failure, initiating an automatic unplanned failover (AUFO) workflow without manual intervention on the original volumes in the CG1 of a first storage cluster of the primary storage site and the original volumes in the CG2 of a second storage cluster when the CG storage expansion process is in progress with a new source volume to be a member of CG1 and a new destination volume to be a member of CG2.

Claims (49)

1 . A computer implemented method performed by one or more processing resources of a distributed storage system, the method comprising:

for each of a plurality of original volumes of a first storage node of a primary storage site of the distributed storage system that are members of a local consistency group (CG1), maintaining information indicative of whether a data replication relationship between a dataset associated with the CG1 and a mirror copy of the dataset stored on a corresponding plurality of original volumes that are members of a remote CG (CG2) of a second storage node of a remote secondary storage site is in an in-synchronization (InSync) state or an out-of-synchronization (OOS) state;

determining whether a CG storage expansion process is in progress including a source expansion to create a new source volume as a member of the CG1 and a destination expansion to create a new destination volume as a member of the CG2;

determining the OOS state for the data replication relationship between original volumes of the CG1 and original volumes of the CG2; and

in response to determining the OOS state and whether the primary storage site has a failure, initiating and performing a dependent write order consistent automatic unplanned failover (AUFO) workflow without manual intervention to switch a primary role for serving input/output (I/O) operations for requesting devices from the original volumes in the CG1 of a first storage cluster of the primary storage site to the original volumes in the CG2 of a second storage cluster without including the new source volume and the new destination volume in the AUFO workflow when the CG storage expansion process is in progress with the new source volume to be a member of the CG1 and the new destination volume to be a member of the CG2.

2 . The computer implemented method of claim 1 , further comprising:

in response to determining the OOS state and whether the primary storage site has a failure, initiating an automatic unplanned (AUFO) failover workflow on all volumes in the CG1 of the first storage cluster and all volumes in the CG2 of the second storage cluster when the CG storage expansion process is completed with a new source volume having already been added as a member of CG1 and a new destination volume having been already added as a member of CG2.

3 . The computer implemented method of claim 2 , further comprising:

setting an indicator on any new volumes being added to CG1 or CG2 during the CG storage expansion process.

4 . The computer implemented method of claim 1 , further comprising:

modifying the CG2 of the second storage cluster to remove details of any new volumes being added to CG2 from a replicating database (RDB) table of the second storage cluster when the AUFO workflow has been initiated;

removing expand in progress state from the RDB table; and

after the first storage cluster returns to operational state, modifying the CG1 of the first storage cluster to remove details of any new volumes being added to CG1 from a replicating database (RDB) table of the first storage cluster.

5 . The computer implemented method of claim 1 , further comprising:

initiating a data replication relationship from the original volumes of CG2 to the original volumes of CG1 while maintaining zero recover point objective (RPO) and Zero recover time objective (RTO).

6 . The computer implemented method of claim 1 , further comprising:

in response to an OOS state between the original volumes of CG1 and the original volumes of CG2, performing resynchronization between the original volumes of CG1 and the original volumes of CG2 while the CG storage expansion process is in progress.

7 . The computer implemented method of claim 1 , further comprising:

automatically resuming CG storage expansion if one or more failures occur during the CG storage expansion.

8 . A non-transitory computer-readable storage medium embodying a set of instructions, which when executed by one or more processing resources of a distributed storage system, cause the distributed storage system to:

for each of a plurality of original volumes of a primary storage site of the distributed storage system that are members of a local consistency group (CG1), maintain information indicative of whether a data replication relationship between a dataset associated with the CG1 and a mirror copy of the dataset stored on a corresponding plurality of original volumes that are members of a remote CG (CG2) of a secondary storage site is in an in-synchronization (InSync) state or an out-of-synchronization (OOS) state;

determine whether a CG container expansion process is in progress including an existing source volume to be added as a member of the CG1;

determine the OOS state for the data replication relationship between original volumes of the CG1 and original volumes of the CG2; and

in response to determining the OOS state and whether the primary storage site has a failure, initiate and perform an automatic unplanned failover (AUFO) workflow without manual intervention to switch a primary role for serving input/output (I/O) operations for requesting devices from the original volumes in the CG1 of a first storage cluster to the original volumes in the CG2 of a second storage cluster without including the existing source volume being added to CG1 in the AUFO workflow when the CG storage expansion process is in progress with the existing source volume to be added as a member of CG1.

9 . The non-transitory computer-readable storage medium of claim 8 , wherein the instructions further cause the distributed storage system to:

in response to determining the OOS state and whether the primary storage site has a failure, initiate an automatic unplanned failover (AUFO) workflow on all volumes in the CG1 of a first storage cluster and all volumes in the CG2 of a second storage cluster when the CG container expansion process is completed with the existing source volume having already been added as a member of CG1 and an existing destination volume having been already added as a member of CG2.

10 . The non-transitory computer-readable storage medium of claim 9 , wherein the instructions further cause the distributed storage system to:

set an indicator on any existing volumes being added to CG1 or CG2 during the CG container expansion process.

11 . The non-transitory computer-readable storage medium of claim 9 , wherein the instructions further cause the distributed storage system to:

modify the CG2 of the second storage cluster to remove details of any existing volumes being added to CG2 from a replicating database (RDB) table when the CG storage expansion process is in progress.

12 . The non-transitory computer-readable storage medium of claim 11 , wherein the instructions further cause the distributed storage system to:

remove container expand in progress state from the RDB table.

13 . The non-transitory computer-readable storage medium of claim 8 , wherein the instructions further cause the distributed storage system to:

initiate a data replication relationship from the original volumes of CG2 to the original volumes of CG1 while maintaining zero recover point objective (RPO) and Zero recover time objective (RTO).

14 . The non-transitory computer-readable storage medium of claim 8 , wherein the instructions further cause the distributed storage system to:

in response to an OOS state between the original volumes of CG1 and the original volumes of CG2, perform resynchronization between the original volumes of CG1 and the original volumes of CG2 while the CG container expansion process is in progress.

15 . The non-transitory computer-readable storage medium of claim 8 , wherein the instructions further cause the distributed storage system to:

automatically resume CG container expansion if one or more failures occur during the CG storage expansion.

16 . A distributed storage system comprising:

one or more processing resource; and

one or more non-transitory computer-readable media, coupled to the one or more processing resources, having stored therein instructions that when executed by the one or more processing resource cause the distributed storage system to:

for each of a plurality of original volumes of a primary storage site of the distributed storage system that are members of a local consistency group (CG1), maintain information indicative of whether a data replication relationship between a dataset associated with the CG1 and a mirror copy of the dataset stored on a corresponding plurality of original volumes that are members of a remote CG (CG2) of a remote secondary storage site is in an in-synchronization (InSync) state or an out-of-synchronization (OOS) state;

determine whether a CG storage expansion process is in progress with a new source volume to be added as a member of CG1 and a new destination volume to be added as a member of CG2;

in response to determining the OOS state and whether the primary storage site has a failure,

automatically resume CG storage expansion without manual intervention when the CG storage expansion process has a failure before completing.

17 . The distributed storage system of claim 16 , wherein the instructions further cause the distributed storage system to in response to determining the OOS state and whether the primary storage site has a failure, initiate and perform an unplanned failover workflow on all volumes in the CG1 of a first storage cluster and all volumes in the CG2 of a second storage cluster when the CG storage expansion process is completed with a new source volume having already been added as a member of CG1 and a new destination volume having been already added as a member of CG2.

18 . The distributed storage system of claim 16 , wherein the instructions further cause the distributed storage system to set an indicator on any new volumes being added to CG1 or CG2 during the CG storage expansion process.

19 . The distributed storage system distributed storage system of claim 16 , wherein the instructions further cause the distributed storage system to initiate a data replication relationship from the original volumes of CG2 to the original volumes of CG1 while maintaining zero recover point objective (RPO) and Zero recover time objective (RTO).

20 . The distributed storage system of claim 16 , wherein the instructions further cause the distributed storage system to in response to an OOS state between the original volumes of CG1 and the original volumes of CG2, performing resynchronization between the original volumes of CG1 and the original volumes of CG2 while the CG storage expansion process is in progress.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2023
From: YADALA, V RAMAKRISHNA RAO; SHETTY, SOHAN; KAUSHIK, AKHIL
To: NETAPP, INC.
Reel/Frame 065495/0317 →
Continuity (1)
Related Publication 20250130985A1 · Apr 24, 2025
References Cited (10)
US 9043567B1 · Modukuri · 2015 [cited by examiner]
US 10481963B1 · Walker · 2019 [cited by applicant]
US 20070283089A1 · Tsurudome · 2007 [cited by applicant]
US 20220067060A1 · O'Halloran et al. · 2022 [cited by applicant]
US 20220317882A1 · Vijayan et al. · 2022 [cited by applicant]
US 20220317897A1 · Subramanian · 2022 [cited by examiner]
US 20250130733A1 · Yadala et al. · 2025 [cited by applicant]
Non Final Office Action mailed on Jun. 4, 2025 for U.S. Appl. No. 18/489,610, filed Oct. 18, 2023, 16 pages. [cited by applicant]
Notice of Allowance mailed on Apr. 10, 2026 for U.S. Appl. No. 18/489,610, filed Oct. 18, 2023, 05 pages. [cited by applicant]
Final Office Action mailed on Dec. 22, 2025 for U.S. Appl. No. 18/489,610, filed Oct. 18, 2023, 22 pages. [cited by applicant]
Cited By (1)
US 12,710,889