IP Library Granted Patent US 12,339,755
Granted Patent B2
US 12,339,755 · App. 18/208,478 · Granted Jun 24, 2025

Failover methods and system in a networked storage environment

Inventors: Ratnesh Gupta (Dublin, CA); Kalaivani Arumugham (Sunnyvale, CA); Ram Kesavan (Los Altos, CA); Ravikanth Dronamraju (Pleasanton, CA)
Assignee: NETAPP, INC.
G06F11/2069G06F11/1662G06F11/2064G06F11/2058G06F11/2071G06F11/2082G06F2201/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,339,755
App. No.
18/208,478
Filed
Jun 12, 2023
Granted
Jun 24, 2025
Kind
B2
Examiner
XU, MICHAEL
Art Unit
2113
USPC
714/6.3
Abstract

Failover methods and systems for a storage environment are provided. During a takeover operation to take over storage of a first storage system node by a second storage system node, the second storage system node copies information from a first storage location to a second storage location. The first storage location points to an active file system of the first storage system node, and the second storage location is assigned to the second storage system node for the takeover operation. The second storage system node quarantines storage space likely to be used by the first storage system node for a write operation, while the second storage system node attempts to take over the storage of the first storage system node. The second storage system node utilizes information stored at the second storage location during the takeover operation to give back control of the storage to the first storage system node.

Claims (50)

1. A method executed by one or more processors, comprising:

reserving storage space in a second storage system node, wherein the storage space is for use by a first storage system node for any write operations that occur while the second storage system node attempts to take over storage of the first storage system node, wherein the first storage system node and the second storage system node are configured to operate as failover partner nodes;

in response to detecting the first storage system node is healthy, copying, by the second storage system node, information from a second storage location assigned to the second storage system node to a first storage location assigned to the first storage system node; and

releasing, by the second storage system node, ownership of the storage space to the first storage system.

2. The method of claim 1 , further comprising:

in response to detecting the first storage system node is unresponsive, copying by the second storage system node configuration information regarding the storage of the first system node from the first storage location to the second storage location.

3. The method of claim 1 , further comprising:

prior to the second storage system node attempting to take over storage of the first storage system node, assigning the first storage location to the first storage system node to point to an active file system of the first storage system node.

4. The method of claim 3 , further comprising:

accessing the active file system of the first storage system node using information copied from the first storage location to the second storage location.

5. The method of claim 1 , further comprising:

upon detecting a failure in the second storage system node while taking over the storage of the first storage system node, using a third storage system node for taking over storage of the second storage system node to complete taking over storage of the first storage system node.

6. The method of claim 1 , further comprising:

allocating a third storage location for the second storage system node to track write requests for storage managed by the second storage system node.

7. The method of claim 1 , further comprising:

allocating, by the second storage system node, storage space for storing data for a write request that would have been written by the first storage system node.

8. A non-transitory, machine readable storage medium having stored thereon instructions comprising machine executable code, which when executed by a machine, causes the machine to:

reserve storage space in a second storage system node, wherein the storage space is for use by a first storage system node for any write operations that occur while the second storage system node attempts to take over storage of the first storage system node, wherein the first storage system node and the second storage system node are configured to operate as failover partner nodes;

in response to detecting the first storage system node is healthy, copy, by the second storage system node, information from a second storage location assigned to the second storage system node to a first storage location assigned to the first storage system node; and

release, by the second storage system node, ownership of the storage space to the first storage system.

9. The non-transitory, machine readable storage medium of claim 8 , wherein the machine executable code further causes the machine to:

in response to detecting the first storage system node is unresponsive, copy by the second storage system node configuration information regarding the storage of the first system node from the first storage location to the second storage location.

10. The non-transitory, machine readable storage medium of claim 8 , wherein the machine executable code further causes the machine to:

prior to the second storage system node attempting to take over storage of the first storage system node, assign the first storage location to the first storage system node to point to an active file system of the first storage system node.

11. The non-transitory, machine readable storage medium of claim 10 , wherein the machine executable code further causes the machine to:

access the active file system of the first storage system node using information copied from the first storage location to the second storage location.

12. The non-transitory, machine readable storage medium of claim 8 , wherein the machine executable code further causes the machine to:

upon detecting a failure in the second storage system node while taking over the storage of the first storage system node, use a third storage system node for taking over storage of the second storage system node to complete taking over storage of the first storage system node.

13. The non-transitory, machine readable storage medium of claim 8 , wherein the machine executable code further causes the machine to:

allocate a third storage location for the second storage system node to track write requests for storage managed by the second storage system node.

14. A system, comprising:

a first storage system node;

a second storage system node, wherein the first storage system node and the second storage system node are configured to operate as failover partner nodes;

a memory containing machine readable medium comprising machine executable code having stored thereon instructions; and

a processor coupled to the memory to execute the machine executable code to cause the second storage system node to:

reserve storage space in the second storage system node, wherein the storage space is for use by the first storage system node for any write operations that occur while the second storage system node attempts to take over storage of the first storage system node;

in response to detecting the first storage system node is healthy, copy, by the second storage system node, information from a second storage location assigned to the second storage system node to a first storage location assigned to the first storage system node; and

release, by the second storage system node, ownership of the storage space to the first storage system.

15. The system of claim 14 , wherein the machine executable code further causes to:

in response to detecting the first storage system node is unresponsive, copy by the second storage system node configuration information regarding the storage of the first system node from the first storage location to the second storage location.

16. The system of claim 14 , wherein the machine executable code further causes to:

prior to the second storage system node attempting to take over storage of the first storage system node, assign the first storage location to the first storage system node to point to an active file system of the first storage system node.

17. The system of claim 16 , wherein the machine executable code further causes to:

access the active file system of the first storage system node using information copied from the first storage location to the second storage location.

18. The system of claim 14 , wherein the machine executable code further causes to:

upon detecting a failure in the second storage system node while taking over the storage of the first storage system node, use a third storage system node for taking over storage of the second storage system node to complete taking over storage of the first storage system node.

19. The system of claim 14 , wherein the machine executable code further causes to:

allocate a third storage location for the second storage system node to track write requests for storage managed by the second storage system node.

20. The system of claim 14 , wherein the machine executable code further causes to:

allocate, by the second storage system node, storage space for storing data for a write request that would have been written by the first storage system node.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2023
From: GUPTA, RATNESH; ARUMUGHAM, KALAIVANI; KESAVAN, RAM; DRONAMRAJU, RAVIKANTH
To: NETAPP, INC.
Reel/Frame 063922/0222 →
Continuity (3)
Continuation 17648531 · Jan 20, 2022
Continuation 17026785 · Sep 21, 2020
Related Publication 20230325289A1 · Oct 12, 2023
References Cited (17)
US 8689043B1 · Bezbaruah et al. · 2014 [cited by applicant]
US 10503612B1 · Wang · 2019 [cited by examiner]
US 11249869B1 · Gupta et al. · 2022 [cited by applicant]
US 20150309892A1 · Ramasubramaniam et al. · 2015 [cited by applicant]
US 20160239437A1 · Le et al. · 2016 [cited by applicant]
US 20170052707A1 · Koppolu · 2017 [cited by examiner]
US 20170235591A1 · Kanada et al. · 2017 [cited by applicant]
US 20170286238A1 · Kesavan et al. · 2017 [cited by applicant]
US 20210073089A1 · Sathavalli et al. · 2021 [cited by applicant]
US 20210081287A1 · Koning et al. · 2021 [cited by applicant]
US 20210157694A1 · Dye et al. · 2021 [cited by applicant]
US 20210286515A1 · Gazit · 2021 [cited by examiner]
US 20210303423A1 · MacCarthaigh · 2021 [cited by examiner]
US 20210334181A1 · Satoyama · 2021 [cited by examiner]
US 20220147428A1 · Gupta et al. · 2022 [cited by applicant]
NetApp, Inc., “Clustered Data ONTAP 8.2,” High-Availability Configuration Guide, updated for 8.2.1, Feb. 2014, 108 pages. [cited by applicant]
Notice of Allowance mailed on Mar. 29, 2023 for U.S. Appl. No. 17/648,531, filed Jan. 20, 2022, 8 pages. [cited by applicant]