IP Library › Granted Patent US 12,282,676
Granted Patent B2
US 12,282,676 · App. 18/116,740 · Granted Apr 22, 2025

Recovery of clustered storage systems

Inventors: Siamak Nazari (Mountain View, CA); David Dejong (Fremont, CA); Srinivasa Murthy (Cupertino, CA); Shayan Askarian Namaghi (San Jose, CA); Roopesh Tamma (San Ramon, CA)
Assignee: Nvidia Corporation
G06F3/065G06F3/0623G06F3/067G06F3/0683
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,282,676
App. No.
18/116,740
Granted
Apr 22, 2025
Kind
B2
Abstract

A cluster storage system takes snapshots that are consistent across all storage nodes. The storage system can nearly instantaneously promote a set of consistent snapshots to their respective base volumes to restore the base volumes to be the same as the snapshots. Given these two capabilities, users can restore the system to a recovery point of the user's choice, by turning off storage service I/O, promoting the snapshots constituting the recovery point, rebooting their servers, and resuming storage service I/O.

Claims (56)

1. A recovery process for a storage system comprising:

creating, for one or more recovery points, a plurality of snapshot sets respectively in a plurality of storage nodes of the storage system, each snapshot set of the snapshot sets corresponding to a recovery point of the one or more recovery points and containing snapshots of every volume that a storage node of the plurality of storage nodes owns or maintains;

assigning, for the one or more recovery points, a generation number to each snapshot in a corresponding snapshot set associated with a respective recovery point;

receiving, from a user of the storage system, a selection of a recovery point from the one or more recovery points;

suspending one or more storage services of the storage system;

in response to receiving the selection of the recovery point, in each storage node of the plurality of storage nodes, promoting the snapshots in the snapshot set corresponding to the recovery point selected by altering metadata associated with each storage node to point to respective data with a generation number in a range between a generation number at creation of the volumes and the generation number assigned to the selected recovery point;

rebooting the storage nodes; and

resuming the one or more storage services of the storage system.

2. The process of claim 1 , wherein creating the snapshot sets comprises:

receiving, from the user of the storage system, a selection of a schedule for creation of the snapshot sets;

synchronizing the storage nodes to identify a first time in the schedule; and

creating, at the first time, the snapshot sets that correspond to a first of the recovery points, the first recovery point corresponding to a state of the storage system at the first time.

3. The process of claim 2 , wherein creating the snapshot set further comprises:

causing the storage nodes to identify a second time in the schedule; and

at the second time, causing the storage nodes to create the snapshot sets that correspond to a second of the recovery points, the second recovery point corresponding to a state of the storage system at the second time.

4. The process of claim 2 , wherein synchronizing the storage nodes comprises executing a Network Time Protocol (NTP) process on the storage nodes through a network interconnecting the storage nodes.

5. The process of claim 1 , further comprising, after the suspending of the storage services, creating a plurality of new snapshot sets respectively in the storage nodes of the storage system, the new snapshot sets corresponding to a new recovery point and capturing a state of the storage system when the recovery process was performed.

6. The process of claim 1 , wherein the creating of the snapshot sets respectively in the storage nodes of the storage system, consists of each of the storage nodes, without copying or moving data of the volumes, creating a metadata structure that prevents the storage node from deleting from the physical storage any of the data required for the snapshots.

7. The process of claim 6 , wherein the promoting of the snapshots consists of each of the storage nodes, without copying or moving the data of the volumes in physical storage, modifying a metadata structure.

8. A cluster storage system comprising:

a plurality of storage nodes; and

a network interconnecting the storage nodes, wherein the storage nodes are configured to perform a process including:

creating, for one or more recovery points, a plurality of snapshot sets respectively in the storage nodes, each snapshot set of the snapshot sets corresponding to a recovery point of the one or more recovery points and containing snapshots of every volume that a storage node of the plurality of storage nodes owns or maintains;

assigning, for the one or more recovery points, a generation number to each snapshot in a corresponding snapshot set associated with a respective recovery point;

receiving a request to roll back the storage system to one of the one or more recovery points selected by a user of the cluster storage system from the one or more recovery points;

suspending storage services of the storage system;

in response to receiving the selection of the recovery point, in each storage node of the plurality of storage nodes, promoting the snapshots in the snapshot set corresponding to the recovery point selected by altering metadata associated with each storage node to point to respective data with a generation number in a range between a generation number at creation of the volumes and the generation number assigned to the selected recovery point;

rebooting the storage nodes; and

resuming the storage service of the storage system.

9. The cluster storage system of claim 8 , wherein each of the storage nodes comprises:

a server;

a backend storage device; and

a storage processing unit resident in the server and connected to control the backend storage to provide the storage services that target any of the volumes that the storage node owns or maintains.

10. The cluster storage system of claim 9 , wherein in each storage node, the storage processing unit comprises a card plugged into a bus of the server.

11. The cluster storage system of claim 8 , wherein the cluster storage system is configured to synchronize the storage nodes by executing a Network Time Protocol (NTP) process through the network interconnecting the storage nodes.

12. A system comprising:

at least one processor; and

a memory coupled to the at least one processor, the memory storing instructions that, when executed by the at least one processor, cause the system to perform steps comprising:

creating, for one or more recovery points, a plurality of snapshot sets respectively in a plurality of storage nodes of the storage system, each snapshot set of the snapshot sets corresponding to a recovery point of the one or more recovery points and containing snapshots of every volume that a storage node of the plurality of storage nodes owns or maintains;

assigning, for the one or more recovery points, a generation number to each snapshot in a corresponding snapshot set associated with a respective recovery point;

receiving, from a user of the storage system, a selection of a recovery point from the one or more recovery points selecting one of the one or more recovery points;

suspending one or more storage services of the storage system;

in response to receiving the selection of the recovery point, in each storage node of the plurality of storage nodes, promoting the snapshots in the snapshot set corresponding to the recovery point selected by altering metadata associated with each storage node to point to respective data with a generation number in a range between a generation number at creation of the volumes and the generation number assigned to the selected recovery point;

rebooting the storage nodes; and

resuming the one or more storage services of the storage system.

13. The system of claim 12 , wherein creating the snapshot sets comprises:

receiving, from the user of the storage system, a selection of a schedule for creation of the snapshot sets;

synchronizing the storage nodes to identify a first time in the schedule; and

creating, at the first time, the snapshot sets that correspond to a first of the recovery points, the first recovery point corresponding to a state of the storage system at the first time.

14. The system of claim 13 , wherein creating the snapshot set further comprises:

causing the storage nodes to identify a second time in the schedule; and

at the second time, causing the storage nodes to create the snapshot sets that correspond to a second of the recovery points, the second recovery point corresponding to a state of the storage system at the second time.

15. The system of claim 13 , wherein synchronizing the storage nodes comprises executing a Network Time Protocol (NTP) process on the storage nodes through a network interconnecting the storage nodes.

16. The system of claim 12 , wherein the instructions cause the system to further perform steps comprising, after the suspending of the storage services, creating a plurality of new snapshot sets respectively in the storage nodes of the storage system, the new snapshot sets corresponding to a new recovery point and capturing a state of the storage system when the recovery process was performed.

17. The system of claim 12 , wherein the creating of the snapshot sets respectively in the storage nodes of the storage system, consists of each of the storage nodes, without copying or moving data of the volumes, creating a metadata structure that prevents the storage node from deleting from the physical storage any of the data required for the snapshots.

18. The system of claim 12 , wherein the promoting of the snapshots consists of each of the storage nodes, without copying or moving the data of the volumes in physical storage, modifying a metadata structure.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2024
From: NEBULON, INC.; NEBULON LTD
To: NVIDIA CORPORATION
Reel/Frame 067338/0929 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 12, 2023
From: NAZARI, SIAMAK; DEJONG, DAVID; MURTHY, SRINIVASA; ASKARIAN NAMAGHI, SHAYAN; TAMA, ROOPESH
To: NEBULON, INC.
Reel/Frame 063306/0442 →
Continuity (6)
Continuation In Part 18115211 · Feb 28, 2023
Provisional Application 63314987 · Feb 28, 2022
Provisional Application 63314996 · Feb 28, 2022
Provisional Application 63314970 · Feb 28, 2022
Provisional Application 63316081 · Mar 3, 2022
Related Publication 20230273742A1 · Aug 31, 2023
References Cited (38)
US 7827350B1 · Jiang · 2010 [cited by examiner]
US 8326803B1 · Stringham · 2012 [cited by examiner]
US 8886995B1 · Pallapothu · 2014 [cited by examiner]
US 9588847B1 · Natanzon · 2017 [cited by examiner]
US 9842117B1 · Zhou · 2017 [cited by examiner]
US 10514989B1 · Borodin · 2019 [cited by examiner]
US 11307935B2 · Keller · 2022 [cited by examiner]
US 11327844B1 · Aquino et al. · 2022 [cited by applicant]
US 11620070B2 · Balasubramanian et al. · 2023 [cited by applicant]
US 20040250033A1 · Prahlad · 2004 [cited by examiner]
US 20090013012A1 · Ichikawa · 2009 [cited by examiner]
US 20100179959A1 · Shoens · 2010 [cited by examiner]
US 20100299309A1 · Maki et al. · 2010 [cited by applicant]
US 20120317079A1 · Shoens · 2012 [cited by examiner]
US 20130166863A1 · Buragohain · 2013 [cited by examiner]
US 20140258239A1 · Amlekar · 2014 [cited by examiner]
US 20150286695A1 · Kadayam · 2015 [cited by examiner]
US 20160132396A1 · Kimmel · 2016 [cited by examiner]
US 20160205182A1 · Lazar · 2016 [cited by examiner]
US 20170031769A1 · Zheng · 2017 [cited by examiner]
US 20170123657A1 · B · 2017 [cited by examiner]
US 20170123700A1 · Sinha · 2017 [cited by examiner]
US 20200272544A1 · Carvelli · 2020 [cited by examiner]
US 20200341855A1 · Tanwer · 2020 [cited by examiner]
US 20200379646A1 · Blau · 2020 [cited by examiner]
US 20210133328A1 · Wu · 2021 [cited by examiner]
US 20210216408A1 · Huskisson · 2021 [cited by examiner]
US 20210224161A1 · Wang · 2021 [cited by applicant]
US 20220263897A1 · Karr · 2022 [cited by examiner]
US 20220350492A1 · Beedu · 2022 [cited by examiner]
US 20230132591A1 · Karr · 2023 [cited by examiner]
US 20230325285A1 · O'Connor et al. · 2023 [cited by applicant]
CN 112416860A · 2021 [cited by examiner]
WO WO2021150563A1 · 2021 [cited by examiner]
WO WO2021150576A1 · 2021 [cited by examiner]
Machine translation of CN 112416860 (Year: 2021). [cited by examiner]
Non-Final Office Action issued in U.S. Appl. No. 18/115,211, dated Jun. 5, 2024. [cited by applicant]
Final Office Action issued in U.S. Appl. No. 18/115,211, dated Oct. 23, 2024. [cited by applicant]