IP Library › Granted Patent US 12,561,436
Granted Patent B2
US 12,561,436 · App. 18/115,211 · Granted Feb 24, 2026

Storage system with cloud assisted recovery

Inventors: Siamak Nazari (Mountain View, CA); David Dejong (Fremont, CA); Srinivasa Murthy (Cupertino, CA); Shayan Askarian Namaghi (San Jose, CA); Roopesh Tamma (San Ramon, CA)
Assignee: Nvidia Corporation
G06F21/566G06F9/4403G06F9/4416G06F11/1446G06F2221/033
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,436
App. No.
18/115,211
Granted
Feb 24, 2026
Kind
B2
Abstract

A cluster storage system including servers containing storage processing units (SPUs) can create synchronized snapshot sets for the volumes that the SPUs maintain and can report the snapshot sets to a cloud-based service. Each snapshot in a set reflects the state a corresponding volume had at a rollback point corresponding to the set. A user of the storage system contacts the cloud-based service about recovery of the storage system, and the cloud-based service may present the user with a list of rollback points corresponding to the synchronized snapshot sets. The user may select to recover the storage system to any of the rollback points, and the SPUs promote the selected snapshots to replace the volumes for storage services.

Claims (55)

1 . A process for operating a storage system including a plurality of servers containing a plurality of service processing units (SPUs), the process comprising:

creating a first snapshot set, the first snapshot set including a plurality of first snapshots respectively of a plurality of volumes that the SPUs maintain for storage services, the first snapshots capturing respective states that the volumes have at a first time;

creating a second snapshot set, the second snapshot set including a plurality of boot volume snapshots respectively of a plurality of boot volumes at a second time when a verified operating system image is written to the plurality of boot volumes;

reporting, by the storage system, the first snapshot set and the second snapshot set to a cloud-based service;

receiving, from a user of the storage system, communication to the cloud-based service about a recovery of the storage system;

providing for presentation to the user a list of recovery points, the recovery points including a first recovery point corresponding to the first time and a second recovery point corresponding to the second time, each recovery point corresponding to one or more snapshots of the plurality of volumes;

receiving from the user a selection to recover the storage system to the first recovery point, the first recovery point corresponding to the first snapshots; and

in response to the selection from the user, promoting the first snapshots to replace the volumes that SPUs present for storage services.

2 . The process of claim 1 , wherein the process further comprises creating a third snapshot set, the third snapshot set including a plurality of third snapshots respectively of the plurality of volumes, the third snapshots capturing respective states that the volumes have at a third time, wherein the list of recovery points presented to the user further includes a third recovery point corresponding to the third time.

3 . The process of claim 2 , further comprising instructing each of the SPUs to perform a snapshot operation at times indicated by a schedule, wherein:

each performance of the snapshot operation including each of the SPUs snapshotting each of the volumes that the SPU exports;

the schedule includes the first time and the third time;

the first snapshot set is results of the SPUs performing the snapshot operation at the first time; and

the third snapshot set is results of the SPUs performing the snapshot operation at the third time.

4 . The process of claim 1 , wherein the process further comprises receiving communication from the user to communicate with the cloud-based service about recovery while the volumes are corrupted or encrypted.

5 . The process of claim 1 , wherein when the cloud-based service presents the list, the cloud-based service also provides to the user information indicating whether the first snapshots appear to be corrupted or encrypted.

6 . The process of claim 1 , wherein creating the first snapshot set comprises performing, by each of the SPUs, a snapshot operation that includes:

suspending incoming storage service requests;

completing pending storage operations for prior storage service requests to the SPU;

creating the first snapshots of the volumes that the SPU maintains; and

resuming acceptance of incoming storage requests to the SPU.

7 . The process of claim 6 , wherein each of the SPUs creates the first snapshots of the volumes without copying data of the volumes.

8 . The process of claim 1 , wherein the promoting of snapshots comprises promoting, by the SPUs, their respective volumes, thereby rolling back an entire cluster of storage nodes associated with the SPUs.

9 . A cluster storage system comprising: a plurality of storage nodes each storage node including:

a server;

a backend storage device; and

an SPU (storage processing unit) connected to the backend storage device, the SPU operating the backend storage to provide storage services for a plurality of virtual volumes, wherein the SPU is configured to execute a first process including:

performing a snapshot operation at times indicated by a cluster-wide schedule, each performance of the snapshot operation including the SPU snapshotting each of the volumes that the SPU exports and thereby creating a synchronized snapshot set associated with the time of the snapshotting, wherein at least one snapshot operation includes creating a plurality of boot volume snapshots at a time when a verified operating system image is written to the plurality of boot volumes;

the SPUs are configured to respond to instructions from a cloud-based service by promoting the snapshots of a selected one of the synchronized snapshot sets to recover the cluster storage system to a state it had at the time associated with the synchronized set, wherein the instructions are based on a user selection from a list of recovery points, the user selection indicating the time as a recovery point.

10 . The cluster storage system of claim 9 , wherein the SPUs are configured to send information regarding the synchronized snapshot sets to the cloud-based service.

11 . A system, comprising at least one processor and memory storing instructions for executing steps, the steps comprising:

creating a first snapshot set, the first snapshot set including a plurality of first snapshots respectively of a plurality of volumes that a plurality of SPUs (storage processing units) maintain for storage services, the first snapshots capturing respective states that the volumes have at a first time;

creating a second snapshot set, the second snapshot set including a plurality of boot volume snapshots respectively of a plurality of boot volumes at a second time when a verified operating system image is written to the plurality of boot volumes;

reporting the first snapshot set and the second snapshot set to a cloud-based service;

receiving, from a user of a storage system, communication to the cloud-based service about a recovery of the storage system;

providing, for presentation to the user, a list of recovery points, the recovery points including a first recovery point corresponding to the first time and a second recovery point corresponding to the second time, each recovery point corresponding to one or more snapshots of the plurality of volumes;

receiving, from the user, a selection to recover the storage system to the first recovery point, the first recovery point corresponding to the first snapshots; and

in response to instructions from the cloud-based service associated with the selection, promoting the first snapshots to replace the volumes that SPUs present for storage services.

12 . The system of claim 11 , wherein the steps further comprise creating, by the SPUs, a third snapshot set, the third snapshot set including a plurality of third snapshots respectively of the plurality of volumes, the third snapshots capturing respective states that the volumes have at a third time, wherein the list of recovery points presented to the user further includes a third recovery point corresponding to the third time.

13 . The system of claim 12 , wherein the instructions further comprise instructing each of the SPUs to perform a snapshot operation at times indicated by a schedule, wherein:

each performance of the snapshot operation includes each of the SPUs snapshotting each of the volumes that the SPU exports;

the schedule includes the first time and the third time;

the first snapshot set is results of the SPUs performing the snapshot operation at the first time; and

the third snapshot set is results of the SPUs performing the snapshot operation at the third time.

14 . The system of claim 12 , wherein the steps further comprise receiving communication from the user to communicate with the cloud-based service about recovery while the volumes are corrupted or encrypted.

15 . The system of claim 11 , wherein when the cloud-based service presents the list, the cloud-based service also provides to the user information indicating whether the first snapshots appear to be corrupted or encrypted.

16 . The system of claim 11 , wherein creating the first snapshot set comprises performing, by each of the SPUs, a snapshot operation that includes:

suspending incoming storage service requests;

completing pending storage operations for prior storage service requests to the SPU;

creating the first snapshots of the volumes that the SPU maintains; and

resuming acceptance of incoming storage requests to the SPU.

17 . The system of claim 16 , wherein each of the SPUs creates the first snapshots of the volumes without copying data of the volumes.

18 . The system of claim 11 , wherein the promoting of snapshots includes all the SPUs promoting their respective volumes in response to the instructions from the cloud-based service, thereby rolling back an entire cluster of storage nodes associated with the SPUs.

19 . The system of claim 11 , wherein the steps further comprise deleting one or more snapshots of volumes that are not boot volumes.

20 . The system of claim 19 , wherein the one or more deleted snapshots are disqualified from providing a possible recovery point.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2024
From: NEBULON, INC.; NEBULON LTD
To: NVIDIA CORPORATION
Reel/Frame 067338/0929 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 12, 2023
From: NAZARI, SIAMAK; DEJONG, DAVID; ASKARIAN NAMAGHI, SHAYAN; MURTHY, SRINIVASA; TAMMA, ROOPESH
To: NEBULON, INC.
Reel/Frame 063306/0932 →
Continuity (5)
Provisional Application 63316081 · Mar 3, 2022
Provisional Application 63314970 · Feb 28, 2022
Provisional Application 63314987 · Feb 28, 2022
Provisional Application 63314996 · Feb 28, 2022
Related Publication 20230273861A1 · Aug 31, 2023
References Cited (40)
US 7827350B1 · Jiang et al. · 2010 [cited by applicant]
US 8326803B1 · Stringham · 2012 [cited by applicant]
US 8886995B1 · Pallapothu et al. · 2014 [cited by applicant]
US 9588847B1 · Natanzon et al. · 2017 [cited by applicant]
US 9842117B1 · Zhou et al. · 2017 [cited by applicant]
US 10514989B1 · Borodin et al. · 2019 [cited by applicant]
US 11307935B2 · Keller et al. · 2022 [cited by applicant]
US 11327844B1 · Aquino · 2022 [cited by examiner]
US 11620070B2 · Balasubramanian · 2023 [cited by examiner]
US 20040019823A1 · Gere · 2004 [cited by examiner]
US 20040250033A1 · Prahlad et al. · 2004 [cited by applicant]
US 20090013012A1 · Ichikawa et al. · 2009 [cited by applicant]
US 20100179959A1 · Shoens · 2010 [cited by applicant]
US 20100299309A1 · Maki · 2010 [cited by examiner]
US 20120317079A1 · Shoens et al. · 2012 [cited by applicant]
US 20130166863A1 · Buragohain et al. · 2013 [cited by applicant]
US 20140258239A1 · Amlekar et al. · 2014 [cited by applicant]
US 20150286695A1 · Kadayam et al. · 2015 [cited by applicant]
US 20160132396A1 · Kimmel et al. · 2016 [cited by applicant]
US 20160205182A1 · Lazar et al. · 2016 [cited by applicant]
US 20170031769A1 · Zheng et al. · 2017 [cited by applicant]
US 20170123657A1 · B · 2017 [cited by applicant]
US 20170123700A1 · Sinha et al. · 2017 [cited by applicant]
US 20200272544A1 · Carvelli et al. · 2020 [cited by applicant]
US 20200341855A1 · Tanwer et al. · 2020 [cited by applicant]
US 20200379646A1 · Blau et al. · 2020 [cited by applicant]
US 20210133328A1 · Wu · 2021 [cited by applicant]
US 20210216408A1 · Huskisson et al. · 2021 [cited by applicant]
US 20210224161A1 · Wang et al. · 2021 [cited by applicant]
US 20220263897A1 · Karr et al. · 2022 [cited by applicant]
US 20220350492A1 · Beedu et al. · 2022 [cited by applicant]
US 20230132591A1 · Karr et al. · 2023 [cited by applicant]
US 20230325285A1 · O'Connor · 2023 [cited by examiner]
CN 112416860A · 2021 [cited by applicant]
WO 2021150563A1 · 2021 [cited by applicant]
WO 2021150576A1 · 2021 [cited by applicant]
Non-Final Office Action issued in U.S. Appl. No. 18/116,740, dated Mar. 25, 2024. [cited by applicant]
Final Office Action issued in U.S. Appl. No. 18/116,740, dated Jul. 29, 2024. [cited by applicant]
Non-Final Office Action issued in U.S. Appl. No. 18/116,740, dated Oct. 22, 2024. [cited by applicant]
Notice of Allowance issued in U.S. Appl. No. 18/116,740, dated Jan. 28, 2025. [cited by applicant]