IP Library Granted Patent US 11,567,660
Granted Patent B2
US 11,567,660 · App. 17/203,452 · Granted Jan 31, 2023

Managing cloud storage for distributed file systems

Inventors: Michael Anthony Chmiel (Seattle, WA); Duncan Robert Fairbanks (Seattle, WA); Stephen Craig Fleischman (Seattle, WA); Daniel Marcos Motles (Seattle, WA); Nicholas Graeme Williams (Seattle, WA)
Assignee: Qumulo, Inc.
G06F3/0604G06F3/067G06F3/0629G06F3/0653G06F3/0683
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,567,660
App. No.
17/203,452
Granted
Jan 31, 2023
Kind
B2
Abstract

Embodiments are directed to managing data in a file system that includes a plurality of storage nodes and a plurality of storage volumes in a cloud computing environment. Metrics associated with each storage volume may be monitored. In response to the metrics exceeding a threshold value, performing further actions, including: determining storage volumes that are unhealthy based on the metrics that exceed the threshold value; updating metadata associated with the storage volumes to indicate that the storage volumes are unhealthy; decoupling the unhealthy storage volumes from storage nodes coupled to the unhealthy storage volumes; determining replacement storage volumes based on the metadata associated with the unhealthy storage volumes; updating other metadata associated with the replacement storage volumes to indicate that the replacement storage volumes are healthy storage volumes; and coupling the healthy storage volumes with the storage nodes that were coupled to the unhealthy storage volumes.

Claims (141)

1. A method for managing data in a file system over a network using one or more processors that execute instructions to perform actions, comprising:

providing the file system that includes a plurality of storage nodes and a plurality of storage volumes, wherein each storage node is coupled to a portion of the plurality of storage volumes, and wherein each storage node is a compute instance in a cloud computing environment and each storage volume is a data store in the cloud computing environment;

monitoring one or more metrics associated with each storage volume; and

in response to the one or more metrics exceeding one or more threshold values, performing further actions, including:

determining one or more storage volumes in the plurality of storage volumes that are disabled or missing based on the one or more metrics that exceed the one or more threshold values and one or more parameters for one or more storage slots associated with each of the disabled or missing one or more storage volumes;

updating metadata associated with the disabled one or more storage volumes to indicate that the disabled one or more storage volumes are also unhealthy, wherein the one or more disabled storage volumes are tagged as disabled;

decoupling the one or more unhealthy storage volumes from one or more storage nodes coupled to the one or more unhealthy storage volumes;

determining one or more replacement storage volumes based on the metadata associated with the one or more unhealthy storage volumes, one or more native queries that associate the metadata with the one or more unhealthy storage volumes, one or more queries for orphaned healthy storage volumes, and the one or more parameters for the one or more storage slots associated with each of the disabled one or more storage volumes, wherein each replacement storage volume matches the one or more parameters that correspond to at least one of the storage slots;

updating other metadata associated with the one or more replacement storage volumes to indicate that the one or more replacement storage volumes are healthy storage volumes;

coupling the one or more healthy storage volumes with the one or more storage nodes that were coupled to the one or more unhealthy storage volumes;

generating replacement metadata for the missing one or more storage volumes, wherein the replacement metadata is employed to generate and provision one or more replacement storage volumes; and

associating the one or more replacement storage volumes with the plurality of storage nodes, wherein the one or more replacement storage volumes are tagged as healthy and coupled to the one or more storage nodes, and wherein one or more portions of the healthy replacement storage volumes are assigned to each storage slot that lacks one or more healthy storage volumes.

2. The method of claim 1 , wherein determining the one or more replacement storage volumes, further comprises:

generating a query based on the metadata associated with the one or more unhealthy storage volumes;

employing the query to determine one or more uncoupled healthy storage volumes in the cloud computing environment, wherein the metadata associated with the one or more unhealthy storage volumes matches the other metadata associated with the one or more uncoupled preexisting healthy storage volumes; and

providing the one or more healthy storage volumes as at least a portion of the one or more replacement storage volumes.

3. The method of claim 1 , wherein determining the one or more replacement storage volumes, further comprises:

in response to a quantity of the one or more unhealthy storage volumes exceeding a quantity of one or more preexisting uncoupled healthy storage volumes, performing further actions, including:

provisioning one or more additional uncoupled storage volumes from the cloud computing environment that match the one or more unhealthy storage volumes;

updating metadata associated with the one or more uncoupled storage volumes to indicate that the one or more additional uncoupled storage volumes are uncoupled healthy storage volumes; and

providing the one or more uncoupled healthy storage volumes as the one or more replacement storage volumes.

4. The method of claim 1 , wherein coupling the one or more healthy storage volumes with the one or more storage nodes, further comprises:

determining the one or more storage slots associated with the one or more storage nodes based on the file system, wherein each storage slot corresponds to a fraction of a storage capacity of the file system; and

associating the one or more healthy storage volumes with the one or more storage slots, wherein each healthy storage volume is employed to provide the fraction of the storage capacity that corresponds to its associated storage slot.

5. The method of claim 1 , wherein monitoring the one or more metrics associated with each storage volume, further comprises, querying each storage node of the plurality of storage nodes for one or more values of the one or more metrics that are associated a portion of the plurality of storage volumes that are coupled to the queried storage node, wherein the one or more values of the one or more metrics are based on error information provided to each storage node by the cloud computing environment.

6. The method of claim 1 , wherein the metadata associated with the one or more storage volumes, further comprises, one or more of a storage cluster identifier, a storage slot identifier, and a field for storing a value that indicates that a storage volume is healthy or unhealthy.

7. The method of claim 1 , wherein determining the one or more replacement storage volumes, further comprises:

determining one or more storage volumes that are missing from the plurality of storage volumes based on the one or more metrics;

providing storage volume information associated with the one or more missing storage volumes based on querying the one or more storage nodes;

generating replacement metadata based on the storage volume information;

provisioning one or more additional uncoupled storage volumes from the cloud computing environment based on the replacement metadata;

updating metadata associated with the one or more additional uncoupled storage volumes to indicate that the one or more additional uncoupled storage volumes are uncoupled healthy storage volumes; and

providing the one or more uncoupled healthy storage volumes as the one or more replacement storage volumes.

8. A system for managing data in a file system comprising:

a network computer, comprising:

a memory that stores at least instructions; and

one or more processors that execute instructions that perform actions, including:

providing the file system that includes a plurality of storage nodes and a plurality of storage volumes, wherein each storage node is coupled to a portion of the plurality of storage volumes, and wherein each storage node is a compute instance in a cloud computing environment and each storage volume is a data store in the cloud computing environment;

monitoring one or more metrics associated with each storage volume; and

in response to the one or more metrics exceeding one or more threshold values, performing further actions, including:

determining one or more storage volumes in the plurality of storage volumes that are disabled or missing based on the one or more metrics that exceed the one or more threshold values and one or more parameters for one or more storage slots associated with each of the disabled or missing one or more storage volumes;

updating metadata associated with the disabled one or more storage volumes to indicate that the disabled one or more storage volumes are also unhealthy;

decoupling the one or more unhealthy storage volumes from one or more storage nodes coupled to the one or more unhealthy storage volumes;

determining one or more replacement storage volumes based on the metadata associated with the one or more unhealthy storage volumes, one or more native queries that associate the metadata with the one or more unhealthy storage volumes, one or more queries for orphaned healthy storage volumes, and the one or more parameters for the one or more storage slots associated with each of the disabled one or more storage volumes, wherein each replacement storage volume matches the one or more parameters that correspond to at least one of the storage slots;

updating other metadata associated with the one or more replacement storage volumes to indicate that the one or more replacement storage volumes are healthy storage volumes;

coupling the one or more healthy storage volumes with the one or more storage nodes that were coupled to the one or more unhealthy storage volumes;

generating replacement metadata for the missing one or more storage volumes, wherein the replacement metadata is employed to generate and provision one or more replacement storage volumes; and

associating the one or more replacement storage volumes with the plurality of storage nodes, wherein the one or more replacement storage volumes are tagged as healthy and coupled to the one or more storage nodes, and wherein one or more portions of the healthy replacement storage volumes are assigned to each storage slot that lacks one or more healthy storage volumes; and

a client computer, comprising:

a memory that stores at least instructions; and

one or more processors that execute instructions that enable performance of actions, including:

providing one or more of the one or more threshold values.

9. The system of claim 8 , wherein determining the one or more replacement storage volumes, further comprises:

generating a query based on the metadata associated with the one or more unhealthy storage volumes;

employing the query to determine one or more uncoupled healthy storage volumes in the cloud computing environment, wherein the metadata associated with the one or more unhealthy storage volumes matches the other metadata associated with the one or more uncoupled preexisting healthy storage volumes; and

providing the one or more healthy storage volumes as at least a portion of the one or more replacement storage volumes.

10. The system of claim 8 , wherein determining the one or more replacement storage volumes, further comprises:

in response to a quantity of the one or more unhealthy storage volumes exceeding a quantity of one or more preexisting uncoupled healthy storage volumes, performing further actions, including:

provisioning one or more additional uncoupled storage volumes from the cloud computing environment that match the one or more unhealthy storage volumes;

updating metadata associated with the one or more uncoupled storage volumes to indicate that the one or more additional uncoupled storage volumes are uncoupled healthy storage volumes; and

providing the one or more uncoupled healthy storage volumes as the one or more replacement storage volumes.

11. The system of claim 8 , wherein coupling the one or more healthy storage volumes with the one or more storage nodes, further comprises:

determining the one or more storage slots associated with the one or more storage nodes based on the file system, wherein each storage slot corresponds to a fraction of a storage capacity of the file system; and

associating the one or more healthy storage volumes with the one or more storage slots, wherein each healthy storage volume is employed to provide the fraction of the storage capacity that corresponds to its associated storage slot.

12. The system of claim 8 , wherein monitoring the one or more metrics associated with each storage volume, further comprises, querying each storage node of the plurality of storage nodes for one or more values of the one or more metrics that are associated a portion of the plurality of storage volumes that are coupled to the queried storage node, wherein the one or more values of the one or more metrics are based on error information provided to each storage node by the cloud computing environment.

13. The system of claim 8 , wherein the metadata associated with the one or more storage volumes, further comprises, one or more of a storage cluster identifier, a storage slot identifier, and a field for storing a value that indicates that a storage volume is healthy or unhealthy.

14. The system of claim 8 , wherein determining the one or more replacement storage volumes, further comprises:

determining one or more storage volumes that are missing from the plurality of storage volumes based on the one or more metrics;

providing storage volume information associated with the one or more missing storage volumes based on querying the one or more storage nodes;

generating replacement metadata based on the storage volume information;

provisioning one or more additional uncoupled storage volumes from the cloud computing environment based on the replacement metadata;

updating metadata associated with the one or more additional uncoupled storage volumes to indicate that the one or more additional uncoupled storage volumes are uncoupled healthy storage volumes; and

providing the one or more uncoupled healthy storage volumes as the one or more replacement storage volumes.

15. A processor readable non-transitory storage media that includes instructions for managing data in a file system over a network, wherein execution of the instructions by one or more processors on one or more network computers performs actions, comprising:

providing the file system that includes a plurality of storage nodes and a plurality of storage volumes, wherein each storage node is coupled to a portion of the plurality of storage volumes, and wherein each storage node is a compute instance in a cloud computing environment and each storage volume is a data store in the cloud computing environment;

monitoring one or more metrics associated with each storage volume; and

in response to the one or more metrics exceeding one or more threshold values, performing further actions, including:

determining one or more storage volumes in the plurality of storage volumes that are disabled or missing based on the one or more metrics that exceed the one or more threshold values and one or more parameters for one or more storage slots associated with each of the disabled or missing one or more storage volumes;

updating metadata associated with the disabled one or more storage volumes to indicate that the disabled one or more storage volumes are also unhealthy;

decoupling the one or more unhealthy storage volumes from one or more storage nodes coupled to the one or more unhealthy storage volumes;

determining one or more replacement storage volumes based on the metadata associated with the one or more unhealthy storage volumes, one or more native queries that associate the metadata with the one or more unhealthy storage volumes, one or more queries for orphaned healthy storage volumes, and the one or more parameters for the one or more storage slots associated with each of the disabled one or more storage volumes, wherein each replacement storage volume matches the one or more parameters that correspond to at least one of the storage slots;

updating other metadata associated with the one or more replacement storage volumes to indicate that the one or more replacement storage volumes are healthy storage volumes;

coupling the one or more healthy storage volumes with the one or more storage nodes that were coupled to the one or more unhealthy storage volumes;

generating replacement metadata for the missing one or more storage volumes, wherein the replacement metadata is employed to generate and provision one or more replacement storage volumes; and

associating the one or more replacement storage volumes with the plurality of storage nodes, wherein the one or more replacement storage volumes are tagged as healthy and coupled to the one or more storage nodes, and wherein one or more portions of the healthy replacement storage volumes are assigned to each storage slot that lacks one or more healthy storage volumes.

16. The media of claim 15 , wherein determining the one or more replacement storage volumes, further comprises:

generating a query based on the metadata associated with the one or more unhealthy storage volumes;

employing the query to determine one or more uncoupled healthy storage volumes in the cloud computing environment, wherein the metadata associated with the one or more unhealthy storage volumes matches the other metadata associated with the one or more uncoupled preexisting healthy storage volumes; and

providing the one or more healthy storage volumes as at least a portion of the one or more replacement storage volumes.

17. The media of claim 15 , wherein determining the one or more replacement storage volumes, further comprises:

in response to a quantity of the one or more unhealthy storage volumes exceeding a quantity of one or more preexisting uncoupled healthy storage volumes, performing further actions, including:

provisioning one or more additional uncoupled storage volumes from the cloud computing environment that match the one or more unhealthy storage volumes;

updating metadata associated with the one or more uncoupled storage volumes to indicate that the one or more additional uncoupled storage volumes are uncoupled healthy storage volumes; and

providing the one or more uncoupled healthy storage volumes as the one or more replacement storage volumes.

18. The media of claim 15 , wherein coupling the one or more healthy storage volumes with the one or more storage nodes, further comprises:

determining the one or more storage slots associated with the one or more storage nodes based on the file system, wherein each storage slot corresponds to a fraction of a storage capacity of the file system; and

associating the one or more healthy storage volumes with the one or more storage slots, wherein each healthy storage volume is employed to provide the fraction of the storage capacity that corresponds to its associated storage slot.

19. The media of claim 15 , wherein monitoring the one or more metrics associated with each storage volume, further comprises, querying each storage node of the plurality of storage nodes for one or more values of the one or more metrics that are associated a portion of the plurality of storage volumes that are coupled to the queried storage node, wherein the one or more values of the one or more metrics are based on error information provided to each storage node by the cloud computing environment.

20. The media of claim 15 , wherein the metadata associated with the one or more storage volumes, further comprises, one or more of a storage cluster identifier, a storage slot identifier, and a field for storing a value that indicates that a storage volume is healthy or unhealthy.

21. The media of claim 15 , wherein determining the one or more replacement storage volumes, further comprises:

determining one or more storage volumes that are missing from the plurality of storage volumes based on the one or more metrics;

providing storage volume information associated with the one or more missing storage volumes based on querying the one or more storage nodes;

generating replacement metadata based on the storage volume information;

provisioning one or more additional uncoupled storage volumes from the cloud computing environment based on the replacement metadata;

updating metadata associated with the one or more additional uncoupled storage volumes to indicate that the one or more additional uncoupled storage volumes are uncoupled healthy storage volumes; and

providing the one or more uncoupled healthy storage volumes as the one or more replacement storage volumes.

22. A network computer for managing data in a file system, comprising:

a memory that stores at least instructions; and

one or more processors that execute instructions that perform actions, including:

providing the file system that includes a plurality of storage nodes and a plurality of storage volumes, wherein each storage node is coupled to a portion of the plurality of storage volumes, and wherein each storage node is a compute instance in a cloud computing environment and each storage volume is a data store in the cloud computing environment;

monitoring one or more metrics associated with each storage volume; and

in response to the one or more metrics exceeding one or more threshold values, performing further actions, including:

determining one or more storage volumes in the plurality of storage volumes that are disabled or missing based on the one or more metrics that exceed the one or more threshold values and one or more parameters for one or more storage slots associated with each of the disabled or missing one or more storage volumes;

updating metadata associated with the disabled one or more storage volumes to indicate that the disabled one or more storage volumes are also unhealthy;

decoupling the one or more unhealthy storage volumes from one or more storage nodes coupled to the one or more unhealthy storage volumes;

determining one or more replacement storage volumes based on the metadata associated with the one or more unhealthy storage volumes, one or more native queries that associate the metadata with the one or more unhealthy storage volumes, one or more queries for orphaned healthy storage volumes, and the one or more parameters for the one or more storage slots associated with each of the disabled one or more storage volumes, wherein each replacement storage volume matches the one or more parameters that correspond to at least one of the storage slots;

updating other metadata associated with the one or more replacement storage volumes to indicate that the one or more replacement storage volumes are healthy storage volumes;

coupling the one or more healthy storage volumes with the one or more storage nodes that were coupled to the one or more unhealthy storage volumes;

generating replacement metadata for the missing one or more storage volumes, wherein the replacement metadata is employed to generate and provision one or more replacement storage volumes; and

associating the one or more replacement storage volumes with the plurality of storage nodes, wherein the one or more replacement storage volumes are tagged as healthy and coupled to the one or more storage nodes, and wherein one or more portions of the healthy replacement storage volumes are assigned to each storage slot that lacks one or more healthy storage volumes.

23. The network computer of claim 22 , wherein determining the one or more replacement storage volumes, further comprises:

generating a query based on the metadata associated with the one or more unhealthy storage volumes;

employing the query to determine one or more uncoupled healthy storage volumes in the cloud computing environment, wherein the metadata associated with the one or more unhealthy storage volumes matches the other metadata associated with the one or more uncoupled preexisting healthy storage volumes; and

providing the one or more healthy storage volumes as at least a portion of the one or more replacement storage volumes.

24. The network computer of claim 22 , wherein determining the one or more replacement storage volumes, further comprises:

in response to a quantity of the one or more unhealthy storage volumes exceeding a quantity of one or more preexisting uncoupled healthy storage volumes, performing further actions, including:

provisioning one or more additional uncoupled storage volumes from the cloud computing environment that match the one or more unhealthy storage volumes;

updating metadata associated with the one or more uncoupled storage volumes to indicate that the one or more additional uncoupled storage volumes are uncoupled healthy storage volumes; and

providing the one or more uncoupled healthy storage volumes as the one or more replacement storage volumes.

25. The network computer of claim 22 , wherein coupling the one or more healthy storage volumes with the one or more storage nodes, further comprises:

determining the one or more storage slots associated with the one or more storage nodes based on the file system, wherein each storage slot corresponds to a fraction of a storage capacity of the file system; and

associating the one or more healthy storage volumes with the one or more storage slots, wherein each healthy storage volume is employed to provide the fraction of the storage capacity that corresponds to its associated storage slot.

26. The network computer of claim 22 , wherein monitoring the one or more metrics associated with each storage volume, further comprises, querying each storage node of the plurality of storage nodes for one or more values of the one or more metrics that are associated a portion of the plurality of storage volumes that are coupled to the queried storage node, wherein the one or more values of the one or more metrics are based on error information provided to each storage node by the cloud computing environment.

27. The network computer of claim 22 , wherein the metadata associated with the one or more storage volumes, further comprises, one or more of a storage cluster identifier, a storage slot identifier, and a field for storing a value that indicates that a storage volume is healthy or unhealthy.

28. The network computer of claim 22 , wherein determining the one or more replacement storage volumes, further comprises:

determining one or more storage volumes that are missing from the plurality of storage volumes based on the one or more metrics;

providing storage volume information associated with the one or more missing storage volumes based on querying the one or more storage nodes;

generating replacement metadata based on the storage volume information;

provisioning one or more additional uncoupled storage volumes from the cloud computing environment based on the replacement metadata;

updating metadata associated with the one or more additional uncoupled storage volumes to indicate that the one or more additional uncoupled storage volumes are uncoupled healthy storage volumes; and

providing the one or more uncoupled healthy storage volumes as the one or more replacement storage volumes.

Assignments (2)
SECURITY INTEREST Recorded Jun 24, 2022
From: QUMULO, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 060439/0967 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2021
From: CHMIEL, MICHAEL ANTHONY; FAIRBANKS, DUNCAN ROBERT; FLEISCHMAN, STEPHEN CRAIG; MOTLES, DANIEL MARCOS; WILLIAMS, NICHOLAS GRAEME
To: QUMULO, INC.
Reel/Frame 055611/0758 →
Continuity (1)
Related Publication 20220300155A1 · Sep 22, 2022
Cited By (9)
US 12,222,903 US 12,292,853 US 12,346,290 US 12,443,559 US 12,443,568 US 12,481,625 US 12,585,563 US 12,619,582 US 12,670,081