IP Library Granted Patent US 11,212,196
Granted Patent B2
US 11,212,196 · App. 17/163,679 · Granted Dec 28, 2021

Proportional quality of service based on client impact on an overload condition

Inventors: David D. Wright (Dacula, GA); Michael Xu (Boulder, CO)
Assignee: NetApp, Inc.
H04L41/5022G06F3/061G06F3/067G06F3/0659G06F11/3433G06F11/3485H04L41/50H04L41/5009H04L41/5067H04L67/1097G06F2201/81
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,212,196
App. No.
17/163,679
Granted
Dec 28, 2021
Kind
B2
Abstract

A distributed storage system monitors one or more system performance metrics and one or more client performance metrics related usage of the distributed storage system, including a read latency metric, a write latency metric, a total input/output (I/O) operations per second (IOPS) metric, a read IOPS metric, a write IOPS metric, an I/O size metric, a total bandwidth metric, a read bandwidth metric, a write bandwidth metric, a read/write ratio metric or statistical measures thereof over a period of time. When the distributed storage system is determined to be in an overload condition (e.g., when a system load value, calculated based on the performance metrics, exceeds a threshold), the distributed storage system independently throttles access to one or more components of the distributed storage system by one or more of multiple clients performing I/O operations to the distributed storage system based on their respective contribution to the overload condition.

Claims (42)

1. A method performed by a processing resource of a distributed storage system, the method comprising:

monitoring one or more system performance metrics of the distributed storage system each representative of an aggregate usage of the distributed storage system by a plurality of clients;

monitoring one or more client performance metrics for each of the plurality of clients performing input/output (I/O) operations to the distributed storage system, wherein each of the one or more client performance metrics represent usage of the distributed storage system by a respective client of the plurality of clients;

determining, based on the one or more system performance metrics, whether the distributed storage system is in an overload condition representing a system load value, calculated based on the one or more client performance metrics and the one or more system performance metrics, exceeding a threshold; and

when said determining is affirmative, independently throttling access to one or more components of the distributed storage system by one or more clients of the plurality of clients based on their respective contribution to the overload condition.

2. The method of claim 1 , wherein said independently throttling access to the distributed storage system comprises:

calculating a scaling factor for a particular client of the plurality of clients based on a ratio between a first client performance metric of the one or more client metrics and a corresponding system performance metric of the one or more system performance metrics;

generating a scaled target performance value based on a target performance value for the first client performance metric within a time period and the scaling factor;

generating a client performance adjustment value for the particular client using a proportional-integral-derivative (PID) controller to match the scaled target performance value based on feedback regarding the one or more monitored client performance metrics; and

causing the first client performance metric to move toward the scaled target performance value by throttling the access by the particular client during the time period.

3. The method of claim 1 , wherein the overload condition is determined by analyzing a plurality of system load values, calculated based on the one or more client performance metrics and the one or more system performance metrics, against respective corresponding thresholds in a prioritized order.

4. The method of claim 1 , wherein the one or more components comprise a particular service or group of services of a plurality of services running on the distributed storage system, a particular resource or group of resources of a plurality of resources associated with the distributed storage system, or a particular volume of a plurality of volumes of the distributed storage system.

5. The method of claim 1 , wherein the one or more system performance metrics include a read latency metric, a write latency metric, a total input/output (I/O) operations per second (IOPS) metric, a read IOPS metric, a write IOPS metric, an I/O size metric, a write cache capacity metric, a dedupe-ability metric, a compressibility metric, a total bandwidth metric, a read bandwidth metric, a write bandwidth metric, a read/write ratio metric or statistical measures thereof over a period of time.

6. The method of claim 1 , wherein the one or more client performance metrics include a read latency metric, a write latency metric, a total input/output (I/O) operations per second (IOPS) metric, a read IOPS metric, a write IOPS metric, an I/O size metric, a total bandwidth metric, a read bandwidth metric, a write bandwidth metric, a read/write ratio metric or statistical measures thereof over a period of time.

7. A distributed storage system comprising:

a processing resource; and

a non-transitory computer-readable medium, coupled to the processing resource, having stored therein instructions that when executed by the processing resource cause the distributed storage system to:

monitor one or more system performance metrics of the distributed storage system each representative of an aggregate usage of the distributed storage system by a plurality of clients;

monitor one or more client performance metrics for each of the plurality of clients performing input/output (I/O) operations to the distributed storage system, wherein each of the one or more client performance metrics represent usage of the distributed storage system by a respective client of the plurality of clients; and

when the distributed storage system is determined to be in an overload condition based on the one or more system performance metrics, reduce a load on the distributed storage system by independently throttling access to one or more components of the distributed storage system by one or more clients of the plurality of clients based on their respective contribution to the overload condition, wherein the overload condition represents a system load value, calculated based on the one or more client performance metrics and the one or more system performance metrics, exceeding a threshold.

8. The distributed storage system of claim 7 , wherein said independently throttling access to the distributed storage system comprises:

calculating a scaling factor for a particular client of the plurality of clients based on a ratio between a first client performance metric of the one or more client metrics and a corresponding system performance metric of the one or more system performance metrics;

generating a scaled target performance value based on a target performance value for the first client performance metric within a time period and the scaling factor;

generating a client performance adjustment value for the particular client using a proportional-integral-derivative (PID) controller to match the scaled target performance value based on feedback regarding the one or more monitored client performance metrics; and

causing the first client performance metric to move toward the scaled target performance value by throttling the access by the particular client during the time period.

9. The distributed storage system of claim 7 , wherein the overload condition is determined by analyzing a plurality of system load values, calculated based on the one or more client performance metrics and the one or more system performance metrics, against respective corresponding thresholds in a prioritized order.

10. The distributed storage system of claim 7 , wherein the one or more components comprise a particular service or group of services of a plurality of services running on the distributed storage system, a particular resource or group of resources of a plurality of resources associated with the distributed storage system, or a particular volume of a plurality of volumes of the distributed storage system.

11. The distributed storage system of claim 7 , wherein the one or more system performance metrics include a read latency metric, a write latency metric, a total input/output (I/O) operations per second (IOPS) metric, a read IOPS metric, a write IOPS metric, an I/O size metric, a write cache capacity metric, a dedupe-ability metric, a compressibility metric, a total bandwidth metric, a read bandwidth metric, a write bandwidth metric, a read/write ratio metric or statistical measures thereof over a period of time.

12. The distributed storage system of claim 7 , wherein the one or more client performance metrics include a read latency metric, a write latency metric, a total input/output (I/O) operations per second (IOPS) metric, a read IOPS metric, a write IOPS metric, an I/O size metric, a total bandwidth metric, a read bandwidth metric, a write bandwidth metric, a read/write ratio metric or statistical measures thereof over a period of time.

13. A non-transitory computer-readable medium having instructions stored thereon, which when executable by a processor of a distributed storage system, cause the distributed storage system to:

monitor one or more system performance metrics of the distributed storage system each representative of an aggregate usage of the distributed storage system by a plurality of clients;

monitor one or more client performance metrics for each of the plurality of clients performing input/output (I/O) operations to the distributed storage system, wherein each of the one or more client performance metrics represent usage of the distributed storage system by a respective client of the plurality of clients;

determine, based on the one or more system performance metrics, whether the distributed storage system is in an overload condition representing a system load value, calculated based on the one or more client performance metrics and the one or more system performance metrics, exceeding a threshold; and

when the distributed storage system is determined to be in the overload condition reduce a load on the distributed storage system by independently throttling access to one or more components of the distributed storage system by one or more clients of the plurality of clients based on their respective contribution to the overload condition.

14. The non-transitory computer-readable medium of claim 13 , wherein said independently throttling access to the distributed storage system comprises:

calculating a scaling factor for a particular client of the plurality of clients based on a ratio between a first client performance metric of the one or more client metrics and a corresponding system performance metric of the one or more system performance metrics;

generating a scaled target performance value based on a target performance value for the first client performance metric within a time period and the scaling factor;

generating a client performance adjustment value for the particular client using a proportional-integral-derivative (PID) controller to match the scaled target performance value based on feedback regarding the one or more monitored client performance metrics; and

causing the first client performance metric to move toward the scaled target performance value by throttling the access by the particular client during the time period.

15. The non-transitory computer-readable medium of claim 13 , wherein the one or more components comprise a particular service or group of services of a plurality of services running on the distributed storage system, a particular resource or group of resources of a plurality of resources associated with the distributed storage system, or a particular volume of a plurality of volumes of the distributed storage system.

16. The non-transitory computer-readable medium of claim 13 , wherein the one or more system performance metrics include a read latency metric, a write latency metric, a total input/output (I/O) operations per second (IOPS) metric, a read IOPS metric, a write IOPS metric, an I/O size metric, a write cache capacity metric, a dedupe-ability metric, a compressibility metric, a total bandwidth metric, a read bandwidth metric, a write bandwidth metric, a read/write ratio metric or statistical measures thereof over a period of time.

17. The non-transitory computer-readable medium of claim 13 , wherein the one or more client performance metrics include a read latency metric, a write latency metric, a total input/output (I/O) operations per second (IOPS) metric, a read IOPS metric, a write IOPS metric, an I/O size metric, a total bandwidth metric, a read bandwidth metric, a write bandwidth metric, a read/write ratio metric or statistical measures thereof over a period of time.

Continuity (8)
Continuation 16588594 · Sep 30, 2019
Continuation 15651438 · Jul 17, 2017
Continuation 14701832 · May 1, 2015
Continuation 13856997 · Apr 4, 2013
Continuation PCTUS2012071844 · Dec 27, 2012
Continuation In Part 13338039 · Dec 27, 2011
Provisional Application 61697905 · Sep 7, 2012
Related Publication 20210160155A1 · May 27, 2021
Cited By (4)
US 12,236,126 US 12,250,129 US 12,373,326 US 12,443,550