IP Library Granted Patent US 10,303,508
Granted Patent B2
US 10,303,508 · App. 15/926,260 · Granted May 28, 2019

Adaptive self-maintenance scheduler

Inventors: Tarang Vaish (Santa Clara, CA); Anirvan Duttagupta (San Jose, CA); Sashi Madduri (Mountain View, CA)
Assignee: Cohesity, Inc.
G06F9/48G06F3/061G06F3/0605G06F3/067G06F3/0659G06F9/4843G06F9/4881G06F9/4887G06F9/50G06F9/5005G06F9/505G06F9/5011G06F9/5016G06F9/5022G06F9/5027G06F9/5033G06F9/5038G06F9/5044G06F9/5055G06F9/5094
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,303,508
App. No.
15/926,260
Granted
May 28, 2019
Kind
B2
Abstract

Embodiments presented herein disclose adaptive techniques for scheduling self-maintenance processes. A load predictor estimates, based on a current state of a distributed storage system, an amount of resources of the system required to perform each of a plurality of self-maintenance processes. A maintenance process scheduler estimates, based on one or more inputs, an amount of resources of the distributed system available to perform one or more of the self-maintenance processes during at least a first time period. The maintenance process scheduler determines a schedule for the one or more of the self-maintenance processes to perform during the first time period, based on the estimated amount of resources required and available.

Claims (46)

1. A method, comprising:

estimating, based on a current state of a distributed storage system comprising a primary storage server and a plurality of secondary storage servers, an amount of computing resources of the primary storage server and the plurality of secondary storage servers required to perform each of a plurality of self-maintenance processes, wherein the plurality of secondary storage servers are configured to perform one or more backup processes and the plurality of self-maintenance processes;

estimating an amount of computing resources of the primary storage server and the plurality of secondary storage servers available to perform one or more of the plurality of self-maintenance processes during at least a first time period of a plurality of time periods;

determining which of the plurality of self-maintenance processes to perform during the at least first time period based on the estimated amount of computing resources of the primary storage server and the plurality of secondary storage servers required to perform each of the plurality of self-maintenance processes and the estimated amount of computing resources of the primary storage server and the plurality of secondary storage servers available to perform one or more of the plurality of self-maintenance processes; and

scheduling the determined one or more self-maintenance processes to perform during the first time period.

2. The method of claim 1 , wherein the amount of computing resources of the primary storage server are estimated based on one or more inputs, wherein the one or more inputs includes at least one of a plurality of current activities of the primary storage server and the plurality of secondary storage servers, the current state of the distributed storage system comprising the primary storage server and the the plurality of secondary storage servers, a plurality of external events, and the estimated amount of computing resources of the primary storage server and the plurality of secondary storage servers required to perform each of the self-maintenance processes.

3. The method of claim 1 , further comprising:

performing each of the scheduled self-maintenance processes during the first time period; and

collecting execution statistics of the performance of each of the scheduled self-maintenance processes.

4. The method of claim 3 , wherein the amount of computing resources of the primary storage server and the plurality of secondary storage servers required and the amount of computing resources of the primary storage server and the plurality of secondary storage servers available are further estimated based on the execution statistics.

5. The method of claim 4 , wherein the amount of computing resources of the primary storage server are estimated based on one or more inputs, the method further comprising:

upon determining that the estimated amount of computing resources available is within a specified range of an actual amount of computing resources available based on the execution statistics, reinforcing the one or more inputs using a machine learning algorithm; and

upon determining that the estimated amount of computing resources available is not within the specified range of the actual amount of computing resources available based on the execution statistics, readjusting the one or more inputs using the machine learning algorithm.

6. The method of claim 1 , wherein the amount of computing resources required and amount of computing resources available include at least one of an amount of I/O resources, network resources, storage capacity, and processing resources.

7. The method of claim 1 , wherein the plurality of self-maintenance processes include at least one of garbage collection, updating free blocks, updating internal statistics, performing compaction methods, cloud spill optimization, or remote replication.

8. A non-transitory computer-readable storage medium having instructions, which, when executed on a processor, perform an operation comprising:

estimating, based on a current state of a distributed storage system comprising a primary storage server and a plurality of secondary storage servers, an amount of computing resources of the primary storage server and the plurality of secondary storage servers required to perform each of a plurality of self-maintenance processes, wherein the plurality of secondary storage servers are configured to perform one or more backup processes and the plurality of self-maintenance processes;

estimating an amount of computing resources of the primary storage server and the plurality of secondary storage servers available to perform one or more of the plurality of self-maintenance processes during at least a first time period of a plurality of time periods;

determining which of the plurality of self-maintenance processes to perform during the at least first time period based on the estimaed amount of computing resources of the primary storage server and the plurality of secondary storage servers required to perform each of the plurality of self-maintenance processes and the estimated amount of computing resources of the primary storage server and the plurality of secondary storage servers available to perform one or more of the plurality of self-maintenance processes; and

scheduling the determined one or more self-maintenance processes to perform during the first time period.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the amount of computing resources of the primary storage server are estimated based on one or more inputs, wherein the one or more inputs includes at least one of a plurality of current activities of the primary storage server and the plurality of secondary storage servers, the current state of the distributed storage system comprising the primary storage server and the plurality of secondary storage servers, a plurality of external events, and the estimated amount of computing resources of the primary storage server and the plurality of secondary storage servers required to perform each of the self-maintenance processes.

10. The non-transitory computer-readable storage medium of claim 8 , wherein the operation further comprises:

performing each of the scheduled self-maintenance processes during the first time period; and

collecting execution statistics of the performance of each of the scheduled self-maintenance processes.

11. The non-transitory computer-readable storage medium of claim 10 , wherein the amount of computing resources of the primary storage server and the plurality of secondary storage servers required and the amount of computing resources of the primary storage server and the plurality of secondary storage servers available are further estimated based on the execution statistics.

12. The non-transitory computer-readable storage medium of claim 11 , wherein the amount of computing resources of the primary storage server are estimated based on one or more inputs, the operation further comprising:

upon determining that the estimated amount of computing resources available is within a specified range of an actual amount of computing resources available based on the execution statistics, reinforcing the one or more inputs using a machine learning algorithm; and

upon determining that the estimated amount of computing resources available is not within the specified range of the actual amount of computing resources available based on the execution statistics, readjusting the one or more inputs using the machine learning algorithm.

13. The non-transitory computer-readable storage medium of claim 8 , wherein the amount of computing resources required and amount of computing resources available include at least one of an amount of I/O resources, network resources, storage capacity, and processing resources.

14. The non-transitory computer-readable storage medium of claim 8 , wherein the plurality of self-maintenance processes include at least one of garbage collection, updating free blocks, updating internal statistics, performing compaction methods, cloud spill optimization, or remote replication.

15. A system, comprising:

a processor; and

a memory containing program code, which, when executed on the processor, performs an operation comprising:

estimating, based on a current state of a distributed storage system comprising a primary storage server and a plurality of secondary storage servers, an amount of computing resources of the primary storage server and the plurality of secondary storage servers required to perform each of a plurality of self-maintenance processes, wherein the plurality of secondary storage servers are configured to perform one or more backup processes and the plurality of self-maintenance processes;

estimating, based on one or more inputs, an amount of computing resources of the primary storage server and the plurality of secondary storage servers available to perform one or more of the plurality of self-maintenance processes during at least a first time period of a plurality of time periods;

determine which of the plurality of self-maintenance processes to perform during the at least first time period based on the estimated amount of computing resources of the primary storage server and the plurality of secondary storage servers required to perform each of the plurality of self-maintenance processes and the estimated amount of computing resources of the primary storage server and the plurality of secondary storage servers available to perform one or more of the plurality of self-maintenance processes; and

scheduling the determined one or more self-maintenance processes to perform during the first time period.

16. The system of claim 15 , wherein the one or more inputs includes at least one of a plurality of current activities of the primary storage server and the plurality of secondary storage servers, the current state of the distributed storage system comprising the primary storage server and the plurality of secondary storage servers, a plurality of external events, and the estimated amount of computing resources of the primary storage server and the plurality of secondary storage servers required to perform each of the self-maintenance processes.

17. The system of claim 15 , wherein the operation further comprises:

performing each of the scheduled self-maintenance processes during the first time period; and

collecting execution statistics of the performance of each of the scheduled self-maintenance processes.

18. The system of claim 17 , wherein the amount of computing resources of the primary storage server and the plurality of secondary storage servers required and the amount of computing resources of the primary storage server and the plurality of secondary storage servers available are further estimated based on the execution statistics.

19. The system of claim 18 , the operation further comprising:

upon determining that the estimated amount of computing resources available is within a specified range of an actual amount of computing resources available based on the execution statistics, reinforcing the one or more inputs using a machine learning algorithm; and

upon determining that the estimated amount of computing resources available is not within the specified range of the actual amount of computing resources available based on the execution statistics, readjusting the one or more inputs using the machine learning algorithm.

20. The system of claim 15 , wherein the amount of computing resources required and amount of computing resources available include at least one of an amount of I/O resources, network resources, storage capacity, and processing resources.

Assignments (4)
TERMINATION AND RELEASE OF INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 10, 2024
From: FIRST-CITIZENS BANK & TRUST COMPANY (AS SUCCESSOR TO SILICON VALLEY BANK)
To: COHESITY, INC.
Reel/Frame 069584/0498 →
SECURITY INTEREST Recorded Dec 9, 2024
From: VERITAS TECHNOLOGIES LLC; COHESITY, INC.
To: JPMORGAN CHASE BANK. N.A.
Reel/Frame 069890/0001 →
SECURITY INTEREST Recorded Sep 23, 2022
From: COHESITY, INC.
To: SILICON VALLEY BANK, AS ADMINISTRATIVE AGENT
Reel/Frame 061509/0818 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2019
From: VAISH, TARANG; DUTTAGUPTA, ANIRVAN; MADDURI, SASHI
To: COHESITY, INC.
Reel/Frame 048255/0730 →
Continuity (2)
Continuation 14852249 · Sep 11, 2015
Related Publication 20180210754A1 · Jul 26, 2018