IP Library Granted Patent US 12,026,072
Granted Patent B2
US 12,026,072 · App. 17/675,263 · Granted Jul 2, 2024

Metering framework for improving resource utilization for a disaster recovery environment

Inventors: Abhishek Gupta (Ghaziabad, IN); Bhushan Pandit (Pune, IN); Pranab Patnaik (Cary, NC)
Assignee: Nutanix, Inc.
G06F11/2035G06F11/1435G06F11/1464G06F11/1469G06F2201/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,026,072
App. No.
17/675,263
Filed
Feb 18, 2022
Granted
Jul 2, 2024
Kind
B2
Examiner
TRUONG, LOAN
Art Unit
2114
USPC
714/4.11
Abstract

A framework is described that improves resource utilization during operations executing within workflows of the distributed data processing system (e.g., having a plurality of interconnected nodes) in a disaster recovery (DR) environment configured to support synchronous and asynchronous (i.e., heterogeneous) DR workflows (e.g., generating snapshots and replicating data) that include synchronous replication, asynchronous replication, nearsync (i.e., short duration snapshots of metadata) replication and migration of data objects associated with the workflows for failover (e.g., replication and/or migration) to a secondary site in the event of failure of the primary site. The framework meters (regulates) execution of the operations directed to the workloads so as to efficiently use the resources in a manner that allows timely progress (completion) of certain (e.g., high-frequency) operations and reduction in blocking (stalling) of other (e.g., low-frequency) operations by avoiding unnecessary resource hoarding/consumption and contention. Notably, the framework also provides metering and tuning of properties during execution of the workflows and maintains their state to provide for recovery.

Claims (45)

1. A method comprising:

estimating a load of a first disaster recovery (DR) workflow based on priority and system resources needed for completion in a multi-site DR environment;

loading an intent-to-create meta-operation into a first level of a queue based on the estimated load of the first DR workflow;

instantiating the first DR workflow from the intent-to-create meta-operation based on a first access policy associated with the first level of the queue;

loading the first DR workflow in the first level of the queue for execution when the system resources needed for completion of the first DR workflow are available; and

metering use of the system resources for the first DR workflow based on the estimated load to permit completion of a shorter duration second DR workflow.

2. The method of claim 1 , further comprising:

borrowing a system resource quota capacity from the first DR workflow of the first level of the queue for use by the second DR workflow.

3. The method of claim 1 , wherein execution of the first DR workflow is delayed without violating a recovery point objective.

4. The method of claim 1 , wherein the first access policy applied to the first level of the queue is different from a second access policy applied to a second level of the queue.

5. The method of claim 1 , wherein the first DR workflow has multiple stages including a snapshot stage and data replication stage, and wherein the first level of the queue applies to the snapshot stage and a second level of the queue applies to the data replication stage.

6. The method of claim 1 , wherein the first DR workflow is an asynchronous replication based on incremental changes between snapshots.

7. The method of claim 1 , wherein the load of the first DR workflow is calculated as relative to loads of other DR workflows.

8. The method of claim 1 , wherein the system resources include network bandwidth between sites of the multi-site DR environment.

9. The method of claim 1 , further comprising employing feedback to determine capacity of the system resources for queuing of the first DR workload.

10. A non-transitory computer readable medium including program instructions for execution on a processor, the program instructions configured to:

estimate a load of a first disaster recovery (DR) workflow based on priority and system resources needed for completion in a multi-site DR environment;

load an intent-to-create meta-operation into a first level of a queue based on the estimated load of the first DR workflow;

instantiate the first DR workflow from the intent-to-create meta-operation based on a first access policy associated with the first level of the queue;

load the first DR workflow in the first level of the queue for execution when the system resources needed for completion of the first DR workflow are available; and

meter use of the system resources for the first DR workflow based on the estimated load to permit completion of a shorter duration second DR workflow.

11. The non-transitory computer readable medium of claim 10 wherein the program instructions for execution on a processor are further configured to:

borrow a system resource quota capacity from the first DR workflow of the first level of the queue for use by the second DR workflow.

12. The non-transitory computer readable medium of claim 10 , wherein execution of the first DR workflow is delayed without violating a recovery point objective.

13. The non-transitory computer readable medium of claim 10 , wherein the first access policy applied to the first level of the queue is different from a second access policy applied to a second level of the queue.

14. The non-transitory computer readable medium of claim 10 , wherein the first DR workflow has multiple stages including a snapshot stage and data replication stage, and wherein the first level of the queue applies to the snapshot stage and a second level of the queue applies to the data replication stage.

15. The non-transitory computer readable medium of claim 10 , wherein the first DR workflow is an asynchronous replication based on incremental changes between snapshots.

16. The non-transitory computer readable medium of claim 10 , wherein the load of the first DR workflow is calculated as relative to loads of other DR workflows.

17. The non-transitory computer readable medium of claim 10 , wherein the program instructions for execution on a processor are further configured to employ feedback to determine capacity of the system resources for queuing of the first DR workload.

18. An apparatus comprising:

a replication manager of a node in a cluster of interconnected nodes of a multi-site DR environment, the replication manager running on the node having a processor configured to execute program instructions to,

estimate a load of a first disaster recovery (DR) workflow based on priority and system resources needed for completion in the DR environment;

load an intent-to-create meta-operation into a first level of a queue based on the estimated load of the first DR workflow;

instantiate the first DR workflow from the intent-to-create meta-operation based on a first access policy associated with the first level of the queue;

load the first DR workflow in the first level of the queue for execution when the system resources needed for completion of the first DR workflow are available; and

meter use of the system resources for the first DR workflow based on the estimated load to permit completion of shorter duration second DR workflow.

19. The apparatus of claim 18 wherein the program instructions further include program instructions to:

borrow a system resource quota capacity from the first DR workflow of the first level of the queue for use by the second DR workflow.

20. The apparatus of claim 18 , wherein execution of the first DR workflow is delayed without violating a recovery point objective.

21. The apparatus of claim 18 , wherein the first access policy applied to the first level of the queue is different from a second access policy applied to a second level of the queue.

22. The apparatus of claim 18 , wherein the first DR workflow has multiple stages including a snapshot stage and data replication stage, and wherein the first level of the queue applies to the snapshot stage and a second level of the queue applies to the data replication stage.

23. The apparatus of claim 18 , wherein the first DR workflow is an asynchronous replication based on incremental changes between snapshots.

24. The apparatus of claim 18 , wherein the load of the first DR workflow is calculated as relative to loads of other DR workflows.

25. The apparatus of claim 18 , wherein the system resources include network bandwidth between sites of the multi-site DR environment.

26. The apparatus of claim 18 , wherein the program instructions further include program instructions to employ feedback to determine capacity of the system resources for queuing of the first DR workload.

Assignments (2)
SECURITY INTEREST Recorded Feb 13, 2025
From: NUTANIX, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 070206/0463 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2022
From: GUPTA, ABHISHEK; PANDIT, BHUSHAN; PATNAIK, PRANAB
To: NUTANIX, INC.
Reel/Frame 059047/0812 →
Priority Claims (1)
IN 202141060697 · Dec 24, 2021 · national
Continuity (1)
Related Publication 20230205653A1 · Jun 29, 2023