IP Library Granted Patent US 10,185,637
Granted Patent B2
US 10,185,637 · App. 14/623,013 · Granted Jan 22, 2019

Preserving management services with distributed metadata through the disaster recovery life cycle

Inventors: Yu Deng (Yorktown Heights, NY); Ruchi Mahindru (Elmsford, NY); HariGovind V. Ramasamy (Ossining, NY); Soumitra Sarkar (Cary, NC); Long Wang (White Plains, NY)
Assignee: International Business Machines Corporation
G06F11/2007G06F11/1458G06F11/1662G06F11/1666G06F11/2069G06F2201/805G06F2201/84G06F2201/85
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,185,637
App. No.
14/623,013
Granted
Jan 22, 2019
Kind
B2
Abstract

For disaster recovery involving a first site and a disaster recovery site, where at least a portion of management service metadata not isolated within the management service, a failover process is initiated, including creating an initial snapshot of the distributed metadata state. In a failback process, a representation is created of state changes for the management service and a delta description is calculated therefrom. The delta description is transmitted to the first site; and a reverse replica is created, at the first site, of all the workload components from the disaster recovery site. The delta description is played back to restore a distributed metadata state that existed in the disaster recovery site and to re-create it in the first site.

Claims (49)

1. A non-transitory computer readable medium comprising computer executable instructions which when executed by a computer cause the computer to perform the method of:

during normal operation, at a first site, of a disaster recovery management unit comprising at least one customer workload machine and at least one management service machine implementing at least one management service, replicating to a remote disaster recovery site said at least one customer workload machine, said at least one management service machine, and metadata for said at least one management service implemented by said at least one management service machine, at least a portion of said metadata not being isolated within said at least one management service;

after a disaster at said first site, initiating a failover process comprising:

bringing up, at said remote disaster recovery site, a replicated version of said at least one customer workload machine;

bringing up, at said remote disaster recovery site, a replicated version of said at least one management service machine;

operating, at said remote disaster recovery site, said replicated version of said at least one customer workload machine and said replicated version of said at least one management service machine, in accordance with said metadata for said at least one management service implemented on said at least one management service machine; and

creating an initial snapshot of a distributed metadata state of said metadata for said at least one management service implemented on said replicated version of said at least one management service machine, wherein said distributed metadata is distributed across at least two of a provisioning service, a customer virtual machine, a hypervisor, a network switch or bridge, a storage system, and said replicated version of said at least one management service machine;

subsequent to initiating said failover process, initiating a failback process comprising:

creating a representation of state changes for said at least one management service implemented on said replicated version of said at least one management service machine made in said remote disaster recovery site since said failover process and calculating therefrom a delta description from said initial snapshot;

transmitting said delta description to said first site; and

creating a reverse replica of all the workload components from the remote disaster recovery site at the first site and playing back the delta description to restore a distributed metadata state that existed in the remote disaster recovery site and re-create it in the first site,

wherein said at least one management service comprises a provisioning service, further comprising computer executable instructions which when executed by the computer cause the computer to perform the additional method steps of:

subsequent to said step of operating said replicated version of said at least one customer workload machine and said replicated version of said at least one management service machine in accordance with said metadata for said at least one management service, carrying out and tracking additional provisioning at said remote disaster recovery site and

subsequent to said additional provisioning, upon said first site coming back up, restoring said first site to reflect said tracked additional provisioning.

2. The non-transitory computer readable medium of claim 1 , wherein said additional provisioning comprises provisioning a new virtual machine.

3. The non-transitory computer readable medium of claim 1 , wherein said additional provisioning comprises deleting an existing virtual machine.

4. An apparatus comprising:

a memory; and

at least one processor, coupled to said memory, and operative to:

during normal operation, at a first site, of a disaster recovery management unit comprising at least one customer workload machine and at least one management service machine implementing at least one management service, replicate to a remote disaster recovery site said at least one customer workload machine, said at least one management service machine, and metadata for said at least one management service, at least a portion of said metadata not being isolated within said at least one management service;

after a disaster at said first site, initiate a failover process comprising:

bringing up, at said remote disaster recovery site, a replicated version of said at least one customer workload machine;

bringing up, at said remote disaster recovery site, a replicated version of said at least one management service machine;

operating, at said remote disaster recovery site, said replicated version of said at least one customer workload machine and said replicated version of said at least one management service machine, in accordance with said metadata for said at least one management service; and

creating an initial snapshot of a distributed metadata state of said metadata for said at least one management service implemented on said replicated version of said at least one management service machine, wherein said distributed metadata is distributed across at least two of a provisioning service, a customer virtual machine, a hypervisor, a network switch or bridge, a storage system, and said replicated version of said at least one management service machine;

subsequent to initiating said failover process, initiate a failback process comprising:

creating a representation of state changes for said at least one management service implemented on said replicated version of said at least one management service machine made in said remote disaster recovery site since said failover process and calculating therefrom a delta description from said initial snapshot;

transmitting said delta description to said first site; and

creating a reverse replica of all the workload components from the remote disaster recovery site at the first site and playing back the delta description to restore a distributed metadata state that existed in the remote disaster recovery site and re-create it in the first site,

wherein said at least one management service comprises a provisioning service, wherein said at least one processor is further operative to:

subsequent to said step of operating said replicated version of said at least one customer workload machine and said replicated version of said at least one management service machine in accordance with said metadata for said at least one management service, carry out and track additional provisioning at said remote disaster recovery site; and

subsequent to said additional provisioning, upon said first site coming back up, restore said first site to reflect said tracked additional provisioning.

5. The apparatus of claim 4 , wherein said additional provisioning comprises provisioning a new virtual machine.

6. The apparatus of claim 4 , wherein said additional provisioning comprises deleting an existing virtual machine.

7. A non-transitory computer readable medium comprising computer executable instructions which when executed by a computer cause the computer to perform the method of:

during normal operation, at a first site, of a disaster recovery management unit comprising at least one customer workload machine and at least one management service machine implementing at least one management service, replicating to a remote disaster recovery site said at least one customer workload machine, said at least one management service machine, and metadata for said at least one management service, at least a portion of said metadata not being isolated within said at least one management service;

after a disaster at said first site, initiating a failover process comprising:

bringing up, at said remote disaster recovery site, a replicated version of said at least one customer workload machine;

bringing up, at said remote disaster recovery site, a replicated version of said at least one management service machine;

operating, at said remote disaster recovery site, said replicated version of said at least one customer workload machine and said replicated version of said at least one management service machine, in accordance with said metadata for said at least one management service;

creating an initial snapshot of a distributed metadata state of said metadata for said at least one management service implemented on said replicated version of said at least one management service machine; and

performing a fix-up process on said distributed metadata state for said replicated version of said at least one management service machine, responsive to differences between said initial snapshot and a state of said replicated version of said at least one customer workload machine;

subsequent to initiating said failover process, initiating a failback process comprising:

creating a representation of state changes for said at least one management service implemented on said replicated version of said at least one management service machine made in said remote disaster recovery site since said failover process and calculating therefrom a delta description from said initial snapshot;

transmitting said delta description to said first site; and

creating a reverse replica of all the workload components from the remote disaster recovery site at the first site and playing back the delta description to restore a distributed metadata state that existed in the remote disaster recovery site and re-create it in the first site,

wherein said at least one management service comprises a provisioning service, further comprising computer executable instructions which when executed by the computer cause the computer to perform the additional method steps of:

subsequent to said step of operating said replicated version of said at least one customer workload machine and said replicated version of said at least one management service machine in accordance with said metadata for said at least one management service, carrying out and tracking additional provisioning at said remote disaster recovery site; and

subsequent to said additional provisioning, upon said first site coming back up, restoring said first site to reflect said tracked additional provisioning.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 16, 2015
From: DENG, YU; MAHINDRU, RUCHI; RAMASAMY, HARIGOVIND V.; SARKAR, SOUMITRA; WANG, LONG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 034966/0201 →
Continuity (1)
Related Publication 20160239392A1 · Aug 18, 2016
Cited By (1)
US 12,475,004