Low impact migration of large data to cloud and virtualized environments
Illustrative methods and systems for migrating a large amount of data from a source site to one or more cloud destination sites with minimal interruptions to the source site during the migration process. Configured media agents initiate specialized backup jobs on source data to create backup copies of a specified size that minimally impacts the source site. Those backup copies are received at storage management site and stored in cache. For each received data blob, the illustrative system restores the backup copies to the destination site in the format of the original application data.
1 . A computer-implemented method, the computer-implemented method comprising:
identifying a dataset within a source system that is to be migrated to a destination system via a staged migration operation;
dividing up the dataset among a plurality of backup jobs, wherein the dividing is based on one or more of: application type, data sensitivity, user type, or available network or computing resources;
assigning priority to each of the plurality of backup jobs according to one or more of:
access frequency,
data sensitivity,
a creation or modification time,
a user type,
a user identification,
an application type;
executing a first backup job of the plurality of backup jobs,
wherein the first backup job comprises:
backing up first source data associated with the first backup job to store a first backup in cache;
determining operating system-specific settings and application-specific settings configurations associated with the first source data;
configuring a first destination client within the destination system according to the operating system-specific settings and application-specific configurations associated with the first source data; and
upon completion of the first backup job, restoring the first backup stored in the cache to the first destination client within the destination system,
wherein the restoring is initiated before remainder of the plurality of backup jobs are completed; and
upon completion of the staged migration operation, initiating one or more incremental backup operations, wherein the one or more incremental backup operations backs up changes to the dataset that occurred after the staged migration operation was initiated.
2 . The computer-implemented method of claim 1 , wherein the dividing of the dataset among the plurality of backup jobs is determined, at least in part, based on a file or data types of data in the dataset.
3 . The computer-implemented method of claim 1 , wherein the first backup job is executed using a storage accelerator.
4 . The computer-implemented method of claim 3 , wherein the storage accelerator writes the first backup directly to the cache of a migration system.
5 . The computer-implemented method of claim 1 , wherein the identification of the dataset to be migrated is made in accordance with an information management policy assigned to the dataset.
6 . The computer-implemented method of claim 1 , the computer-implemented method further comprising initiating creation or initiation of one or more computing devices at the destination system for hosting data to be migrated to the destination system.
7 . The computer-implemented method of claim 1 , wherein the dividing of the dataset among the plurality of backup jobs is, at least in part, based on network or computing resources of the source system.
8 . The computer-implemented method of claim 1 , wherein the first backup is stored in a backup format and comprises metadata to facilitate the staged migration operation.
9 . The computer-implemented method of claim 1 , the computer-implemented method further comprising: upon completion of the first backup job, initiating a replication job to replicate the first backup stored in the cache to create a first secondary copy.
10 . The computer-implemented method of claim 9 , wherein the replication job is executed simultaneously as the restoration of the first backup to the destination system.
11 . A computer-implemented system, the computer-implemented system configured to:
with one or more processors:
identify a dataset within a source system that is to be migrated to a destination system via a staged migration operation;
divide up the dataset among a plurality of backup jobs, wherein the dividing is based on one or more of: application type, data sensitivity, user type, or available network or computing resources;
assign priority to each of the plurality of backup jobs according to one or more of:
access frequency,
data sensitivity,
a creation or modification time,
a user type,
a user identification,
an application type; and
execute a first backup job of the plurality of backup jobs,
wherein the first backup job comprises:
backing up first source data associated with the first backup job to store a first backup in cache;
determine operating system-specific settings and application-specific settings configurations associated with the first source data;
configure a first destination client within the destination system according to the operating system-specific settings and application-specific configurations associated with the first source data; and
upon completion of the first backup job, restore the first backup stored in the cache to the first destination client within the destination system,
wherein the restoring is initiated before remainder of the plurality of backup jobs are completed; and
upon completion of the staged migration operation, initiate one or more incremental backup operations, wherein the one or more incremental backup operations backs up changes to the dataset that occurred after the staged migration operation was initiated.
12 . The computer-implemented system of claim 11 , wherein the dividing of the dataset among the plurality of backup jobs is determined, at least in part, based on a file or data types of data in the dataset.
13 . The computer-implemented system of claim 11 , wherein the first backup job is executed using a storage accelerator.
14 . The computer-implemented system of claim 13 , wherein the storage accelerator writes the first backup directly to the cache of a migration system.
15 . The computer-implemented system of claim 11 , wherein the identification of the dataset to be migrated is made in accordance with an information management policy assigned to the dataset.
16 . The computer-implemented system of claim 11 , wherein the computer-implemented system is further configured to create or initiate one or more computing devices at the destination system for hosting data to be migrated to the destination system.
17 . The computer-implemented system of claim 11 , wherein the dividing of the dataset among the plurality of backup jobs is, at least in part, based on network or computing resources of the source system.
18 . The computer-implemented system of claim 11 , wherein the first backup is stored in a backup format and comprises metadata to facilitate the staged migration operation.
19 . The computer-implemented system of claim 11 , wherein the computer-implemented system is further configured to: upon completion of the first backup job, initiate a replication job to replicate the first backup stored in the cache to create a first secondary copy.
20 . The computer-implemented system of claim 19 , wherein the replication job is executed simultaneously as the restoration of the first backup to the destination system.