IP Library › Granted Patent US 11,481,616
Granted Patent B2
US 11,481,616 · App. 16/198,642 · Granted Oct 25, 2022

Framework for providing recommendations for migration of a database to a cloud computing system

Inventors: Mitchell Gregory Spryn (Seattle, WA); Intaik Park (Bellevue, WA); Felipe Vieira Frujeri (Vancouver, CA); Vijay Govind Panjeti (Bothell, WA); Ashok Sai Madala (Redmond, WA); Ajay Kumar Karanam (Redmond, WA)
Assignee: MICROSOFT TECHNOLOGY LICENSING, LLC
G06N3/08G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,481,616
App. No.
16/198,642
Granted
Oct 25, 2022
Kind
B2
Abstract

To obtain one or more recommendations for the migration of a database to a cloud computing system, information about performance of the database operating under a workload may be obtained. A first machine learning model (e.g., a neural network-based autoencoder) may be used to generate a compressed representation of characteristics of the database operating under the workload. The compressed representation may then be provided as input to a second machine learning model (e.g., a neural network-based classifier), which outputs a recommendation regarding a characteristic (e.g., size, configuration, level of service) of the cloud database to which the database should be migrated. This type of recommendation may be made prior to migration, thereby making it easier to properly estimate the cost of running the cloud database and plan the migration accordingly.

Claims (58)

1. A method for providing one or more recommendations for migration of a database to a cloud computing system, comprising:

executing a plurality of different workloads against a plurality of different databases, wherein executing the plurality of different workloads against the plurality of different databases comprises:

executing a first workload against a first plurality of different instances of a first database running on different computer systems having different hardware configurations to generate first sets of information corresponding to the first database; and

executing a second workload against a second plurality of different instances of a second database running on the different computer systems having the different hardware configurations to generate second sets of performance counter data and metadata corresponding to the second database;

generating a plurality of different compressed representations of characteristics of the plurality of different databases operating under the plurality of different workloads using a first machine learning model, wherein generating the plurality of different compressed representations comprises:

repeatedly executing an autoencoder using different combinations of the first sets of information corresponding to the first database as input layer values and output layer values of the autoencoder; and

repeatedly executing the autoencoder using a second plurality of combinations of the second sets of information corresponding to the second database as the input layer values and the output layer values of the autoencoder;

determining at least one recommended characteristic for a cloud database using the plurality of different compressed representations and a second machine learning model; and

outputting the at least one recommended characteristic.

2. The method of claim 1 , wherein:

the first machine learning model comprises a neural network-based autoencoder; and

the second machine learning model comprises a neural network-based classifier.

3. The method of claim 1 , wherein:

the method further comprises training the second machine learning model to map a particular compressed representation that is generated by the first machine learning model to a particular characteristic of a cloud database; and

training the second machine learning model comprises using compressed representations from the first machine learning model as inputs to the second machine learning model, the compressed representations corresponding to databases with one or more known characteristics.

4. The method of claim 1 , further comprising:

determining a plurality of recommended characteristics corresponding to different categories of cloud databases using the compressed representation and a plurality of instances of the second machine learning model; and

outputting the plurality of recommended characteristics.

5. The method of claim 1 , wherein:

the compressed representation is generated for a subset of the information; and

the method further comprises selecting the subset by identifying a defined number of percentile values for time series data that is included in the information.

6. The method of claim 1 , further comprising generating synthetic workloads to increase an amount of data available for training.

7. A system for providing one or more recommendations for migration of a database to a cloud computing system, comprising:

one or more processors; and

memory comprising instructions that are executable by the one or more processors to perform operations comprising:

executing a plurality of different workloads against a plurality of different databases, wherein executing the plurality of different workloads against the plurality of different databases comprises:

executing a first workload against a first plurality of different instances of a first database running on different computer systems having different hardware configurations to generate first sets of information corresponding to the first database; and

executing a second workload against a second plurality of different instances of a second database running on the different computer systems having the different hardware configurations to generate second sets of performance counter data and metadata corresponding to the second database;

generating a plurality of different compressed representations of characteristics of the plurality of different databases operating under the plurality of different workloads using a first machine learning model, wherein generating the plurality of different compressed representations comprises:

repeatedly executing an autoencoder using different combinations of the first sets of information corresponding to the first database as input layer values and output layer values of the autoencoder; and

repeatedly executing the autoencoder using a second plurality of combinations of the second sets of information corresponding to the second database as the input layer values and the output layer values of the autoencoder;

determining at least one recommended characteristic for a cloud database using the plurality of different compressed representations and a second machine learning model; and

outputting the at least one recommended characteristic.

8. The system of claim 7 , wherein:

the first machine learning model comprises a neural network-based autoencoder; and

the second machine learning model comprises a neural network-based classifier.

9. The system of claim 7 , wherein:

the operations further comprise training the second machine learning model to map a particular compressed representation that is generated by the first machine learning model to a particular characteristic of a cloud database; and

training the second machine learning model comprises using compressed representations from the first machine learning model as inputs to the second machine learning model, the compressed representations corresponding to databases with one or more known characteristics.

10. The system of claim 7 , wherein the operations further comprise:

determining a plurality of recommended characteristics corresponding to different categories of cloud databases using the compressed representation and a plurality of instances of the second machine learning model; and

outputting the plurality of recommended characteristics.

11. The system of claim 7 , wherein:

the compressed representation is generated for a subset of the information; and

the operations further comprise selecting the subset by identifying a defined number of percentile values for time series data that is included in the information.

12. The system of claim 7 , wherein the operations further comprise generating synthetic workloads to increase an amount of data available for training.

13. A non-transitory computer-readable medium having computer-executable instructions stored thereon that, when executed, cause one or more processors to perform operations comprising:

executing a plurality of different workloads against a plurality of different databases, wherein executing the plurality of different workloads against the plurality of different databases comprises:

executing a first workload against a first plurality of different instances of a first database running on different computer systems having different hardware configurations to generate first sets of information corresponding to the first database; and

executing a second workload against a second plurality of different instances of a second database running on the different computer systems having the different hardware configurations to generate second sets of performance counter data and metadata corresponding to the second database;

generating, based at least on a subset of the information, a plurality of different compressed representations of characteristics of the plurality of different databases operating under the plurality of different workloads using a first machine learning model, wherein generating the plurality of different compressed representations comprises:

repeatedly executing an autoencoder using different combinations of the first sets of information corresponding to the first database as input layer values and output layer values of the autoencoder; and

repeatedly executing the autoencoder using a second plurality of combinations of the second sets of information corresponding to the second database as the input layer values and the output layer values of the autoencoder;

determining at least one recommended characteristic for a cloud database using the plurality of different compressed representations and a second machine learning model; and

outputting the at least one recommended characteristic.

14. The computer-readable medium of claim 13 , wherein:

the operations further comprise training the second machine learning model to map a particular compressed representation that is generated by the first machine learning model to a particular characteristic of a cloud database; and

training the second machine learning model comprises using compressed representations from the first machine learning model as inputs to the second machine learning model, the compressed representations corresponding to databases with one or more known characteristics.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 27, 2018
From: SPRYN, MITCHELL GREGORY; PARK, INTAIK; VIEIRA FRUJERI, FELIPE; PANJETI, VIJAY GOVIND; MADALA, ASHOK SAI; KARANAM, AJAY KUMAR
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 047595/0631 →
Continuity (2)
Provisional Application 62692505 · Jun 29, 2018
Related Publication 20200005136A1 · Jan 2, 2020
Cited By (2)
US 12,602,260 US 12,748,850