IP Library › Granted Patent US 12,737,375
Granted Patent B2
US 12,737,375 · App. 17/590,358 · Granted Sep 15, 2026

Optimization of virtual warehouse computing resource allocation

Inventors: Syed Shamaz Salim (North Potomac, MD); Ganesh Bharathan (Henrico, VA)
Assignee: Capital One Services, LLC
G06F16/254G06F16/2455G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,375
App. No.
17/590,358
Granted
Sep 15, 2026
Kind
B2
Abstract

Methods, systems, and apparatuses for optimizing the configuration of virtual warehouses for execution of queries on one or more data warehouses are described herein. A plurality of different events associated with a data sharing platform may be logged. The data sharing platform may enable users to access one or more databases managed by the data sharing platform. The data sharing platform may be configured to provide access to the data stored by the data sharing platform via one or more of a plurality of virtual warehouses. A testing database may be generated. An optimized virtual warehouse configuration may be predicted for a first virtual warehouse by selecting a plurality of different warehouse configurations for the first virtual warehouse, measuring performance parameters of each of the plurality of different warehouse configurations by emulating, and selecting the optimized virtual warehouse configuration based on the performance parameters.

Claims (79)

1 . A computing device comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the computing device to:

log a plurality of different queries executed by a data sharing platform, wherein the data sharing platform enables users to access one or more databases managed by the data sharing platform, wherein the data sharing platform is configured to provide access to the data stored by the data sharing platform via one or more of a plurality of virtual warehouses, wherein each of the plurality of virtual warehouses comprises a respective set of computing resources configured to:

execute one or more queries with respect to at least a portion of a plurality of data warehouses;

collect results from the one or more queries; and

provide access to the collected results;

generate a testing database by duplicating at least one of the one or more databases;

predict an optimized virtual warehouse configuration for a first virtual warehouse by:

selecting a plurality of different warehouse configurations for the first virtual warehouse, wherein the plurality of different warehouse configurations each correspond to a different set of computing resources available to the first virtual warehouse;

emulating, via the first virtual warehouse, the plurality of different queries at the testing database by, for each query of the plurality of different queries, for each warehouse configuration of the plurality of different warehouse configurations, and at different times, causing a virtual warehouse configured based on the warehouse configuration to execute the query, collect results based on the query, and provide access to the results based on the query;

monitoring the emulating of the plurality of different queries to measure performance parameters of each of the plurality of different warehouse configurations; and

selecting the optimized virtual warehouse configuration based on the performance parameters; and

output the optimized virtual warehouse configuration.

2 . The computing device of claim 1 , wherein the instructions, when executed by the one or more processors, cause the computing device to emulate the plurality of different queries at the testing database by causing the computing device to:

cause each of the plurality of different queries to be executed by the virtual warehouse configured based on the warehouse configuration at a different time.

3 . The computing device of claim 1 , wherein the plurality of different queries comprise one or more of:

a write action to a database of the one or more databases; or

a read action to the database of the one or more databases.

4 . The computing device of claim 1 , wherein the instructions, when executed by the one or more processors, cause the computing device to log the plurality of different queries during a time period.

5 . The computing device of claim 1 , wherein the instructions, when executed by the one or more processors, cause the computing device to log the plurality of different queries by causing the computing device to log:

at least one first query that satisfies a processing time threshold; and

at least one second query that does not satisfy the processing time threshold.

6 . The computing device of claim 1 , wherein the instructions, when executed by the one or more processors, cause the computing device to predict the optimized virtual warehouse configuration by causing the computing device to:

train, using training data, a machine learning model to output a recommended virtual warehouse configuration, wherein the training data comprises a history of different virtual warehouse configurations and a history of different query processing times;

provide, to the trained machine learning model, input comprising the performance parameters; and

receive, from the trained machine learning model, output indicating the optimized virtual warehouse configuration.

7 . The computing device of claim 1 , wherein the performance parameters indicate a processing time corresponding to each of the plurality of different queries.

8 . A method comprising:

logging, by a computing device, a plurality of different queries executed by a data sharing platform, wherein the data sharing platform enables users to access one or more databases managed by the data sharing platform, wherein the data sharing platform is configured to provide access to the data stored by the data sharing platform via one or more of a plurality of virtual warehouses, wherein each of the plurality of virtual warehouses comprises a respective set of computing resources configured to:

execute one or more queries with respect to at least a portion of a plurality of data warehouses;

collect results from the one or more queries; and

provide access to the collected results;

generating a testing database by duplicating at least one of the one or more databases;

predicting an optimized virtual warehouse configuration for a first virtual warehouse by:

selecting a plurality of different warehouse configurations for the first virtual warehouse, wherein the plurality of different warehouse configurations each correspond to a different set of computing resources available to the first virtual warehouse;

emulating, via the first virtual warehouse, the plurality of different queries at the testing database by, for each query of the plurality of different queries, for each warehouse configuration of the plurality of different warehouse configurations, and at different times, causing a virtual warehouse configured based on the warehouse configuration to execute the query, collect results based on the query, and provide access to the results based on the query;

monitoring the emulating of the plurality of different queries to measure performance parameters of each of the plurality of different warehouse configurations; and

selecting the optimized virtual warehouse configuration based on the performance parameters; and

outputting the optimized virtual warehouse configuration.

9 . The method of claim 8 , wherein emulating the plurality of different queries at the testing database comprises:

causing each of the plurality of different queries to be executed by the virtual warehouse configured based on the warehouse configuration at a different time.

10 . The method of claim 8 , wherein the plurality of different queries comprise one or more of:

a write action to a database of the one or more databases; or

a read action to the database of the one or more databases.

11 . The method of claim 8 , wherein logging the plurality of different queries comprises logging the plurality of different events during a time period.

12 . The method of claim 8 , wherein logging the plurality of different queries comprises logging:

at least one first query that satisfies a processing time threshold; and

at least one second query that does not satisfy the processing time threshold.

13 . The method of claim 8 , wherein predicting the optimized virtual warehouse configuration comprises:

training, using training data, a machine learning model to output a recommended virtual warehouse configuration, wherein the training data comprises a history of different virtual warehouse configurations and a history of different query processing times;

providing, to the trained machine learning model, input comprising the performance parameters; and

receiving, from the trained machine learning model, output indicating the optimized virtual warehouse configuration.

14 . The method of claim 8 , wherein the performance parameters indicate a processing time corresponding to each of the plurality of different queries.

15 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a computing device, cause the computing device to:

log a plurality of different queries executed by a data sharing platform, wherein the data sharing platform enables users to access one or more databases managed by the data sharing platform, wherein the data sharing platform is configured to provide access to the data stored by the data sharing platform via one or more of a plurality of virtual warehouses, wherein each of the plurality of virtual warehouses comprises a respective set of computing resources configured to:

execute one or more queries with respect to at least a portion of a plurality of data warehouses;

collect results from the one or more queries; and

provide access to the collected results;

generate a testing database by duplicating at least one of the one or more databases;

predict an optimized virtual warehouse configuration for a first virtual warehouse by:

selecting a plurality of different warehouse configurations for the first virtual warehouse, wherein the plurality of different warehouse configurations each correspond to a different set of computing resources available to the first virtual warehouse;

emulating, via the first virtual warehouse, the plurality of different queries at the testing database by, for each query of the plurality of different queries, for each warehouse configuration of the plurality of different warehouse configurations, and at different times, causing a virtual warehouse configured based on the warehouse configuration to execute the query, collect results based on the query, and provide access to the results based on the query;

monitoring the emulating of the plurality of different queries to measure performance parameters of each of the plurality of different warehouse configurations; and

selecting the optimized virtual warehouse configuration based on the performance parameters; and

output the optimized virtual warehouse configuration.

16 . The non-transitory computer-readable media of claim 15 , wherein the instructions, when executed by the one or more processors, cause the computing device to emulate the plurality of different queries at the testing database by causing the computing device to:

cause each of the plurality of different queries to be executed by the virtual warehouse configured based on the warehouse configuration at a different time.

17 . The non-transitory computer-readable media of claim 15 , wherein the plurality of different queries comprise one or more of:

a write action to a database of the one or more databases; or

a read action to the database of the one or more databases.

18 . The non-transitory computer-readable media of claim 15 , wherein the instructions, when executed by the one or more processors, cause the computing device to log the plurality of different queries during a time period.

19 . The non-transitory computer-readable media of claim 15 , wherein the instructions, when executed by the one or more processors, cause the computing device to log the plurality of different queries by causing the computing device to log:

at least one first query that satisfies a processing time threshold; and

at least one second query that does not satisfy the processing time threshold.

20 . The non-transitory computer-readable media of claim 15 , wherein the instructions, when executed by the one or more processors, cause the computing device to predict the optimized virtual warehouse configuration by causing the computing device to:

train, using training data, a machine learning model to output a recommended virtual warehouse configuration, wherein the training data comprises a history of different virtual warehouse configurations and a history of different query processing times;

provide, to the trained machine learning model, input comprising the performance parameters; and

receive, from the trained machine learning model, output indicating the optimized virtual warehouse configuration.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2022
From: SALIM, SYED SHAMAZ; BHARATHAN, GANESH
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 058848/0775 →
Continuity (1)
Related Publication 20230244687A1 · Aug 3, 2023
References Cited (28)
US 10970303B1 · Denton · 2021 [cited by examiner]
US 11537575B1 · McNair et al. · 2022 [cited by applicant]
US 20100257513A1 · Thirumalai et al. · 2010 [cited by applicant]
US 20120131591A1 · Moorthi · 2012 [cited by examiner]
US 20130117305A1 · Varakin · 2013 [cited by examiner]
US 20140006384A1 · Jerzak · 2014 [cited by examiner]
US 20140089495A1 · Akolkar et al. · 2014 [cited by applicant]
US 20170316078A1 · Funke et al. · 2017 [cited by applicant]
US 20190163795A1 · Lai · 2019 [cited by examiner]
US 20190280918A1 · Hermoni · 2019 [cited by examiner]
US 20210089560A1 · Funke et al. · 2021 [cited by applicant]
US 20210097076A1 · Rajaperumal · 2021 [cited by examiner]
US 20210334283A1 · Gladwin et al. · 2021 [cited by applicant]
US 20210342360A1 · Cseri · 2021 [cited by examiner]
WO 2020228378A1 · 2020 [cited by applicant]
May 9, 2023 (WO) International Search Report and Written Opinion—App PCT/US2023/011421, 20 pages. [cited by applicant]
Yan, et al., “Snowtrail: Testing with Production Queries on a Cloud Database,” Proceedings of the Workshop on Testing Database Systems, Jun. 15, 2018, pp. 1-6. [cited by applicant]
Aug. 1, 20216, VentureBeat, Lazzaro, “How to migrate to Snowflake without getting ‘data drunk’,” <<https://venturebeat.com/2021/08/16/how-to-migrate-to-snowflake-without-getting-data-drunk/>>, 8 pages. [cited by applicant]
Nadilytics—The Intelligence Platform on top of Snowflake. Retreived from https://www.nadilytics.com/, retrieved online May 14, 2021, at 4:25:15 pm. [cited by applicant]
Snowflake for Data Sharing, Snowflake Workloads, <https://www.snowflake.com/workloads/data-sharing/>>, date of publication unknown but, prior to Jul. 13, 2021, 3 pages, published Aug. 2020. [cited by applicant]
Data Governance is Worth the Time and Trouble, <<https://www.snowflake.com/trending/data-governance-framework>>, date of publication unknown but, prior to Jul. 13, 2021, 14 pages. [cited by applicant]
Working with Secuve Views, Snowflake Documention, <<https://docs.snowflake.com/en/user-guide/views-secure.html>>, date of publication unknown but, prior to Jul. 13, 2021, 3 pages, published Aug. 2020. [cited by applicant]
Data governance with Snowflake: 3 things you need to know, <<https://www.talend.com/resources/data-governance-snowflake-3-things-to-know/, date of publication unknown but, prior to Jul. 13, 2021, 11 pages. [cited by applicant]
GPDR: a quick way to reduce scope, <<https://www.talend.com/resources/anonymize-data/>>, date of publication unknown but, prior to Jul. 13, 2021, 6 pages. [cited by applicant]
Jun. 14, 2019, Mushtaq, Data preprocessing in detail, <<https://developer.ibm.com/technologies/data-science/articles/data-preprocessing-in-detail/>>, 9 pages. [cited by applicant]
Wikipedia, Data pre-processing, <<https://en.wikipedia.org/wiki/Data_pre-processing>>, date of publication but, prior to Jul. 13, 2021, 4 pages. [cited by applicant]
Snowflake Data Exchange, <<https://resources.snowflake.com/solution-briefs/data-exchange-solution-brief>>, date of publication unknown but, prior to Jul. 13, 2021, 2 pages. [cited by applicant]
Zhao, et al., “SLA-based Profit Optimization Resource Scheduling for Big Data Analytics-as-a-Service Platforms in Cloud Computing Environments,” IEEE Transactions on Cloud Computing, Dec. 27, 2018. [cited by applicant]