IP Library Granted Patent US 11,789,782
Granted Patent B2
US 11,789,782 · App. 18/160,290 · Granted Oct 17, 2023

Techniques for modifying cluster computing environments

Inventors: Sandeep Akinapelli (Fremont, CA); Devaraj Das (Fremont, CA); Devarajulu Kavali (Santa Clara, CA); Puneet Jaiswal (Milpitas, CA); Velimir Radanovic (Half Moon Bay, CA)
Assignee: Oracle International Corporation
G06F9/505G06F11/3433G06F18/214G06N20/00H04L67/1059
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,789,782
App. No.
18/160,290
Granted
Oct 17, 2023
Kind
B2
Abstract

Systems, devices, and methods discussed herein are directed to intelligently adjusting the set of worker nodes within a computing cluster. By way of example, a computing device (or service) may monitor performance metrics of a set of worker nodes of a computing cluster. When a performance metric is detected that is below a performance threshold, the computing device may perform a first adjustment (e.g., an increase or decrease) to the number of nodes in the cluster. Training data may be obtained based at least in part on the first adjustment and utilized with supervised learning techniques to train a machine-learning model to predict future performance changes in the cluster. Subsequent performance metrics and/or cluster metadata may be provided to the machine-learning model to obtain output indicating a predicted performance change. An additional adjustment to the number of worker nodes may be performed based at least in part on the output.

Claims (49)

1. A computer-implemented method, comprising:

detecting, by a computing device, a performance degradation related to a set of worker nodes of a computing cluster, the performance degradation being detected based at least in part on identifying that a performance metric associated with the computing cluster is below a performance threshold;

in response to detecting the performance degradation, performing, by the computing device, a first adjustment to a number of worker nodes of the set of worker nodes;

generating, by the computing device, a training data example based at least in part on performing the first adjustment, the training data example comprising: 1) first performance metrics and first cluster metadata associated with a first time period that occurs prior to the first adjustment, 2) second performance metrics and second cluster metadata associated with a second time period that follows the first adjustment, and 3) an indication of whether the first adjustment resolved the performance degradation;

obtaining, by the computing device, a machine-learning model that has been trained utilizing the training data example and a supervised machine-learning algorithm, the machine-learning model being trained to identify a particular adjustment to the set of worker nodes from current performance metrics and current cluster metadata provided as input; and

performing, by the computing device, a subsequent adjustment to the set of worker nodes based at least in part on output generated by the machine-learning model, the output being generated by the machine-learning model based at least in part on providing, to the machine-learning model as input, subsequent performance metrics and subsequent cluster metadata of the computing cluster.

2. The computer-implemented method of claim 1 , further comprising:

in response to detecting the performance degradation, performing a second adjustment to the number of worker nodes in the set of worker nodes of the computing cluster, the second adjustment being performed prior to the first adjustment, the second adjustment being similar to the first adjustment; and

determining that performing the second adjustment has caused the performance metric to approach the performance threshold, wherein the first adjustment is performed further based at least in part on determining that the second adjustment has caused the performance metric to approach the performance threshold.

3. The computer-implemented method of claim 1 , further comprising:

in response to detecting the performance degradation, performing a second adjustment to the number of worker nodes in the set of worker nodes of the computing cluster, the second adjustment being performed prior to the first adjustment, the second adjustment dissimilar to the first adjustment; and

determining that performing the second adjustment has caused performance of the computing cluster to worsen, wherein the first adjustment is performed further based at least in part on determining that the second adjustment has caused performance of the computing cluster to worsen.

4. The computer-implemented method of claim 1 , wherein the output indicates how many worker nodes are to be added or removed from the computing cluster.

5. The computer-implemented method of claim 1 , wherein performing the first adjustment comprises provisioning a number of additional worker nodes to the set of worker nodes of the computing cluster.

6. The computer-implemented method of claim 5 , further comprising determining that provisioning the number of additional worker nodes has resulted in a subsequent performance metric that exceeds the performance threshold, wherein the training data example is generated in response to determining the number of additional worker nodes has resulted in the subsequent performance metric.

7. The computer-implemented method of claim 1 , wherein the performance metric is one of a set of performance metrics comprising one or more of: a number of pending queries, a number of pending tasks, a latency measurement, a processing utilization, or a memory utilization.

8. A computing device, comprising:

one or more processing devices communicatively coupled to a computer-readable medium; and

a computer-readable medium storing non-transitory computer-executable program instructions that, when executed by the one or more processing devices, causes the computing device to perform operations comprising:

detecting a performance degradation related to a set of worker nodes of a computing cluster, the performance degradation being detected based at least in part on identifying that a performance metric associated with the computing cluster is below a performance threshold;

in response to detecting the performance degradation, performing a first adjustment to a number of worker nodes of the set of worker nodes;

generating a training data example based at least in part on performing the first adjustment, the training data example comprising: 1) first performance metrics and first cluster metadata associated with a first time period that occurs prior to the first adjustment, 2) second performance metrics and second cluster metadata associated with a second time period that follows the first adjustment, and 3) an indication of whether the first adjustment resolved the performance degradation;

obtaining a machine-learning model that has been trained utilizing the training data example and a supervised machine-learning algorithm, the machine-learning model being trained to identify a particular adjustment to the set of worker nodes from current performance metrics and current cluster metadata provided as input; and

performing a subsequent adjustment to the set of worker nodes based at least in part on output generated by the machine-learning model, the output being generated by the machine-learning model based at least in part on providing, to the machine-learning model as input, subsequent performance metrics and subsequent cluster metadata of the computing cluster.

9. The computing device of claim 8 , wherein executing the computer-executable program instructions causes the computing device to perform additional operations comprising:

in response to detecting the performance degradation, performing a second adjustment to the number of worker nodes in the set of worker nodes of the computing cluster, the second adjustment being performed prior to the first adjustment, the second adjustment being similar to the first adjustment; and

determining that performing the second adjustment has caused the performance metric to approach the performance threshold, wherein the first adjustment is performed further based at least in part on determining that the second adjustment has caused the performance metric to approach the performance threshold.

10. The computing device of claim 8 , wherein executing the computer-executable program instructions causes the computing device to perform additional operations comprising:

in response to detecting the performance degradation, performing a second adjustment to the number of worker nodes in the set of worker nodes of the computing cluster, the second adjustment being performed prior to the first adjustment, the second adjustment dissimilar to the first adjustment; and

determining that performing the second adjustment has caused performance of the computing cluster to worsen, wherein the first adjustment is performed further based at least in part on determining that the second adjustment has caused performance of the computing cluster to worsen.

11. The computing device of claim 8 , wherein the output indicates how many worker nodes are to be added or removed from the computing cluster.

12. The computing device of claim 8 , wherein performing the first adjustment comprises provisioning a number of additional worker nodes to the set of worker nodes of the computing cluster.

13. The computing device of claim 12 , wherein executing the computer-executable program instructions causes the computing device to perform additional operations comprising determining that provisioning the number of additional worker nodes has resulted in a subsequent performance metric that exceeds the performance threshold, and wherein the training data example is generated in response to determining the number of additional worker nodes has resulted in the subsequent performance metric.

14. The computing device of claim 8 , wherein the performance metric is one of a set of performance metrics comprising one or more of: a number of pending queries, a number of pending tasks, a latency measurement, a processing utilization, or a memory utilization.

15. A non-transitory computer-readable medium storing computer-executable program instructions that, when executed by a processing device of a computing device, causes the computing device to perform operations comprising:

detecting a performance degradation related to a set of worker nodes of a computing cluster, the performance degradation being detected based at least in part on identifying that a performance metric associated with the computing cluster is below a performance threshold;

in response to detecting the performance degradation, performing a first adjustment to a number of worker nodes of the set of worker nodes;

generating a training data example based at least in part on performing the first adjustment, the training data example comprising: 1) first performance metrics and first cluster metadata associated with a first time period that occurs prior to the first adjustment, 2) second performance metrics and second cluster metadata associated with a second time period that follows the first adjustment, and 3) an indication of whether the first adjustment resolved the performance degradation;

obtaining a machine-learning model that has been trained utilizing the training data example and a supervised machine-learning algorithm, the machine-learning model being trained to identify a particular adjustment to the set of worker nodes from current performance metrics and current cluster metadata provided as input; and

performing a subsequent adjustment to the set of worker nodes based at least in part on output generated by the machine-learning model, the output being generated by the machine-learning model based at least in part on providing, to the machine-learning model as input, subsequent performance metrics and subsequent cluster metadata of the computing cluster.

16. The non-transitory computer-readable medium of claim 15 , wherein executing the computer-executable program instructions causes the computing device to perform additional operations comprising:

in response to detecting the performance degradation, performing a second adjustment to the number of worker nodes in the set of worker nodes of the computing cluster, the second adjustment being performed prior to the first adjustment, the second adjustment being similar to the first adjustment; and

determining that performing the second adjustment has caused the performance metric to approach the performance threshold, wherein the first adjustment is performed further based at least in part on determining that the second adjustment has caused the performance metric to approach the performance threshold.

17. The non-transitory computer-readable medium of claim 15 , wherein executing the computer-executable program instructions causes the computing device to perform additional operations comprising:

in response to detecting the performance degradation, performing a second adjustment to the number of worker nodes in the set of worker nodes of the computing cluster, the second adjustment being performed prior to the first adjustment, the second adjustment dissimilar to the first adjustment; and

determining that performing the second adjustment has caused performance of the computing cluster to worsen, wherein the first adjustment is performed further based at least in part on determining that the second adjustment has caused performance of the computing cluster to worsen.

18. The non-transitory computer-readable medium of claim 15 , wherein the output indicates how many worker nodes are to be added or removed from the computing cluster.

19. The non-transitory computer-readable medium of claim 15 , wherein performing the first adjustment comprises provisioning a number of additional worker nodes to the set of worker nodes of the computing cluster, wherein executing the computer-executable program instructions causes the computing device to determine that provisioning the number of additional worker nodes has resulted in a subsequent performance metric that exceeds the performance threshold, and wherein the training data example is generated in response to determining the number of additional worker nodes has resulted in the subsequent performance metric.

20. The non-transitory computer-readable medium of claim 15 , wherein the performance metric is one of a set of performance metrics comprising one or more of: a number of pending queries, a number of pending tasks, a latency measurement, a processing utilization, or a memory utilization.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2023
From: AKINAPELLI, SANDEEP; DAS, DEVARAJ; KAVALI, DEVARAJULU; JAISWAL, PUNEET; RADANOVIC, VELIMIR
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 064379/0098 →
Continuity (2)
Continuation 17094715 · Nov 10, 2020
Related Publication 20230222002A1 · Jul 13, 2023