IP Library › Granted Patent US 12,332,761
Granted Patent B1
US 12,332,761 · App. 18/582,114 · Granted Jun 17, 2025

Metric management in multi-computing cluster environment

Inventors: Seep Goel (Bengaluru, IN); Kavya Govindarajan (Chennai, IN); Chander Govindarajan (Chennai, IN); Priyanka Prakash Naik (Mumbai, IN); Praveen Jayachandran (Bangalore, IN); Aishwariya Chakraborty (Bankura, IN)
Assignee: International Business Machines Corporation
G06F11/3409H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,332,761
App. No.
18/582,114
Filed
Feb 20, 2024
Granted
Jun 17, 2025
Kind
B1
Art Unit
2441
USPC
709/224
Abstract

Metric management techniques in a multi-cluster computing environment are disclosed. For example, a method obtains a set of metrics collected from one or more computing clusters of a distributed computing environment, wherein the set of metrics are associated with one or more processes that are executable on the one or more computing clusters. The method computes one or more metric importance values for the set of metrics based on one or more processing criteria. The method sends the one or more metric importance values to at least one computing cluster of the one or more computing clusters to enable the at least one computing cluster to adapt a collection frequency of at least one metric of the set of metrics. Such adaptation, by way of example, can be based on the metric importance and resource availability.

Claims (35)

1. A computer-implemented method, comprising:

obtaining a set of metrics collected from one or more computing clusters of a distributed computing environment, wherein the set of metrics are associated with one or more processes that are executable on the one or more computing clusters;

computing one or more metric importance values for the set of metrics based on one or more processing criteria; and

sending the one or more metric importance values to at least one computing cluster of the one or more computing clusters to enable the at least one computing cluster to adapt a collection frequency of at least one metric of the set of metrics;

wherein the collection frequency of the at least one metric of the set of metrics accounts for a total frequency of obtaining a set of important metrics collected from the one or more computing clusters of the distributed computing environment, and is proportional to the one or more metric importance values for the set of metrics; and

wherein the computer-implemented method is performed by a processing platform when executing program code, the processing platform comprising one or more processing devices, each of the one or more processing devices comprising a processor coupled to a memory.

2. The computer-implemented method of claim 1 , wherein the one or more processes comprise at least one of one or more applications, one or more microservices, and one or more network slices.

3. The computer-implemented method of claim 1 , further comprising computing one or more metric importance values for the set of metrics based on a violation of the one or more processing criteria.

4. The computer-implemented method of claim 3 , wherein the computing is performed with application-specific logic.

5. The computer-implemented method of claim 1 , wherein computing the one or more metric importance values for the set of metrics is further based on correlating the set of metrics to a corresponding set of key performance indicators of the one or more processes following an analysis of the one or more processing criteria.

6. The computer-implemented method of claim 3 , wherein the computing is performed by a machine learning algorithm.

7. The computer-implemented method of claim 6 , wherein the machine learning algorithm is a supervised or a semi-supervised algorithm and wherein the machine learning algorithm is trained on fault detection data.

8. The computer-implemented method of claim 1 , wherein computing the one or more metric importance values for the set of metrics is further based on a set of respective weights assigned to the set of metrics.

9. The computer-implemented method of claim 1 , wherein computing the one or more metric importance values for the set of metrics is further based on a binary value assigned to one or more metrics of the set of metrics.

10. The computer-implemented method of claim 1 , wherein computing the one or more metric importance values for the set of metrics is further based on a set of weights assigned to the set of metrics based on domain knowledge.

11. The computer-implemented method of claim 1 , further comprising determining a resource allocation for a computing cluster of the one or more computing clusters based on at least one of a current resource availability of the computing cluster and one or more monitoring limits of the computing cluster; and

wherein the current resource availability of the computing cluster is determined by monitoring real-time cluster resource utilization and a number of the one or more processes that are executable on the one or more computing clusters.

12. The computer-implemented method of claim 11 , wherein the collection frequency of at least one metric of the set of metrics is based on at least one of the resource allocation for the computing cluster and the one or more metric importance values for the set of metrics, and wherein the collection frequency of at least one metric of the set of metrics is determined by a machine learning model.

13. An apparatus comprising:

a processing platform comprising a processor set, a set of one or more computer-readable storage media, and program instructions, collectively stored in the set of one or more storage media, for causing the processor to perform computer operations comprising:

obtaining a set of metrics collected from one or more computing clusters of a distributed computing environment, wherein the set of metrics are associated with one or more processes that are executable on the one or more computing clusters;

computing one or more metric importance values for the set of metrics based on one or more processing criteria; and

sending the one or more metric importance values to at least one computing cluster of the one or more computing clusters to enable the at least one computing cluster to adapt a collection frequency of at least one metric of the set of metrics;

wherein the collection frequency of the at least one metric of the set of metrics accounts for a total frequency of obtaining a set of important metrics collected from the one or more computing clusters of the distributed computing environment, and is proportional to the one or more metric importance values for the set of metrics.

14. The apparatus of claim 13 , wherein the collection frequency of at least one metric of the set of metrics is based on at least one of a resource allocation for a given computing cluster and the one or more metric importance values for the set of metrics.

15. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, wherein the program instructions are executable by a processing platform comprising one or more processing devices, each of the one or more processing device comprising a processor coupled to a memory, to cause the processors of the one or more processing devices to:

obtain one or more metric importance values for a set of metrics, wherein the set of metrics is associated with one or more processes that are executable on one or more computing clusters of a distributed computing environment; and

adapt a collection frequency of at least one metric of the set of metrics;

wherein the adapted collection frequency of the at least one metric of the set of metrics is based on at least the one or more metric importance values for the set of metrics; and

wherein the adapted collection frequency of the at least one metric of the set of metrics accounts for a total frequency of obtaining a set of important metrics collected from the one or more computing clusters of the distributed computing environment, and is proportional to the one or more metric importance values for the set of metrics.

16. The computer program product of claim 15 , wherein the program instructions executable by the processing platform further cause the processors of the one or more processing devices to determine a resource allocation for a given computing cluster based on at least one of a current resource availability of the given computing cluster and one or more monitoring limits of the given computing cluster.

17. The apparatus of claim 13 , wherein the one or more processes comprise at least one of one or more applications, one or more microservices, and one or more network slices.

18. The apparatus of claim 13 , wherein the collection frequency of at least one metric of the set of metrics is determined by a machine learning model.

19. The computer program product of claim 15 , wherein the one or more processes comprise at least one of one or more applications, one or more microservices, and one or more network slices.

20. The computer program product of claim 16 , wherein the program instructions executable by the processing platform further cause the processors of the one or more processing devices to determine the current resource availability of the computing cluster by monitoring real-time cluster resource utilization and a number of the one or more processes that are executable on the one or more computing clusters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2024
From: GOEL, SEEP; GOVINDARAJAN, KAVYA; GOVINDARAJAN, CHANDER; NAIK, PRIYANKA PRAKASH; JAYACHANDRAN, PRAVEEN; CHAKRABORTY, AISHWARIYA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 066509/0214 →
References Cited (14)
US 9755925B2 · Poola · 2017 [cited by examiner]
US 10467036B2 · Anwar et al. · 2019 [cited by applicant]
US 10678602B2 · Manglik et al. · 2020 [cited by applicant]
US 11671312B2 · Puri · 2023 [cited by examiner]
US 20160080226A1 · Poola · 2016 [cited by examiner]
US 20160105335A1 · Choudhary · 2016 [cited by examiner]
US 20170351715A1 · Cudak et al. · 2017 [cited by applicant]
US 20180241649A1 · Mazzitelli · 2018 [cited by examiner]
US 20190196929A1 · Megahed et al. · 2019 [cited by applicant]
M. A. Aleisa et al., “Examining the Performance of Fog-Aided, Cloud-Centered IoT in a Real-World Environment,” Sensors, Oct. 20, 2021, 32 pages, vol. 21, No. 6950. [cited by applicant]
W. Wang et al., “Closed-Loop Network Performance Monitoring and Diagnosis with SpiderMon,” 19th USENIX Symposium on Networked Systems Design and Implementation, Apr. 2022, pp. 267-285. [cited by applicant]
T. Yang et al., “Elastic Sketch: Adaptive and Fast Network-wide Measurements,” Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication, Aug. 2018, pp. 561-575. [cited by applicant]
N. Yaseen et al., “Towards a Cost vs. Quality Sweet Spot for Monitoring Networks,” arXiv:2110.05554v1, Oct. 11, 2021, 9 pages. [cited by applicant]
W. Zhao et al., “Scheduling Sensor Data Collection with Dynamic Traffic Patterns,” IEEE Transactions on Parallel and Distributed Systems, May 29, 2012, 14 pages. [cited by applicant]