IP Library › Granted Patent US 11,416,364
Granted Patent B2
US 11,416,364 · App. 17/119,462 · Granted Aug 16, 2022

Methods and systems that identify dimensions related to anomalies in system components of distributed computer systems using clustered traces, metrics, and component-associated attribute values

Inventors: Naira Movses Grigoryan (Yerevan, AM); Arnak Poghosyan (Yerevan, AM); Ashot Nshan Harutyunyan (Yerevan, AM); Clement Pang (Palo Alto, CA); Dev Nag (Palo Alto, CA)
Assignee: VMware, Inc.
G06F11/3006G06F11/3075G06F11/323G06F11/3476
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,416,364
App. No.
17/119,462
Granted
Aug 16, 2022
Kind
B2
Abstract

The current document is directed to methods and systems that employ distributed-computer-system metrics collected by one or more distributed-computer-system metrics-collection services, call traces collected by one or more call-trace services, and attribute values for distributed-computer-system components to identify attribute dimensions related to anomalous behavior of distributed-computer-system components. In a described implementation, nodes correspond to particular types of system components and node instances are individual components of the component type corresponding to a node. Node instances are associated with attribute values and node are associated with attribute-value spaces defined by attribute dimensions. A set of call traces is partitioned, by clustering. Using attribute values and call traces, attribute dimensions that are likely related to particular anomalous behaviors of distributed-computer-system components are determined by decision-tree-related analyses for each partition and are reported to one or more computational entities to facilitate resolution of the anomalous behaviors.

Claims (79)

1. A system that determines attribute dimensions correlated with anomalous operational behaviors of components of a distributed computer system, the system comprising:

one or more processors;

one or more memories; and

computer instructions, stored in one or more of the one or more memories that, when executed by one or more of the one or more processors, control the system to

collect metric data comprising a series of timestamped metric values associated with each metric of multiple metrics, wherein each metric of the multiple metrics is associated with a component or component type of the distributed computer system,

identify components of the distributed computer system which exhibit anomalous operational behaviors using the collected metric data,

access collected call traces from a call-tracing service,

access attribute values for one or more components of the distributed computer system,

cluster the collected call traces into multiple subsets of related call traces, wherein each subset of related call traces of the multiple subsets of related call traces corresponds to a different trace type,

apply decision-tree-based analysis to each subset of related call traces of the multiple subsets of related call traces to determine attribute dimensions of component types that are correlated with the identified components of the distributed computer system which exhibit anomalous operational behaviors, and

transmit the determined attribute dimensions of the component types to a computational entity to facilitate amelioration of the anomalous operational behaviors.

2. The system of claim 1 wherein the one or more components of the distributed computer system are selected from among:

a distributed service-oriented application;

service nodes of the distributed service-oriented application;

service instances of the service nodes of the distributed service-oriented application;

servers;

mass-storage devices and appliances; and

networking components.

3. The system of claim 1 wherein the collected call traces each encodes a series of component types related to execution of a requested task or service.

4. The system of claim 3 wherein the collected call traces each encodes a series of service calls to service nodes within a distributed service-oriented application related to a service call made by a remote client to the distributed service-oriented application.

5. The system of claim 1 wherein the attribute values for the one or more components of the distributed computer system are points within an attribute-value space, for which attributes are dimensions, that is associated with component types of the one or more components of the distributed computer system.

6. The system of claim 5 wherein the attribute values for the one or more components of the distributed computer system are collected from one or more of an attribute-value store and call traces that include component types of the identified components of the distributed computer system which exhibit anomalous operational behaviors.

7. The system of claim 6 wherein the decision-tree-based analysis applied to the subset of related call traces of the multiple subsets of related call traces determine the attribute dimensions of the component types in which the attribute values for the one or more components of the distributed computer system are localized, rather than distributed across the attribute dimensions of the component types.

8. The system of claim 7 wherein the decision-tree-based analysis determine attributes and attribute values that partition the collected call traces into a subset that contains call traces that include components of the distributed computer system, and only call traces that include components of the distributed computer system which exhibit anomalous operational behaviors, and one or more additional subsets.

9. The system of claim 1 wherein the collected call traces are clustered into the multiple subsets of related call traces by:

vectorizing the collected call traces to generate an initial set of call-trace vectors;

clustering the initial set of call-trace vectors;

choosing a provisional set of clusters; and

verifying the provisional set of clusters.

10. The system of claim 9 wherein a call trace is vectorized by:

identifying unique service calls in the call trace;

sorting the identified unique service calls to produce an ordered set of call traces;

for each identified unique service call in the ordered set of call traces,

collecting attribute values for service-call instances invoked during execution of a service entrypoint represented by the call trace; and

mapping the ordered set of call traces and the collected attribute values for the service call instances to a call-trace vector.

11. The system of claim 10

wherein the call-trace vector is a bit vector; and

a unique bit in the call-trace vector corresponds to each different collected attribute-value/service-call pair.

12. The system of claim 10

wherein the call-trace vector is a bit vector; and

a unique bit in the call-trace vector corresponds to each different attribute-value-combination/service-call pair.

13. The system of claim 9 wherein the initial set of call-trace vectors is clustered by:

initially assigning each call-trace vector to a unique single-vector cluster; and

iteratively merging a closest pair of clusters into a new cluster, where a distance between pairs of clusters is determined using a cluster-distance metric.

14. The system of claim 9 wherein the provisional set of clusters is chosen by:

selecting a cut-off clustering distance at a clustering distance greater than a clustering distance of a prominent knee of a cluster-distance-versus-clustering-sequence graph; and

selecting, as the provisional set of clusters, clusters formed from pairs of clusters closer than the cut-off clustering distance that were subsequently merged from pairs of clusters further from one another than the cut-off clustering distance.

15. The system of claim 9 wherein the provisional set of clusters is verified by:

calculating a cophenetic correlation coefficient for the clustering of the initial set of call-trace vectors and determining that the cophenetic correlation coefficient for the clustering of the initial set of call-trace vectors is greater than a first threshold value;

determining that a ratio of an average sparsity of the initial set of call-trace vectors produced by re-vectorizing the collected call traces in each cluster of the provisional set of clusters to a sparsity of the initial set of call-trace vectors is less than a second threshold value;

determining a number of call-trace vectors in each cluster of the provisional set of clusters;

determining a percentage of relevant call-trace vectors specified to be relevant in each cluster of the provisional set of clusters; and

when the number of call-trace vectors in any cluster of the provisional set of clusters is less than a third threshold value or the percentage of relevant call-trace vectors in any cluster of the provisional set of clusters is less than a fourth threshold value or greater than a fifth threshold value, determining that the provisional set of clusters can be adjusted to produce an adjusted set of clusters that does not include any clusters with a percentage of relevant call-trace vectors less than the fourth threshold value or greater than the fifth threshold value and that does not include any clusters with a number of call-trace vectors less than the third threshold value.

16. A method that determines attribute dimensions correlated with anomalous operational behaviors of components of a distributed computer system, the method comprising:

collecting metric data comprising a series of timestamped metric values associated with each metric of multiple metrics, wherein each metric of the multiple metrics is associated with a component or component type of the distributed computer system;

identifying components of the distributed computer system which exhibit anomalous operational behaviors using the collected metric data;

accessing collected call traces from a call-tracing service;

accessing attribute values for one or more components of the distributed computer system;

clustering the collected call traces into multiple subsets of related call traces, wherein each subset of related call traces of the multiple subsets of related call traces corresponds to a different trace type,

applying decision-tree-based analysis to each subset of related call traces of the multiple subsets of related call traces to determine attribute dimensions of component types that are correlated with the identified components of the distributed computer system which exhibit anomalous operational behaviors; and

transmitting the determined attribute dimensions of the component types to a computational entity to facilitate amelioration of the anomalous operational behaviors.

17. The method of claim 16 wherein the collected call traces are clustered into the multiple subsets of related call traces by:

vectorizing the collected call traces to generate an initial set of call-trace vectors;

clustering the initial set of call-trace vectors;

choosing a provisional set of clusters; and

verifying the provisional set of clusters.

18. A physical data-storage device that stores computer instructions that, when executed by one or more processors of a system that includes one or more memories and one or more mass-storage devices, controls the system to determine attribute dimensions correlated with anomalous operational behaviors of components of a distributed computer system by:

collecting metric data comprising a series of timestamped metric values associated with each metric of multiple metrics, wherein each metric of the multiple metrics is associated with a component or component type of the distributed computer system;

identifying components of the distributed computer system which exhibit anomalous operational behaviors using the collected metric data;

accessing collected call traces from a call-tracing service;

accessing attribute values for one or more components of the distributed computer system;

clustering the collected call traces into multiple subsets of related call traces, wherein each subset of related call traces of the multiple subsets of related call traces corresponds to a different trace type;

applying decision-tree-based analysis to each subset of related call traces of the multiple subsets of related call traces to determine attribute dimensions of component types that are correlated with the identified components of the distributed computer system which exhibit anomalous operational behaviors; and

transmitting the determined attribute dimensions of the component types to a computational entity to facilitate amelioration of the anomalous operational behaviors.

19. The physical data-storage device of claim 18 wherein the collected call traces are clustered into the multiple subsets of related call traces by:

vectorizing the collected call traces to generate an initial set of call-trace vectors;

clustering the initial set of call-trace vectors;

choosing a provisional set of clusters; and

verifying the provisional set of clusters.

Assignments (2)
CHANGE OF NAME Recorded Apr 15, 2024
From: VMWARE, INC.
To: VMWARE LLC
Reel/Frame 067102/0395 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2020
From: GRIGORYAN, NAIRA; POGHOSYAN, ARNAK; HARUTYUNYAN, ASHOT; PANG, CLEMENT; NAG, DEV
To: VMWARE, INC.
Reel/Frame 054621/0054 →
Continuity (2)
Continuation In Part 16833102 · Mar 27, 2020
Related Publication 20210303431A1 · Sep 30, 2021