IP Library Patent Application 13646978
Patent Application
App. No. 13/646,978

METHOD AND APPARATUS FOR ANALYZING A ROOT CAUSE OF A SERVICE IMPACT IN A VIRTUALIZED ENVIRONMENT

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
13/646,978
Abstract

A dependency graph includes nodes representing states of infrastructure elements in a managed system, and impacts and events among the infrastructure elements in a managed system that are related to delivery of a service by the managed system. Events are received that cause change among the states in the dependency graph. An event occurs in relation to one of the infrastructure elements of the dependency graph. Each individual node that was affected by the event is analyzed and ranked based on (i) states of the nodes which impact the individual node, and (ii) the states of the nodes which are impacted by the individual node, to provide a score for event(s) which is associated with the individual node. Plural events are ranked based on the scores. The root cause of the events with respect to the service is provided based on the events which were ranked.

Claims (80)

1 . A computer-implemented system that determines a root cause of a service impact, comprising:

a dependency graph data storage configured to store a dependency graph that includes nodes which represent states of infrastructure elements in a managed system, and impacts and events among the infrastructure elements in a managed system that are related to delivery of a service by the managed system; and

a processor that is configured to

receive events that can cause change among the states in the dependency graph, wherein an event occurs in relation to one of the infrastructure elements in a managed system;

for each of the events, execute an analyzer that analyzes and ranks each individual node in the dependency graph that was affected by the event based on (i) states of the nodes which impact the individual node, and (ii) the states of the nodes which are impacted by the individual node, to provide a score for each of at least one event which is associated with the individual node;

rank all of the events based on the scores; and

provide the rank as indicating a root cause of the events with respect to the service.

2 . The computer-implemented system of claim 1 , wherein the dependency graph represents relationships among all infrastructure elements in the managed system that are related to delivery of the service by the managed system, and how the infrastructure elements interact with each other in a delivery of said service, and a state of an infrastructure element is impacted only by states among its immediately dependent infrastructure elements of the dependency tree; and

the processor is configured to determines the state of the service by checking current states of infrastructure elements in the dependency tree that immediately depend from the service.

3 . The computer-implemented system of claim 1 , wherein the individual node in the dependency graph is ranked consistent with the formula (ra/n+1)+w, to provide the score for each of the at least one event which is associated with the individual node, wherein:

r=an integer value of the state caused by the at least one event;

a=an average of the integer values of the states of nodes impacted, directly or indirectly, by the node affected by the at least one event;

n=number of nodes with states affected by other events impacting the node affected by the at least one event; and

w=an optional adjustment that can be provided to influence the score for the at least one event.

4 . The computer-implemented system of claim 1 , wherein

states indicated for the infrastructure element include availability states of at least: up, down, at risk, and degraded,

“up” indicates a normally functional state, “down” indicates a non-functional state, “at risk” indicates a state at risk for being “down”, and “degraded” indicates a state which is available and not fully functional.

5 . The computer-implemented system of claim 1 , wherein

states indicated for the infrastructure element include performance states of at least up, degraded, and down,

“up” indicates a normally functional state, “down” indicates a non-functional state, and “degraded” indicates a state which is available and not fully functional.

6 . The computer-implemented system of claim 1 , wherein the infrastructure elements include:

the service;

a physical element that generates an event caused by a pre-defined physical change in the physical element;

a logical element that generates an event when it has a pre-defined characteristic as measured through a synthetic transaction;

a virtual element that generates an event when a predefined condition occurs; and

a reference element that is a pre-defined collection of other different elements among the same dependency tree, for which a single policy is defined for handling an event that occurs within the reference element.

7 . The computer-implemented system of claim 1 , wherein

the processor determines the state of the infrastructure element according an absolute calculation specified in a policy assigned to the infrastructure element.

8 . A computer-implemented method that determines a root cause of a service impact, comprising:

storing, in a dependency graph data storage that stores a dependency graph that includes nodes which represent states of infrastructure elements in a managed system, and impacts and events among the infrastructure elements in a managed system that are related to delivery of a service by the managed system;

receiving, in a processor, events that can cause change among the states in the dependency graph, wherein an event occurs in relation to one of the infrastructure elements in a managed system;

for each of the events, executing, in the processor, an analyzer that analyzes and ranks each individual node in the dependency graph that was affected by the event based on (i) states of the nodes which impact the individual node, and (ii) the states of the nodes which are impacted by the individual node, to provide a score for each of at least one event which is associated with the individual node;

ranking, in the processor, all of the events based on the scores; and

providing, in the processor, the rank as indicating a root cause of the events with respect to the service.

9 . The method of claim 8 , wherein the dependency graph represents relationships among all infrastructure elements in the managed system that are related to delivery of the service by the managed system, and how the infrastructure elements interact with each other in a delivery of said service, and a state of an infrastructure element is impacted only by states among its immediately dependent infrastructure elements of the dependency tree; and further comprising

determining, in the processor, the state of the service by checking current states of infrastructure elements in the dependency tree that immediately depend from the service.

10 . The method of claim 8 , wherein the individual node in the dependency graph is ranked consistent with the formula (ra/n+1)+w, to provide the score for each of the at least one event which is associated with the individual node, wherein:

r=an integer value of the state caused by the at least one event;

a=an average of the integer values of the states of nodes impacted, directly or indirectly, by the node affected by the at least one event;

n=number of nodes with states affected by other events impacting the node affected by the at least one event; and

w=an optional adjustment that can be provided to influence the score for the at least one event.

11 . The method of claim 8 , wherein

states indicated for the infrastructure element include availability states of at least: up, down, at risk, and degraded,

“up” indicates a normally functional state, “down” indicates a non-functional state, “at risk” indicates a state at risk for being “down”, and “degraded” indicates a state which is available and not fully functional.

12 . The method of claim 8 , wherein

states indicated for the infrastructure element include performance states of at least up, degraded, and down,

“up” indicates a normally functional state, “down” indicates a non-functional state, and “degraded” indicates a state which is available and not fully functional.

13 . The method of claim 8 , wherein the infrastructure elements include:

the service;

a physical element that generates an event caused by a pre-defined physical change in the physical element;

a logical element that generates an event when it has a pre-defined characteristic as measured through a synthetic transaction;

a virtual element that generates an event when a predefined condition occurs; and

a reference element that is a pre-defined collection of other different elements among the same dependency tree, for which a single policy is defined for handling an event that occurs within the reference element.

14 . The method of claim 8 , further comprising determining, in the processor, the state of the infrastructure element according an absolute calculation specified in a policy assigned to the infrastructure element.

15 . A non-transitory computer-readable medium comprising instructions being executed by a computer, the instructions including a computer-implemented method that determines a root cause of a service impact, the instructions implement:

storing, in a dependency graph data storage that stores a dependency graph that includes nodes which represent states of infrastructure elements in a managed system, and impacts and events among the infrastructure elements in a managed system that are related to delivery of a service by the managed system;

receiving events that can cause change among the states in the dependency graph, wherein an event occurs in relation to one of the infrastructure elements in a managed system;

for each of the events, executing an analyzer that analyzes and ranks each individual node in the dependency graph that was affected by the event based on (i) states of the nodes which impact the individual node, and (ii) the states of the nodes which are impacted by the individual node, to provide a score for each of at least one event which is associated with the individual node;

ranking all of the events based on the scores; and

providing the rank as indicating a root cause of the events with respect to the service.

16 . The non-transitory computer-readable medium of claim 15 , wherein the dependency graph represents relationships among all infrastructure elements in the managed system that are related to delivery of the service by the managed system, and how the infrastructure elements interact with each other in a delivery of said service, and a state of an infrastructure element is impacted only by states among its immediately dependent infrastructure elements of the dependency tree; and further comprising

determining the state of the service by checking current states of infrastructure elements in the dependency tree that immediately depend from the service.

17 . The non-transitory computer-readable medium of claim 15 , wherein the individual node in the dependency graph is ranked consistent with the formula (ra/n+1)+w, to provide the score for each of the at least one event which is associated with the individual node, wherein:

r=an integer value of the state caused by the at least one event;

a=an average of the integer values of the states of nodes impacted, directly or indirectly, by the node affected by the at least one event;

n=number of nodes with states affected by other events impacting the node affected by the at least one event; and

w=an optional adjustment that can be provided to influence the score for the at least one event.

18 . The non-transitory computer-readable medium of claim 15 , wherein

states indicated for the infrastructure element include availability states of at least: up, down, at risk, and degraded,

“up” indicates a normally functional state, “down” indicates a non-functional state, “at risk” indicates a state at risk for being “down”, and “degraded” indicates a state which is available and not fully functional.

19 . The non-transitory computer-readable medium of claim 15 , wherein

states indicated for the infrastructure element include performance states of at least up, degraded, and down,

“up” indicates a normally functional state, “down” indicates a non-functional state, and “degraded” indicates a state which is available and not fully functional.

20 . The non-transitory computer-readable medium of claim 15 , wherein the infrastructure elements include:

the service;

a physical element that generates an event caused by a pre-defined physical change in the physical element;

a logical element that generates an event when it has a pre-defined characteristic as measured through a synthetic transaction;

a virtual element that generates an event when a predefined condition occurs; and

a reference element that is a pre-defined collection of other different elements among the same dependency tree, for which a single policy is defined for handling an event that occurs within the reference element.

21 . The non-transitory computer-readable medium of claim 15 , further comprising determining the state of the infrastructure element according an absolute calculation specified in a policy assigned to the infrastructure element.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Nov 5, 2021
From: COMERICA BANK
To: ZENOSS, INC.
Reel/Frame 058035/0917 →
SECURITY INTEREST Recorded Jan 10, 2019
From: ZENOSS, INC.
To: COMERICA BANK
Reel/Frame 047952/0294 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2012
From: MCCRACKEN, IAN C.
To: ZENOSS, INC.
Reel/Frame 029090/0893 →