SYSTEM AND METHOD FOR MANAGING DATA CENTER ALARMS
Systems, methods, architectures, mechanisms and/or apparatus to manage alarm generation associated with event-sourcing objects or entities at a data center in accordance with a hierarchy of failure relationships of the event-sourcing objects or entities, wherein alarms normally generated in response to a received event are suppressed (i.e., not generated) if the source of the event is not the root cause of the event.
1 . A method of managing alarms at a data center (DC), comprising:
defining a hierarchy of failure relationships of DC entities, each of said failure relationships comprising a higher-level DC entity and a lower level DC entity, each lower level DC entity necessarily failing in response to failure of a corresponding higher-level DC entity;
in response to received events indicative of failed DC entities, correlating failed higher-level entities corresponding to failed lower level DC entities to identify thereby root cause failed DC entities; and
generating failure alarms associated with said root cause failed DC entities.
2 . The method of claim 1 , wherein an alarm normally generated in response to a received event is suppressed unless the event source comprises a root cause of the received event.
3 . The method of claim 1 , further comprising suppressing failure alarm generation associated with failed DC entities that do not comprise root cause failed DC entities.
4 . The method of claim 1 , further comprising for each generated failure alarm associated with a root cause failed DC entity, suppressing failure alarm generation associated with failed DC entities in a corresponding failure relationship with the root cause failed DC entity.
5 . The method of claim 1 , further comprising generating failure alarms associated with failed DC entities comprising priority DC entities.
6 . The method of claim 1 , further comprising generating failure alarms associated with a failed DC entities comprising priority entity types.
7 . The method of claim 1 , further comprising generating failure alarms associated with a failed DC entities associated with a priority service.
8 . The method of claim 1 , further comprising generating failure alarms associated with a failed DC entities associated with a priority customer.
9 . The method of claim 1 , wherein received events are included within event streams processed by a rules engine.
10 . The method of claim 1 , wherein said root cause failed DC entities are selected as those higher-level failed DC entities corresponding to lower-level failed entities.
11 . The method of claim 10 , wherein selecting root cause failed DC entities comprises identifying a minimum number of higher order failed DC entities hierarchically corresponding to a group of failed DC entities.
12 . The method of claim 10 , wherein said relational graph is formed as a directed tree structure.
13 . The method of claim 12 , wherein a first directed tree represents a data center object failure hierarchy, a second directed tree represents a Border Gateway Protocol (BGP) failure hierarchy, and a third directed tree represents and Interior Gateway Protocol (IGP) failure hierarchy.
14 . The method of claim 12 , wherein the entities comprise data center objects, wherein a first directed tree represents a data center object hard failure hierarchy and a second directed tree represents a data center object soft failure hierarchy.
15 . The method of claim 12 , wherein the BC entities comprise Border Gateway Protocol (BGP) objects, wherein a first directed tree represents a BGP object hard failure hierarchy and a second directed tree represents a BGP object soft failure hierarchy.
16 . The method of claim 1 , wherein the DC entities comprise any virtual or nonvirtual event-sourcing entity in the data center.
17 . The method of claim 1 , wherein the entities comprise any of a virtual machine (VM), a VM-based appliance, a virtual router (VR) and a virtual service.
18 . An apparatus for managing alarms at a data center, the apparatus comprising:
a processor configured for:
defining a hierarchy of failure relationships of DC entities, each of said failure relationships comprising a higher-level DC entity and a lower level DC entity, each lower level DC entity necessarily failing in response to failure of a corresponding higher-level DC entity;
in response to received events indicative of failed DC entities, correlating failed higher-level entities corresponding to failed lower level DC entities to identify thereby root cause failed DC entities; and
generating failure alarms associated with said root cause failed DC entities.
19 . A tangible and non-transient computer readable storage medium storing instructions which, when executed by a computer, adapt the operation of the computer to perform a method for managing alarms at a data center, the method comprising:
defining a hierarchy of failure relationships of DC entities, each of said failure relationships comprising a higher-level DC entity and a lower level DC entity, each lower level DC entity necessarily failing in response to failure of a corresponding higher-level DC entity;
in response to received events indicative of failed DC entities, correlating failed higher-level entities corresponding to failed lower level DC entities to identify thereby root cause failed DC entities; and
generating failure alarms associated with said root cause failed DC entities.
20 . A computer program product wherein computer instructions, when executed by a processor in a network element, adapt the operation of the network element to provide a method for managing alarms at a data center, the method comprising:
defining a hierarchy of failure relationships of DC entities, each of said failure relationships comprising a higher-level DC entity and a lower level DC entity, each lower level DC entity necessarily failing in response to failure of a corresponding higher-level DC entity;
in response to received events indicative of failed DC entities, correlating failed higher-level entities corresponding to failed lower level DC entities to identify thereby root cause failed DC entities; and
generating failure alarms associated with said root cause failed DC entities.