Agent functionality evaluation in managed endpoints
An embodiment includes a method of health and functionality evaluation of an agent on a managed endpoint. The method includes receiving an agent event message that includes data representing platform health indicators and capacity health indicators. The platform health indicators include quantifications of functionality of communication channels and components that implement the agent. The capacity health indicators include quantifications of functionality of engines that are configured to implement a management operation. The method includes examining the agent event message for a change in status of the health of the agent. Responsive to the agent event message indicating the change, the method includes emitting an updated agent event. The method includes triggering generation a health score for the agent based on the updated agent event and historical agent health data. The method includes communicating to a webhost the health score where it is caused to be displayed.
1 . A method of health and functionality evaluation of an agent on a managed endpoint, the method comprising:
receiving, at a service bus, an agent event message, wherein the agent event message includes data representative of a current status of the agent loaded on the managed endpoint, the agent event message is communicated when the agent checks in with a cloud management device, the agent event message includes first data representing platform health indicators and second data representing capacity health indicators, the platform health indicators are quantifications of functionality of communication channels and components that implement the agent on the managed endpoint, the capacity health indicators are quantifications of functionality of engines that are configured to implement a management operation at the agent, and the first data and the second data are communicated separately;
examining, by a series of detectors, the agent event message for a change in status of at least one aspect of health of the agent;
responsive to the agent event message indicating the change, emitting, by the series of detectors, an updated agent event;
triggering generation of a health score of the agent based on the updated agent event and historical agent health data, the health score representing a current level of functionality degradation of the agent that is based on an agent platform health and health of the engines delivered on the managed endpoint;
communicating to a webhost the health score; and
causing display on a web-hosted user interface, the health score of the agent.
2 . The method of claim 1 , wherein the platform health indicators represent one or more or a combination of:
connection to a distribution server;
an ability to download a manifest;
validity of a manifest;
age of a manifest;
ability to download one of the components;
a trust status of one or more of the engines;
accessibility of one or more of the engines;
installation status of one of the components; and
installation status of a prerequisite.
3 . The method of claim 1 , wherein:
the engines include a first engine and a second engine;
the capacity health indicators include a plurality of status statements;
each of the plurality of status statements includes a name-equal-value pair from each of the engines;
the plurality of status statements includes a first status statement and a second status statement;
the first status statement represents a first status of the first engine and the second status statement represents a second status of the second engine; and
a first weight is associated with the first status statement and a second weight is associated with the second status statement such that a significance of the first status statement is greater than a significance of the second status statement.
4 . The method of claim 3 , further comprising receiving a registration event that specifies a list of package names and key names with associated weights, wherein the registration event populates a status item dictionary, wherein the name-equal-value pair is provided in the status item dictionary.
5 . The method of claim 1 , wherein the capacity health indicators are communicated from an update plugin that is installed in at least one of the engines.
6 . The method of claim 1 , wherein:
the health score is based on an agent health equation:
AgentScore=1.0−(0.5×PlatformScore+0.5×AggCapacityScore),
in which:
AgentScore represents the health score of the agent,
PlatformScore represents a value for the platform health indicators,
AggCapacityScore represents a value for the capacity health indicators.
7 . The method of claim 6 , wherein:
the value for the platform health indicators is based on a platform health equation:
PlatformScore=min(1.0,sum(PlatformIndicators), in which:
min (X,Y) represents a function that returns a lower of variables X or Y,
sum (X) represents a summation function, and
PlatformIndicators represents individual numerical values from each of the platform health indicators; and
the platform health equation is configured such that a platform health score is capped at a value of 1.
8 . The method of claim 6 , wherein the value for the capacity health indicators is based on capacity health indicators equations:
for i=1 to NumCap:
capacityIndicator
(
i
)
=
(
1
N
u
m
C
a
p
)
*
capacityindUnweighted
(
i
)
*
weight
(
i
)
;
AggCapacityScore=min(1.0,sum(capabilityInciators( i ))), in which:
NumCap represents a number of capabilities;
capacityIndicator(i) represents a weighted capacity indicator score for a capacity assigned an indicator i;
capacityindUnweighted(i) represents an unweighted capacity score for a capacity assigned an indicator i;
weight(i) represents a weight assigned to a capacity assigned an indicator i;
min (X,Y) represents a function that returns a lower of variables X or Y; and
sum (X) represents a summation function.
9 . The method of claim 1 , wherein:
a first indicator of the platform health indicators includes a connection stability indicator; and
the connection stability indicator is a quantification of stability of a connection between the agent and the cloud management device.
10 . The method of claim 9 , further comprising:
responsive to a failure to connect, generating a forensic diagnosis of the connection between the agent and the cloud management device; and
retaining the forensic diagnosis at the managed endpoint until a check-in by the agent with the cloud management device, wherein the forensic diagnosis includes one or more or a combination of:
a test of domain name system (DNS) resolution of the cloud management device;
a route trace to the cloud management device;
a ping to the cloud management device;
an adapter configuration; and
a list of products at the managed endpoint that are out of date.
11 . The method of claim 1 , wherein:
a first indicator of the platform health indicators includes an agent check-in indicator;
the agent check-in indicator is a quantification of whether or not the agent has checked-in with the cloud management device; and
the agent check-in indicator is based on a scheduled background task that is configured to examine multiple agents that includes the agent deployed on a plurality of managed endpoints that includes the managed endpoint.
12 . The method of claim 1 , wherein:
a first indicator of the platform health indicators includes a software component update indicator;
the software component update indicator includes a list of products at the managed endpoint that are out of date;
the software component update indicator is based on a package status change event combined with a policy indicating a prescribed version of software components at the managed endpoint; and
the list of products includes manifests published per distribution ring in a managed network.
13 . A non-transitory computer-readable medium having encoded therein programming code executable by one or more processors to perform or control performance of operations of health and functionality evaluation of an agent on a managed endpoint, the operations comprising:
receiving, at a service bus, an agent event message, wherein the agent event message includes data representative of a current status of the agent loaded on the managed endpoint, the agent event message is communicated when the agent checks in with a cloud management device, the agent event message includes first data representing platform health indicators and second data representing capacity health indicators, the platform health indicators are quantifications of functionality of communication channels and components that implement the agent on the managed endpoint, the capacity health indicators are quantifications of functionality of engines that are configured to implement a management operation at the agent, and the first data and the second data are communicated separately;
examining, by a series of detectors, the agent event message for a change in status of at least one aspect of health of the agent;
responsive to the agent event message indicating the change, emitting, by the series of detectors, an updated agent event;
triggering generation of a health score of the agent based on the updated agent event and historical agent health data, the health score representing a current level of functionality degradation of the agent that is based on an agent platform health and health of the engines delivered on the managed endpoint;
communicating to a webhost the health score; and
causing display, on a web-hosted user interface, the health score of the agent.
14 . The non-transitory computer-readable medium of claim 13 , wherein the platform health indicators represent one or more or a combination of:
connection to a distribution server;
an ability to download a manifest;
validity of a manifest;
age of a manifest;
ability to download one of the components;
a trust status of one or more of the engines;
accessibility of one or more of the engines;
installation status of one of the components; and
installation status of a prerequisite.
15 . The non-transitory computer-readable medium of claim 13 , wherein:
the engines include a first engine and a second engine;
the capacity health indicators include a plurality of status statements;
each of the plurality of status statements includes a name-equal-value pair from each of the engines;
the plurality of status statements includes a first status statement and a second status statement;
the first status statement represents a first status of the first engine and the second status statement represents a second status of the second engine; and
a first weight is associated with the first status statement and a second weight is associated with the second status statement such that a significance of the first status statement is greater than a significance of the second status statement.
16 . The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise receiving a registration event that specifies a list of package names and key names with associated weights, wherein the registration event populates a status item dictionary, wherein the name-equal-value pair is provided in the status item dictionary.
17 . The non-transitory computer-readable medium of claim 13 , wherein the capacity health indicators are communicated from an update plugin that is installed in at least one of the engines.
18 . The non-transitory computer-readable medium of claim 13 , wherein:
the health score is based on an agent health equation:
AgentScore=1.0−(0.5×PlatformScore+0.5×AggCapacityScore),
in which:
AgentScore represents the health score of the agent,
PlatformScore represents a value for the platform health indicators,
AggCapacityScore represents a value for the capacity health indicators.
19 . The non-transitory computer-readable medium of claim 18 , wherein:
the value for the platform health indicators is based on a platform health equation:
PlatformScore=min(1.0,sum(PlatformIndicators), in which:
min (X,Y) represents a function that returns a lower of variables X or Y,
sum (X) represents a summation function, and
PlatformIndicators represents individual numerical values from each of the platform health indicators; and
the platform health equation is configured such that a platform health score is capped at a value of 1.
20 . The non-transitory computer-readable medium of claim 18 , wherein the value for the capacity health indicators is based on capacity health indicators equations:
for i=1 to NumCap:
for
i
=
1
to
NumCap
:
capacityIndicator
(
i
)
=
(
1
NumCap
)
*
capacityindUnweighted
(
i
)
*
weight
(
i
)
;
AggCapacityScore
=
min
(
1.
,
sum
(
capabilityInciators
(
i
)
)
)
,
in which:
NumCap represents a number of capabilities;
capacityIndicator(i) represents a weighted capacity indicator score for a capacity assigned an indicator i;
capacityindUnweighted(i) represents an unweighted capacity score for a capacity assigned an indicator i;
weight(i) represents a weight assigned to a capacity assigned an indicator i;
min (X,Y) represents a function that returns a lower of variables X or Y; and
sum (X) represents a summation function.
21 . The non-transitory computer-readable medium of claim 13 , wherein:
a first indicator of the platform health indicators includes a connection stability indicator; and
the connection stability indicator is a quantification of stability of a connection between the agent and the cloud management device.
22 . The non-transitory computer-readable medium of claim 21 , wherein the operations further comprise:
responsive to a failure to connect, generating a forensic diagnosis of the connection between the agent and the cloud management device; and
retaining the forensic diagnosis at the managed endpoint until a check-in by the agent with the cloud management device, wherein the forensic diagnosis includes one or more or a combination of:
a test of domain name system (DNS) resolution of the cloud management device;
a route trace to the cloud management device;
a ping to the cloud management device;
an adapter configuration; and
a list of products at the managed endpoint that are out of date.
23 . The non-transitory computer-readable medium of claim 13 , wherein:
a first indicator of the platform health indicators includes an agent check-in indicator;
the agent check-in indicator is a quantification of whether or not the agent has checked-in with the cloud management device; and
the agent check-in indicator is based on a scheduled background task that is configured to examine multiple agents that includes the agent deployed on a plurality of managed endpoints that includes the managed endpoint.
24 . The non-transitory computer-readable medium of claim 13 , wherein:
a first indicator of the platform health indicators includes a software component update indicator;
the software component update indicator includes a list of products at the managed endpoint that are out of date;
the software component update indicator is based on a package status change event combined with a policy indicating a prescribed version of software components at the managed endpoint; and
the list of products includes manifests published per distribution ring in a managed network.