Detecting network events having adverse user impact
A method includes receiving, by a network management system, network data from a plurality of network devices configured to provide a network at a site; receiving, by the processing circuitry, user impact data from a plurality of client devices that access the network at the site; determining, based on the network data, a pattern of one or more network events occurring over time; correlating in time the pattern of the one or more network events to an adverse user impact event indicated by the user impact data received from the plurality of client devices; and determining, in response to the correlating, an instance of overwhelming network traffic having an adverse user impact. In some examples, the network data includes network traffic impact data, such as a number of packets dropped at a switch port due to congestion.
1 . A network management system (NMS) comprising processing circuitry in communication with storage media, the processing circuitry configured to:
determine a pattern of network events within a rolling time window based on network telemetry data obtained from a plurality of network devices that provide network access to one or more client devices;
determine an impact event based on user data obtained from the one or more client devices for one or more users, the user data indicating feedback related to a quality of an application session;
correlate in time the pattern of network events within the rolling time window to the impact event; and
determine, based at least in part on the correlation, that the pattern of network events is indicative of network behavior that is a cause of the impact event, wherein the network behavior comprises an instance of a high volume of discovery messages per unit time exchanged by the one or more client devices.
2 . The NMS of claim 1 , wherein the processing circuitry is further configured to:
identify a root cause of the network behavior that is the cause of the impact event; and
initiate a remedial action to remedy the root cause of the network behavior that is the cause of the impact event.
3 . The NMS of claim 1 , wherein, to determine the pattern of network events based on the network telemetry data, the processing circuitry is configured to apply the network telemetry data as input to a machine learning system trained to perform anomaly detection based on the network telemetry data, the machine learning system configured to output an indication of the pattern of network events.
4 . The NMS of claim 1 , wherein the processing circuitry is further configured to generate, for output to a user, a real-time alert indicating that the pattern of network events is indicative of the network behavior that is the cause of the impact event.
5 . The NMS of claim 1 , wherein the network telemetry data for the plurality of network devices comprises network telemetry data indicating a number of packets dropped at each network device of the plurality of network devices due to congestion.
6 . The NMS of claim 1 , wherein the network telemetry data for the plurality of network devices comprises network telemetry data indicating a number of reflected packets received by each network device of the plurality of network devices.
7 . The NMS of claim 1 , wherein the processing circuitry is further configured to obtain the user data indicating the feedback related to the quality of the application session as a response to a prompt presented to a user of the one or more users by a client device of the one or more client devices.
8 . The NMS of claim 1 ,
wherein to determine that the pattern of network events is indicative of the network behavior that is the cause of the impact event, the processing circuitry is configured to determine that the pattern of network events is indicative of a worsening trend of the network behavior that is the cause of the impact event;
identify a root cause of the worsening trend of the network behavior; and
initiate a remedial action to remedy the root cause of the worsening trend of the network behavior.
9 . A method comprising:
determining, by a network management system (NMS) executed by processing circuitry, a pattern of network events within a rolling time window based on network telemetry data obtained from a plurality of network devices that provide network access to one or more client devices;
determining, by the NMS, an impact event based on user data obtained from the one or more client devices for one or more users, the user data indicating feedback related to a quality of an application session;
correlating, by the NMS, in time the pattern of network events within the rolling time window to the impact event; and
determining, by the NMS, based at least in part on the correlation, that the pattern of network events is indicative of network behavior that is a cause of the impact event, wherein the network behavior comprises an instance of a high volume of discovery messages per unit time exchanged by the one or more client devices.
10 . The method of claim 9 , further comprising:
identifying, by the NMS, a root cause of the network behavior that is the cause of the impact event; and
initiating, by the NMS, a remedial action to remedy the root cause of the network behavior that is the cause of the impact event.
11 . The method of claim 9 , wherein determining the pattern of network events based on the network telemetry data comprises applying, by the NMS, the network telemetry data as input to a machine learning system trained to perform anomaly detection based on the network telemetry data, the machine learning system configured to output an indication of the pattern of network events.
12 . The method of claim 9 , wherein the network telemetry data for the plurality of network devices comprises network telemetry data indicating a number of packets dropped at each network device of the plurality of network devices due to congestion.
13 . The method of claim 9 , wherein the network telemetry data for the plurality of network devices comprises network telemetry data indicating a number of reflected packets received by each network device of the plurality of network devices.
14 . The method of claim 9 , further comprising obtaining, by the NMS, the user data indicating the feedback related to the quality of the application session as a response to a prompt presented to a user of the one or more users by a client device of the one or more client devices.
15 . Non-transitory, computer-readable media comprising instructions that, when executed, cause processing circuitry to:
execute a network management system (NMS) configured to:
determine a pattern of network events within a rolling time window based on network telemetry data obtained from a plurality of network devices that provide network access to one or more client devices;
determine an impact event based on user data obtained from the one or more client devices for one or more users, the user data indicating feedback related to a quality of an application session;
correlate in time the pattern of network events within the rolling time window to the impact event; and
determine, based at least in part on the correlation, that the pattern of network events is indicative of network behavior that is a cause of the impact event, wherein the network behavior comprises an instance of a high volume of discovery messages per unit time exchanged by the one or more client devices.