IP Library Granted Patent US 11,863,921
Granted Patent B2
US 11,863,921 · App. 18/313,255 · Granted Jan 2, 2024

Application performance monitoring and management platform with anomalous flowlet resolution

Inventors: Ashutosh Kulshreshtha (Cupertino, CA); Omid Madani (San Carlos, CA); Vimal Jeyakumar (Sunnyvale, CA); Navindra Yadav (Cupertino, CA); Ali Parandehgheibi (Sunnyvale, CA); Andy Sloane (Pleasanton, CA); Kai Chang (San Francisco, CA); Khawar Deen (Sunnyvale, CA); Shih-Chun Chang (San Jose, CA); Hai Vu (San Jose, CA)
Assignee: Cisco Technology, Inc.
H04Q9/02G06F11/3495H04L13/04H04L41/064H04L41/0681H04L43/026H04L67/12H04L41/14H04L43/16H04Q2209/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,863,921
App. No.
18/313,255
Granted
Jan 2, 2024
Kind
B2
Abstract

An application and network analytics platform can capture telemetry from servers and network devices operating within a network. The application and network analytics platform can determine an application dependency map (ADM) for an application executing in the network. Using the ADM, the application and network analytics platform can resolve flows into flowlets of various granularities, and determine baseline metrics for the flowlets. The baseline metrics can include transmission times, processing times, and/or data sizes for the flowlets. The application and network analytics platform can compare new flowlets against the baselines to assess availability, load, latency, and other performance metrics for the application. In some implementations, the application and network analytics platform can automate remediation of unavailability, load, latency, and other application performance issues.

Claims (38)

1. A method comprising:

processing telemetry data for a plurality of flows associated with a set of service instances in a network, the set of service instances residing in a data center, the telemetry data received from one or more sensors installed in the data center;

generating, based on the processed telemetry data, an application dependency map for an application executing in the network, the application dependency map indicating dependencies between the set of service instances in the network, each service instance implementing one or more processes associated with the application;

determining one or more metrics associated with requests and responses transmitted between at least a first service instance and a second service instance of the application dependency map;

comparing the determined one or more metrics to respective ranges; and

responsive to detecting a deviation of at least one of the determined one or more metrics from a corresponding respective range, initiating one or more remediation actions, at least one of the one or more remediation actions comprising instantiating one or more new service instances associated with the application in a public cloud remote from the data center.

2. The method of claim 1 , wherein at least a second one of the one or more remediation actions comprises load balancing among the set of service instances associated with the application dependency map.

3. The method of claim 1 , wherein at least one of the determined one or more metrics comprises central processing unit (CPU) utilization of the first service instance or the second service instance.

4. The method of claim 3 , wherein at least a second one of the one or more remediation actions comprises instantiating one or more new service instances associated with the application dependency map in the data center.

5. The method of claim 1 , wherein at least a second one of the one or more remediation actions comprises disabling network connectivity for one or more problematic servers.

6. The method of claim 1 , wherein at least a second one of the one or more remediation actions comprises migrating the first service instance from a first location to a second location, wherein the migrating reduces a distance between the first service instance and a third location of the second service instance.

7. The method of claim 1 , wherein at least one of the respective ranges is based on one or more baseline metrics, the one or more baseline metrics determined from analysis of the telemetry data over a period of time associated with the requests and responses transmitted between at least the first service instance and the second service instance of the application dependency map.

8. The method of claim 1 , wherein at least one sensor of the one or more sensors is installed on a network device in the network, and wherein at least a second sensor of the one or more sensors is installed on a server device of the network.

9. A system comprising:

one or more processors; and

memory including instructions that, upon being executed by the one or more processors, cause the system to:

process telemetry data for a plurality of flows associated with a set of service instances in a network, the set of service instances residing in a data center, the telemetry data received from one or more sensors installed in the data center;

generate, based on the processed telemetry data, an application dependency map for an application executing in the network, the application dependency map indicating dependencies between the set of service instances in the network, each service instance implementing one or more processes associated with the application;

determine one or more metrics associated with requests and responses transmitted between at least a first service instance and a second service instance of the application dependency map;

compare the determined one or more metrics to respective ranges; and responsive to detecting a deviation of at least one of the determined one or more metrics from a corresponding respective range, initiate one or more remediation actions, at least one of the one or more remediation actions comprising instantiating one or more new service instances associated with the application in a public cloud remote from the data center.

10. The system of claim 9 , wherein at least a second one of the one or more remediation actions comprises load balancing among the set of service instances associated with the application dependency map.

11. The system of claim 9 , wherein at least one of the determined one or more metrics comprises central processing unit (CPU) utilization of the first service instance or the second service instance.

12. The system of claim 11 , wherein at least a second one of the one or more remediation actions comprises instantiating one or more new service instances associated with the application dependency map in the data center.

13. The system of claim 9 , wherein at least a second one of the one or more remediation actions comprises disabling network connectivity for one or more problematic servers.

14. The system of claim 9 , wherein at least a second one of the one or more remediation actions comprises migrating the first service instance from a first location to a second location, wherein the migrating reduces a distance between the first service instance and a third location of the second service instance.

15. The system of claim 9 , wherein at least one of the respective ranges is based on one or more baseline metrics, the one or more baseline metrics determined from analysis of the telemetry data over a period of time associated with the requests and responses transmitted between at least the first service instance and the second service instance of the application dependency map.

16. The system of claim 9 , wherein at least one sensor of the one or more sensors is installed on a network device in the network, and wherein at least a second sensor of the one or more sensors is installed on a server device of the network.

17. A non-transitory computer-readable medium having instructions that, upon being executed by one or more processors, cause the one or more processors to:

process telemetry data for a plurality of flows associated with a set of service instances in a network, the set of service instances residing in a data center, the telemetry data received from one or more sensors installed in the data center;

generate, based on the processed telemetry data, an application dependency map for an application executing in the network, the application dependency map indicating dependencies between the set of service instances in the network, each service instance implementing one or more processes associated with the application;

determine one or more metrics associated with requests and responses transmitted between at least a first service instance and a second service instance of the application dependency map;

compare the determined one or more metrics to respective ranges; and

responsive to detecting a deviation of at least one of the determined one or more metrics from a corresponding respective range, initiate one or more remediation actions, at least one of the one or more remediation actions comprising instantiating one or more new service instances associated with the application in a public cloud remote from the data center.

18. The non-transitory computer-readable medium of claim 17 , wherein at least a second one of the one or more remediation actions comprises load balancing among the set of service instances associated with the application dependency map.

19. The non-transitory computer-readable medium of claim 17 , wherein at least one of the determined one or more metrics comprises central processing unit (CPU) utilization of the first service instance or the second service instance.

20. The non-transitory computer-readable medium of claim 19 , wherein at least a second one of the one or more remediation actions comprises instantiating one or more new service instances associated with the application dependency map in the data center.

21. The non-transitory computer-readable medium of claim 17 , wherein at least a second one of the one or more remediation actions comprises disabling network connectivity for one or more problematic servers.

22. The non-transitory computer-readable medium of claim 17 , wherein at least a second one of the one or more remediation actions comprises migrating the first service instance from a first location to a second location, wherein the migrating reduces a distance between the first service instance and a third location of the second service instance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2023
From: KULSHRESHTHA, ASHUTOSH; MADANI, OMID; JEYAKUMAR, VIMAL; YADAV, NAVINDRA; PARANDEHGHEIBI, ALI; SLOANE, ANDY; CHANG, KAI; DEEN, KHAWAR; CHANG, SHIH-CHUN; VU, HAI
To: CISCO TECHNOLOGY, INC.
Reel/Frame 063556/0802 →
Continuity (4)
Continuation 17529727 · Nov 18, 2021
Continuation 17094815 · Nov 11, 2020
Continuation 15471183 · Mar 28, 2017
Related Publication 20230276152A1 · Aug 31, 2023
Cited By (2)
US 12,284,087 US 12,309,039