Distributed application call path performance analysis
In general, techniques are described for managing a distributed application based on call paths among the multiple services of the distributed application that traverse underlying network infrastructure. In an example, a method comprises determining, by a computing system, and for a distributed application implemented with a plurality of services, a call path from an entry endpoint service of the plurality of services to a terminating endpoint service of the plurality of services; determining, by the computing system, a corresponding network path for each pair of adjacent services from a plurality of pairs of services that communicate for the call path; and based on a performance indicator for a network device of the corresponding network path meeting a threshold, performing, by the computing system, one or more of: reconfiguring the network; or redeploying one of the plurality of services to a different compute node of the compute nodes.
1 . A computing system, comprising:
memory; and
processing circuitry in communication with the memory, and configured to:
determine, for a distributed application implemented with a plurality of services executing on compute nodes interconnected by a network of network devices, a call path from an entry endpoint service of the plurality of services to a terminating endpoint service of the plurality of services;
determine a corresponding network path traversed by corresponding calls for each pair of one or more pairs of adjacent services from a plurality of pairs of services that communicate for the call path; and
for a pair of the one or more pairs of adjacent services, based on a performance indicator for a network device, of the network devices, of the corresponding network path traversed by the corresponding calls for the pair of the one or more pairs of adjacent services meeting a threshold, perform one or more of:
reconfigure the network; or
redeploy one of the plurality of services to a different compute node of the compute nodes.
2 . The computing system of claim 1 , wherein to determine the corresponding network path, the processing circuitry is configured to:
determine, for the pair of the one or more pairs of adjacent services, a source IP address and a destination IP address, wherein the source IP address is the IP address of a caller service of the pair of the one or more pairs of adjacent services, and wherein the destination IP address is the IP address of a callee service of the pair of the one or more pairs of adjacent services; and
correlate, based on the source IP address and the destination IP address, a flow to the corresponding network path.
3 . The computing system of claim 2 , wherein to correlate the flow to the corresponding network path the processing circuitry is configured to determine, based on flow data, one or more of the network devices that have processed flows that include the source IP address and the destination IP address.
4 . The computing system of claim 2 , wherein the source IP address is one of a physical IP address for a compute node that hosts the caller service or a virtual IP address for a workload for the caller service.
5 . The computing system of claim 1 , wherein the network devices comprise at least one top-of-rack switch and at least one chassis switch.
6 . The computing system of claim 1 , wherein to reconfigure the network, the processing circuitry is further configured to:
identify a replacement service for either service of the pair of the one or more pairs of adjacent services, wherein the replacement service is another instance of a first service of the pair of the one or more pairs of adjacent services executing on a different compute node; and
reconfigure the distributed application to execute using the replacement service.
7 . The computing system of claim 1 , wherein the call path is a first call path, wherein the entry endpoint is a first entry endpoint, wherein the terminating endpoint is a first terminating endpoint, and wherein the processing circuitry is further configured to:
based on determining that no performance indicators for the network device meet the threshold, determine, for the distributed application implemented with the plurality of services, a second call path from a second entry endpoint service of the plurality of services to a second terminating endpoint service of the plurality of services;
determine a second corresponding network path traversed by corresponding calls for a different pair of the one or more pairs of adjacent services from the plurality of pairs of services that communicate for the second call path;
identify a second network device on one of the second corresponding network path; and
based on a performance indicator for the second network device meeting the threshold, perform one or more of:
reconfigure the network; or
redeploy one of the plurality of services to a different compute node of the compute nodes.
8 . The computing device of claim 1 , wherein the performance indicator is one or more of:
a packet loss rate,
a transmission time,
a resource utilization, or
a latency.
9 . The computing device of claim 1 , wherein the call path is a call path of a plurality of call paths for the distributed application that has a highest end-to-end latency of the plurality of call paths.
10 . A method comprising:
determining, by a computing system, for a distributed application implemented with a plurality of services executing on compute nodes interconnected by a network of network devices, a call path from an entry endpoint service of the plurality of services to a terminating endpoint service of the plurality of services;
determining, by the computing system, a corresponding network path traversed by corresponding calls for each pair of one or more pairs of adjacent services from a plurality of pairs of services that communicate for the call path; and
for a pair of the one or more pairs of adjacent services, based on a performance indicator for a network device of the network devices, of the corresponding network path traversed by the corresponding calls for the pair of the one or more pairs of adjacent services meeting a threshold, performing, by the computing system, one or more of:
reconfiguring the network; or
redeploying one of the plurality of services to a different compute node of the compute nodes.
11 . The method of claim 10 , wherein determining a corresponding network path further comprises:
determining, for the pair of the one or more pairs of adjacent services, a source IP address and a destination IP address, wherein the source IP address is the IP address of a caller service of the pair of the one or more pairs of adjacent services, and wherein the destination IP address is the IP address of a callee service of the pair of the one or more pairs of adjacent services; and
correlating, based on the source IP address and the destination IP address, a flow to the corresponding network path.
12 . The method of claim 11 , wherein correlating the flow to the corresponding network path further comprises determining, based on flow data, one or more of the network devices that have processed flows that include the source IP address and the destination IP address.
13 . The method of claim 12 , wherein the source IP address is one of a physical IP address for a compute node that hosts the caller service or a virtual IP address for a workload for the caller service.
14 . The method of claim 10 , wherein the network devices comprise at least one top-of-rack switch and at least one chassis switch.
15 . The method of claim 10 , wherein reconfiguring further comprises:
identifying a replacement service for either service of the pair of the one or more pairs of adjacent services, wherein the replacement service is another instance of a first service of the pair of the one or more pairs of adjacent services executing on a different compute node; and
reconfiguring the distributed application to execute using the replacement service.
16 . The method of claim 15 , wherein the call path is a first call path, wherein the entry endpoint is a first entry endpoint, wherein the terminating endpoint is a first terminating endpoint, and further comprising:
based on determining that no performance indicators for the network device meet the threshold, determining, by the computing system and for the distributed application implemented with the plurality of services, a second call path from a second entry endpoint service of the plurality of services to a second terminating endpoint service of the plurality of services;
determining, by the computing system, a second corresponding network path traversed by corresponding calls for a different pair of the one or more pairs of adjacent services from the plurality of pairs of services that communicate for the second call path;
identifying, by the computing system, a second network device on one of the second corresponding network path; and
based on a performance indicator for the second network device meeting a threshold, perform one or more of:
reconfiguring the network; or
redeploying one of the plurality of services to a different compute node of the compute nodes.
17 . The method of claim 10 , wherein the performance indicator is one or more of:
a packet loss rate,
a transmission time,
a resource utilization, or
a latency.
18 . The method of claim 10 , wherein the call path is a call path of a plurality of call paths for the distributed application that has a highest end-to-end latency of the plurality of call paths.
19 . Non-transitory computer-readable storage media comprising instructions that, when executed, cause one or more processors to:
determine, for a distributed application implemented with a plurality of services executing on compute nodes interconnected by a network of network devices, a call path from an entry endpoint service of the plurality of services to a terminating endpoint service of the plurality of services;
determine a corresponding network path traversed by corresponding calls for each pair of one or more pairs of adjacent services from a plurality of pairs of services that communicate for the call path; and
for a pair of the one or more pairs of adjacent services, based on a performance indicator for a network device, of the network devices, of the corresponding network path traversed by the corresponding calls for the pair of the one or more pairs of adjacent services meeting a threshold, perform one or more of:
reconfigure the network; or
redeploy one of the plurality of services to a different compute node of the compute nodes.
20 . The non-transitory computer-readable storage media of claim 19 , wherein the instructions further cause the one or more processors to:
determine, for the pair of the one or more pairs of adjacent services, a source IP address and a destination IP address, wherein the source IP address is the IP address of a caller service of the pair of the one or more pairs of adjacent services, and wherein the destination IP address is the IP address of a callee service of the pair of the one or more pairs of adjacent services; and
correlate, based on the source IP address and the destination IP address, a flow to the corresponding network path.