USER INTERFACE FOR HEALTH MONITORING OF MULTI-SERVICE SYSTEM
Some embodiments provide a method for providing health status for a system implemented in a network. The method displays, in a graphical user interface (GUI), representations of health status for multiple different services of the system. Each representation for a respective service shows health status for the respective service over a first particular time period. Upon receiving selection of a particular service, the method displays representations of health status data for each of multiple different aspects of the particular service. Each representation for a respective aspect of the service shows operational status for the respective aspect of the service over a second particular time period.
1 . A method for providing health status for a system implemented in a network, the method comprising:
displaying, in a graphical user interface (GUI), representations of health status for a plurality of different services of the system, each representation for a respective service showing health status for the respective service over a first particular time period;
upon receiving selection of a particular service, displaying representations of health status data for each of a plurality of different aspects of the particular service, each representation for a respective aspect of the service showing operational status for the respective aspect of the service over a second particular time period.
2 . The method of claim 1 , wherein:
the system is a network management system that manages a plurality of groups of datacenters for a plurality of different tenants; and
the plurality of different services comprises (i) a set of common multi-tenant services and (ii) a set of services belonging to tenant-specific service instances.
3 . The method of claim 2 , wherein:
the set of common multi-tenant services comprises services accessed by the plurality of different tenants; and
each tenant-specific service instance performs a respective service of the network management system for a respective group of datacenters of a respective tenant.
4 . The method of claim 1 , wherein:
the system is implemented in a Kubernetes cluster within a public cloud; and
the plurality of different services are Kubernetes microservices.
5 . The method of claim 1 , wherein the health status data for the plurality of different services specifies, over the first particular time period, whether each of the different services is operational, degraded, or nonoperational.
6 . The method of claim 5 , wherein the particular service is nonoperational when at least one aspect of the particular service is nonoperational.
7 . The method of claim 5 , wherein the particular service is nonoperational when at least one of a first subset of the aspects of the particular service is nonoperational and is degraded when at least one of a second subset of the aspects of the particular service is nonoperational.
8 . The method of claim 1 , wherein the representations of health status for at least a set of the services comprises representations of health status for each of a plurality of replicas of the services of the set of services executing in the network over the first particular time period.
9 . The method of claim 8 , wherein, for a specific service, different replicas of the specific service have different health statuses at a particular time.
10 . The method of claim 1 , wherein the representations of the health status data for the plurality of different aspects of the particular service provides operational status for each of the different aspects on each of a plurality of different replicas of the particular service.
11 . The method of claim 10 , wherein the system is implemented in a Kubernetes cluster, wherein each of the different replicas of the particular service executes on a different node of the cluster.
12 . A non-transitory machine-readable medium storing a program which when executed by at least one processing unit provides health status for a system implemented in a network, the program comprising sets of instructions for:
displaying, in a graphical user interface (GUI), representations of health status for a plurality of different services of the system, each representation for a respective service showing health status for the respective service over a first particular time period;
upon receiving selection of a particular service, displaying representations of health status data for each of a plurality of different aspects of the particular service, each representation for a respective aspect of the service showing operational status for the respective aspect of the service over a second particular time period.
13 . The non-transitory machine-readable medium of claim 12 , wherein:
the system is a network management system that manages a plurality of groups of datacenters for a plurality of different tenants; and
the plurality of different services comprises (i) a set of common multi-tenant services and (ii) a set of services belonging to tenant-specific service instances.
14 . The non-transitory machine-readable medium of claim 13 , wherein:
the set of common multi-tenant services comprises services accessed by the plurality of different tenants; and
each tenant-specific service instance performs a respective service of the network management system for a respective group of datacenters of a respective tenant.
15 . The non-transitory machine-readable medium of claim 12 , wherein:
the system is implemented in a Kubernetes cluster within a public cloud; and
the plurality of different services are Kubernetes microservices.
16 . The non-transitory machine-readable medium of claim 12 , wherein the health status data for the plurality of different services specifies, over the first particular time period, whether each of the different services is operational, degraded, or nonoperational.
17 . The non-transitory machine-readable medium of claim 16 , wherein the particular service is nonoperational when at least one aspect of the particular service is nonoperational.
18 . The non-transitory machine-readable medium of claim 16 , wherein the particular service is nonoperational when at least one of a first subset of the aspects of the particular service is nonoperational and is degraded when at least one of a second subset of the aspects of the particular service is nonoperational.
19 . The non-transitory machine-readable medium of claim 12 , wherein the representations of health status for at least a set of the services comprises representations of health status for each of a plurality of replicas of the services of the set of services executing in the network over the first particular time period.
20 . The non-transitory machine-readable medium of claim 12 , wherein the representations of the health status data for the plurality of different aspects of the particular service provides operational status for each of the different aspects on each of a plurality of different replicas of the particular service.