HIGH AVAILABILITY MANAGEMENT FOR A HIERARCHY OF RESOURCES IN AN SDDC
Some embodiments provide a hierarchical data service (HDS) that manages many resource clusters that are in a resource cluster hierarchy. In some embodiments, each resource cluster has its own cluster manager, and the cluster managers are in a cluster manager hierarchy that mimics the hierarchy of the resource clusters. In some embodiments, both the resource cluster hierarchy and the cluster manager hierarchy are tree structures, e.g., a directed acyclic graph (DAG) structure that has one root node with multiple other nodes in a hierarchy, with each other node having only one parent node and one or more possible child nodes.
1 . A method of managing resources arranged in a hierarchy in at least one datacenter, the method comprising:
defining different resource managers for different resource cluster at different level of the resource hierarchy in the datacenter;
associating each of a plurality of child resource managers to have one parent resource manager, each child or parent resource manager managing a resource cluster, each child resource manager providing state data regarding the child resource manager's resource cluster to its associated parent resource manager, and
configuring each particular resource manager to send notification to remove data regarding a child resource manager of the particular resource manager to the parent resource manager of the particular resource manager when the particular resource manager loses connection with the child resource manager.
2 . The method of claim 1 , wherein the particular resource cluster loses connection with the child resource cluster when the child resource cluster has an operational failure or a network connectivity between the particular resource cluster and the child resource cluster fails.
3 . The method of claim 1 further comprising configuring each pair of parent and child resource cluster managers to exchange control messages to ensure that the connection between the pair is maintained, said particular resource manager detecting that the connection with the child resource manager has failed after.
4 . The method of claim 1 further comprising associating each particular child resource manager of a set of child resource managers to have at least one child resource manager of its own that manages a resource cluster that is a child resource cluster of the resource cluster managed by the particular child resource manager, the particular child resource manager receiving state data from each of its child resource manager regarding the resource cluster managed by that child resource manager.
5 . The method of claim 1 further comprising configuring each particular resource manager of a plurality of resource managers to pass along the state data provided by a first set of progeny resource managers of the particular resource manager to the parent resource manager of the particular resource manager without aggregating the provided state data, and to aggregate state data provided by a second set of progeny resource managers of the particular resource manager to the parent resource manager as each progeny resource manager in the second set has reached a maximum upstream state propagation level.
6 . The method of claim 1 , wherein each parent resource manager to provide commands to each of its child resource managers to direct the child resource manager to effectuate a change in state of the resource cluster managed by the child resource manager.
7 . The method of claim 1 further comprising:
providing each resource manager of a set of resource managers with a list of potential parent resource managers; and
configuring each particular resource manager of the set of resource managers to detect that a connection with a parent first resource manager has failed, to select a second resource manager from the list, and to establish a connection with the second resource manager as the parent resource manager of the particular resource manager.
8 . The method of claim 7 , wherein
before the connection to the first parent resource manager fails, each particular resource manager of the set of resource managers provides state data regarding the resource cluster managed by the particular resource manager to the first parent resource manager; and
after the connection to the first parent resource manager fails, each particular resource manager of the set of resource managers provides state data regarding the resource cluster managed by the particular resource manager to the second parent resource manager.
9 . The method of claim 1 further comprising configuring each particular resource manager to send notification to remove data regarding a progeny resource manager of the particular resource manager to the parent resource manager of the particular resource manager when the particular resource manager receives notification from a child resource manager that state data associated with one of its progeny resource managers should be updated due to loss of connectivity to at least one of its progeny resource managers.
10 . The method of claim 1 , wherein the resource clusters comprise compute clusters and network clusters, and said managers are machines executing on the datacenter, said machines being one of containers, Pods, virtual machines and standalone computers.
11 . A system for managing resources arranged in a hierarchy in at least one datacenter, the system comprising:
different resource managers for different resource cluster at different level of the resource hierarchy in the datacenter;
each of a plurality of child resource managers associated with one parent resource manager, each child or parent resource manager managing a resource cluster, each child resource manager providing state data regarding the child resource manager's resource cluster to its associated parent resource manager, and
each particular resource manager configured to send notification to remove data regarding a child resource manager of the particular resource manager to the parent resource manager of the particular resource manager when the particular resource manager loses connection with the child resource manager.
12 . The system of claim 11 , wherein the particular resource cluster loses connection with the child resource cluster when the child resource cluster has an operational failure or a network connectivity between the particular resource cluster and the child resource cluster fails.
13 . The system of claim 11 , wherein each pair of parent and child resource cluster managers is configured to exchange control messages to ensure that the connection between the pair is maintained, said particular resource manager detecting that the connection with the child resource manager has failed after.
14 . The system of claim 11 , wherein each particular child resource manager of a set of child resource managers is associated with at least one child resource manager of its own that manages a resource cluster that is a child resource cluster of the resource cluster managed by the particular child resource manager, the particular child resource manager receiving state data from each of its child resource manager regarding the resource cluster managed by that child resource manager.
15 . The system of claim 11 , wherein each particular resource manager of a plurality of resource managers is configured to pass along the state data provided by a first set of progeny resource managers of the particular resource manager to the parent resource manager of the particular resource manager without aggregating the provided state data, and to aggregate state data provided by a second set of progeny resource managers of the particular resource manager to the parent resource manager as each progeny resource manager in the second set has reached a maximum upstream state propagation level.
16 . The system of claim 11 , wherein each parent resource manager to provide commands to each of its child resource managers to direct the child resource manager to effectuate a change in state of the resource cluster managed by the child resource manager.
17 . The system of claim 11 , wherein
each resource manager of a set of resource managers is provided with a list of potential parent resource managers; and
each particular resource manager of the set of resource managers is configured to detect that a connection with a parent first resource manager has failed, to select a second resource manager from the list, and to establish a connection with the second resource manager as the parent resource manager of the particular resource manager.
18 . The system of claim 17 , wherein
before the connection to the first parent resource manager fails, each particular resource manager of the set of resource managers provides state data regarding the resource cluster managed by the particular resource manager to the first parent resource manager; and
after the connection to the first parent resource manager fails, each particular resource manager of the set of resource managers provides state data regarding the resource cluster managed by the particular resource manager to the second parent resource manager.
19 . The system of claim 11 , wherein each particular resource manager is configured to send notification to remove data regarding a progeny resource manager of the particular resource manager to the parent resource manager of the particular resource manager when the particular resource manager receives notification from a child resource manager that state data associated with one of its progeny resource managers should be updated due to loss of connectivity to at least one of its progeny resource managers.
20 . The system of claim 11 , wherein the resource clusters comprise compute clusters and network clusters, and said managers are machines executing on the datacenter, said machines being one of containers, Pods, virtual machines and standalone computers.