IP Library › Granted Patent US 12,321,790
Granted Patent B2
US 12,321,790 · App. 17/718,419 · Granted Jun 3, 2025

Distributed control plane for handling worker node failures of a distributed storage architecture

Inventors: Praveen Kumar Hasti (Acton, MA); Christopher Alan Busick (Littleton, MA)
Assignee: NetApp, Inc.
G06F9/5072G06F3/0607G06F3/0614G06F3/0616G06F3/0635G06F3/0653G06F3/067G06F9/48G06F9/4806G06F9/4843G06F9/4881G06F9/50G06F9/5005G06F9/5061G06F9/5077G06F9/5083G06F9/541G06F11/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,321,790
App. No.
17/718,419
Granted
Jun 3, 2025
Kind
B2
Abstract

Techniques are provided for implementing a distributed control plane to facilitate communication between a container orchestration platform and a distributed storage architecture. The distributed storage architecture hosts worker nodes that manage distributed storage that can be made accessible to applications within the container orchestration platform through the distributed control plane. The distributed control plane includes control plane controllers that are each paired with a single worker node of the distributed storage architecture. The distributed control plane is configured to selectively route commands to control plane controllers that are paired with worker nodes that are current owners of objects targeted by the commands. If a worker node fails and ownership of an object has changed from the failed worker node to another worker node, then subsequent commands are re-routed to a control plane controller paired with the other worker node now owning the object in place of the failed worker node.

Claims (58)

1. A system, comprising:

a distributed storage architecture including a plurality of worker nodes managing distributed storage comprised of storage devices hosted by the plurality of worker nodes;

a container orchestration platform hosting applications running through containers;

a distributed control plane hosted within the container orchestration platform, wherein the distributed control plane comprises a plurality of pods hosting control plane controllers paired with the worker nodes, wherein the distributed control plane is configured to:

route a first command originating from an application running as a container within the container orchestration platform to a first control plane controller paired with a first worker node based upon the first command targeting an object owned by the first worker node; and

in response to detecting a failure of the first worker node, route a second command to a second control plane controller paired with a second worker node based upon the second command targeting the object whose ownership changed from the first worker node to the second worker node in response to the failure; and

the first control plane controller configured to:

translate the first command from being formatted according to a first model supported by the container orchestration platform into a first reformatted command formatted according to a second model supported by the distributed storage architecture; and

transmit the first reformatted command to the first worker node for execution through the distributed storage architecture.

2. The system of claim 1 , further comprising:

the second control plane controller configured to:

translate the second command from being formatted according to the first model into a second reformatted command formatted according to the second model; and

transmit the second reformatted command to the second worker node for execution through the distributed storage architecture.

3. The system of claim 1 , the distributed control plane is further configured to:

execute a polling thread to detect object ownership changes of objects amongst the worker nodes due to worker node failures.

4. The system of claim 1 , the distributed storage architecture is further configured to:

track health status information related to the plurality of worker nodes; and

provide the health status information to the distributed control plane for tracking ownership of objects by worker nodes, wherein ownership of a set of objects owned by a worker node is transferred to a different worker node based upon the health status information.

5. The system of claim 1 , the distributed storage architecture is further configured to:

track health status information related to the plurality of worker nodes; and

in response to the health status information indicating that the first worker node failed, transferring ownership of objects owned by the first worker node to the second worker node, wherein commands targeting the objects are routed to the second control plane controller paired with the second worker node based upon the ownership being transferred.

6. The system of claim 1 , the distributed storage architecture is further configured to:

change ownership of a first object from a worker node to a different worker node while retaining current storage locations of data of the first object within the distributed storage.

7. The system of claim 1 , the distributed control plane is further configured to:

in response to detecting the distributed storage architecture adding a new worker node to take over for a failed worker node, host a new control plane controller paired with the new worker node; and

route commands to the new control plane controller based upon the commands targeting objects owned by the new worker node.

8. The system of claim 1 , the distributed control plane is further configured to:

in response to detecting the distributed storage architecture removing a worker node, remove at least one of a control plane controller paired with the worker node or a pod hosting the control plane controller, wherein ownership of objects owned by the worker node are transferred to a different worker node.

9. The system of claim 1 , the distributed control plane is further configured to:

in response to detecting the failure of the first worker node, remove at least one of the first control plane controller paired with the first worker node or a pod hosting the first control plane controller, wherein ownership of objects owned by the first worker node are transferred to a different worker node.

10. A method, comprising:

receiving, by a distributed control plane hosted within a container orchestration platform hosting applications running through containers, a first command originating from an application running as a container within the container orchestration platform, wherein the command is formatted according to a first model supported by the container orchestration platform;

routing the first command from the application a first control plane controller paired with a first worker node of a distributed storage architecture based upon the first command targeting an object owned by the first worker node; and

in response to detecting a failure of the first worker node, routing a second command originating from the application to a second control plane controller paired with a second worker node of the distributed storage architecture based upon the second command targeting the object whose ownership changed from the first worker node to the second worker node based upon the failure.

11. The method of claim 10 , the method further comprising:

translating, by the first control plane controller, the first command from being formatted according to the first model into a first reformatted command formatted according to a second model supported by the distributed storage architecture; and

transmitting, by the first control plane controller, the first reformatted command to the first worker node for execution through the distributed storage architecture.

12. The method of claim 10 , the method further comprising:

translating, by the second control plane controller, the second command from being formatted according to the first model into a second reformatted command formatted according to a second model supported by the distributed storage architecture; and

transmitting, by the second control plane controller, the second reformatted command to the second worker node for execution through the distributed storage architecture.

13. The method of claim 10 , the method further comprising:

polling the distributed storage architecture to detect object ownership changes of objects amongst the worker nodes due to worker node failures, wherein ownership of objects of a failed worker node are transferred to a different worker node.

14. The method of claim 10 , the method further comprising:

providing, by the distributed storage architecture, health status information of the worker nodes to the distributed control plane, wherein the distributed control plane utilizes the health status information to determine which worker nodes are to be paired with which control plane controllers.

15. The method of claim 10 , the method further comprising:

tracking, by the first control plane controller, health status information of the first worker node, wherein ownership of objects owned by the first worker node are transferred to a different worker node based upon the health status information indicating a degraded health state of the first worker node.

16. The method of claim 10 , the method further comprising:

in response to health status information indicating that the first worker node failed, transferring ownership of objects owned by the first worker node to the second worker node, wherein commands targeting the objects are rerouted to the second control plane controller.

17. The method of claim 10 , the method further comprising:

in response to detecting the distributed storage architecture adding a new worker node to take over for a failed worker node, hosting a new control plane controller paired with the new worker node, wherein ownership of objects owned by the failed worker node are transferred to the new worker node.

18. The method of claim 10 , the method further comprising:

in response to detecting the distributed storage architecture removing a worker node, removing a control plane controller paired with the worker node, wherein ownership of objects owned by the worker node are transferred to a different worker node.

19. The method of claim 10 , the method further comprising:

in response to detecting the failure of the first worker node, removing the first control plane controller paired with the first worker node, wherein ownership of objects owned by the first worker node are transferred to a different worker node.

20. A non-transitory machine readable medium comprising instructions, which when executed by a machine, causes the machine to:

receive, by a distributed control plane hosted within a container orchestration platform hosting applications running through containers, a first command originating from an application running as a container within the container orchestration platform, wherein the command is formatted according to a first model supported by the container orchestration platform;

route the first command from the application to a first control plane controller paired with a first worker node of a distributed storage architecture based upon the first command targeting an object owned by the first worker node, wherein the first control plane controller reformats the first command according to a second model supported by the distributed storage architecture and transmits the first command to the first worker node; and

in response to detecting a failure of the first worker node, route a second command to a second control plane controller paired with a second worker node based upon the second command targeting the object whose ownership changed from the first worker node to the second worker node based upon the failure, wherein the second control plane controller reformats the second command according to the second model and transmits the second command to the second worker node.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 12, 2022
From: HASTI, PRAVEEN KUMAR; BUSICK, CHRISTOPHER ALAN
To: NETAPP INC.
Reel/Frame 059570/0043 →
Continuity (1)
Related Publication 20230325254A1 · Oct 12, 2023
References Cited (56)
US 9438665B1 · Vasanth et al. · 2016 [cited by applicant]
US 10871922B2 · East · 2020 [cited by applicant]
US 11245748B1 · Hannon · 2022 [cited by applicant]
US 11650886B2 · Mathew et al. · 2023 [cited by applicant]
US 11748030B1 · East · 2023 [cited by applicant]
US 11775204B1 · Hasti · 2023 [cited by examiner]
US 11789660B1 · Hasti · 2023 [cited by examiner]
US 11822370B2 · Patel et al. · 2023 [cited by applicant]
US 20190065096A1 · Sterin et al. · 2019 [cited by applicant]
US 20190173793A1 · Liu et al. · 2019 [cited by applicant]
US 20200026625A1 · Konka · 2020 [cited by examiner]
US 20200083909A1 · Kusters et al. · 2020 [cited by applicant]
US 20200099610A1 · Heron et al. · 2020 [cited by applicant]
US 20200104275A1 · Sen et al. · 2020 [cited by applicant]
US 20200213879A1 · Kim et al. · 2020 [cited by applicant]
US 20200280592A1 · Ithal et al. · 2020 [cited by applicant]
US 20210034423A1 · Hallur et al. · 2021 [cited by applicant]
US 20210311764A1 · Rosoff et al. · 2021 [cited by applicant]
US 20210349767A1 · Asayag · 2021 [cited by examiner]
US 20220012095A1 · Ylinen et al. · 2022 [cited by applicant]
US 20220391138A1 · Dronamraju et al. · 2022 [cited by applicant]
US 20230121460A1 · Banerjee et al. · 2023 [cited by applicant]
US 20230325118A1 · Hasti et al. · 2023 [cited by applicant]
US 20240028255A1 · Hasti et al. · 2024 [cited by applicant]
US 20240036770A1 · Hasti et al. · 2024 [cited by applicant]
US 20240419362A1 · Hasti et al. · 2024 [cited by applicant]
EP 3792760A1 · 2021 [cited by applicant]
KR 20200108228A · 2020 [cited by applicant]
Alibaba Cloud; Distributed Cloud Container Platform for Kubernetes (ACK One); 8 Pgs. [cited by applicant]
Google; GKE cluster architecture; 5 Pgs. [cited by applicant]
Notice of Allowance mailed on May 7, 2024 for U.S. Appl. No. 18/487,366, filed Oct. 16, 2023, 08 pages. [cited by applicant]
Co-pending U.S. Appl. No. 17/718,382, inventors Praveen; Kumar Hasti et al., filed Apr. 12, 2022. [cited by applicant]
Co-pending U.S. Appl. No. 17/718,395, inventors Praveen; Kumar Hasti et al., filed Apr. 12, 2022. [cited by applicant]
Co-pending U.S. Appl. No. 17/718,403, inventors Praveen; Kumar Hasti et al., filed Apr. 12, 2022. [cited by applicant]
Notice of Allowance mailed on Apr. 14, 2023 for U.S. Appl. No. 17/718,403, filed Apr. 12, 2022, 7 pages. [cited by applicant]
International Search Report cited in PCT/US2023/065353, mailed Jul. 13, 2023, 14 pages. [cited by applicant]
Notice of Allowance mailed on Jun. 7, 2023 for U.S. Appl. No. 17/718,403, filed Apr. 12, 2022, 7 pages. [cited by applicant]
Notice of Allowance mailed on May 24, 2023 for U.S. Appl. No. 17/718,395, filed Apr. 12, 2022, 8 pages. [cited by applicant]
U.S. Appl. No. 17/718,395, filed Apr. 12, 2022, Hasti et al. [cited by applicant]
U.S. Appl. No. 17/718,403, filed Apr. 12, 2022, Hasti et al. [cited by applicant]
U.S. Appl. No. 17/718,382, filed Apr. 12, 2022, Hasti et al. [cited by applicant]
Notice of Allowance mailed Aug. 14, 2023 for U.S. Appl. No. 17/718,395, filed Apr. 12, 2022, 05 pages. [cited by applicant]
Notice of Allowance mailed Aug. 18, 2023 for U.S. Appl. No. 17/718,403, filed Apr. 12, 2022, 04 pages. [cited by applicant]
Notice of Allowance malled on Aug. 30, 2023 for U.S. Appl. No. 17/718,395, filed Apr. 12, 2022, 5 pages. [cited by applicant]
Notice of Allowance mailed on Aug. 2, 2023 for U.S. Appl. No. 17/718,403, filed Apr. 12, 2022, 04 pages. [cited by applicant]
“Google Kubernetes Engine: Ultimate Quick Start Guide”, Jun. 2021, Cloud Central, Kubernetes Storage, Reprinted from the Internet at: https://cloud.netapp.com/blog/gcp-cvo-blg-google-kubernetes-engine-ultimate-quick-sta… [cited by applicant]
“Kubernetes Autoscaling: 3 Methods and How to Make Them Great”, Jan. 2022, Spot by NetApp, Reprinted from the Internet at: https://spot.io/resources/kubernetes-autoscaling-3-methods-and-how-to-make-them-great/, 10 pgs. [cited by applicant]
“Use Cases”, 2022, NetApp Cloud Volumes Service for Google Cloud, Cloud Architecture Center, Reprinted from the Internet at: https://cloud.google.com/architecture/partners/netapp-cloud-volumes/use-cases#:˜:text=Cloud%20… [cited by applicant]
“Using the Kubernetes Vertical Pod Autoscaler”, Oracle cloud Infrastructure Documentation, Reprinted from the Internet at: https://docs.oracle.com/en-us/iaas/Content/ContEng/Tasks/contengusingverticalpodautoscaler.htm, … [cited by applicant]
“Vertical Autoscaling in Kubernetes”, May 2021, Pua Abbassi, Giant Swarm, Reprinted from the Internet at: https://docs.oracle.com/en-us/iaas/Content/ContEng/Tasks/contengusingverticalpodautoscaler.htm, 10 pgs. [cited by applicant]
“Vertical Pod Autoscaler”, Amazon EKS User Guide, Reprinted from the Internet at: https://docs.aws.amazon.com/eks/latest/userguide/vertical-pod-autoscaler.html, 5 pgs. [cited by applicant]
“Vertical Pod Autoscaling”, Kubernetes Engine Documentation, Google Cloud, Reprinted from the Internet at: https://docs.aws.amazon.com/eks/latest/userguide/vertical-pod-autoscaler.html, 9 pgs. [cited by applicant]
“What Can NetApp's Cloud Volumes Services Do?”, Volta, Reprinted from the Internet at: https://voltainc.com/cloud-volumes-services/, 3 pgs. [cited by applicant]
International Preliminary Report on Patentability for Application No. PCT/US2023/065353 mailed on Oct. 24, 2024, 09 pages. [cited by applicant]
Non-Final Office Action mailed on Feb. 12, 2025 for U.S. Appl. No. 18/479,195, filed Oct. 2, 2023, 8 pages. [cited by applicant]
Notice of Allowance mailed on Apr. 15, 2025 for U.S. Appl. No. 18/479,195, filed Oct. 2, 2023, 08 pages. [cited by applicant]
Cited By (2)
US 12,717,521 US 12,743,319