IP Library Granted Patent US 10,133,619
Granted Patent B1
US 10,133,619 · App. 15/174,977 · Granted Nov 20, 2018

Cluster-wide virtual machine health monitoring

Inventors: Abhinay Nagpal (San Jose, CA); Alexander J. Kaufmann (San Jose, CA); Himanshu Shukla (San Jose, CA); Jason Sims (Redwood City, CA); Varun Kumar Arora (Mountain View, CA); Venkata Vamsi Krishna Kothuri (San Jose, CA)
Assignee: Nutanix, Inc.
G06F11/079G06F9/45558G06F11/0709G06F11/0712G06F11/0751G06F11/0787G06F2009/45575
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,133,619
App. No.
15/174,977
Filed
Jun 6, 2016
Granted
Nov 20, 2018
Kind
B1
Art Unit
2113
USPC
714/37
Abstract

Systems for self-configuring health monitoring instrumentation for clustered storage platforms. Master and slave health modules implement a health monitoring system in a clustered virtualization environment comprising a plurality of nodes of the cluster with an installed health module instance running on the nodes. The health module system may gather and analyze data on a node level and at a cluster level to manage the cluster. The cluster health module system observes I/O commands issued to, and I/O command responses returned from, a common storage pool. Health data is stored in the storage pool.

Claims (37)

1. A computer-implemented method for managing a clustered virtualization environment, comprising:

identifying a cluster of computing nodes that are interconnected by at least one communication path, wherein at least two of the computing nodes are further interconnected to a storage pool comprising at least one node-local storage device and at least one networked storage device;

invoking, on a first node, a first health module comprising at least one first node data collection unit that accesses the storage pool;

invoking, on a second node, a second health module comprising at least one second data collection unit that accesses the storage pool;

storing, in the storage pool, a first set of collected data generated by the first health module, the first set of collected data tracking a health status for the first node or the cluster of computing nodes by receiving observations taken at the first node; and

storing, in the storage pool, a second set of collected data generated by the second health module, the second set of collected data tracking a health status for the second node or the cluster of computing nodes by receiving observations taken at the second node.

2. The method of claim 1 , further comprising performing an analysis, by the second node, of the first set of collected data to determine a health status of the first node.

3. The method of claim 1 , further comprising detecting, by the second node, an interruption of service of the first node and invoking an instance of a new health module.

4. The method of claim 3 , wherein the new health module accesses the storage pool and reads at least a portion of the first set of collected data gathered by the first health module.

5. The method of claim 1 , further comprising issuing one or more alerts based on results of an analysis of at least a portion of the collected data.

6. The method of claim 1 , further comprising logging an occurrence of an alert based at least in part on one or more observed events.

7. The method of claim 6 , further comprising updating an alert database based at least in part on the one or more observed events.

8. The method of claim 6 , further comprising initiating a corrective action based at least in part on the one or more observed events.

9. The method of claim 6 , further comprising modifying a data collection schedule.

10. The method of claim 1 , further comprising determining a version identifier of the first node health module and replacing the first node health module with a new instance of a health module having a newer version.

11. A computer readable medium, embodied in a non-transitory computer readable medium, the non-transitory computer readable medium having stored thereon a sequence of instructions which, when stored in memory and executed by a processor causes the processor to perform a set of acts for managing a clustered virtualization environment, the acts comprising:

identifying a cluster of computing nodes that are interconnected by at least one communication path, wherein at least two of the computing nodes are further interconnected to a storage pool comprising at least one node-local storage device and at least one networked storage device;

invoking, on a first node, a first health module comprising at least one first node data collection unit that accesses the storage pool;

invoking, on a second node, a second health module comprising at least one second data collection unit that accesses the storage pool;

storing, in the storage pool, a first set of collected data generated by the first health module, the first set of collected data tracking a health status for the first node or the cluster of computing nodes by receiving observations taken at the first node; and

storing, in the storage pool, a second set of collected data generated by the second health module, the second set of collected data tracking a health status for the second node or the cluster of computing nodes by receiving observations taken at the second node.

12. The computer readable medium of claim 11 , further comprising instructions which, when stored in memory and executed by the processor causes the processor to perform acts of performing an analysis, by the second node, of the first set of collected data to determine a health status of the first node.

13. The computer readable medium of claim 11 , further comprising instructions which, when stored in memory and executed by the processor causes the processor to perform acts of detecting, by the second node, an interruption of service of the first node and invoking an instance of a new health module.

14. The computer readable medium of claim 13 , wherein the new health module accesses the storage pool and reads at least a portion of the first set of collected data gathered by the first health module.

15. The computer readable medium of claim 11 , further comprising instructions which, when stored in memory and executed by the processor causes the processor to perform acts of issuing one or more alerts based on results of an analysis of at least a portion of the collected data.

16. The computer readable medium of claim 11 , further comprising instructions which, when stored in memory and executed by the processor causes the processor to perform acts of logging an occurrence of an alert based at least in part on one or more observed events.

17. The computer readable medium of claim 16 , further comprising instructions which, when stored in memory and executed by the processor causes the processor to perform acts of updating an alert database based at least in part on the one or more observed events.

18. The computer readable medium of claim 11 , further comprising instructions which, when stored in memory and executed by the processor causes the processor to perform acts of determining a version identifier of the first node health module and replacing the first node health module with a new instance of a health module having a newer version.

19. A system for managing a clustered virtualization environment comprising:

a storage medium having stored thereon a sequence of instructions; and

a processor or processors that execute the instructions to cause the processor or processors to perform a set of acts, the acts comprising,

identifying a cluster of computing nodes that are interconnected by at least one communication path, wherein at least two of the computing nodes are further interconnected to a storage pool comprising at least one node-local storage device and at least one networked storage device;

invoking, on a first node, a first health module comprising at least one first node data collection unit that accesses the storage pool;

invoking, on a second node, a second health module comprising at least one second data collection unit that accesses the storage pool;

storing, in the storage pool, a first set of collected data generated by the first health module, the first set of collected data tracking a health status for the first node or the cluster of computing nodes by receiving observations taken at the first node; and

storing, in the storage pool, a second set of collected data generated by the second health module, the second set of collected data tracking a health status for the second node or the cluster of computing nodes by receiving observations taken at the second node.

20. The system of claim 19 , wherein a new health module accesses the storage pool and reads at least a portion of the first set of collected data gathered by the first health module.

Assignments (2)
SECURITY INTEREST Recorded Feb 13, 2025
From: NUTANIX, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 070206/0463 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 19, 2018
From: NAGPAL, ABHINAY; KAUFMANN, ALEXANDER J.; SHUKLA, HIMANSHU; SIMS, JASON; ARORA, VARUN KUMAR; KOTHURI, VENKATA VAMSI KRISHNA
To: NUTANIX, INC.
Reel/Frame 046400/0378 →
Continuity (1)
Provisional Application 62172738 · Jun 8, 2015
Cited By (14)
US 12,190,140 US 12,197,398 US 12,242,455 US 12,248,434 US 12,248,435 US 12,248,709 US 12,314,753 US 12,367,072 US 12,367,108 US 12,461,834 US 12,481,531 US 12,513,221 US 12,517,874 US 12,627,681