IP Library › Granted Patent US 10,318,392
Granted Patent B2
US 10,318,392 · App. 15/517,031 · Granted Jun 11, 2019

Management system for virtual machine failure detection and recovery

Inventors: Lei Sun (Tokyo, JP); Shinya Miyakawa (Tokyo, JP); Masaki Kan (Tokyo, JP); Jun Suzuki (Tokyo, JP); Yuki Hayashi (Tokyo, JP)
Assignee: NEC CORPORATION
G06F11/203G06F9/46G06F11/1484G06F11/30G06F13/10G06F11/202G06F11/3409G06F2201/815
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,318,392
App. No.
15/517,031
Granted
Jun 11, 2019
Kind
B2
Abstract

A Management system 10 includes: resource pools 11 1 - 11 4 which act as the hardware components on which multiple virtual machines are running; an inter-connecting network 12 which connects various resource pools; and a HA manager 13 which snoops all traffic of the inter-connecting network 12 to detect failure of a target VM and triggers corresponded actions when failure is detected.

Claims (54)

1. A management system for detecting failure of virtual machines and triggering corresponded actions when failure is found in resource disaggregation data center architecture, the management system comprising:

resource pools which act as hardware components on which multiple virtual machines are running;

an inter-connecting network which connects various resource pools; and

a high availability (HA) manager which snoops all traffic of the inter-connecting network to detect failure of a target virtual machine (VM) and triggers corresponded actions when failure is detected,

wherein the HA manager comprises:

a snooping module which snoops all traffic of the inter-connecting network;

a packet parsing module which parses a snooped packet and extracts info from header and payload of a parsed packet;

a VM-manager which provides basic operation of VMs on a computing device which is connected to the VM-manager;

an action module which sends predefined commands to a local VM-manager; and

an HA-database (DB) which stores records of all target VMs,

wherein the packet parsing module determines whether a heartbeat message from a VM arrives on time or delayed or missing, determines whether there is I/O traffic from a VM or not, and determines whether current status is follow predefined normal patterns or not,

wherein the VM-manager starts a new instance of a specific VM, and gathers further info of a specific VM, such system resource utility and system availability, and

wherein the action module starts a new instance of a target VM when a VM is believed unavailable, and sends diagnosis command to gather more information of the target VM.

2. A management system according to claim 1 , wherein the records of all target VMs stored in the HA-DB comprise:

Node_Id: ID of a CPU pool;

VM_Id: ID of the VM;

Device_Id: ID of a device;

Image_d: ID of an image used by VM;

NW_Address: network address used by VM;

NW_Id: ID of a network used by VM;

Heartbeat_state: a state of a heartbeat message;

Traffic_state: a state of I/O traffic;

Heartbeat_timeout: a default value of heartbeat timeout;

Traffic_timeout: a default value of I/O traffic timeout,

wherein the NW_Address may be MAC address when ExpEther is used,

wherein the Heartbeat_state may be either healthy or delayed, and

wherein the Traffic_state may be either healthy or delayed.

3. A management system according to claim 2 , wherein the packet parsing module determines whether the heartbeat timeout expires or not, and determines whether the I/O traffic timeout expires or not.

4. A management system according to claim 3 , wherein the action module just updates a corresponded timer if there is neither heartbeat timeout nor I/O traffic timeout;

wherein the action module requires system resource info for further diagnosis if I/O traffic timeout occurs; and

wherein the action module triggers a recovery action if both heartbeat timeout and I/O traffic timeout occur.

5. A management system according to claim 2 , wherein the packet parsing module extracts corresponded info from a heartbeat message, extracts corresponded info from a normal I/O traffic message, and extracts corresponded info from diagnosis information.

6. A management system according to claim 1 , wherein the packet parsing module extracts corresponded info from a heartbeat message, extracts corresponded info from a normal I/O traffic message, and extracts corresponded info from diagnosis information.

7. A management method executed in a device included in a virtualization system including resource pools acting as hardware components on which multiple virtual machines are running and an inter-connecting network connecting various resource pools for detecting failure of virtual machines and triggering corresponded actions when failure is found in resource disaggregation data center architecture, the management method comprising:

snooping all traffic of the inter-connecting network;

parsing a snooped packet;

extracting info from header and payload of a parsed packet;

determining whether a heartbeat message from a virtual machine (VM) arrives on time or delayed or missing;

determining whether there is I/O traffic from a VM or not;

determining whether current status is follow predefined normal patterns or not;

providing basic operation of VMs on a computing device which is connected to the inter-connecting network for starting a new instance of a specific VM;

gathering further info of a specific VM, such system resource utility and system availability;

sending predefined commands to a local VM-manager for starting a new instance of a target VM when a VM is believed unavailable; and

sending diagnosis command to gather more information of the target VM.

8. A non-transitory computer-readable recording medium having recorded therein a management program for detecting failure of virtual machines and triggering corresponded actions when failure is found in resource disaggregation data center architecture, the management program causing a computer included in a virtualization system including resource pools acting as hardware components on which multiple virtual machines are running and an inter-connecting network connecting various resource pools, to execute:

a snooping process of snooping all traffic of the inter-connecting network;

a parsing process of parsing a snooped packet;

an extracting process of extracting info from header and payload of a parsed packet;

a determining process of determining whether a heartbeat message from a VM arrives on time or delayed or missing; a determining process of determining whether there is I/O traffic from a VM or not;

a determining process of determining whether current status is follow predefined normal patterns or not;

a providing process of providing basic operation of VMs on a computing device which is connected to the inter-connecting network for starting a new instance of a specific VM;

a gathering process of gathering further info of a specific VM, such system resource utility and system availability;

a sending process of sending predefined commands to a local VM-manager for starting a new instance of a target VM when a VM is believed unavailable; and

a sending process of sending diagnosis command to gather more information of the target VM.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2017
From: SUN, LEI; MIYAKAWA, SHINYA; KAN, MASAKI; SUZUKI, JUN; HAYASHI, YUKI
To: NEC CORPORATION
Reel/Frame 042075/0672 →
Continuity (1)
Related Publication 20170293537A1 · Oct 12, 2017