IP Library Granted Patent US 8,862,727
Granted Patent B2
US 8,862,727 · App. 13/470,589 · Granted Oct 14, 2014

Problem determination and diagnosis in shared dynamic clouds

Inventors: Praveen Jayachandran (Bangalore, IN); Bikash Sharma (State College, PA); Akshat Verma (New Delhi, IN)
Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,862,727
App. No.
13/470,589
Granted
Oct 14, 2014
Kind
B2
Abstract

An apparatus and an article of manufacture for problem determination and diagnosis in a shared dynamic cloud environment include monitoring each virtual machine and physical server in the shared dynamic cloud environment for at least one metric, identifying a symptom of a problem and generating an event based on said monitoring, analyzing the event to determine a deviation from normal behavior, and classifying the event as a cloud-based anomaly or an application fault based on existing knowledge.

Claims (34)

1. An article of manufacture comprising a computer readable storage device having computer readable instructions for problem determination and diagnosis in a shared dynamic cloud environment during live virtual machine migration tangibly embodied thereon which, when implemented, cause a computer to carry out a plurality of method steps comprising:

monitoring each virtual machine and each physical server in a shared dynamic cloud environment for at least one metric;

identifying a symptom of a problem within the shared dynamic cloud environment and generating an event based on said monitoring, wherein said event corresponds to said symptom;

analyzing the event to determine a deviation from normal behavior, wherein said normal behavior is determined based on said monitoring; and

classifying the event as a cloud-based anomaly or an application fault based on a comparison of the event with multiple fault signatures, wherein said multiple fault signatures capture a set of deviations from normal behavior associated with (i) one or more cloud-based anomalies and (ii) one or more application faults, and wherein said set of deviations comprises information pertaining to a given virtual machine and a physical server hosting the given virtual machine in the shared dynamic cloud environment, and wherein said classifying comprises:

matching the deviation from normal behavior associated with the event with one of the one or more fault signatures based on the at least one monitored metric monitored for each virtual machine and for each physical server corresponding to said event.

2. The article of manufacture of claim 1 , wherein said monitoring comprises outputting a stream of data-points, each data-point being derived from a monitoring data time series, corresponding to each system and application metric in the shared dynamic cloud environment.

3. The article of manufacture of claim 1 , wherein the method steps comprise:

building a model of normal application behavior by leveraging a machine learning technique; and

using the model to detect a deviation in the monitored data from normal behavior.

4. The article of manufacture of claim 1 , wherein the at least one metric comprises a metric pertaining to at least one of central processing unit, memory, cache, network resources and disk resources.

5. The article of manufacture of claim 1 , wherein said identifying comprises identifying a trend from time series data.

6. The article of manufacture of claim 1 , wherein said analyzing comprises using statistical correlation across virtual machines and resources to pinpoint the location of the deviation to an affected resource and virtual machine.

7. A system for problem determination and diagnosis in a shared dynamic cloud environment during live virtual machine migration, comprising:

at least one distinct software module, each distinct software module being embodied on a tangible computer-readable medium;

a memory; and

at least one processor coupled to the memory and operative for:

monitoring each virtual machine and each physical server in a shared dynamic cloud environment for at least one metric;

identifying a symptom of a problem within the shared dynamic cloud environment and generating an event based on said monitoring, wherein said event corresponds to said symptom;

analyzing the event to determine a deviation from normal behavior, wherein said normal behavior is determined based on said monitoring; and

classifying the event as a cloud-based anomaly or an application fault based on a comparison of the event with multiple fault signatures, wherein said multiple fault signatures capture a set of deviations from normal behavior associated with (i) one or more cloud-based anomalies and (ii) one or more application faults, and wherein said set of deviations comprises information pertaining to a given virtual machine and a physical server hosting the given virtual machine in the shared dynamic cloud environment, and wherein said classifying comprises:

matching the deviation from normal behavior associated with the event with one of the one or more fault signatures based on the at least one monitored metric monitored for each virtual machine and for each physical server corresponding to said event.

8. The system of claim 7 , wherein the at least one metric comprises a metric pertaining to at least one of central processing unit, memory, cache, network resources and disk resources.

9. The system of claim 7 , wherein said identifying comprises identifying a trend from time series data.

10. The system of claim 7 , wherein said analyzing comprises using statistical correlation across virtual machines and resources to pinpoint the location of the deviation to an affected resource and virtual machine.

11. A system for problem determination and diagnosis in a shared dynamic cloud environment during live virtual machine migration, comprising:

a memory;

at least one processor coupled to the memory; and

at least one distinct software module, each distinct software module being embodied on a tangible computer-readable medium, the at least one distinct software module comprising:

a monitoring engine module, executing on the processor, for monitoring each virtual machine and each physical server in a shared dynamic cloud environment for at least one metric and outputting a monitoring data time-series corresponding to each metric;

an event generation engine module, executing on the processor, for identifying a symptom of a problem within the shared dynamic cloud environment and generating an event based on said monitoring, wherein said event corresponds to said symptom;

a problem determination engine module, executing on the processor, for analyzing the event to determine and locate a deviation from normal behavior, wherein said normal behavior is determined based on said monitoring; and

a diagnosis engine module, executing on the processor, for classifying the event as a cloud-based anomaly or an application fault based on a comparison of the event with multiple fault signatures, wherein said multiple fault signatures capture a set of deviations from normal behavior associated with (i) one or more cloud-based anomalies and (ii) one or more application faults, and wherein said set of deviations comprises information pertaining to a given virtual machine and a physical server hosting the given virtual machine in the shared dynamic cloud environment, and wherein said classifying comprises:

matching the deviation from normal behavior associated with the event with one of the one or more fault signatures based on the at least one monitored metric monitored for each virtual machine and for each physical server corresponding to said event.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2017
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: SINOEAST CONCEPT LIMITED
Reel/Frame 041388/0557 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2012
From: JAYACHANDRAN, PRAVEEN; SHARMA, BIKASH; VERMA, AKSHAT
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 028203/0623 →
Continuity (1)
Related Publication 20130305092A1 · Nov 14, 2013