IP Library Granted Patent US 8,230,262
Granted Patent B2
US 8,230,262 · App. 12/830,175 · Granted Jul 24, 2012

Method and apparatus for dealing with accumulative behavior of some system observations in a time series for Bayesian inference with a static Bayesian network model

Assignee: Oracle International Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,230,262
App. No.
12/830,175
Granted
Jul 24, 2012
Kind
B2
Abstract

A method and apparatus are provided for determining the probability that one or more problems have occurred within a complex multi-host system. A probabilistic model representing the cause/effect relationships among potential system problems identifies the probability that a problem occurred in the system based at least on system measure states that are input into the probabilistic model. System measure states may be determined based on an aggregation of system measurement values taken periodically. Aggregating system measurement values may be performed over system measurement values that were taken during a recent time interval. A rolling count aggregation function may be used for this purpose. A rolling count function counts the number of system measurement values taken within the recent time interval that lie within a particular range of values. A system measure state may be determined based on whether the rolling count exceeds a threshold associated with the system measure.

Claims (80)

1. A method comprising steps of:

receiving a set of system measurement values periodically over time, wherein each system measurement value of the set of system measurement values is a value for a particular system measure of a set of system measures;

wherein each system measurement value of the set of system measurement values is taken at a measurement time point within a time interval associated with the particular system measure;

wherein the time interval includes a plurality of recent successive measurement time points;

identifying a particular set of system measurement values that were taken at measurement time points included in the time interval;

determining a count of system measurement values of the particular set of system measurement values, wherein a system measurement value is counted if said system measurement value exceeds a threshold value

wherein the particular system measure is associated with the threshold value and a number of measurement time points included in the time interval;

determining a system measure state for the particular system measure based on the count of system measurement values, wherein the system measure state of the particular system measure is provided to a probabilistic model, wherein the probabilistic model determines a probability of a failure value based on the system measure state of the particular system measure in a complex multi-host system;

wherein each step is performed by one or more computing devices.

2. The method of claim 1 , wherein a first count is determined for a first set of measurement values taken within a first time interval for a system measure, wherein the steps further include:

determining a first system measure state for the system measure based on the first count;

a probabilistic model determining a probability of failure value based on the first system measure state;

receiving a second set of system measurement values periodically over time, wherein each system measurement value of the second set of system measurement values is a value for said system measure;

wherein each system measurement value of the second set of system measurement values is taken at a measurement time point within a second time interval associated with said system measure;

determining a second count for the second set of measurement values taken within the second time interval for said system measure;

determining a second system measure state for said system measure based on the second count;

in response to determining that the second system measure state is different from the first system measure state, the probabilistic model determining a second probability of failure value based on the second system measure state.

3. The method of claim 2 , wherein there is at least one measurement time point that is included in both the first time interval and the second time interval, wherein the oldest time point included in the first time interval is not included in the second time interval, and the newest time point in the second interval is not included in the first time interval; and

wherein there is a plurality of system measurement values that correspond to the plurality of measurement time points, wherein each system measurement value of the plurality of system measurement values corresponds to a measurement time point of the plurality of measurement time points; and

wherein both the first system measure state and the second system measure state are based at least in part on a count of system measurement values in the plurality of system measurement values, wherein a system measurement value in the plurality of system measurement values is counted if it exceeds a threshold value associated with said system measure.

4. The method of claim 1 , wherein:

the probabilistic model is a Bayesian network;

an observation node of the Bayesian network performs determining the count of system measurement values of the particular set of system measurement values, wherein the system measurement value is counted if it exceeds the threshold value; and

wherein the steps further include determining the system measure state for the particular system measure is performed by the observation node mapping the count of system measurement values to the system measure state for the particular system measure.

5. The method of claim 4 , wherein

the system measure state is provided as input into a root cause node in the Bayesian network; and

the root cause node maps the system measure state to a probability of failure value associated with the root cause node.

6. A non-transitory computer-readable storage medium storing one or more sequences of instructions, said one or more sequences of instructions, which, when executed by one or more processors, causes the one or more processors to perform steps of:

receiving a set of system measurement values periodically over time, wherein each system measurement value of the set of system measurement values is a value for a particular system measure of a set of system measures;

wherein each system measurement value of the set of system measurement values is taken at a measurement time point within a time interval associated with the particular system measure;

wherein the time interval includes a plurality of recent successive measurement time points;

identifying a particular set of system measurement values that were taken at measurement time points included in the time interval;

determining a count of system measurement values of the particular set of system measurement values, wherein a system measurement value is counted if said system measurement value exceeds a threshold value

wherein the particular system measure is associated with the threshold value and a number of measurement time points included in the time interval;

determining a system measure state for the particular system measure based on the count of system measurement values, wherein the system measure state of the particular system measure is provided to a probabilistic model, wherein the probabilistic model determines a probability of a failure value based on the system measure state of the particular system measure in a complex multi-host system;

wherein each step is performed by one or more computing devices.

7. The non-transitory computer-readable storage medium of claim 6 , wherein a first count is determined for a first set of measurement values taken within a first time interval for a system measure, wherein the steps further include:

determining a first system measure state for the system measure based on the first count;

a probabilistic model determining a probability of failure value based on the first system measure state;

receiving a second set of system measurement values periodically over time, wherein each system measurement value of the second set of system measurement values is a value for said system measure;

wherein each system measurement value of the second set of system measurement values is taken at a measurement time point within a second time interval associated with said system measure;

determining a second count for the second set of measurement values taken within the second time interval for said system measure;

determining a second system measure state for said system measure based on the second count;

in response to determining that the second system measure state is different from the first system measure state, the probabilistic model determining a second probability of failure value based on the second system measure state.

8. The non-transitory computer-readable storage medium of claim 7 , wherein there is at least one measurement time point that is included in both the first time interval and the second time interval, wherein the oldest time point included in the first time interval is not included in the second time interval, and the newest time point in the second interval is not included in the first time interval; and

wherein there is a plurality of system measurement values that correspond to the plurality of measurement time points, wherein each system measurement value of the plurality of system measurement values corresponds to a measurement time point of the plurality of measurement time points; and

wherein both the first system measure state and the second system measure state are based at least in part on a count of system measurement values in the plurality of system measurement values, wherein a system measurement value in the plurality of system measurement values is counted if it exceeds a threshold value associated with said system measure.

9. The non-transitory computer-readable storage medium of claim 6 , wherein:

the probabilistic model is a Bayesian network;

an observation node of the Bayesian network performs determining the count of system measurement values of the particular set of system measurement values, wherein the system measurement value is counted if it exceeds the threshold value; and

wherein the steps further include determining the system measure state for the particular system measure is performed by the observation node mapping the count of system measurement values to the system measure state for the particular system measure.

10. The non-transitory computer-readable storage medium of claim 9 , wherein the system measure state is provided as input into a root cause node in the Bayesian network; and

the root cause node maps the system measure state to a probability of failure value associated with the root cause node.

11. A multi-host computing system, wherein said multi-host computing system is configured to:

to generate a set of system measurement values periodically over time, wherein each system measurement value of the set of system measurement values is a value for particular system measure of a set of system measures;

wherein each system measurement value of the set of system measurement values is taken at a measurement time point within a time interval associated with the particular system measure;

wherein the time interval includes a plurality of recent successive measurement time points;

to identify a particular set of system measurement values that were taken at measurement time points included in the time interval;

to determine a count of system measurement values of the particular set of system measurement values, wherein a system measurement value is counted if said system measurement value exceeds a threshold value;

wherein the particular system measure is associated with the threshold value and a number of measurement time points included in the time interval; and

to determine a system measure state for the particular system measure based on the count of system measurement values, wherein the system measure state of the particular system measure is provided to a probabilistic model, wherein the probabilistic model determines a probability of a failure value based on the system measure state of the particular system measure in a complex multi-host system.

12. The multi-host computing system of claim 11 , the multi-host computing system being further configured to:

to determine a first system measure state for the system measure based on the first count;

to use a probabilistic model to determine a probability of failure value based on the first system measure state;

to receive a second set of system measurement values periodically over time, wherein each system measurement value of the second set of system measurement values is a value for said system measure;

wherein each system measurement value of the second set of system measurement values is taken at a measurement time point within a second time interval associated with said system measure;

to determine a second count for the second set of measurement values taken within the second time interval for said system measure;

to determine a second system measure state for said system measure based on the second count; and

to use the probabilistic model to determine a second probability of failure value based on the second system measure state, in response to determining that the second system measure state is different from the first system measure state.

13. The multi-host computing system of claim 11 ,

wherein the oldest time point included in the first time interval is not included in the second time interval, and the newest time point in the second interval is not included in the first time interval; and

wherein there is a plurality of system measurement values that correspond to the plurality of measurement time points, wherein each system measurement value of the plurality of system measurement values corresponds to a measurement time point of the plurality of measurement time points; and

wherein both the first system measure state and the second system measure state are based at least in part on a count of system measurement values in the plurality of system measurement values, wherein a system measurement value in the plurality of system measurement values is counted if it exceeds a threshold value associated with said system measure.

14. The multi-host computing system of claim 11 , wherein:

the probabilistic model is a Bayesian network;

an observation node of the Bayesian network is configured to determine the count of system measurement values of the particular set of system measurement values, wherein the system measurement value is counted if it exceeds the threshold value; and

wherein the observation node is configured to map the count of system measurement values to the system measure state for the particular system measure to determine the system measure state for the particular system measure.

15. The multi-host computing system of claim 11 , wherein

the system measure state is provided as input into a root cause node in the Bayesian network; and

the root cause node maps the system measure state to a probability of failure value associated with the root cause node.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2010
From: LI, FULU; BEG, MOHSIN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 024655/0486 →
Continuity (1)
Related Publication 20120005534A1 · Jan 5, 2012