IP Library Granted Patent US 7,870,440
Granted Patent B2
US 7,870,440 · App. 12/048,922 · Granted Jan 11, 2011

Method and apparatus for detecting multiple anomalies in a cluster of components

Assignee: Oracle America, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,870,440
App. No.
12/048,922
Granted
Jan 11, 2011
Kind
B2
Abstract

A system that detects multiple anomalies in a cluster of components is presented. During operation, the system monitors derivatives obtained from one or more inferential variables which are received from sensors in the cluster of components. The system then determines whether one or more components within the cluster have experienced an anomalous event based on the monitored derivatives. If so, the system performs one or more remedial actions.

Claims (219)

1. A method for detecting multiple anomalies in a cluster of components, comprising:

monitoring one or more inferential variables that are received from sensors in the cluster of components;

obtaining derivatives from the one or more inferential variables using a moving-window numerical derivative technique to calculate the rate-of-change in the one or more inferential variables over a specified time interval;

monitoring the derivatives obtained from the one or more inferential variables;

determining whether one or more components within the cluster have experienced an anomalous event based on the monitored derivatives; and

if so, performing one or more remedial actions.

2. The method of claim 1 , wherein determining whether one or more components within the cluster have experienced an anomalous event involves:

determining whether the monitored derivatives indicate that the value of the one or more inferential variables is changing with a specified polarity; and

if so, determining that a component in the cluster of components has experienced an anomalous event.

3. The method of claim 1 , wherein monitoring derivatives for one or more inferential variables involves using a Sequential Probability Ratio Test (SPRT).

4. The method of claim 1 , wherein performing the one or more remedial actions involves one or more of:

generating warnings;

replacing failed components;

reporting the actual number of components that have failed;

reporting the probability that a specified number of components have failed;

monitoring additional variables monitored by sensors in the cluster of components;

adjusting the frequency at which the one or more variables are polled by the sensors;

adjusting test conditions during an accelerated-life study of the components;

scheduling the failed components to be replaced at the next scheduled maintenance interval; and

estimating the remaining useful life of the components.

5. The method of claim 1 , wherein the one or more inferential variables received from the sensors include one or more of:

hardware variables; and

software variables.

6. The method of claim 5 , wherein the hardware variables include one or more of:

voltage;

current;

temperature;

vibration;

optical power;

optical wavelength;

air velocity;

measures of signal integrity; and

fan speed.

7. The method of claim 5 , wherein the software variables include one or more of:

throughput;

transaction latencies;

queue lengths;

central processing unit load;

memory load;

cache load;

I/O traffic;

bus saturation metrics;

FIFO overflow statistics;

network traffic; and

disk-related metrics.

8. A computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method for detecting multiple anomalies in a cluster of components, wherein the method comprises:

monitoring one or more inferential variables that are received from sensors in the cluster of components;

obtaining derivatives from the one or more inferential variables using a moving-window numerical derivative technique to calculate the rate-of-change in the one or more inferential variables over a specified time interval;

monitoring the derivatives obtained from the one or more inferential variables;

determining whether one or more components within the cluster have experienced an anomalous event based on the monitored derivatives; and

if so, performing one or more remedial actions.

9. The computer-readable storage medium of claim 8 , wherein determining whether one or more components within the cluster have experienced an anomalous event involves:

determining whether the monitored derivatives indicate that the value of the one or more inferential variables is changing with a specified polarity; and

if so, determining that a component in the cluster of components has experienced an anomalous event.

10. The computer-readable storage medium of claim 8 , wherein monitoring derivatives for one or more inferential variables involves using a Sequential Probability Ratio Test (SPRT).

11. The computer-readable storage medium of claim 8 , wherein performing the one or more remedial actions involves one or more of:

generating warnings;

replacing failed components;

reporting the actual number of components that have failed;

reporting the probability that a specified number of components have failed;

monitoring additional variables monitored by sensors in the cluster of components;

adjusting the frequency at which the one or more variables are polled by the sensors;

adjusting test conditions during an accelerated-life study of the components;

scheduling the failed components to be replaced at the next scheduled maintenance interval; and

estimating the remaining useful life of the components.

12. The computer-readable storage medium of claim 8 , wherein the one or more inferential variables received from the sensors include one or more of:

hardware variables; and

software variables.

13. The computer-readable storage medium of claim 12 , wherein the hardware variables include one or more of:

voltage;

current;

temperature;

vibration;

optical power;

optical wavelength;

air velocity;

measures of signal integrity; and

fan speed.

14. The computer-readable storage medium of claim 12 , wherein the software variables include one or more of:

throughput;

transaction latencies;

queue lengths;

central processing unit load;

memory load;

cache load;

I/O traffic;

bus saturation metrics;

FIFO overflow statistics;

network traffic; and

disk-related metrics.

15. An apparatus that detects multiple anomalies in a cluster of components, comprising:

a monitoring mechanism configured to:

monitor one or more inferential variables that are received from sensors in the cluster of components;

obtain derivatives from the one or more inferential variables using a moving-window numerical derivative technique to calculate the rate-of-change in the one or more inferential variables over a specified time interval; and

monitor the derivatives obtained from the one or more inferential variables;

an analysis mechanism configured to determine whether one or more components within the cluster have experienced an anomalous event based on the monitored derivatives; and

a remedial action mechanism, wherein if the analysis mechanism determines that one or more components within the cluster has failed, the remedial action mechanism is configured to perform one or more remedial actions.

16. A method for detecting multiple anomalies in a cluster of components, comprising:

monitoring derivatives obtained from one or more inferential variables which are received from sensors in the cluster of components;

determining whether one or more specified events occurred within the cluster of components based on the monitored derivatives; and

if so, determining the probability that the cluster of components is in a specified state based on the one or more specified events that occurred by

determining the number of first events x and second events y that occurred; and

determining the probability p i,j that the cluster of components is in a specified state |i,j> as:

p

i

,

j

=

a

i

,

j

2

k

,

l

a

k

,

l

2

;

wherein a′ i,j is the coefficient associated with the specified state |i, j>, and wherein −y≦i≦x, −x≦j≦y, −y≦k≦x, −x≦l≦y and

i

,

j

p

i

,

j

=

1.

17. The method of claim 16 , wherein after determining the probability that the cluster of components is in a specified state, the method further comprises:

determining whether the probability that the cluster of components is in the specified state meets specified criteria; and

if so, performing one or more remedial actions.

18. The method of claim 17 , wherein performing the one or more remedial actions involves one or more of:

generating warnings;

replacing failed components;

reporting the actual number of components that have failed;

reporting the probability that a specified number of components have failed;

monitoring additional variables monitored by sensors in the cluster of components;

adjusting the frequency at which the one or more variables are polled by the sensors;

adjusting test conditions during an accelerated-life study of the components;

scheduling the failed components to be replaced at the next scheduled maintenance interval; and

estimating the remaining useful life of the components.

19. The method of claim 16 , wherein the one or more specified events include one or more of:

a first event wherein the value of the inferential variable is increasing; and

a second event wherein the value of the inferential variable is decreasing.

20. A computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method for detecting multiple anomalies in a cluster of components, wherein the method comprises:

monitoring derivatives obtained from one or more inferential variables which are received from sensors in the cluster of components;

determining whether one or more specified events occurred within the cluster of components based on the monitored derivatives; and

if so, determining the probability that the cluster of components is in a specified state based on the one or more specified events that occurred by

determining the number of first events x and second events y that occurred; and

determining the probability p i,j that the cluster of components is in a specified state |i,j> as:

p

i

,

j

=

a

i

,

j

2

k

,

l

a

k

,

l

2

;

wherein a′ i,j is the coefficient associated with the specified state |i, j>, and wherein −y≦i≦x, −x≦j≦y, −y≦k≦x, −x≦l≦y, and

i

,

j

p

i

,

j

=

1.

21. The computer-readable storage medium of claim 20 , wherein after determining the probability that the cluster of components is in a specified state, the method further comprises:

determining whether the probability that the cluster of components is in the specified state meets specified criteria; and

if so, performing one or more remedial actions.

22. The computer-readable storage medium of claim 21 , wherein performing the one or more remedial actions involves one or more of:

generating warnings;

replacing failed components;

reporting the actual number of components that have failed;

reporting the probability that a specified number of components have failed;

monitoring additional variables monitored by sensors in the cluster of components;

adjusting the frequency at which the one or more variables are polled by the sensors;

adjusting test conditions during an accelerated-life study of the components;

scheduling the failed components to be replaced at the next scheduled maintenance interval; and

estimating the remaining useful life of the components.

23. The computer-readable storage medium of claim 20 , wherein the one or more specified events include one or more of:

a first event wherein the value of the inferential variable is increasing; and

a second event wherein the value of the inferential variable is decreasing.

Assignments (2)
MERGER AND CHANGE OF NAME Recorded Dec 16, 2015
From: ORACLE USA, INC.; SUN MICROSYSTEMS, INC.; ORACLE AMERICA, INC.
To: ORACLE AMERICA, INC.
Reel/Frame 037306/0556 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2008
From: VACAR, DAN; MCELFRESH, DAVID K.; GROSS, KENNY C.; LOPEZ, LEONCIO D.
To: SUN MICROSYSTEMS, INC.
Reel/Frame 020867/0263 →
Continuity (1)
Related Publication 20090234484A1 · Sep 17, 2009