IP Library › Granted Patent US 8,479,048
Granted Patent B2
US 8,479,048 · App. 13/211,694 · Granted Jul 2, 2013

Root cause analysis method, apparatus, and program for IT apparatuses from which event information is not obtained

Inventors: Tomohiro Morimura (Kawasaki, JP); Takayuki Nagai (Machida, JP); Kiminori Sugauchi (Yokohama, JP); Takaki Kuroda (Machida, JP); Yoshihiro Arato (Tokyo, JP)
Assignee: Hitachi, Ltd.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,479,048
App. No.
13/211,694
Granted
Jul 2, 2013
Kind
B2
Abstract

In the system management server, an information processing apparatus that is an event-information acquisition target is registered as a monitored apparatus in configuration information; event information that complies with a rule stored in advance is identified from among a plurality of pieces of event information stored in the system management server; a server apparatus for a network service related to the event information is identified; and a message is displayed which indicates that the cause of the event that occurred in a client information processing apparatus which has generated event information is an event related to the network service, which occurred in the server apparatus.

Claims (80)

1. A system comprising:

a plurality of information processing apparatuses; and

a management computer,

wherein the management computer stores correlation analysis rule information, indicating that an event of a second event type is a root cause of an event of a first event type for a network service,

wherein the management computer stores configuration information including at least information about the network service of a plurality of monitored apparatuses,

wherein the plurality of monitored apparatuses are included in the plurality of information processing apparatuses,

wherein the management computer obtains event information from the plurality of monitored apparatuses,

wherein the management computer identifies, from the event information, a first event of the first event type,

wherein the management computer identifies a first monitored apparatus in which the first event occurs, and

wherein the management computer identifies a root cause apparatus which is a server of the network service, based on the correlation analysis rule information and the configuration information, even if the root cause apparatus is not included in the plurality of monitored apparatuses.

2. The system according to claim 1 , wherein the management server selects the plurality of monitored apparatuses, each of the plurality of monitoried apparatuses having an IP (Internet Protocol) address in a predetermined IP address range.

3. The system according to claim 1 ,

wherein the root cause apparatus is a storage apparatus,

wherein the network service provides a logical volume of the storage apparatus, and

wherein the second event type is an occurrence of a fault in the storage apparatus, and the first event type is a failure of accessing the logical volume by a computer.

4. The system according to claim 1 ,

wherein the root cause apparatus is a DNS (Domain Name Service) server,

wherein the network service is a DNS,

wherein the second event type is a fault in the DNS server, and

wherein the first event type is a disconnection of communication for a DNS.

5. The system according to claim 1 ,

wherein the root cause apparatus is a file server computer,

wherein the network service is a file sharing service,

wherein the second event type is a fault in the file server computer, and

wherein the first event type is an access failure of a file provided by the file sharing service.

6. The system according to claim 1 , wherein the management computer identifies the first monitored apparatus, the first event type, the root cause apparatus, and the second event type, and sends information identifying the first monitored apparatus, the first event type, the root cause apparatus, and the second event type to the screen output apparatus for displaying a root cause of the first event of the first event type that occurred in the first monitored apparatus and is estimated to be caused by a not obtained event of the second event type that occurred in the root cause apparatus.

7. The system according to claim 2 , wherein the management computer suggests obtaining event information from the root cause apparatus, after checking whether or not the management server is able to obtain information from the root cause apparatus.

8. A management computer comprising:

a memory storing a management program; and

a CPU (Central Processing Unit) that executes the management program,

wherein when executed, the management program causes the CPU to:

store correlation analysis rule information, indicating that an event of a second event type is a root cause of an event of a first event type for a network service;

store configuration information including at least information about the network service of a plurality of monitored apparatuses;

obtain event information from the plurality of monitored apparatuses;

identify, from the event information, a first event of the first event type;

identify a first monitored apparatus in which the first event occurs; and

identify a root cause apparatus which is a server of the network service, based on the correlation analysis rule information and the configuration information, even if the root cause apparatus is not included in the plurality of monitored apparatuses.

9. The management computer according to claim 8 , wherein the management program further causes the CPU to select the plurality of monitored apparatuses, each of the plurality of monitored apparatuses having an IP (Internet Protocol) address in a predetermined IP address range.

10. The management computer according to claim 8 ,

wherein the root cause apparatus is a storage apparatus,

wherein the network service provides a logical volume of the storage apparatus, and

wherein the second event type is an occurrence of a fault in the storage apparatus, and the first event type is a failure of accessing the logical volume by a computer.

11. The management computer according to claim 8 ,

wherein the root cause apparatus is a DNS (Domain Name Service) server,

wherein the network service is a DNS,

wherein the second event type is a fault in the DNS server, and

wherein the first event type is a disconnection of communication for a DNS.

12. The management computer according to claim 8 ,

wherein the root cause apparatus is a file server computer,

wherein the network service is a file sharing service,

wherein the second event type is a fault in the file server computer, and

wherein the first event type is a access failure of a file provided by the file sharing service.

13. The management computer according to claim 8 ,

wherein the management program further causes the CPU to: identify the first monitored apparatus, the first event type, the root cause apparatus, and the second event type; and

send information identifying the first monitored apparatus, the first event type, the root cause apparatus, and the second event type to the screen output apparatus for displaying a root cause of the first event of the first event type that occurred in the first monitored apparatus and is estimated to be caused by a not obtained event of the second event type that occurred in the root cause apparatus.

14. The management computer according to claim 9 , wherein the management computer suggests obtaining event information from the root cause apparatus, after checking whether or not the CPU is able to obtain information from the root cause apparatus.

15. A non-transitory machine-readable storage medium tangibly embodying a program for execution on a management computer, the program comprising code causing the management computer to:

store correlation analysis rule information, indicating that an event of a second event type is a root cause of an event of a first event type for a network service;

store configuration information including at least information about the network service of a plurality of monitored apparatuses;

obtain event information from the plurality of monitored apparatuses;

identify, from the event information, a first event of the first event type;

identify a first monitored apparatus in which the first event occurs; and

identify a root cause apparatus which is a server of the network service, based on the correlation analysis rule information and the configuration information, even if the root cause apparatus is not included in the plurality of monitored apparatuses.

16. The non-transitory machine-readable storage medium according to claim 15 , wherein the program further causes the management computer to select the plurality of monitored apparatuses, each of the plurality of monitored apparatuses having an IP (Internet Protocol) address in a predetermined IP address range.

17. The non-transitory machine-readable storage medium according to claim 15 ,

wherein the root cause apparatus is a storage apparatus,

wherein the network service provides a logical volume of the storage apparatus, and

wherein the second event type is an occurrence of a fault in the storage apparatus, and the first event type is a failure of accessing the logical volume by a computer.

18. The non-transitory machine-readable storage medium according to claim 15 ,

wherein the root cause apparatus is a DNS (Domain Name Service) server,

wherein the network service is a DNS,

wherein the second event type is a fault in the DNS server, and

wherein the first event type is a disconnection of communication for a DNS.

19. The non-transitory machine-readable storage medium according to claim 15 ,

wherein the root cause apparatus is a file server computer,

wherein the network service is a file sharing service,

wherein the second event type is a fault in the file server computer, and

wherein the first event type is a access failure of a file provided by the file sharing service.

20. The non-transitory machine-readable storage medium according to claim 15 , wherein the program causes the management computer to identify the first monitored apparatus, the first event type, the root cause apparatus, and the second event type, and send information identifying the first monitored apparatus, the first event type, the root cause apparatus, and the second event type to the screen output apparatus for displaying a root cause of the first event of the first event type that occurred in the first monitored apparatus and is estimated to be caused by a not obtained event of the second event type that occurred in the root cause apparatus.

21. The non-transitory machine-readable storage medium according to claim 16 , wherein the program causes the management computer to suggest obtaining event information from the root cause apparatus, after checking whether or not the management computer is able to obtain information from the root cause apparatus.

Priority Claims (1)
JP 2008-252093 · Sep 30, 2008 · national
Continuity (2)
Continuation 12444398
Related Publication 20110302305A1 · Dec 8, 2011