IP Library Granted Patent US 8,726,065
Granted Patent B2
US 8,726,065 · App. 13/275,394 · Granted May 13, 2014

Managing failover operations on a cluster of computers

Inventors: Travis M. Drucker (Rochester, MN); Joel C. Dubbels (Eyota, MN); Thomas J. Eggebraaten (Rochester, MN); Janice R. Glowacki (Rochester, MN); Richard J. Stevens (Rochester, MN); David A. Wall (Rochester, MN)
Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,726,065
App. No.
13/275,394
Granted
May 13, 2014
Kind
B2
Abstract

Managing failover operations on a cluster of computers, including: identifying, by a failover hold module, a failure to access data storage in the cluster of computers; preventing the execution of all read operations directed to the data storage that were received after the failure to access data storage was identified; executing all write operations directed to the data storage that were received after the failure to access data storage was identified, including writing data to a cache; identifying that a failover to alternative data storage is complete; executing the held read operations, including reading data from the alternative data storage; and copying, from cache to the alternative data storage, the data written to the cache as part of the write operations.

Claims (26)

1. An apparatus for managing failover operations on a cluster of computers, the apparatus comprising:

a computer processor; and

a computer memory operatively coupled to the computer processor, the computer memory having disposed within it computer program instructions that, when executed by the computer processor, cause the computer processor to function by:

identifying, by a failover hold module, a failure to access data storage in the cluster of computers;

in response to identifying the failure to access data storage:

holding all read operations directed to the data storage; and

executing all write operations directed to the data storage by redirecting the write operations to a cache;

identifying that a failover to an alternative data storage is complete;

in response to identifying that the failover to the alternative data storage is complete:

copying, from cache to the alternative data storage, the data written to the cache as part of execution of the write operations; and

executing held read operations by reading data from the alternative data storage instead of the data storage to which the read operation is directed.

2. The apparatus of claim 1 wherein identifying a failure to access data storage includes identifying that a read operation failed.

3. The apparatus of claim 1 wherein identifying a failure to access data storage includes identifying that a write operation failed.

4. The apparatus of claim 1 wherein identifying a failure to access data storage includes identifying a failure to access a file system.

5. The apparatus of claim 1 wherein identifying a failure to access data storage includes identifying a failure to access a database.

6. An apparatus for managing maintenance of a cluster of computers, the apparatus comprising:

a computer processor; and

a computer memory operatively coupled to the computer processor, the computer memory having disposed within it computer program instructions that, when executed by the computer processor, cause the computer processor to function by:

identifying, by a failover module, one or more scheduled maintenance operations to be executed on the cluster of computers;

initiating the execution of the scheduled maintenance operations;

in response to initiating the execution of the scheduled maintenance operations, holding all received data storage access requests;

determining that the scheduled maintenance operations are complete; and

in response to determining that the scheduled maintenance operations are complete, executing all held data storage access requests including executing all write requests that were received after initiating the execution of the scheduled maintenance operations.

7. The apparatus of claim 6 wherein holding all received data storage access requests further comprises storing, in a cache, all data storage access requests that were received after initiating the execution of the scheduled maintenance operations.

8. The apparatus of claim 6 wherein executing all held data storage access requests includes executing all read requests that were received after initiating the execution of the scheduled maintenance operations.

9. The apparatus of claim 6 wherein the one or more scheduled maintenance operations to be executed on the cluster of computers includes backing up data storage on the cluster of computers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2011
From: DRUCKER, TRAVIS M.; DUBBELS, JOEL C.; EGGEBRAATEN, THOMAS J.; GLOWACKI, JANICE R.; STEVENS, RICHARD J.; WALL, DAVID A.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 027075/0967 →
Continuity (1)
Related Publication 20130097456A1 · Apr 18, 2013