IP Library Granted Patent US 7,523,341
Granted Patent B2
US 7,523,341 · App. 10/844,796 · Granted Apr 21, 2009

Methods, apparatus and computer programs for recovery from failures in a computing environment

Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,523,341
App. No.
10/844,796
Granted
Apr 21, 2009
Kind
B2
Abstract

Provided are methods, apparatus and computer programs for recovery from failures affecting a server in a data processing environment in which a set of servers controls a client's access to a set of resource instances. Independent of any server failure, the client or a gateway is provided with an identification of both a primary server for accessing the resource and at least one secondary server for use as a backup server for accessing the same resource instance (for example, the same physical storage disk). The client or gateway connects to the primary server to perform resource access operations. Following a failure that affects availability of the primary server, the client or gateway connects to the previously identified secondary server to access the same resource instance. Provision of the identification of at least one backup secondary server (without requiring the ‘trigger’ of a failure) avoids the need to discover a new server as part of the recovery operation following a failure. Release of existing reservations using a reset operation, and re-reservation by the original initiator via a backup server, deals with any dangling reservations.

Claims (20)

1. A method for failover recovery, comprising:

receiving an identification at an initiating computer of a primary server and, independently of any failure of the primary server, of a backup server for providing access to a data storage resource;

responsively to the identification, submitting a first request from an initiating computer to the primary server to reserve the data storage resource, while maintaining a record of a reservation of the data storage resource at the initiating computer;

after submitting the first request, detecting a failure of the primary server;

responsively to detecting the failure, submitting a second request from the initiating computer to the backup server, responsively to the record, to reset the data storage resource on which the reservation was made; and

after resetting the data storage resource, accessing the data storage resource via the backup server.

2. The method according to claim 1 , wherein accessing the data storage resource comprises submitting a third request to the backup server to place a new reservation on the data storage resource, and accessing the data storage resource via the backup server using the new reservation.

3. The method according to claim 1 , wherein the data storage resource is a data storage device.

4. The method according to claim 3 , wherein the primary and backup servers comprise storage access controllers within a storage area network (SAN).

5. The method according to claim 3 , and comprising sending a reset instruction, by reference to the record maintained by the initiating computer, from the backup server to the data storage device on which the reservation was made.

6. The method according to claim 5 , wherein the reset instructions is implemented as a Logical Unit Reset operation sent in an iSCSI Task management Request from the backup server to the data storage device.

7. The method according to claim 1 , wherein the primary and backup servers comprise iSCSI target servers.

8. The method according to claim 1 , wherein the initiating computer is an iSCSI initiator.

9. The method according to claim 1 , wherein the initiating computer is an iSCSI gateway.

10. The method according to claim 1 , wherein detecting the failure comprises waiting to receive a response from the primary server for a preset time period after submitting the first request, and identifying the failure when the time period expires prior to receipt of the response.

11. The method according to claim 10 , wherein detecting the failure comprises, upon identifying the failure, verifying that the primary server is unavailable before submitting the second request.

12. The method according to claim 1 , wherein receiving the identification of the primary server and the backup server comprises requesting discovery of a server, within a set of servers, that is available to provide access to the data storage resource, and initiating a discovery operation by one of the servers in the set in order to identify at least one of the primary server and the backup server.

13. The method according to claim 12 , wherein the discovery operation is performed by an iSNS server.

14. The method according to claim 12 , wherein requesting the discovery comprises sending a Service Level Protocol request as a multicast to the servers in the set.

15. The method according to claim 1 , wherein resetting the data storage resource comprises removing dangling reservations from the data storage resource.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2004
From: HUFFERD, JOHN; METH, KALMAN Z.; SATRAN, JULIAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 015117/0473 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 20, 2004
From: HUFFERD, JOHN; METH, KALMAN; SATRAN, JULIAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 014870/0208 →
Continuity (1)
Related Publication 20050268145A1 · Dec 1, 2005