IP Library Granted Patent US 9,454,435
Granted Patent B2
US 9,454,435 · App. 14/014,852 · Granted Sep 27, 2016

Write performance in fault-tolerant clustered storage systems

Inventors: Wendy A. Belluomini (San Jose, CA); Karan Gupta (San Jose, CA); Dean Hildebrand (San Jose, CA); Anna S. Povzner (San Jose, CA); Himabindu Pucha (San Jose, CA); Renu Tewari (San Jose, CA)
Assignee: International Business Machines Corporation
G06F11/1415G06F11/1666G06F17/30132G06F17/30174G06F17/30362G06F17/30371G06F17/30581H04L67/1095H04L67/2842G06F11/20G06F2201/825
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,454,435
App. No.
14/014,852
Granted
Sep 27, 2016
Kind
B2
Abstract

Embodiments of the invention relate to supporting transaction data committed to a stable storage. Committed data in the cluster is stored in the persistent cache layer and replicated and stored in the cache layer of one or more secondary nodes. One copy is designated as a master copy and all other copies are designated as replica, with an exclusive write lock assigned to the master and a shared write lock extended to the replica. An acknowledgement of receiving the data is communicated following confirmation that the data has been replicated to each node designated to receive the replica. Managers and a director are provided to support management of the master copy and the replicas within the file system, including invalidation of replicas, fault tolerance associated with failure of a node holding a master copy, recovery from a failed node, recovered of the file system from a power failure, and transferring master and replica copies within the file system.

Claims (18)

1. A method comprising:

integrating a stable memory layer with a page cache layer in a file system to temporarily hold committed data in distributed non-volatile memory of nodes in a cluster;

in response to receiving a synchronous write transaction in the file system, placing data associated with the received write transaction in the page cache layer and replicating the received data within the page cache layer of one or more remote nodes in the cluster;

distinguishing between a master copy of the received data and a replica of the received data, including applying existing cache policies to the master copy of the received data; and

in response to flushing the master copy to persistent storage, invalidating each replica on the one or more remote nodes.

2. The method of claim 1 , further comprising the master copy holding an exclusive cluster-wide write lock and the replica holding a shared cluster-wide write lock.

3. The method of claim 2 , further comprising maintaining fault tolerance in response to a failure of a node holding the master copy, including electing a remote node to acquire the master copy by acquiring the exclusive cluster-wide write lock for the data.

4. The method of claim 3 , further comprising recovering content on the recovered failed node in response to the recovery of the failed node, including reading data from non-volatile memory of a surviving node, wherein the surviving node is selected from the group consisting of: an available node and the master node.

5. The method of claim 2 , further comprising transferring designation of the master copy to a requestor node, including revoking the exclusive write lock from the master copy, and transferring the exclusive write lock to the requestor node.

6. The method of claim 1 , further comprising in response to recovery of the file system from a power failure, recovering data from non-volatile memory content on each node, and identifying master and replica copies from a characteristic of data byte-range and validating master and replica copies by re-acquiring cluster-wide write locks.

7. A method comprising:

integrating a stable memory layer with a page cache layer in a file system to temporarily hold data committed to a stable storage in non-volatile memory of nodes in a cluster;

receiving a synchronous write transaction in the file system, including placing data associated with the transaction in the page cache layer of the file system and replicating received data within the page cache layer;

maintaining a master copy and a replica of the received data, the master copy having distinguishing characteristics from the replica; and

applying existing cache policies to the master copy of the received data; and

invalidating the replica following flushing of the master copy to persistent storage.

8. The method of claim 7 , further comprising the master copy holding an exclusive cluster-wide write lock and the replica holding a shared cluster-wide write lock.

9. The method of claim 7 , further comprising recovering data from non-volatile memory content on each node, identifying master and replica copies from a characteristic of data byte-range, and validating master and replica copies by re-acquiring cluster-wide write locks, responsive to recovery from a power failure.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2016
From: BELLUOMINI, WENDY A.; GUPTA, KARAN; HILDEBRAND, DEAN; POVZNER, ANNA S.; PUCHA, HIMABINDU; TEWARI, RENU
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 039292/0358 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2016
From: BELLUOMINI, WENDY A.; GUPTA, KARAN; HILDEBRAND, DEAN; POVZNER, ANNA S.; PUCHA, HIMABINDU; TEWARI, RENU
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 039113/0525 →
Continuity (2)
Continuation 13719590 · Dec 19, 2012
Related Publication 20140173185A1 · Jun 19, 2014