IP Library Granted Patent US 8,386,834
Granted Patent B1
US 8,386,834 · App. 12/772,006 · Granted Feb 26, 2013

Raid storage configuration for cached data storage

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,386,834
App. No.
12/772,006
Granted
Feb 26, 2013
Kind
B1
Abstract

A storage server receives a notification indicating a failure of a mass storage device in a storage array. The storage server determines whether a number of failures exceeds a fault tolerance level of the array and if the number of failures exceeds the fault tolerance level, recovers an address space corresponding to the failed storage device. When recovering the address space, the storage server replaces the failed storage device with a spare storage device having an identifiable pattern stored thereon and determines whether a file system on the storage system can automatically invalidate cached data blocks on the failed storage device.

Claims (38)

1. A method comprising:

receiving, by a storage system, a notification indicating a failure of a storage device in a cache storage array for caching data;

determining whether a number of failures exceeds a fault tolerance level of the cache storage array; and

if the number of failures exceeds the fault tolerance level, recovering an address space corresponding to the failed storage device, wherein recovering the address space comprises:

replacing the failed storage device with a spare storage device having an identifiable pattern stored thereon;

determining whether a file system on the storage system can automatically invalidate cached data blocks on the failed storage device;

if the file system cannot automatically invalidate cached data blocks on the failed storage device, setting an indicator for each of a plurality of data blocks on the spare storage device, wherein the indicator indicates that an associated data block has an unrecoverable error, and wherein the associated data block contains the identifiable pattern; and

updating a parity calculation for a stripe of data blocks in the storage system, wherein the stripe of data blocks includes a data block on the spare storage device, and wherein the data block on the spare storage device contains the identifiable pattern.

2. The method of claim 1 , further comprising:

if the number of failures exceeds the fault tolerance level, returning an error message in response to an input/output request for the address space corresponding to the failed mass storage device.

3. A method, comprising:

receiving, by a storage system, a data access request for a data block in an array of cache storage devices connected to the storage system, the cache storage device containing data replicated in a primary storage device;

determining a data block has an unrecoverable error, wherein the data block has an unrecoverable error if a number of errors in a stripe of data blocks across the array of cache storage devices exceeds a number of errors that can be corrected by an underlying RAID protection level;

setting a cache-miss indicator bit associated with the data block in response to determining the data block has an unrecoverable error;

writing an identifiable pattern to the data block in response to determining the data block has an unrecoverable error;

updating a parity calculation for a stripe of data blocks in the storage system, wherein the parity calculation includes the data block with the identifiable pattern; and

continuing to service data access requests without initiating a consistency check operation in the storage system.

4. The method of claim 3 , wherein the array of storage devices comprises an array of solid state drives (SSDs).

5. A system comprising:

an array of cache storage devices;

a storage server coupled to the array of cache storage devices, the storage server comprising;

a storage access module configured to detect a failure of a storage device in the array of cache storage devices and to determine whether a number of failures exceeds a fault tolerance level of the array; and

a RAID-C management module configured to recover an address space corresponding to the failed storage device if the number of failures exceeds the fault tolerance level, wherein recovering the address space comprises:

replacing the failed storage device with a spare storage device having an identifiable pattern stored thereon;

determining whether a file system on the storage system can automatically invalidate cached data blocks on the failed storage device;

if the file system cannot automatically invalidate cached data blocks on the failed storage device, setting an indicator for each of a plurality of data blocks on the spare storage device, wherein the indicator indicates that an associated data block has an unrecoverable error, and wherein the associated data block contains the identifiable pattern; and

updating a parity calculation for a stripe of data blocks in the storage system, wherein the stripe of data blocks includes a data block on the spare storage device, and wherein the data block on the spare storage device contains the identifiable pattern.

6. The system of claim 5 , wherein the RAID-C management module is further configured to:

return an error message in response to an input/output request for the address space corresponding to the failed mass storage device if the number of failures exceeds the fault tolerance level.

7. A system, comprising:

a primary array of mass storage devices;

a cache array of mass storage devices containing data replicated in the primary array of mass storage devices; and

a storage server coupled to the primary and cache arrays of mass storage devices, the storage server comprising:

a storage access module that maintains a data protection mechanism, wherein the storage access module is configured to identify, during a data access request, a data block in the cache array of mass storage devices having an unrecoverable error condition, wherein the data block has an unrecoverable error if a number of errors in a stripe of data blocks across the cache array of mass storage devices exceeds a number of errors that can be corrected by an underlying RAID protection level; and

a RAID-C management module configured to:

write an identifiable pattern to the data block having the unrecoverable error;

set an indicator bit for th data block having the unrecoverable error condition, wherein the storage access module continues to service data access requests even if the indicator bit is set; and

updating a parity calculation for a stripe of data blocks in the storage array, wherein the stripe of data blocks includes the block with the identifiable pattern.

Assignments (2)
CHANGE OF NAME Recorded Jan 15, 2014
From: NETWORK APPLIANCE, INC
To: NETAPP, INC.
Reel/Frame 031978/0526 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2010
From: GOEL, ATUL; STRANGE, STEPHEN H.
To: NETAPP, INC., A CORPORATION OF DELAWARE
Reel/Frame 024501/0236 →