Method and system for disk drive exercise and maintenance of high-availability storage systems
Methods and apparatuses for maintaining a particular disk drive that is powered off in a storage system are disclosed. The method includes powering on the particular disk drive and executing a test on the particular disk drive. The method further includes powering off the particular disk drive on completing the test.
1 . A method for maintaining a particular disk drive in a storage system, wherein the storage system includes a plurality of disk drives and the particular disk drive that is powered-off, the method comprising:
powering-on the particular disk drive;
executing a test on the particular disk drive; and
powering-off the particular disk drive.
2 . The method of claim 1 , wherein the test includes a buffer test.
3 . The method of claim 2 , wherein the buffer test includes a write/read/compare test of sector buffer random access memory in the particular disk drive.
4 . The method of claim 1 , wherein the test includes a write test on a plurality of heads in the particular disk drive.
5 . The method of claim 4 , wherein the write test includes a write/read/compare operation on each head in the particular disk drive.
6 . The method of claim 5 , wherein write the test includes accessing non-user accessible sectors on the particular disk drive.
7 . The method of claim 1 , wherein the test includes a random read test.
8 . The method of claim 7 , wherein the random read test includes a read operation on randomly selected logical block addresses.
9 . The method of claim 7 , wherein the random read test includes auto defect reallocation.
10 . The method of claim 1 , wherein the test includes a read scan test.
11 . The method of claim 10 , wherein the read scan test includes auto defect reallocation.
12 . The method of claim 11 , wherein the read scan test is performed over an entire surface of the particular disk drive.
13 . The method of claim 1 , wherein the plurality of disk drives is power-managed according to a power budget, the method further comprising:
determining that powering-on the particular disk drive would cause the power budget to be exceeded; and
postponing powering-on of the particular disk drive and executing the test.
14 . The method of claim 1 , wherein the plurality of disk drives is power-managed according to a power budget, the method further comprising:
determining after powering-on the particular disk drive that a request to power-on an additional disk drive would cause the power budget to be exceeded;
suspending the test; and
powering-off the particular disk drive.
15 . The method of claim 1 , wherein the test is performed after a predetermined interval.
16 . The method of claim 15 , wherein the predetermined interval is approximately 30 days.
17 . The method of claim 1 , further comprising:
detecting a request for access of the particular drive;
suspending the test to fulfill the request for access; and
resuming the test.
18 . The method of claim 1 , further comprising:
determining whether the test has failed; and
indicating whether the particular disk drive should be replaced in response to determining whether the test has failed.
19 . A method for maintaining data in a disk drive, the method comprising:
performing a check on the disk drive; and
if a predetermined criterion is not met as a result of the test then performing a recovery action.
20 . The method of claim 19 , wherein the check includes a seek error rate.
21 . The method of claim 19 , wherein the check includes an RSC rate.
22 . The method of claim 19 , wherein the check includes a timeout error.
23 . The method of claim 19 , wherein the check includes a buffer test.
24 . The method of claim 19 , wherein the check includes a write test on a plurality of heads in the disk drive.
25 . The method of claim 19 , wherein the check includes a random read test.
26 . The method of claim 25 , wherein the random read test includes auto defect reallocation.
27 . The method of claim 19 , wherein the check includes a read scan test.
28 . The method of claim 27 , wherein the read scan test includes auto defect reallocation.
29 . An apparatus for maintaining a particular disk drive in a storage system, wherein the storage system includes a plurality of disk drives and the particular disk drive that is powered-off, the apparatus comprising:
a power controller for controlling power to the disk drives and the particular disk drive; and
a test-moderator for executing a test on the particular disk drive;
whereby, the particular disk is powered-on by the power controller before the test is to be executed and the particular disk drive is powered-off after the test is executed.
30 . A machine-readable medium including instructions executable by a processor for maintaining a particular disk drive in a storage system, wherein the storage system includes a plurality of disk drives and the particular disk drive that is powered-off, the machine readable medium comprising:
one or more instructions for powering-on the particular disk drive;
one or more instructions for executing a test on the particular disk drive; and
one or more instructions for powering-off the particular disk drive.
31 . An apparatus for maintaining a particular disk drive in a storage system, wherein the storage system includes a plurality of disk drives and the particular disk drive that is powered-off, the apparatus comprising:
a processor for executing instructions; and
a machine-readable medium including:
one or more instructions for powering-on the particular disk drive;
one or more instructions for executing a test on the particular disk drive; and
one or more instructions for powering-off the particular disk drive.