IP Library › Granted Patent US 10,229,121
Granted Patent B2
US 10,229,121 · App. 15/070,227 · Granted Mar 12, 2019

Detection of file corruption in a distributed file system

Inventors: James C. Davis (Blacksburgh, VA); Willard A. Davis (Rosendale, NY); Felipe Knop (Poughkeepsie, NY)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F17/3007G06F11/0727G06F11/0766G06F11/3688G06F17/30194
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,229,121
App. No.
15/070,227
Granted
Mar 12, 2019
Kind
B2
Abstract

Aspects include testing distributed file systems by selecting a file in a multiple writer environment and selecting an offset of a block in the file. Test data is generated for the block by randomly selecting a starting value from a plurality of possible starting values. A test header that includes the starting value and a test data sequence that starts with the starting value is created. A file system that is being tested writes the test header and the test data sequence to the block. Contents of the block are read by the file system that is being tested, and expected contents of the data sequence are determined based on contents of the read header. The expected contents of the data sequence are compared to the read data sequence and an error indication is output based on the expected contents not being equal to the read contents.

Claims (19)

1. A computer-implemented method for testing distributed file systems, the method comprising:

injecting an error at an offset of a block in a file in a distributed file system;

performing a recovery action that includes the offset of the block in the file system;

generating test data for the block that will detect silent write failure and stale read failures, the generating including:

randomly selecting a starting value from a plurality of possible starting values;

creating a test header that includes the starting value; and

creating a test data sequence that starts with the starting value;

writing, by the distributed file system, the test header and the test data sequence to the block in the file;

reading, by the file system that is being tested, contents of the block from the file subsequent to the writing, the read contents including a read header and a read data sequence;

determining expected contents of the data sequence based on contents of the read header;

comparing the expected contents of the data sequence to the read data sequence; and

outputting an error indication based on the expected contents not being equal to the read contents, the error indication including human-readable content including a time of the writing and the expected contents.

2. The method of claim 1 , wherein the test data sequence is divided into a plurality of chunks characterized by a chunk size, the plurality of chunks including a first chunk of the chunk size, at least one intermediate chunk of the chunk size, and a last chunk that is smaller than or equal to the chunk size, wherein a portion of the test data sequence in the first chunk starts with the starting value and a portion of the test data sequence in the at least one intermediate chunk and the last chunk start with a zero.

3. The method of claim 2 , further comprising selecting the chunk size, wherein the chunk size of the block is different than a chunk size of at least one other block in the file.

4. The method of claim 1 , wherein the writing is atomic within the file.

5. The method of claim 1 , wherein the generating, writing, reading, determining, comparing and outputting are performed by multiple nodes simultaneously on the plurality of files.

6. The method of claim 5 , wherein the generating, writing, reading, determining, comparing and outputting continue to be performed by a node in the multiple nodes subsequent to a failure of an other node in the multiple nodes.

7. The method of claim 1 , further comprising writing the test header to a journal file corresponding to the block synchronously with writing the test header to the block in the file.

8. The method of claim 7 , wherein the reading further comprises reading the journal file corresponding to the block, and the method further comprises comparing contents of the journal file to the read header and outputting an error indication based on the contents of the journal file not matching the contents of the read header.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2016
From: DAVIS, JAMES C.; DAVIS, WILLARD A.; KNOP, FELIPE
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 037979/0539 →
Continuity (2)
Continuation 14869091 · Sep 29, 2015
Related Publication 20170091086A1 · Mar 30, 2017