IP Library Granted Patent US 11,210,170
Granted Patent B2
US 11,210,170 · App. 16/937,308 · Granted Dec 28, 2021

Failed storage device rebuild method

Inventors: Vladislav Bolkhovitin (San Jose, CA); Siva Munnangi (San Jose, CA)
Assignee: Western Digital Technologies, Inc.
G06F11/1092G06F11/1076G06F11/2094
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,210,170
App. No.
16/937,308
Granted
Dec 28, 2021
Kind
B2
Abstract

Methods and systems for rebuilding a failed storage device in a data storage system. For example, a method including identifying a first garbage collection group (GCG) in a storage array for garbage collection; extracting valid data and redundancy information from functioning storage devices in the storage array associated with the first GCG; reconstructing data of a failed storage device associated with the first GCG based on the extracted valid data and redundancy information from the functioning storage devices associated with the first GCG; consolidating the extracted valid data from the functioning storage devices and the reconstructed data of the failed storage device associated with the first GCG; writing the consolidated extracted valid data from the functioning storage devices and the reconstructed data of the failed storage device associated with the first GCG to a second GCG in the storage array; and reclaiming the first GCG identified for garbage collection.

Claims (57)

1. A computer-implemented method comprising:

identifying a first garbage collection group in a storage array for garbage collection;

reconstructing, based on extracted valid data and redundancy information from functioning storage devices in the storage array, partial data of a failed storage device associated with the first garbage collection group;

determining that a predetermined condition concerning the storage array has been met;

performing, responsive to determining that the predetermined condition concerning the storage array has been met, a manual rebuild of additional rebuild data of the failed storage device; and

writing, based on the partial data and the additional rebuild data of the failed storage device, reconstructed data of the failed storage device to the storage array.

2. The computer-implemented method of claim 1 , further comprising:

extracting the extracted valid data and the redundancy information from one or more functioning storage devices in the storage array that are associated with the first garbage collection group; and

consolidating the extracted valid data and the reconstructed partial data of the failed storage device that is associated with the first garbage collection group.

3. The computer-implemented method of claim 2 , further comprising:

writing the consolidated extracted valid data and the reconstructed partial data of the failed storage device that is associated with the first garbage collection group to a second garbage collection group in the storage array; and

reclaiming the first garbage collection group identified for garbage collection.

4. The computer-implemented method of claim 1 , wherein performing the manual rebuild of additional rebuild data of the failed storage device includes:

identifying a stripe in the storage array to be rebuilt;

extracting, from outside of the first garbage collection group, additional valid data and additional redundancy information from functioning storage devices associated with the identified stripe in the storage array; and

reconstructing, based on the additional valid data and the additional redundancy information, the additional rebuild data.

5. The computer-implemented method of claim 4 , wherein the identified stripe is not reconstructed solely from the first garbage collection group.

6. The computer-implemented method of claim 1 , wherein the predetermined condition concerning the storage array is selected from a group comprising:

a rebuild timeout threshold for the failed storage device has been exceeded; and

one or more garbage collection groups in the storage array have not been written to within a predetermined amount of time.

7. The computer-implemented method of claim 1 , wherein the storage array comprises one or more solid-state drives.

8. The computer-implemented method of claim 1 , wherein the storage array is configured as a redundant array of independent disks (RAID) array.

9. The computer-implemented method of claim 1 , wherein the storage array is configured to support an erasure coding scheme.

10. The computer-implemented method of claim 1 , wherein the storage array includes overprovisioned capacity that is configured as spare space to temporarily store reconstructed data of the failed storage device.

11. A data storage system, comprising:

a storage array including a plurality of storage devices;

one or more processors; and

logic executable by the one or more processors to perform operations comprising:

identifying a first garbage collection group in the storage array for garbage collection;

reconstructing, based on extracted valid data and redundancy information from functioning storage devices in the storage array, partial data of a failed storage device associated with the first garbage collection group;

determining that a predetermined condition concerning the storage array has been met;

performing, responsive to determining that the predetermined condition concerning the storage array has been met, a manual rebuild of additional rebuild data of the failed storage device; and

writing, based on the partial data and the additional rebuild data of the failed storage device, reconstructed data of the failed storage device to the storage array.

12. The data storage system of claim 11 , wherein the operations further comprise:

extracting the extracted valid data and the redundancy information from one or more functioning storage devices in the storage array that are associated with the first garbage collection group; and

consolidating the extracted valid data and the reconstructed partial data of the failed storage device that is associated with the first garbage collection group.

13. The data storage system of claim 12 , wherein the operations further comprise:

writing the consolidated extracted valid data and the reconstructed partial data of the failed storage device that is associated with the first garbage collection group to a second garbage collection group in the storage array; and

reclaiming the first garbage collection group identified for garbage collection.

14. The data storage system of claim 11 , wherein performing the manual rebuild of the additional rebuild data of the failed storage device includes:

identifying a stripe in the storage array to be rebuilt;

extracting, from outside of the first garbage collection group, additional valid data and additional redundancy information from functioning storage devices associated with the identified stripe in the storage array; and

reconstructing, based on the additional valid data and the additional redundancy information, the additional rebuild data.

15. The data storage system of claim 14 , wherein the identified stripe is not reconstructed solely from the first garbage collection group.

16. The data storage system of claim 11 , wherein the predetermined condition concerning the storage array is selected from a group comprising:

a rebuild timeout threshold for the failed storage device has been exceeded; and

one or more garbage collection groups in the storage array have not been written to within a predetermined amount of time.

17. The data storage system of claim 11 , wherein the storage array comprises one or more solid-state drives.

18. The data storage system of claim 11 , wherein the storage array is configured as a redundant array of independent disks (RAID) array.

19. The data storage system of claim 11 , wherein the storage array is configured to support an erasure coding scheme.

20. A system, comprising:

a storage array including a plurality of storage devices;

means for identifying a first garbage collection group in the storage array for garbage collection;

means for reconstructing, based on extracted valid data and redundancy information from functioning storage devices in the storage array, partial data of a failed storage device associated with the first garbage collection group;

means for determining that a predetermined condition concerning the storage array has been met;

means for performing, responsive to determining that the predetermined condition concerning the storage array has been met, a manual rebuild of additional rebuild data of the failed storage device; and

means for writing, based on the partial data and the additional rebuild data of the failed storage device, reconstructed data of the failed storage device to the storage array.

Assignments (5)
PATENT COLLATERAL AGREEMENT - A&R LOAN AGREEMENT Recorded Aug 21, 2023
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064715/0001 →
PATENT COLLATERAL AGREEMENT - DDTL LOAN AGREEMENT Recorded Aug 21, 2023
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 067045/0156 →
RELEASE OF SECURITY INTEREST AT REEL 053926 FRAME 0446 Recorded Feb 8, 2022
From: JPMORGAN CHASE BANK, N.A.
To: WESTERN DIGITAL TECHNOLOGIES, INC.
Reel/Frame 058966/0321 →
SECURITY INTEREST Recorded Sep 29, 2020
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS AGENT
Reel/Frame 053926/0446 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 25, 2020
From: BOLKHOVITIN, VLADISLAV; MUNNANGI, SIVA
To: WESTERN DIGITAL TECHNOLOGIES, INC.
Reel/Frame 053593/0565 →
Continuity (2)
Continuation 15913910 · Mar 6, 2018
Related Publication 20200356440A1 · Nov 12, 2020