IP Library Granted Patent US 11,989,449
Granted Patent B2
US 11,989,449 · App. 17/313,960 · Granted May 21, 2024

Method for full data reconstruction in a raid system having a protection pool of storage units

Inventors: Paul Nehse (Livermore, CA); Michael B. Thiels (San Martin, CA); Devendra V. Kulkarni (Santa Clara, CA)
Assignee: EMC IP HOLDING COMPANY LLC
G06F3/0659G06F3/0619G06F3/0631G06F3/0689G06F11/1092G06F11/2094G06F2201/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,989,449
App. No.
17/313,960
Granted
May 21, 2024
Kind
B2
Abstract

A method of performing a full data reconstruction in a redundant array of independent disks (RAID) system with a protection pool of storage units includes determining that a physical disk of a storage cluster has been removed from service. The physical disk includes a set of physical extents and at least one physical extent of the set of physical extents is associated with an array of physical extents distributed across physical disks of the storage cluster. The method further includes transmitting a message to one or more array groups of the physical disks, to allocate replacement physical extents and assign the replacement physical extents to the array of physical extents and initiating reconstruction of data from the set of physical extents of the physical disk to the replacement physical extents.

Claims (53)

1. A method comprising:

determining that a physical disk of a plurality of physical disks of a storage cluster has been removed from service due to failure, the physical disk comprising a plurality of physical extents, wherein at least one physical extent of the plurality of physical extents is associated with an array of physical extents distributed across a plurality of physical disks of the storage cluster, wherein the plurality of physical disks are associated with one or more array groups, wherein the one or more array groups each comprise a plurality of arrays, wherein each of the plurality of arrays of the one or more array groups comprises physical extents of the plurality of physical extents from a group of the plurality of physical disks, and wherein the plurality of physical extents of the physical disk are allocated to a plurality of the arrays;

transmitting a message to the one or more array groups of the plurality of physical disks, the message including a notification to the one or more array groups of each of the physical extents that are failed, wherein the one or more array groups provides notification to each array of the one or more array groups that is associated with the physical extents that are failed, to perform a successful failure of the physical extents that are failed,

in response to receiving a response from the one or more array groups indicating the successful failure of the physical extents that are failed, allocating replacement physical extents and assigning the replacement physical extents to the array of physical extents distributed across the plurality of physical disks of the storage cluster; and

initiating reconstruction of data from the plurality of physical extents of the physical disk to the replacement physical extents.

2. The method of claim 1 , wherein determining that the physical disk has been removed from service is in response to determining that a disk write failure on the physical disk has occurred.

3. The method of claim 1 , wherein determining that a physical disk of a storage cluster has been removed from service comprises:

receiving, from an administrator, a command to remove the physical disk from service and to invoke a disk failure of the physical disk.

4. The method of claim 1 , wherein determining that a physical disk of a storage cluster has been removed from service comprises:

monitoring an error counter associated with the physical disk; and

in response to determining that the error counter has reached a threshold error count, initiating a disk failure of the physical disk.

5. The method of claim 1 , wherein determining that a physical disk of a storage cluster has been removed from service comprises:

determining that a write error has occurred during an input/output (I/O) request to an array of the physical disk.

6. The method of claim 1 , further comprising:

providing, by a local disk manager associated with the physical disk to one or more physical extent arrays that include at least one physical extent of the plurality of physical extents of the physical disk, a notification that the physical disk is failed.

7. The method of claim 1 , wherein the reconstruction of the data from the physical extents of the physical disk comprises:

retrieving remaining data from each remaining physical extent in an extent row; and

reconstructing the data from the physical extents based on the remaining data from each remaining physical extent in the extent row.

8. A system comprising:

a processor; and

a memory to store instructions, which when executed by the processor, cause the processor to perform operations comprising:

determining that a physical disk of a storage cluster has been removed from service due to failure, the physical disk comprising a plurality of physical extents, wherein at least one physical extent of the plurality of physical extents is associated with an array of physical extents distributed across a plurality of physical disks of the storage cluster, wherein the plurality of physical disks are associated with one or more array groups, and wherein the one or more array groups each comprise a plurality of arrays, wherein each of the plurality of arrays of the one or more array groups comprises physical extents of the plurality of physical extents from a group of the plurality of physical disks, and wherein the plurality of physical extents of the physical disk are allocated to a plurality of the arrays;

transmitting a message to the one or more array groups of the plurality of physical disks the message including a notification to the one or more array groups of each of the physical extents that are failed, wherein the one or more array groups provides notification to each array of the one or more array groups that is associated with the physical extents that are failed, to perform a successful failure of the physical extents that are failed,

in response to receiving a response from the one or more array groups indicating the successful failure of the physical extents that are failed, allocating replacement physical extents and assigning the replacement physical extents to the array of physical extents distributed across the plurality of physical disks of the storage cluster; and

initiating reconstruction of data from the plurality of physical extents of the physical disk to the replacement physical extents.

9. The system of claim 8 , wherein determining that the physical disk has been removed from service is in response to determining that a disk write failure on the physical disk has occurred.

10. The system of claim 8 , wherein determining that a physical disk of a storage cluster has been removed from service comprises:

receiving, from an administrator, a command to remove the physical disk from service and to invoke a disk failure of the physical disk.

11. The system of claim 8 , wherein determining that a physical disk of a storage cluster has been removed from service comprises:

monitoring an error counter associated with the physical disk; and

in response to determining that the error counter has reached a threshold error count, initiating a disk failure of the physical disk.

12. The system of claim 8 , wherein determining that a physical disk of a storage cluster has been removed from service comprises:

determining that a write error has occurred during an input/output (I/O) request to an array of the physical disk.

13. The system of claim 8 , further comprising:

providing, by a local disk manager associated with the physical disk to one or more physical extent arrays that include at least one physical extent of the plurality of physical extents of the physical disk, a notification that the physical disk is failed.

14. The system of claim 8 , wherein the reconstruction of the data from the physical extents of the physical disk comprises:

retrieving remaining data from each remaining physical extent in an extent row; and

reconstructing the data from the physical extents based on the remaining data from each remaining physical extent in the extent row.

15. A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations comprising:

determining that a physical disk of a storage cluster has been removed from service due to failure, the physical disk comprising a plurality of physical extents, wherein at least one physical extent of the plurality of physical extents is associated with an array of physical extents distributed across a plurality of physical disks of the storage cluster, wherein the plurality of physical disks are associated with one or more array groups, and wherein the one or more array groups each comprise a plurality of arrays, wherein each of the plurality of arrays of the one or more array groups comprises physical extents of the plurality of physical extents from a group of the plurality of physical disks, and wherein the plurality of physical extents of the physical disk are allocated to a plurality of the arrays;

transmitting a message to the one or more array groups of the plurality of physical disks the message including a notification to the one or more array groups of each of the physical extents that are failed, wherein the one or more array groups provides notification to each array of the one or more array groups that is associated with the physical extents that are failed, to perform a successful failure of the physical extents that are failed,

in response to receiving a response from the one or more array groups indicating the successful failure of the physical extents that are failed, allocating replacement physical extents and assigning the replacement physical extents to the array of physical extents distributed across the plurality of physical disks of the storage cluster; and

initiating reconstruction of data from the plurality of physical extents of the physical disk to the replacement physical extents.

16. The non-transitory machine-readable medium of claim 15 , wherein determining that the physical disk has been removed from service is in response to determining that a disk write failure on the physical disk has occurred.

17. The non-transitory machine-readable medium of claim 15 , wherein determining that a physical disk of a storage cluster has been removed from service comprises:

receiving, from an administrator, a command to remove the physical disk from service and to invoke a disk failure of the physical disk.

18. The non-transitory machine-readable medium of claim 15 , wherein determining that a physical disk of a storage cluster has been removed from service comprises:

monitoring an error counter associated with the physical disk; and

in response to determining that the error counter has reached a threshold error count, initiating a disk failure of the physical disk.

19. The non-transitory machine-readable medium of claim 15 , wherein determining that a physical disk of a storage cluster has been removed from service comprises:

determining that a write error has occurred during an input/output (I/O) request to an array of the physical disk.

20. The non-transitory machine-readable medium of claim 15 , further comprising:

providing, by a local disk manager associated with the physical disk to one or more physical extent arrays that include at least one physical extent of the plurality of physical extents of the physical disk, a notification that the physical disk is failed.

Assignments (8)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (058014/0560) Recorded Jun 10, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
Reel/Frame 062022/0473 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (057931/0392) Recorded Jun 10, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
Reel/Frame 062022/0382 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (057758/0286) Recorded Jun 10, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
Reel/Frame 061654/0064 →
SECURITY INTEREST Recorded Oct 6, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 058014/0560 →
SECURITY INTEREST Recorded Oct 6, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 057758/0286 →
SECURITY INTEREST Recorded Oct 6, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 057931/0392 →
SECURITY AGREEMENT Recorded Oct 1, 2021
From: DELL PRODUCTS, L.P.; EMC IP HOLDING COMPANY LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 057682/0830 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2021
From: NEHSE, PAUL; THIELS, MICHAEL B.; KULKARNI, DEVENDRA V.
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 056164/0284 →
Continuity (1)
Related Publication 20220357881A1 · Nov 10, 2022