IP Library Granted Patent US 10,216,591
Granted Patent B1
US 10,216,591 · App. 15/198,216 · Granted Feb 26, 2019

Method and apparatus of a profiling algorithm to quickly detect faulty disks/HBA to avoid application disruptions and higher latencies

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,216,591
App. No.
15/198,216
Granted
Feb 26, 2019
Kind
B1
Abstract

One embodiment is related to a method for determining a faulty hardware component within a data storage system, comprising: collecting data relating to a plurality of input/output (IO) errors associated with a first storage processor within the data storage system; compiling IO error statistics based on the data relating to the plurality of IO errors; and determining a faulty hardware component based on the IO error statistics, wherein the determining of the faulty hardware component comprises utilizing a second storage processor of the data storage system independent from the first storage processor.

Claims (47)

1. A method for determining a faulty hardware component within a data storage system, comprising:

collecting, by a processor, data relating to a plurality of input/output (IO) errors associated with a first storage processor within the data storage system, wherein the data storage system includes a plurality of disk array enclosures (DAEs), each DAE having one or more disk drives;

compiling, by the processor, IO error statistics based on the data relating to the plurality of IO errors, the IO error statistics being related to a first one of the DAEs of the data storage system; and

determining, by the processor, a faulty hardware component based on the IO error statistics, wherein the determining of the faulty hardware component comprises utilizing a second storage processor of the data storage system independent from the first storage processor, including examining IO access statistics of the second storage processor for accessing the first DAE through a different path, and

wherein the plurality of DAEs are connected to the first storage processor and the second storage processor through an independent first path and an independent second path, and each of the one or more disk drives has a first port connected to the first storage processor through the first path and a second port connected to the second storage processor through the second path.

2. The method of claim 1 , wherein the data relating to the plurality of IO errors associated with the first storage processor comprise timestamps of the plurality of IO errors and drives of the data storage system associated with each of the plurality of IO errors.

3. The method of claim 1 , wherein a count of the plurality of IO errors is above a third threshold, the plurality of IO errors take place over a period of time longer than a second time period threshold, and a time difference between any two consecutive IO errors within the plurality of IO errors is below a first time period threshold.

4. The method of claim 1 , further comprising:

determining drives implicated in the plurality of IO errors and a count of IO errors associated with each implicated drive, wherein each implicated drive is associated with more IO errors than a fourth threshold.

5. The method of claim 4 , further comprising:

determining the drives implicated in the plurality of IO errors as faulty based on a fact that 1) IO operations cannot be successfully performed on the drives implicated in the plurality of IO errors with the second storage processor, 2) a count of the drives implicated in the plurality of IO errors is below a fifth threshold, or a combination of both.

6. The method of claim 4 , further comprising:

determining the first storage processor or its associated hardware data path as faulty based on a fact that 1) a count of the drives implicated in the plurality of IO errors is above a fifth threshold, 2) the first storage processor or at least one connected entity within the hardware data path is associated with an error count above a sixth threshold, or a combination of both, wherein the error count is one of Invalid DWord Count, Running Disparity Error Count, or Loss of DWord Sync Count.

7. The method of claim 6 , further comprising:

switching all future IO operations to the second storage processor in response to the first storage processor or its associated hardware data path having been determined as faulty.

8. A non-transitory machine-readable medium having instructions stored therein which, when executed by a processor, cause the processor to perform troubleshooting operations, the operations comprising:

collecting data relating to a plurality of input/output (IO) errors associated with a first storage processor within a data storage system, wherein the data storage system includes a plurality of disk array enclosures (DAEs), each DAE having one or more disk drives;

compiling IO error statistics based on the data relating to the plurality of IO errors, the IO error statistics being related to a first one of the DAEs of the data storage system; and

determining a faulty hardware component based on the IO error statistics, wherein the determining of the faulty hardware component comprises utilizing a second storage processor of the data storage system independent from the first storage processor, including examining IO access statistics of the second storage processor for accessing the first DAE through a different path, and

wherein the plurality of DAEs are connected to the first storage processor and the second storage processor through an independent first path and an independent second path, and each of the one or more disk drives has a first port connected to the first storage processor through the first path and a second port connected to the second storage processor through the second path.

9. The non-transitory machine-readable medium of claim 8 , wherein the data relating to the plurality of IO errors associated with the first storage processor comprise timestamps of the plurality of IO errors and drives of the data storage system associated with each of the plurality of IO errors.

10. The non-transitory machine-readable medium of claim 8 , wherein a count of the plurality of IO errors is above a third threshold, the plurality of IO errors take place over a period of time longer than a second time period threshold, and a time difference between any two consecutive IO errors within the plurality of IO errors is below a first time period threshold.

11. The non-transitory machine-readable medium of claim 8 , the operations further comprising:

determining drives implicated in the plurality of IO errors and a count of IO errors associated with each implicated drive, wherein each implicated drive is associated with more IO errors than a fourth threshold.

12. The non-transitory machine-readable medium of claim 11 , the operations further comprising:

determining the drives implicated in the plurality of IO errors as faulty based on a fact that 1) IO operations cannot be successfully performed on the drives implicated in the plurality of IO errors with the second storage processor, 2) a count of the drives implicated in the plurality of IO errors is below a fifth threshold, or a combination of both.

13. The non-transitory machine-readable medium of claim 11 , the operations further comprising:

determining the first storage processor or its associated hardware data path as faulty based on a fact that 1) a count of the drives implicated in the plurality of IO errors is above a fifth threshold, 2) the first storage processor or at least one connected entity within the hardware data path is associated with an error count above a sixth threshold, or a combination of both, wherein the error count is one of Invalid DWord Count, Running Disparity Error Count, or Loss of DWord Sync Count.

14. The non-transitory machine-readable medium of claim 13 , the operations further comprising:

switching all future IO operations to the second storage processor in response to the first storage processor or its associated hardware data path having been determined as faulty.

15. A data processing system, comprising:

a processor; and

a memory coupled to the processor storing instructions which, when executed by the processor, cause the processor to perform troubleshooting operations, the operations including

collecting data relating to a plurality of input/output (IO) errors associated with a first storage processor within the data processing system, wherein the data storage system includes a plurality of disk array enclosures (DAEs), each DAE having one or more disk drives;

compiling IO error statistics based on the data relating to the plurality of IO errors, the IO error statistics being related to a first one of the DAEs of the data storage system; and

determining a faulty hardware component based on the IO error statistics, wherein the determining of the faulty hardware component comprises utilizing a second storage processor of the data processing system independent from the first storage processor, including examining IO access statistics of the second storage processor for accessing the first DAE through a path, and

wherein the plurality of DAEs are connected to the first storage processor and the second storage processor through an independent first path and an independent second path, and each of the one or more disk drives has a first port connected to the first storage processor through the first path and a second port connected to the second storage processor through the second path.

16. The system of claim 15 , wherein the data relating to the plurality of IO errors associated with the first storage processor comprise timestamps of the plurality of IO errors and drives of the data processing system associated with each of the plurality of IO errors.

17. The system of claim 15 , wherein a count of the plurality of IO errors is above a third threshold, the plurality of IO errors take place over a period of time longer than a second time period threshold, and a time difference between any two consecutive IO errors within the plurality of IO errors is below a first time period threshold.

18. The system of claim 15 , the operations further comprising:

determining drives implicated in the plurality of IO errors and a count of IO errors associated with each implicated drive, wherein each implicated drive is associated with more IO errors than a fourth threshold.

19. The system of claim 18 , the operations further comprising:

determining the drives implicated in the plurality of IO errors as faulty based on a fact that 1) IO operations cannot be successfully performed on the drives implicated in the plurality of IO errors with the second storage processor, 2) a count of the drives implicated in the plurality of IO errors is below a fifth threshold, or a combination of both.

20. The system of claim 18 , the operations further comprising:

determining the first storage processor or its associated hardware data path as faulty based on a fact that 1) a count of the drives implicated in the plurality of IO errors is above a fifth threshold, 2) the first storage processor or at least one connected entity within the hardware data path is associated with an error count above a sixth threshold, or a combination of both, wherein the error count is one of Invalid DWord Count, Running Disparity Error Count, or Loss of DWord Sync Count.

21. The system of claim 20 , the operations further comprising:

switching all future IO operations to the second storage processor in response to the first storage processor or its associated hardware data path having been determined as faulty.

Assignments (8)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (050724/0466) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO WYSE TECHNOLOGY L.L.C.)
Reel/Frame 060753/0486 →
RELEASE OF SECURITY INTEREST AT REEL 050405 FRAME 0534 Recorded Nov 2, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.
Reel/Frame 058001/0001 →
PATENT SECURITY AGREEMENT (NOTES) Recorded Oct 15, 2019
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 050724/0466 →
SECURITY AGREEMENT Recorded Sep 17, 2019
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 050405/0534 →
SECURITY AGREEMENT Recorded Mar 21, 2019
From: CREDANT TECHNOLOGIES, INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 049452/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2017
From: EMC CORPORATION
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 041872/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2016
From: RAVINDRANATH, ANIL; GUDIPATI, KRISHNA
To: EMC CORPORATION
Reel/Frame 039169/0509 →