IP Library Granted Patent US 10,785,294
Granted Patent B1
US 10,785,294 · App. 14/813,516 · Granted Sep 22, 2020

Methods, systems, and computer readable mediums for managing fault tolerance of hardware storage nodes

Inventors: Ryan Joseph Andersen (Belmont, MA); Donald Edward Norbeck, Jr. (Lafayette Hill, PA); Jonathan Peter Streete (San Jose, CA); Seamus Patrick Kerrigan (Ballincollig, IE)
Assignee: EMC IP HOLDING COMPANY LLC
H04L67/1095G06F3/065G06F3/0619G06F3/0683H04L67/1097
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,785,294
App. No.
14/813,516
Granted
Sep 22, 2020
Kind
B1
Abstract

Methods, systems, and computer readable mediums for managing fault tolerance. A method includes receiving a request to establish a pool of data storage for an application of a distributed computing system. The distributed computing system includes hardware storage nodes integrated with compute nodes. The method includes receiving a target level of fault tolerance for the pool of data storage. The method includes establishing the pool of data storage by specifying, for each hardware storage node, a mirror hardware storage node for mirroring data stored on the hardware storage node so that the hardware storage node and the mirror hardware storage node do not share one or more pieces of physical equipment as specified in a physical layout of the hardware storage nodes to meet the target level of fault tolerance.

Claims (23)

1. A method for managing fault tolerance, the method comprising:

receiving, by one or more computers, a request to establish a pool of data storage for an application of a distributed computing system, the distributed computing system comprising a plurality of hardware storage nodes each integrated with a respective compute node, wherein the distributed computing system comprises a hyper-converged system comprising a plurality of hyper-converged nodes, wherein the hyper-converged system is configured to perform virtualization and includes a hyper-converged storage manager configured to present the pool of data storage to the application as a single logical storage volume;

receiving, by the one or more computers, a target level of fault tolerance for the pool of data storage;

determining a physical layout of the hardware storage nodes by querying the hyper-converged nodes;

establishing, by the one or more computers, the pool of data storage by specifying, for each hardware storage node, a mirror hardware storage node for mirroring data stored on the hardware storage node so that the hardware storage node and the mirror hardware storage node do not share one or more pieces of physical equipment as specified in the physical layout of the hardware storage nodes to meet the target level of fault tolerance; and

auditing the pool of data storage by determining, using the physical layout of the hardware storage nodes, whether any of the hardware storage nodes shares a piece of physical equipment with a mirror hardware storage node in violation of the target level of fault tolerance and generating a list of hardware storage nodes that do not meet the target level of fault tolerance for the pool of data storage;

wherein the distributed computing system comprises a plurality of equipment racks housing the hardware storage nodes in a plurality of chassis, wherein the physical layout of the hardware storage nodes specifies a chassis identifier for each hardware storage node indicating which chassis houses the hardware storage node and a rack identifier for each hardware storage node indicating which equipment rack houses the hardware storage node, and wherein specifying a mirror hardware storage node for each hardware storage node comprises finding, for each hardware storage node, an unallocated hardware storage node having a chassis identifier different from the chassis identifier of the hardware storage node and a rack identifier different from the rack identifier of the hardware storage node, thereby achieving chassis-level and rack-level fault tolerance for the pool of data storage.

2. A system comprising:

one or more physical computers; and

a virtual storage aggregator implemented on the one or more physical computers for performing operations comprising:

receiving a request to establish a pool of data storage for an application of a distributed computing system, the distributed computing system comprising a plurality of hardware storage nodes each integrated with a respective compute node, wherein the distributed computing system comprises a hyper-converged system comprising a plurality of hyper-converged nodes, wherein the hyper-converged system is configured to perform virtualization and includes a hyper-converged storage manager configured to present the pool of data storage to the application as a single logical storage volume;

receiving a target level of fault tolerance for the pool of data storage;

determining a physical layout of the hardware storage nodes by querying the hyper-converged nodes;

establishing the pool of data storage by specifying, for each hardware storage node, a mirror hardware storage node for mirroring data stored on the hardware storage node so that the hardware storage node and the mirror hardware storage node do not share one or more pieces of physical equipment as specified in the physical layout of the hardware storage nodes to meet the target level of fault tolerance; and

auditing the pool of data storage by determining, using the physical layout of the hardware storage nodes, whether any of the hardware storage nodes shares a piece of physical equipment with a mirror hardware storage node in violation of the target level of fault tolerance and generating a list of hardware storage nodes that do not meet the target level of fault tolerance for the pool of data storage;

wherein the distributed computing system comprises a plurality of equipment racks housing the hardware storage nodes in a plurality of chassis, wherein the physical layout of the hardware storage nodes specifies a chassis identifier for each hardware storage node indicating which chassis houses the hardware storage node and a rack identifier for each hardware storage node indicating which equipment rack houses the hardware storage node, and wherein specifying a mirror hardware storage node for each hardware storage node comprises finding, for each hardware storage node, an unallocated hardware storage node having a chassis identifier different from the chassis identifier of the hardware storage node and a rack identifier different from the rack identifier of the hardware storage node, thereby achieving chassis-level and rack-level fault tolerance for the pool of data storage.

3. A non-transitory computer readable medium having stored thereon executable instructions which, when executed by one or more physical computers, cause the one or more physical computers to perform operations comprising:

receiving a request to establish a pool of data storage for an application of a distributed computing system, the distributed computing system comprising a plurality of hardware storage nodes each integrated with a respective compute node, wherein the distributed computing system comprises a hyper-converged system comprising a plurality of hyper-converged nodes, wherein the hyper-converged system is configured to perform virtualization and includes a hyper-converged storage manager configured to present the pool of data storage to the application as a single logical storage volume;

receiving a target level of fault tolerance for the pool of data storage;

determining a physical layout of the hardware storage nodes by querying the hyper-converged nodes;

establishing the pool of data storage by specifying, for each hardware storage node, a mirror hardware storage node for mirroring data stored on the hardware storage node so that the hardware storage node and the mirror hardware storage node do not share one or more pieces of physical equipment as specified in the physical layout of the hardware storage nodes to meet the target level of fault tolerance; and

auditing the pool of data storage by determining, using the physical layout of the hardware storage nodes, whether any of the hardware storage nodes shares a piece of physical equipment with a mirror hardware storage node in violation of the target level of fault tolerance and generating a list of hardware storage nodes that do not meet the target level of fault tolerance for the pool of data storage;

wherein the distributed computing system comprises a plurality of equipment racks housing the hardware storage nodes in a plurality of chassis, wherein the physical layout of the hardware storage nodes specifies a chassis identifier for each hardware storage node indicating which chassis houses the hardware storage node and a rack identifier for each hardware storage node indicating which equipment rack houses the hardware storage node, and wherein specifying a mirror hardware storage node for each hardware storage node comprises finding, for each hardware storage node, an unallocated hardware storage node having a chassis identifier different from the chassis identifier of the hardware storage node and a rack identifier different from the rack identifier of the hardware storage node, thereby achieving chassis-level and rack-level fault tolerance for the pool of data storage.

Assignments (6)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
SECURITY AGREEMENT Recorded Apr 22, 2020
From: CREDANT TECHNOLOGIES INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 053546/0001 →
MERGER Recorded Mar 26, 2020
From: VCE IP HOLDING COMPANY LLC
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 052236/0497 →
SECURITY AGREEMENT Recorded Mar 21, 2019
From: CREDANT TECHNOLOGIES, INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 049452/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2016
From: VCE COMPANY, LLC
To: VCE IP HOLDING COMPANY LLC
Reel/Frame 040576/0161 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2015
From: ANDERSEN, RYAN JOSEPH; NORBECK, DONALD EDWARD, JR.; STREETE, JONATHAN PETER; KERRIGAN, SEAMUS PATRICK
To: VCE COMPANY, LLC
Reel/Frame 036246/0552 →