IP Library Granted Patent US 9,158,540
Granted Patent B1
US 9,158,540 · App. 13/676,019 · Granted Oct 13, 2015

Method and apparatus for offloading compute resources to a flash co-processing appliance

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,158,540
App. No.
13/676,019
Granted
Oct 13, 2015
Kind
B1
Abstract

Solid-State Drive (SSD) burst buffer nodes are interposed into a parallel supercomputing cluster to enable fast burst checkpoint of cluster memory to or from nearby interconnected solid-state storage with asynchronous migration between the burst buffer nodes and slower more distant disk storage. The SSD nodes also perform tasks offloaded from the compute nodes or associated with the checkpoint data. For example, the data for the next job is preloaded in the SSD node and very fast uploaded to the respective compute node just before the next job starts. During a job, the SSD nodes perform fast visualization and statistical analysis upon the checkpoint data. The SSD nodes can also perform data reduction and encryption of the checkpoint data.

Claims (30)

1. A parallel supercomputing cluster comprising:

compute nodes interconnected in a mesh of data links for executing a Message Passing Interface (MPI) job using MPI data transfer between the computer nodes over the mesh of data links; and

solid-state storage nodes each linked to a respective group of the compute nodes for receiving checkpoint data from the respective compute nodes, and

magnetic disk storage linked to each of the solid-state storage nodes for asynchronous migration of the checkpoint data from the solid-state storage nodes to the magnetic disk storage;

wherein each solid-state storage node includes a data processor coupled to the respective group of compute nodes for receiving the checkpoint data from the respective group of compute nodes and coupled to the magnetic disk storage for transmitting the checkpoint data to the magnetic disk storage, solid state storage coupled to the data processor for buffering the checkpoint data, and non-transitory computer readable storage medium storing computer instructions that, when executed by the data processor, perform the steps of:

(a) presenting a file system interface to the MPI job, and multiple MPI processes of the MPI job writing the checkpoint data to a shared file in the solid-state storage in a strided fashion in a first data layout;

(b) asynchronously migrating the checkpoint data from the shared file in the solid-state storage to the magnetic disk storage and writing the checkpoint data to the magnetic disk storage in a sequential fashion in a second data layout; and

(c) performing additional tasks offloaded from the compute nodes or associated with the checkpoint data.

2. The parallel supercomputing cluster as claimed in claim 1 , wherein the additional tasks include preloading data into said each solid-state storage node for processing by the MPI job in the respective group of compute nodes.

3. The parallel supercomputing cluster as claimed in claim 2 , which further includes a control station computer and a network coupling the control station computer to each of the solid-state storage nodes, and the pre-loaded data is transmitted from the control station computer to the solid-state storage nodes over the network coupling the control station computer to each of the solid-state storage nodes.

4. The parallel supercomputing cluster as claimed in claim 3 , wherein the pre-loaded data is transmitted from the control station computer to the solid-state storage nodes over the network coupling the control station computer to each of the solid-state storage nodes and preloaded into said each of the solid-state storage nodes without any impact on MPI traffic on the mesh of data links.

5. The parallel supercomputing cluster as claimed in claim 3 , wherein all of the pre-loaded data is transmitted from the control station computer to the solid-state storage nodes over the network coupling the control station computer to each of the solid-state storage nodes for the MPI job before completion of execution of a previous MPI job.

6. The parallel supercomputing cluster as claimed in claim 1 , wherein the additional tasks include processing of the checkpoint data to produce a visualization of the checkpoint data presented in real time to a user.

7. The parallel supercomputing cluster as claimed in claim 1 , wherein the additional tasks include performing a statistical analysis of the checkpoint data presented in real time to a user.

8. The parallel supercomputing cluster as claimed in claim 1 , wherein the additional tasks include performing an analysis of the checkpoint data in order to terminate a simulation upon detection of a simulation error.

9. The parallel supercomputing cluster as claimed in claim 1 , wherein the additional tasks include performing data reduction operations upon the checkpoint data to reduce the magnetic disk storage capacity needed to store the checkpoint data.

10. The parallel supercomputing cluster as claimed in claim 1 , wherein the additional tasks include encrypting the checkpoint data so that the encrypted checkpoint data is stored in the magnetic disk storage.

11. A method of operating a parallel supercomputing cluster, the parallel supercomputing cluster including compute nodes, solid-state storage nodes, and magnetic disk storage, the compute nodes being interconnected in a mesh of data links for executing a Message Passing Interface (MPI) job using MPI data transfer between the computer nodes over the mesh of data links, each of the solid-state storage nodes being linked to a respective group of the compute nodes for receiving checkpoint data from the respective compute nodes, and the magnetic disk storage being linked to each of the solid-state storage nodes for asynchronous migration of the checkpoint data from the solid-state storage nodes to the magnetic disk storage, and each of the solid-state storage nodes including a data processor coupled to the respective group of compute nodes for receiving the checkpoint data from the respective group of compute nodes and coupled to the magnetic disk storage for transmitting the checkpoint data to the magnetic disk storage, and solid-state storage coupled to the data processor for buffering the checkpoint data, and non-transitory computer readable storage medium storing computer instructions, said method comprising the data processor executing the computer instructions to perform the steps of:

(a) presenting a file system interface to the MPI job, and multiple MPI processes of the MPI job writing the checkpoint data to a shared file in the solid-state storage in a strided fashion in a first data layout;

(b) asynchronously migrating the checkpoint data from the shared file in the solid-state storage to the magnetic disk storage and writing the checkpoint data to the magnetic disk storage in a sequential fashion in a second data layout; and

(c) performing additional tasks offloaded from the compute nodes or associated with the checkpoint data.

12. The method as claimed in claim in claim 11 , wherein the additional tasks include preloading data into said each solid-state storage node for processing by the MPI job in the respective group of compute nodes.

13. The method as claimed in claim 12 , which the parallel supercomputing cluster further includes a control station computer and a network coupling the control station computer to each of the solid-state storage nodes, and the method further includes transmitting the pre-loaded data from the control station computer to the solid-state storage nodes over the network coupling the control station computer to each of the solid-state storage nodes.

14. The method as claimed in claim 13 , wherein the pre-loaded data is transmitted from the control station computer to the solid-state storage nodes over the network coupling the control station computer to each of the solid-state storage nodes and preloaded into said each of the solid-state storage nodes without any impact on MPI traffic on the mesh of data links.

15. The method as claimed in claim 13 , wherein all of the pre-loaded data is transmitted from the control station computer to the solid-state storage nodes over the network coupling the control station computer to each of the solid-state storage nodes for the MPI job before completion of execution of a previous MPI job.

16. The method as claimed in claim 11 , wherein the additional tasks include processing of the checkpoint data to produce a visualization of the checkpoint data presented in real time to a user.

17. The method as claimed in claim 11 , wherein the additional tasks include performing a statistical analysis of the checkpoint data presented in real time to a user.

18. The method as claimed in claim 11 , wherein the additional tasks include performing an analysis of the checkpoint data in order to terminate a simulation upon detection of a simulation error.

19. The method as claimed in claim 11 , wherein the additional tasks include performing data reduction operations upon the checkpoint data to reduce the magnetic disk storage capacity needed to store the checkpoint data.

20. The method as claimed in claim 11 , wherein the additional tasks include encrypting the checkpoint data so that the encrypted checkpoint data is stored in the magnetic disk storage.

Assignments (13)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (045455/0001) Recorded May 20, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO ASAP SOFTWARE EXPRESS, INC.); DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL USA L.P.; DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC CORPORATION (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO MAGINATICS LLC); EMC IP HOLDING COMPANY LLC (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO MOZY, INC.); SCALEIO LLC
Reel/Frame 061753/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (040136/0001) Recorded Apr 26, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO ASAP SOFTWARE EXPRESS, INC.); DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL USA L.P.; DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC CORPORATION (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO MAGINATICS LLC); EMC IP HOLDING COMPANY LLC (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO MOZY, INC.); SCALEIO LLC
Reel/Frame 061324/0001 →
RELEASE OF SECURITY INTEREST Recorded Nov 3, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: ASAP SOFTWARE EXPRESS, INC.; AVENTAIL LLC; CREDANT TECHNOLOGIES, INC.; DELL USA L.P.; DELL INTERNATIONAL, L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL SOFTWARE INC.; DELL SYSTEMS CORPORATION; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; FORCE10 NETWORKS, INC.; MAGINATICS LLC; MOZY, INC.; SCALEIO LLC; WYSE TECHNOLOGY L.L.C.
Reel/Frame 058216/0001 →
SECURITY AGREEMENT Recorded Apr 22, 2020
From: CREDANT TECHNOLOGIES INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 053546/0001 →
SECURITY AGREEMENT Recorded Mar 21, 2019
From: CREDANT TECHNOLOGIES, INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 049452/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2018
From: LOS ALAMOS NATIONAL SECURITY, LLC
To: TRIAD NATIONAL SECURITY, LLC
Reel/Frame 048007/0874 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2016
From: EMC CORPORATION
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 040203/0001 →
SECURITY AGREEMENT Recorded Sep 21, 2016
From: ASAP SOFTWARE EXPRESS, INC.; AVENTAIL LLC; CREDANT TECHNOLOGIES, INC.; DELL USA L.P.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL SOFTWARE INC.; DELL SYSTEMS CORPORATION; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; FORCE10 NETWORKS, INC.; MAGINATICS LLC; MOZY, INC.; SCALEIO LLC; SPANNING CLOUD APPS LLC; WYSE TECHNOLOGY L.L.C.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 040134/0001 →
SECURITY AGREEMENT Recorded Sep 21, 2016
From: ASAP SOFTWARE EXPRESS, INC.; AVENTAIL LLC; CREDANT TECHNOLOGIES, INC.; DELL USA L.P.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL SOFTWARE INC.; DELL SYSTEMS CORPORATION; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; FORCE10 NETWORKS, INC.; MAGINATICS LLC; MOZY, INC.; SCALEIO LLC; SPANNING CLOUD APPS LLC; WYSE TECHNOLOGY L.L.C.
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 040136/0001 →
CONFIRMATORY LICENSE Recorded Mar 6, 2014
From: LOS ALAMOS NATIONAL SECURITY
To: U.S. DEPARTMENT OF ENERGY
Reel/Frame 032361/0308 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2013
From: TZELNIC, PERCY; FAIBISH, SORIN; GUPTA, UDAY K.; BENT, JOHN
To: EMC CORPORATION
Reel/Frame 029658/0370 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2013
From: GRIDER, GARY ALAN; CHEN, HSING-BUNG
To: LOS ALAMOS NATIONAL SECURITY, LLC
Reel/Frame 029659/0097 →