Efficient data delivery for artificial intelligence systems in distributed storage environments
A method is disclosed for managing transformed datasets in a compute cluster environment. The method includes identifying, based on one or more machine learning models to be executed on a compute cluster comprising a plurality of GPU servers, one or more transformations to apply to a dataset. The method further includes generating a transformed dataset based on the one or more transformations, storing the transformed dataset, receiving a request to transmit the transformed dataset to at least one GPU server of the plurality of GPU servers, and, responsive to the request, transmitting the stored transformed dataset to the at least one GPU server without re-performing the one or more transformations on the dataset.
1 . A method comprising:
identifying, based on one or more machine learning models to be executed on a compute cluster comprising a plurality of graphics processing unit (GPU) servers, one or more transformations to apply to a dataset;
generating multiple transformed datasets having different levels of precision or format;
storing the multiple transformed datasets;
receiving a request to transmit a transformed dataset to at least one GPU server of the plurality of GPU servers;
selecting one of the multiple transformed datasets for transmission based on a resource availability of the at least one GPU server; and
responsive to the request, transmitting, to the at least one GPU server without re-performing the one or more transformations on the dataset, the selected one of the multiple transformed datasets.
2 . The method of claim 1 , wherein transmitting the selected one of the multiple transformed datasets to the at least one GPU server comprises performing a remote direct memory access from a storage node of a distributed storage system to a memory of the at least one GPU server.
3 . The method of claim 1 , wherein storing the multiple transformed datasets comprises persisting one or more intermediate states of the multiple transformed datasets in a non-volatile memory tier of a distributed storage system.
4 . The method of claim 1 , wherein identifying the one or more transformations comprises analyzing metadata describing at least one of a schema, structure, or historical usage of the dataset.
5 . The method of claim 1 , wherein storing the multiple transformed datasets comprises storing the multiple transformed datasets in a distributed storage system that provides a single global namespace accessible via multiple storage protocols.
6 . The method of claim 1 , further comprising:
maintaining lineage metadata linking the dataset, the one or more transformations, and the multiple transformed datasets; and
using the lineage metadata to avoid redundant performance of the one or more transformations.
7 . An apparatus comprising:
a memory;
a processing device, operatively coupled to the memory, configured to:
identify, based on one or more machine learning models to be executed on a compute cluster comprising a plurality of graphics processing unit (GPU) servers, one or more transformations to apply to a dataset;
generate multiple transformed datasets having different levels of precision or format;
store the multiple transformed datasets;
receive a request to transmit a transformed dataset to at least one GPU server of the plurality of GPU servers;
select one of the multiple transformed datasets for transmission based on a resource availability of the at least one GPU server; and
responsive to the request, transmit, to the at least one GPU server without re-performing the one or more transformations on the dataset, the selected one of the multiple transformed datasets.
8 . The apparatus of claim 7 , wherein, to transmit the selected one of the multiple transformed datasets to the at least one GPU server, the processing device is configured to perform a remote direct memory access from a storage node of a distributed storage system to a memory of the at least one GPU server.
9 . The apparatus of claim 7 , wherein, to store the multiple transformed datasets, the processing device is configured to persist one or more intermediate states of the transformed dataset in a non-volatile memory tier of a distributed storage system.
10 . The apparatus of claim 7 , wherein, to identify the one or more transformations, the processing devices is configured to analyze metadata describing at least one of a schema, structure, or historical usage of the dataset.
11 . The apparatus of claim 7 , wherein, to store the multiple transformed datasets, the processing device is configured to store the multiple transformed datasets in a distributed storage system that provides a single global namespace accessible via multiple storage protocols.
12 . The apparatus of claim 8 , wherein the processing device is further configured to:
maintain lineage metadata linking the dataset, the one or more transformations, and the multiple transformed datasets; and
use the lineage metadata to avoid redundant performance of the one or more transformations.
13 . A non-transitory computer readable storage medium storing instructions that, when executed, cause a processing device to:
identify, based on one or more machine learning models to be executed on a compute cluster comprising a plurality of graphics processing unit (GPU) servers, one or more transformations to apply to a dataset;
generate multiple transformed datasets having different levels of precision or format;
store the multiple transformed datasets;
receive a request to transmit a transformed dataset to at least one GPU server of the plurality of GPU servers;
select one of the multiple transformed datasets for transmission based on a resource availability of the at least one GPU server; and
responsive to the request, transmit, to the at least one GPU server without re-performing the one or more transformations on the dataset, the selected one of the multiple transformed datasets.
14 . The non-transitory computer readable storage medium of claim 13 , wherein, to transmit the selected one of the multiple transformed datasets to the at least one GPU server, the instructions, when executed, cause the processing device to perform a remote direct memory access from a storage node of a distributed storage system to a memory of the at least one GPU server.
15 . The non-transitory computer readable storage medium of claim 13 , wherein, to store the multiple transformed datasets, the instructions, when executed, cause the processing device to persist one or more intermediate states of the transformed dataset in a non-volatile memory tier of a distributed storage system.
16 . The non-transitory computer readable storage medium of claim 13 , wherein, to identify the one or more transformations, the instructions, when executed, cause the processing device to analyze metadata describing at least one of a schema, structure, or historical usage of the dataset.
17 . The non-transitory computer readable storage medium of claim 13 , wherein, to store the multiple transformed datasets, the instructions, when executed, cause the processing device to store the multiple transformed datasets in a distributed storage system that provides a single global namespace accessible via multiple storage protocols.
18 . The non-transitory computer readable storage medium of claim 13 , wherein instructions, when executed, further cause the processing device to:
maintain lineage metadata linking the dataset, the one or more transformations, and the multiple transformed datasets; and
use the lineage metadata to avoid redundant performance of the one or more transformations.