Identifying dataset data blocks in a chunk to apply tiering/protection updates
A system can maintain a group of data sets, a vector database that comprises respective first identifiers of respective data sets of the group of data sets stored on respective data chunks, respective second identifiers in respective chunk descriptors of the respective data chunks, wherein the respective second identifiers identify the respective first identifiers in the vector database. The system can, based on determining to adjust a characteristic of a data set, determine, from a second identifier of the second identifiers that is stored in a chunk descriptor of the chunk descriptors, that the vector database indicates that a data chunk of the data chunks identifies that the data chunk stores at least part of the data set, and adjust the characteristic of the data set as applicable to at least the part of the data set that is stored in the data chunk.
1 . A system, comprising:
at least one processor; and
at least one memory that stores executable instructions that, when executed by the at least one processor, facilitate performance of operations, comprising:
maintaining a group of data sets on a cluster file system;
maintaining a vector database that comprises respective first identifiers of respective data sets of the group of data sets stored on respective data chunks of the cluster file system;
maintaining respective second identifiers in respective chunk descriptors of the respective data chunks, wherein the respective second identifiers identify the respective first identifiers in the vector database; and
based on determining to adjust a characteristic of a data set of the group of data sets,
determining, from a second identifier of the second identifiers that is stored in a chunk descriptor of the chunk descriptors, that the vector database indicates that a data chunk of the data chunks, which corresponds to the chunk descriptor, identifies that the data chunk stores at least part of the data set, and
adjusting the characteristic of the data set as applicable to at least the part of the data set that is stored in the data chunk, wherein the adjusting comprises locating the at least the part of the data set in the data chunk based on a first pointer in the chunk descriptor that points to a virtual chunk extent, wherein the adjusting comprises adjusting a data tiering applicable to the data set or adjusting a data protection applicable to the data set, wherein the virtual chunk extent comprises an indication of the data set, wherein a chain of virtual chunk extents comprises the virtual chunk extent, and wherein respective virtual chunk extents of the chain of virtual chunk extents comprise respective third identifiers to respective portions of the vector database that indicate which data sets in the data chunk are pointed to by the respective virtual chunk extents.
2 . The system of claim 1 , wherein the adjusting of the characteristic of the data set as applicable to at least the part of the data set that is stored in the data chunk comprises:
locating the at least the part of the data set in the data chunk based on the first pointer in the chunk descriptor that points to the virtual chunk extent, wherein the virtual chunk extent comprises a second pointer that points to a location of at least the part of the data set in the data chunk, and wherein the respective virtual chunk extents comprise respective second pointers to respective locations of at least parts of data sets in the data chunk.
3 . The system of claim 1 , wherein the respective first identifiers comprise respective vectors, wherein respective elements of the respective vectors comprise the respective first identifiers of the respective data sets stored on respective data chunks, and wherein the respective vectors correspond to the respective data chunks.
4 . The system of claim 1 , wherein the adjusting of the characteristic of the data set is performed as a background process relative to a foreground process that facilitates reads and writes of data of the cluster file system.
5 . The system of claim 1 , wherein at least the part of the data set is at least a first part of the data set, and wherein the adjusting of the characteristic of the data set comprises:
parsing respective chunk descriptors of the chunk descriptors to determine, from the vector database, whether the corresponding respective chunks store at least a second part of the data set that is distinct from at least the first part.
6 . The system of claim 1 , wherein the cluster file system is stored across a group of volumes, wherein at least the part of the data set is at least a first part of the data set, and wherein the adjusting of the characteristic of the data set comprises:
for respective volumes of the group of volumes, parsing respective chunk descriptors of the chunk descriptors to determine, from the vector database, whether the corresponding respective chunks store at least a second part of the data set that is distinct from at least the first part.
7 . The system of claim 1 , wherein the adjusting of the characteristic of the data set is performed independently of a complete parsing of a name space of the cluster file system.
8 . A method, comprising:
storing, in a vector database by a system comprising at least one processor, respective first identifiers of respective data sets stored on respective chunks of a cluster file system, and respective second identifiers in respective chunk descriptors of the respective chunks, wherein the respective second identifiers identify the respective first identifiers in the vector database; and
based on determining to adjust a characteristic of a data set of the respective data sets,
determining, by the system and from a second identifier of the second identifiers that is stored in a chunk descriptor of the chunk descriptors, that the vector database indicates that a chunk of the chunks, which corresponds to the chunk descriptor, identifies that the chunk stores at least part of the data set, and
adjusting, by the system, the characteristic of the data set as applied to at least the part of the data set that is stored in the chunk, wherein the adjusting comprises locating the at least the part of the data set in the chunk based on a first pointer in the chunk descriptor that points to a virtual chunk extent, wherein the adjusting comprises adjusting a data tiering applicable to the data set or adjusting a data protection applicable to the data set, wherein the virtual chunk extent comprises an indication of the data set, wherein a chain of virtual chunk extents comprises the virtual chunk extent, and wherein respective virtual chunk extents of the chain of virtual chunk extents comprise respective third identifiers to respective portions of the vector database that indicate which data sets in the data chunk are pointed to by the respective virtual chunk extents.
9 . The method of claim 8 , wherein the adjusting of the characteristic of the data set as applied to at least the part of the data set that is stored in the chunk comprises:
identifying the at least the part of the data set based on the first pointer in the chunk descriptor that points to the virtual chunk extent, wherein the virtual chunk extent comprises a second pointer that points to a location of at least the part of the data set in the chunk.
10 . The method of claim 8 , wherein the respective first identifiers comprise respective vectors, wherein the respective vectors comprise respective elements that store the respective first identifiers, and wherein the respective vectors correspond to the respective chunks.
11 . The method of claim 8 , wherein the adjusting of the characteristic of the data set is performed as a background process relative to a foreground process that executes read and write operations on data from and to the cluster file system.
12 . The method of claim 8 , wherein the adjusting of the data set as applied to at least the part of the data set that is stored in the data chunk comprises:
locating the at least the part of the data set in the data chunk based on the first pointer in the chunk descriptor that points to the virtual chunk extent, wherein the virtual chunk extent comprises a second pointer that points to a location of at least the part of the data set in the data chunk, and wherein the respective virtual chunk extents comprise respective second pointers to respective locations of at least parts of data sets in the data chunk.
13 . The method of claim 8 , wherein the respective first identifiers comprise respective vectors, wherein respective elements of the respective vectors comprise the respective first identifiers of the respective data sets stored on respective data chunks, and wherein the respective vectors correspond to the respective data chunks.
14 . A non-transitory computer-readable medium comprising instructions that, in response to execution, cause a system comprising at least one processor to perform operations, comprising:
maintaining a data store that identifies respective data sets stored on respective chunks of a cluster file system, and respective chunk descriptors of the respective chunks that identify the respective first identifiers in the data store; and
based on determining to adjust a data set of the respective data sets,
determining, from a second identifier of the second identifiers that is stored in a chunk descriptor of the chunk descriptors, that the data store indicates that a data chunk of the data chunks, which corresponds to the chunk descriptor, identifies that the data chunk stores at least part of the data set, and
adjusting the data set as applied to at least the part of the data set that is stored in the data chunk, wherein the adjusting comprises locating the at least the part of the data set in the data chunk based on a first pointer in the chunk descriptor that points to a virtual chunk extent, wherein the adjusting comprises adjusting a data tiering applicable to the data set or adjusting a data protection applicable to the data set, wherein the virtual chunk extent comprises an indication of the data set, wherein a chain of virtual chunk extents comprises the virtual chunk extent, and wherein respective virtual chunk extents of the chain of virtual chunk extents comprise respective third identifiers to respective portions of the data store that indicate which data sets in the data chunk are pointed to by the respective virtual chunk extents.
15 . The non-transitory computer-readable medium of claim 14 , wherein at least the part of the data set is at least a first part of the data set, and wherein the adjusting of the data set comprises:
parsing respective chunk descriptors of the chunk descriptors to determine from the data store whether the corresponding respective chunks store at least a second part of the data set.
16 . The non-transitory computer-readable medium of claim 14 , wherein the cluster file system is stored across a group of volumes, wherein at least the part of the data set is at least a first part of the data set, and wherein the adjusting of the data set comprises:
for respective volumes of the group of volumes, parsing respective chunk descriptors of the chunk descriptors to determine, from the data store, whether the corresponding respective chunks store at least a second part of the data set.
17 . The non-transitory computer-readable medium of claim 14 , wherein the adjusting of the data set is performed independently of a parsing of a name space of the cluster file system.
18 . The non-transitory computer-readable medium of claim 14 , wherein the data store comprises a vector database.
19 . The non-transitory computer-readable medium of claim 14 , wherein the adjusting of the data set as applied to at least the part of the data set that is stored in the data chunk comprises:
locating the at least the part of the data set in the data chunk based on the first pointer in the chunk descriptor that points to the virtual chunk extent, wherein the virtual chunk extent comprises a second pointer that points to a location of at least the part of the data set in the data chunk, and wherein the respective virtual chunk extents comprise respective second pointers to respective locations of at least parts of data sets in the data chunk.
20 . The non-transitory computer-readable medium of claim 14 , wherein the respective first identifiers comprise respective vectors, wherein respective elements of the respective vectors comprise the respective first identifiers of the respective data sets stored on respective data chunks, and wherein the respective vectors correspond to the respective data chunks.