Using stateless nodes to process data of catalog objects
A system and method of using a stateless node to process data of a catalog object. The method includes accessing a catalog object comprising metadata associated with a dataset. The method includes distributing, by one or more processors, a task to a stateless node to cause the stateless node to process the dataset without storing information indicative of a particular state of the stateless node.
1 . A system comprising:
a memory, and
one or more processors, operatively coupled to the memory, the one or more processors to:
access a catalog object comprising metadata associated with a dataset;
create a duplicate catalog object of the catalog object by copying the metadata associated with the dataset without copying a table of the dataset;
distribute a task to a stateless node to cause the stateless node to process the dataset without storing information indicative of a particular state of the stateless node; and
replace the stateless node with a different node without recreating a particular state.
2 . The system of claim 1 , wherein the duplicate catalog object comprises a duplicate hierarchy of one or more generations of children.
3 . The system of claim 1 , wherein to copy the metadata, the one or more processors are further to copy an inventory of the dataset.
4 . The system of claim 1 , further comprising:
copy information regarding the dataset that enables identification of the dataset without requiring access to the dataset.
5 . The system of claim 1 , wherein the one or more processors are further to:
modify the dataset associated with the catalog object to generate modified data that is not visible to the duplicate catalog object of the catalog object.
6 . The system of claim 1 , the one or more processors are further to:
delete the duplicate catalog object of the catalog object in response to modifying the dataset associated with the catalog object.
7 . The system of claim 1 , wherein the one or more processors are further to:
identify the catalog object in a database based on a logical grouping of the dataset in the database.
8 . A method comprising:
accessing a catalog object comprising metadata associated with a dataset;
creating a duplicate catalog object of the catalog object by copying the metadata associated with the dataset without copying a table of the dataset;
distributing, by one or more processors, a task to a stateless node to cause the stateless node to process the dataset without storing information indicative of a particular state of the stateless node; and
replacing the stateless node with a different node without recreating a particular state.
9 . The method of claim 8 , wherein the duplicate catalog object comprises a duplicate hierarchy of one or more generations of children.
10 . The method of claim 9 , further comprising:
copying information regarding the dataset that enables identification of the dataset without requiring access to the dataset.
11 . The method of claim 8 , wherein to copy the metadata comprises:
copying an inventory of the dataset.
12 . The method of claim 8 , further comprising:
modifying the dataset associated with the catalog object to generate modified data that is not visible to the duplicate catalog object of the catalog object.
13 . The method of claim 8 , further comprising:
deleting the duplicate catalog object of the catalog object in response to modifying the dataset associated with the catalog object.
14 . The method of claim 8 , further comprising:
identifying the catalog object in a database based on a logical grouping of the dataset in the database.
15 . A non-transitory computer-readable storage medium comprising instructions which, when executed by one or more processors, cause the one or more processors to:
access a catalog object comprising metadata associated with a dataset;
create a duplicate catalog object of the catalog object by copying the metadata associated with the dataset without copying a table of the dataset;
distribute, by the one or more processors, a task to a stateless node to cause the stateless node to process the dataset without storing information indicative of a particular state of the stateless node; and
replace the stateless node with a different node without recreating a particular state.
16 . The non-transitory computer-readable storage medium of claim 15 , wherein the duplicate catalog object comprises a duplicate hierarchy of one or more generations of children.
17 . The non-transitory computer-readable storage medium of claim 15 , wherein to copy the metadata the one or more processors are further to copy an inventory of the dataset.
18 . The non-transitory computer-readable storage medium of claim 15 , further comprising:
copy information regarding the dataset that enables identification of the dataset without requiring access to the dataset.
19 . The non-transitory computer-readable storage medium of claim 15 , wherein to copy the metadata the one or more processors are further to copy an inventory of the dataset.
20 . The non-transitory computer-readable storage medium of claim 15 , wherein the one or more processors are further to:
modify the dataset associated with the catalog object to generate modified data that is not visible to the duplicate catalog object of the catalog object.
21 . The non-transitory computer-readable storage medium of claim 15 , the one or more processors are further to:
delete the duplicate catalog object of the catalog object in response to modifying the dataset associated with the catalog object.
22 . The non-transitory computer-readable storage medium of claim 15 , wherein the one or more processors are further to:
identify the catalog object in a database based on a logical grouping of the dataset in the database.