Transport of non-standardized data between relational database operations
A method for processing non-standardized data in a relational database may include identifying, within a sequence of operations forming a query pipeline for executing a query, a first operation ingesting a non-standardized data. In response to identifying the first operation ingesting the non-standardized data, a second operation may be inserted before the first operation. The non-standardized data may be output by a third operation preceding the first operation or a source external to the query pipeline. The second operation may serialize the non-standardized data for ingestion by the first operation, for example, by generating a relational table populated by the non-standardized data. The query may be executed by performing the sequence of operations included in the query pipeline. Related systems and computer program products are also provided.
1. A system, comprising:
at least one data processor; and
at least one memory storing instructions which, when executed by the at least one data processor, cause operations comprising:
identifying, within a sequence of operations comprising a query pipeline for executing a query, a first operation ingesting non-standardized data;
in response to identifying the first operation ingesting the non-standardized data, inserting, before the first operation, a second operation to serialize at least a portion of the non-standardized data for ingestion by the first operation; and
executing the query by at least performing the sequence of operations comprising the query pipeline.
2. The system of claim 1 , wherein the second operation serializes at least the portion of the non-standardized data by generating a relational table populated by at least the portion of the non-standardized data.
3. The system of claim 2 , wherein each row of the relational table is populated with a different type of the non-standardized data.
4. The system of claim 2 , wherein each column of the relational table is populated with a different type of the non-standardized data table.
5. The system of claim 1 , wherein the non-standardized data includes one or more data statistics.
6. The system of claim 5 , wherein the one or more data statistics include a row count, a column count, and/or a datatype.
7. The system of claim 1 , wherein the non-standardized data includes synchronization information.
8. The system of claim 7 , wherein the synchronization information includes a versioning timestamp and/or a table identifier.
9. The system of claim 1 , wherein the non-standardized data includes a type of algorithm applied to execute the query, temporary and/or auxiliary data structures used during the executing of the query, information about parallelization or scheduling, and/or status information.
10. The system of claim 1 , wherein the non-standardized data is output by a third operation preceding the first operation, and wherein the second operation is inserted between the first operation and the third operation.
11. The system of claim 1 , wherein the non-standardized data is output by a source external to the query pipeline.
12. The system of claim 1 , wherein the operations further comprise:
in response to determining that the first operation operates on the non-standardized data in its original format, inserting, between the second operation and the first operation, a third operation to convert the portion of the non-standardized, which is serialized, data back to the original format for use by the first operation.
13. A computer-implemented method, comprising:
identifying, within a sequence of operations comprising a query pipeline for executing a query, a first operation ingesting non-standardized data;
in response to identifying the first operation ingesting non-standardized data, inserting, before the first operation, a second operation to serialize at least a portion of the non-standardized data for ingestion by the first operation; and
executing the query by at least performing the sequence of operations comprising the query pipeline.
14. The method of claim 13 , wherein the second operation serializes at least the portion of the non-standardized data by generating a relational table in which each row or column of the relational table is populated by a different type of the non-standardized data.
15. The method of claim 13 , wherein the non-standardized data includes one or more data statistics comprising at least one of a row count, a column count, or a datatype.
16. The method of claim 13 , wherein the non-standardized data includes synchronization information comprising at least one of a versioning timestamp or a table identifier.
17. The method of claim 13 , wherein the non-standardized data includes a type of algorithm applied to execute the query, temporary and/or auxiliary data structures used during the executing of the query, information about parallelization or scheduling, and/or status information.
18. The method of claim 13 , wherein the non-standardized data output is output by a third operation preceding the first operation or a source external to the query pipeline.
19. The method of claim 13 , further comprising:
in response to determining that the first operation operates on the non-standardized data in its original format, inserting, between the second operation and the first operation, a third operation to convert the portion of the non-standardized, which is serialized, data back to the original format for use by the first operation.
20. A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising:
identifying, within a sequence of operations comprising a query pipeline for executing a query, a first operation ingesting non-standardized data;
in response to identifying the first operation ingesting the non-standardized data, inserting, before the first operation, a second operation to serialize at least a portion of the non-standardized data for ingestion by the first operation; and
executing the query by at least performing the sequence of operations comprising the query pipeline.