Systems and methods for versioning a graph database
Embodiments of the present invention provide methods, systems, and/or the like for versioning a graph representation in a graph data structure. In accordance with one embodiment, a method is provided comprising: conducting a plurality of iterations involving: validating a first data source comprising a new version of data based on a schema from a plurality of schemas in which each schema corresponds to a graph representation found in a graph data structure; and identifying errors in the first source based on the validating of the source; identifying an applicable schema as a schema producing fewer errors than at least one other schema; comparing the first source with a second source comprising a previous version of the data to identify a difference; generating a query for the difference based on the applicable schema; and providing the query for execution to migrate the difference into the graph representation.
1 . A method comprising:
conducting, by computing hardware, a plurality of iterations, wherein an iteration of the plurality of iterations comprises:
validating a first data source comprising a new version of data based on a schema from a plurality of schemas in which each schema in the plurality of schemas corresponds to a graph representation found in a graph data structure;
identifying errors in the first data source based on the validating of the first data source;
identifying, by the computing hardware, an applicable schema from the plurality of schemas, wherein the applicable schema produces fewer of the errors than at least one other schema of the plurality of schemas;
parsing, by the computing hardware, the first data source to identify a difference between the first data source and a second data source comprising a previous version of the data, wherein the difference comprises at least a data update, a data deletion, or a data addition;
generating, by the computing hardware, a change set from the identified difference comprising a first query for the difference based on the applicable schema to implement a version update to the graph representation found in the graph data structure corresponding to the applicable schema, wherein the version update comprises at least one of a new node, a new edge, a deleted node, a deleted edge, an updated node, or an updated edge of the graph representation found in the graph data structure corresponding to the applicable schema;
saving, by the computing hardware, the change set in a change set repository;
querying, by the computing hardware, the change set repository to retrieve an unapplied change set;
determining, by the computing hardware, an applicable modification to make to the graph representation by utilizing a machine-learning model to analyze data of the graph representation and the unapplied change set from the identified difference; and
providing, by the computing hardware, a second query to migrate the unapplied change set into the graph representation found in the graph data structure corresponding to the applicable schema by modifying the graph representation found in the graph data structure corresponding to the applicable schema based on the determined applicable modification; and
wherein, for at least one iteration of the plurality of iterations, the computing hardware causes, responsive to identifying an issue with at least one change set within the change set repository and generating the change set from the identified difference, a rollback of the at least one change set within the change set repository.
2 . The method of claim 1 , wherein the applicable schema produces a least number of the errors.
3 . The method of claim 1 , wherein the first data source comprises a matrix and the applicable schema comprises a script indicating a type of data for each column of the matrix.
4 . The method of claim 1 , wherein validating the first data source based on the schema comprises applying at least one of a linear cost function or a least squares cost function.
5 . The method of claim 1 further comprising at least one of:
providing, by the computing hardware, the errors produced by the applicable schema for display on a graphical user interface; or
generating, by the computing hardware, a communication for the errors produced by the applicable schema, wherein the errors produced by the applicable schema are at least one of displayed or communicated for error correction prior to comparing the first data source with the second data source.
6 . The method of claim 1 further comprising determining, by the computing hardware, the applicable modification to make to the graph representation by utilizing the machine-learning model to analyze a matrix representation conversion of the data of the graph representation.
7 . The method of claim 1 , wherein the machine-learning model comprises at least one of a multi-label classification model.
8 . The method of claim 1 , further comprising identifying the difference between the first data source and the second data source utilizing a first vector representation of the first data source and a second vector representation of the second data source.
9 . The method of claim 1 further comprising generating, by the computing hardware, the change set from the identified difference to comprise metadata for the applicable schema and the first query for the difference.
10 . The method of claim 9 , further comprising:
identifying, by the computing hardware, multiple differences between the first data source and the second data source;
generating, by the computing hardware, the change set from the identified multiple differences comprising multiple queries for the multiple differences based on the applicable schema; and
providing, by the computing hardware, the change set to execute the multiple queries to migrate the multiple differences into the graph representation.
11 . A method comprising:
processing, by computing hardware, data found in a first data source comprising a new version of the data using a machine-learning model to identify an applicable schema from a plurality of schemas in which each schema of the plurality of schemas corresponds to a graph representation found in a graph data structure;
parsing, by the computing hardware, the first data source to identify a difference between the first data source and a second data source comprising a previous version of the data, wherein the difference comprises at least a data update, a data deletion, or a data addition;
generating, by the computing hardware, a change set from the identified difference comprising a first query for the difference based on the applicable schema to implement a version update to the graph representation found in the graph data structure corresponding to the applicable schema, wherein the version update comprises at least one of a new node, a new edge, a deleted node, a deleted edge, an updated node, or an updated edge of the graph representation found in the graph data structure corresponding to the applicable schema;
causing, by the computing hardware and responsive to detecting an issue in a prior change set of a change set repository, a rollback of the prior change set, such that the change set can correct the issue upon migration;
saving, by the computing hardware, the change set in the change set repository;
querying, by the computing hardware, the change set repository to retrieve an unapplied change set;
determining, by the computing hardware, an applicable modification to make to the graph representation by utilizing an additional machine-learning model to analyze data of the graph representation and the unapplied change set from the identified difference; and
providing, by the computing hardware, a second query to migrate the unapplied change set into the graph representation found in the graph data structure corresponding to the applicable schema by modifying the graph representation found in the graph data structure corresponding to the applicable schema based on the determined applicable modification.
12 . The method of claim 11 further comprising validating the first data source using the applicable schema to identify errors in the first data source, wherein the errors in the first data source are corrected prior to comparing the first data source with the second data source.
13 . The method of claim 11 , wherein the machine-learning model comprises at least one of a multi-label classification model or an ensemble of multiple classification models that provides a prediction for each schema in the plurality of schemas that represents a likelihood of the schema being applicable to the first data source, and processing the data found in the first data source using the machine-learning model to identify the applicable schema comprises selecting the applicable schema based on the corresponding prediction for the applicable schema being higher than predictions for other schemas in the plurality of schemas.
14 . A system comprising:
a non-transitory computer-readable medium storing instructions; and
a processing device communicatively coupled to the non-transitory computer-readable medium, wherein, the processing device is configured to execute the instructions and thereby perform operations comprising:
conducting a plurality of iterations, wherein an iteration of the plurality of iterations comprises validating a first data source comprising a new version of data based on a schema from a plurality of schemas in which each schema in the plurality of schemas corresponds to a graph representation found in a graph data structure;
identifying, based on the plurality of iterations, an applicable schema from the plurality of schemas;
parsing the first data source to identify a difference between the first data source and a second data source comprising a previous version of the data, wherein the difference comprises at least a data update, a data deletion, or a data addition;
generating a change set from the identified difference comprising a first query for the difference based on the applicable schema to implement a version update to the graph representation found in the graph data structure corresponding to the applicable schema, wherein the version update comprises at least one of a new node, a new edge, a deleted node, a deleted edge, an updated node, or an updated edge of the graph representation found in the graph data structure corresponding to the applicable schema;
causing, by the computing hardware and responsive to detecting an issue in a prior change set of a change set repository, a rollback of the a prior change set, such that the change set can correct the issue upon migration;
saving the change set in the change set repository;
querying the change set repository to retrieve an unapplied change set;
determining an applicable modification to make to the graph representation by utilizing a machine-learning model to analyze data of the graph representation and the unapplied change set from the identified difference; and
providing a second query to migrate the unapplied change set into the graph representation found in the graph data structure corresponding to the applicable schema by modifying the graph representation found in the graph data structure corresponding to the applicable schema based on the determined applicable modification.
15 . The system of claim 14 , wherein each iteration of the plurality of iterations further comprises identifying errors in the first data source based on the validating of the first data source, wherein the identified errors for the applicable schema are fewer than additional errors of at least one other schema of the plurality of schemas.
16 . The system of claim 15 , wherein validating the first data source based on the schema comprises applying at least one of a linear cost function or a least squares cost function.
17 . The system of claim 15 , wherein the operations further comprise at least one of:
providing the identified errors produced by the applicable schema for display on a graphical user interface; or
generating an electronic communication for the identified errors produced by the applicable schema.
18 . The system of claim 14 , wherein the first data source comprises a matrix and the applicable schema comprises a script indicating a data type for each column of the matrix.
19 . The system of claim 14 , wherein the machine-learning model comprises an ensemble of multiple classification models that provides a prediction for each available modification in a plurality of available modifications that represents a likelihood of the applicable modification being applicable to the graph representation.
20 . The system of claim 14 , wherein the operations further comprise:
generating a communication providing the applicable modification as an applicable recommendation; and
sending the communication to an electronic address associated with the graph data structure.