Systems and methods for formatting data using a recurrent neural network
Systems and methods for formatting data are disclosed. For example, a system may include at least one memory storing instructions and one or more processors configured to execute the instructions to perform operations. The operations may include receiving data comprising a plurality of sequences of data values and training a recurrent neural network model to output conditional probabilities of subsequent data values based on preceding data values in the data value sequences. The operations may include generating conditional probabilities using the trained recurrent neural network model and the received data. The operations may include determining a data format of a subset of the data value sequences, based on the generated conditional probabilities, and reformatting at least one of the data value sequences according to the determined data format.
1 . A system for formatting data, the system comprising:
at least one memory storing instructions; and
one or more processors configured to execute the instructions to perform operations comprising:
receiving data comprising a plurality of data value sequences;
generating conditional probabilities using a machine learning model based on the plurality of data value sequences;
determining a data format of at least a subset of the plurality of data value sequences based on the generated conditional probabilities;
generating reformatted data based on the received data by reformatting at least one of the plurality of data value sequences according to the determined data format; and
training a synthetic data model to generate synthetic data based on the reformatted data after the reformatting of the at least one of the plurality of data value sequences according to the determined data format.
2 . The system of claim 1 , wherein the machine learning model is configured to output conditional probabilities of subsequent data values based on preceding data values in the received data.
3 . The system of claim 1 , the operations further comprising displaying a probabilistic graph of the generated conditional probabilities.
4 . The system of claim 1 , the operations further comprising determining a frequency of the determined data format.
5 . The system of claim 4 , the operations further comprising displaying the frequency of the determined data format in a probabilistic graph of the generated conditional probabilities.
6 . The system of claim 1 , the operations further comprising at least one of storing or transmitting the reformatted data.
7 . The system of claim 1 , wherein determining the data format comprises using the machine learning model.
8 . The system of claim 1 , the operations further comprising:
generating a first probabilistic graph, the first probabilistic graph including a set of nodes corresponding to positions in received data value sequences, by iteratively:
determining conditional counts of occurrences of received data values at a subsequent node in the set of nodes, the conditional counts being based on counting instances of received data values at one or more preceding nodes in the set of nodes; and
determining conditional probabilities based on the conditional counts.
9 . The system of claim 8 , the operations further comprising determining a similarity metric of a second probabilistic graph and the first probabilistic graph, the second probabilistic graph being generated by the machine learning model or another machine learning model.
10 . The system of claim 8 , wherein:
generating the first probabilistic graph further includes determining total counts including a total of occurrences of the received data values; and
determining conditional probabilities is further based on the total counts.
11 . The system of claim 8 , the operations further comprising pruning the first probabilistic graph, the pruning comprising deleting a null value conditional probability of a node.
12 . The system of claim 1 , the operations further comprising generating synthetic data to replace sensitive information in the received data.
13 . The system of claim 1 , the operations further comprising determining at least one data schema of the received data value sequences.
14 . The system of claim 1 , the operations further comprising determining one or more foreign keys within the received data value sequences.
15 . The system of claim 1 , wherein the machine learning model is configured to learn relationships between sub-sequences in the received data value sequences.
16 . The system of claim 1 , wherein determining the conditional probabilities is further based on relationships between nonconsecutive data values in the received data value sequences.
17 . The system of claim 1 , wherein the machine learning model is a recurrent neural network.
18 . The system of claim 1 , the operations further comprising generating an expression for determining or reformatting data based on the conditional probabilities.
19 . A method for formatting data, the method comprising:
receiving data comprising a plurality of data value sequences;
generating conditional probabilities using a machine learning model based on the plurality of data value sequences, wherein the machine learning model is configured to output conditional probabilities of subsequent data values based on preceding data values in the received data;
determining a data format of at least a subset of the plurality of data value sequences based on the generated conditional probabilities;
generating reformatted data based on the determined data format; and
training a synthetic data model to generate synthetic data based on the reformatted data after generating the reformatted data based on the determined data format.
20 . A non-transitory computer-readable medium including instructions that are executable by one or more processors to perform operations comprising:
receiving data comprising a plurality of data value sequences;
generating conditional probabilities using a machine learning model based on the plurality of data value sequences, wherein the machine learning model is configured to output conditional probabilities of subsequent data values based on preceding data values in the received data;
generating reformatted data based on the received data by reformatting at least one of the plurality of data value sequences according to a standard format; and
training a synthetic data model to generate synthetic data based on the reformatted data after generating the reformatted data based on the determined data format.