IP Library Granted Patent US 11,727,031
Granted Patent B2
US 11,727,031 · App. 17/833,147 · Granted Aug 15, 2023

Systems and methods for formatting data using a recurrent neural network

Inventors: Anh Truong (Champaign, IL); Reza Farivar (Champaign, IL); Austin Walters (Savoy, IL); Jeremy Goodsitt (Champaign, IL)
Assignee: Capitai One Services, LLC
G06F16/258G06F16/9024G06F17/18G06N3/045G06N3/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,727,031
App. No.
17/833,147
Granted
Aug 15, 2023
Kind
B2
Abstract

Systems and methods for formatting data are disclosed. For example, a system may include at least one memory storing instructions and one or more processors configured to execute the instructions to perform operations. The operations may include receiving data comprising a plurality of sequences of data values and training a recurrent neural network model to output conditional probabilities of subsequent data values based on preceding data values in the data value sequences. The operations may include generating conditional probabilities using the trained recurrent neural network model and the received data. The operations may include determining a data format of a subset of the data value sequences, based on the generated conditional probabilities, and reformatting at least one of the data value sequences according to the determined data format.

Claims (45)

1. A system for formatting data, the system comprising:

at least one memory storing instructions; and

one or more processors configured to execute the instructions to perform operations comprising:

generating a direct probabilistic-graph, the direct probabilistic-graph including a set of nodes corresponding to positions in received data value sequences, by iteratively:

determining conditional counts of occurrences of received data values at a subsequent node in the set of nodes, the conditional counts being based on counting instances of received data values at one or more preceding nodes in the set of nodes;

determining total counts including a total of occurrences of the received data values; and

determining conditional probabilities based on the conditional counts and the total counts;

determining a similarity metric of a modeled probabilistic-graph and the direct probabilistic-graph, the modeled probabilistic-graph being generated by a machine learning model, wherein the direct probabilistic-graph is generated based on a known data format and the modeled probabilistic-graph is generated based on an unknown data format; and

training the machine learning model to output conditional probabilities based on the similarity metric.

2. The system of claim 1 , wherein the operations further comprise encoding the data value sequences.

3. The system of claim 2 , wherein determining the conditional counts and the total counts is further based on the encoded data value sequences.

4. The system of claim 1 , wherein the operations further comprise generating the modeled probabilistic-graph prior to determining the similarity metric.

5. The system of claim 1 , wherein the modeled probabilistic-graph is a previously generated modeled probabilistic-graph.

6. The system of claim 1 , wherein the similarity metric comprises a percent overlap reflecting a percent of nodes that include same conditional probabilities in both the modeled probabilistic-graph and the direct probabilistic-graph.

7. The system of claim 1 , wherein the similarity metric comprises an average relative difference between nodes.

8. The system of claim 1 , wherein the similarity metric comprises a measure of a statistical distribution of differences between nodes.

9. The system of claim 1 , wherein the operations further comprise pruning the direct probabilistic-graph, the pruning comprising at least one of deleting a null value or deleting a conditional probability of a node.

10. The system of claim 1 , wherein the operations further comprise pruning the direct probabilistic-graph, the pruning comprising deleting data associated with a node.

11. The system of claim 1 , wherein the conditional probabilities comprise reformatted data.

12. The system of claim 1 , wherein training the machine learning model is based on a threshold similarity metric value.

13. The system of claim 1 , wherein the operations further comprise transmitting the direct probabilistic-graph to a remote system.

14. The system of claim 1 , wherein the direct probabilistic-graph includes at least one of a Bayesian network or a Markov network.

15. The system of claim 1 , wherein the operations further comprise generating an expression for determining a data format based on the direct probabilistic-graph.

16. The system of claim 1 , wherein the operations further comprise generating an expression for reformatting data based on the direct probabilistic-graph.

17. The system of claim 1 , wherein:

the operations further comprise generating embedded data based on the data value sequences; and

training the machine learning model comprises using the embedded data as training data.

18. The system of claim 17 , wherein generating embedded data comprises implementing at least one of a one-hot encoding method or a glove method.

19. A method for formatting data, comprising:

generating a direct probabilistic-graph, the direct probabilistic-graph including a set of nodes corresponding to positions in received data value sequences, by iteratively:

determining conditional counts of occurrences of received data values at a subsequent node in the set of nodes, the conditional counts being based on counting instances of received data values at one or more preceding nodes in the set of nodes;

determining total counts including a total of occurrences of the received data values; and

determining conditional probabilities based on the conditional counts and the total counts;

determining a similarity metric of a modeled probabilistic-graph and the direct probabilistic-graph, the modeled probabilistic-graph being generated by a machine learning model, wherein the direct probabilistic-graph is generated based on a known data format and the modeled probabilistic-graph is generated based on an unknown data format; and

training the machine learning model to output conditional probabilities based on the similarity metric.

20. A system for formatting data, the system comprising:

at least one memory storing instructions; and

one or more processors configured to execute the instructions to perform operations comprising:

generating a direct probabilistic-graph, the direct probabilistic-graph including a set of nodes corresponding to positions in received data value sequences, by iteratively:

determining conditional counts of occurrences of received data values at a subsequent node in the set of nodes, the conditional counts being based on counting instances of received data values at one or more preceding nodes in the set of nodes;

determining total counts including a total of occurrences of the received data values;

determining conditional probabilities based on the conditional counts and the total counts; and

generating embedded data based on the data value sequences by implementing at least one of a one-hot encoding method or a glove method;

determining a similarity metric of a modeled probabilistic-graph and the direct probabilistic-graph, the similarity metric including at least one of a percent overlap reflecting a percent of nodes that include same conditional probabilities in both the modeled probabilistic-graph and the direct probabilistic-graph, an average relative difference between nodes, or a measure of a statistical distribution of differences between nodes, the modeled probabilistic-graph being generated by a machine learning model, wherein the direct probabilistic-graph is generated based on a known data format and the modeled probabilistic-graph is generated based on an unknown data format; and

training the machine learning model to output conditional probabilities based on the similarity metric and the embedded data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2022
From: TRUONG, ANH; FARIVAR, REZA; WALTERS, AUSTIN; GOODSITT, JEREMY
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 060110/0877 →
Continuity (3)
Continuation 17078775 · Oct 23, 2020
Continuation 16810230 · Mar 5, 2020
Related Publication 20220300526A1 · Sep 22, 2022