IP Library Granted Patent US 12,147,447
Granted Patent B2
US 12,147,447 · App. 18/340,166 · Granted Nov 19, 2024

Systems and methods for formatting data using a recurrent neural network

Inventors: Anh Truong (Champaign, IL); Reza Farivar (Champaign, IL); Austin Walters (Savoy, IL); Jeremy Goodsitt (Champaign, IL)
Assignee: Capital One Services, LLC
G06F16/258G06F16/9024G06F17/18G06N3/045G06N3/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,147,447
App. No.
18/340,166
Granted
Nov 19, 2024
Kind
B2
Abstract

Systems and methods for formatting data are disclosed. For example, a system may include at least one memory storing instructions and one or more processors configured to execute the instructions to perform operations. The operations may include receiving data comprising a plurality of sequences of data values and training a recurrent neural network model to output conditional probabilities of subsequent data values based on preceding data values in the data value sequences. The operations may include generating conditional probabilities using the trained recurrent neural network model and the received data. The operations may include determining a data format of a subset of the data value sequences, based on the generated conditional probabilities, and reformatting at least one of the data value sequences according to the determined data format.

Claims (40)

1. A system for formatting data, the system comprising:

at least one memory storing instructions; and

one or more processors configured to execute the instructions to perform operations comprising:

generating a first probabilistic graph, the first probabilistic graph including a set of nodes corresponding to positions in received data value sequences, by iteratively:

determining conditional counts of occurrences of received data values at a subsequent node in the set of nodes, the conditional counts being based on counting instances of received data values at one or more preceding nodes in the set of nodes; and

determining conditional probabilities based on the conditional counts;

determining a similarity metric of a second probabilistic graph and the first probabilistic graph, the second probabilistic graph being generated by a machine learning model; and

training the machine learning model based on the similarity metric.

2. The system of claim 1 , wherein the operations further comprise classifying the received data value sequences.

3. The system of claim 2 , wherein classifying the received data value sequences includes clustering one or more datasets.

4. The system of claim 2 , wherein classifying the received data value sequences is based on at least one of a data profile, a data schema, a statistical profile, a foreign key, or a relationship between datasets.

5. The system of claim 1 , wherein the operations further comprise training the machine learning model to generate synthetic data to replace sensitive information.

6. The system of claim 1 , wherein the operations further comprise determining at least one data schema of the received data value sequences.

7. The system of claim 1 , wherein the operations further comprise determining one or more foreign keys within the received data value sequences.

8. The system of claim 1 , wherein the operations further comprise determining at least one data format associated with the received data value sequences.

9. The system of claim 8 , wherein the operations further comprise reformatting additional data according to the determined at least one data format.

10. The system of claim 1 , wherein:

the operations further comprise generating embedded data based on the received data value sequences; and

training the machine learning model comprises using the embedded data as training data.

11. The system of claim 10 , wherein generating the embedded data comprises implementing at least one of a one-hot encoding method or a glove method.

12. The system of claim 10 , wherein the operations further comprise clustering the generated embedded data and training the machine learning model comprises using clustered embedded data as training data.

13. The system of claim 10 , wherein determining the conditional counts of data values is further based on the embedded data.

14. The system of claim 1 , wherein determining the conditional probabilities is further based on relationships between nonconsecutive data values in the received data value sequences.

15. The system of claim 1 , wherein the machine learning model is a recurrent neural network.

16. The system of claim 15 , wherein the operations further comprise training the recurrent neural network to learn a relationship between sub-sequences in the received data value sequences.

17. The system of claim 1 , wherein the operations further comprise pruning the first probabilistic graph, the pruning comprising of deleting a null value conditional probability of a node.

18. The system of claim 1 , wherein the operations further comprise pruning the first probabilistic graph, the pruning comprising of deleting data associated with a node.

19. A method for formatting data, comprising:

generating a first probabilistic graph, the first probabilistic graph including a set of nodes corresponding to positions in received data value sequences, by iteratively:

determining conditional counts of occurrences of received data values at a subsequent node in the set of nodes, the conditional counts being based on counting instances of received data values at one or more preceding nodes in the set of nodes; and

determining conditional probabilities based on the conditional counts;

determining a similarity metric of a second probabilistic graph and the first probabilistic graph, the second probabilistic graph being generated by a machine learning model;

generating embedded data based on the received data value sequences; and

training the machine learning model based on the similarity metric and the embedded data.

20. A non-transitory computer-readable medium including instructions that are executable by one or more processors to perform operations comprising:

generating a first direct probabilistic-graph, the first probabilistic graph including a set of nodes corresponding to positions in received data value sequences, by iteratively:

determining conditional counts of occurrences of received data values at a subsequent node in the set of nodes, the conditional counts being based on counting instances of received data values at one or more preceding nodes in the set of nodes; and

determining conditional probabilities based on the conditional counts;

determining a similarity metric of a second probabilistic graph and the first probabilistic graph, the second probabilistic graph being generated by a machine learning model; and

training the machine learning model based on the similarity metric.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2024
From: TRUONG, ANH; FARIVAR, REZA; WALTERS, AUSTIN; GOODSITT, JEREMY
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 068141/0451 →
Continuity (4)
Continuation 17833147 · Jun 6, 2022
Continuation 17078775 · Oct 23, 2020
Continuation 16810230 · Mar 5, 2020
Related Publication 20230334063A1 · Oct 19, 2023