IP Library › Granted Patent US 12,639,330
Granted Patent B2
US 12,639,330 · App. 18/922,033 · Granted May 26, 2026

Systems and methods for formatting data using a recurrent neural network

Inventors: Anh Truong (Champaign, IL); Reza Farivar (Champaign, IL); Austin Walters (Savoy, IL); Jeremy Goodsitt (Champaign, IL)
Assignee: Capital One Services, LLC
G06F16/258G06F16/9024G06F17/18G06N3/045G06N3/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,330
App. No.
18/922,033
Granted
May 26, 2026
Kind
B2
Abstract

Systems and methods for formatting data are disclosed. For example, a system may include at least one memory storing instructions and one or more processors configured to execute the instructions to perform operations. The operations may include receiving data comprising a plurality of sequences of data values and training a recurrent neural network model to output conditional probabilities of subsequent data values based on preceding data values in the data value sequences. The operations may include generating conditional probabilities using the trained recurrent neural network model and the received data. The operations may include determining a data format of a subset of the data value sequences, based on the generated conditional probabilities, and reformatting at least one of the data value sequences according to the determined data format.

Claims (41)

1 . A system for formatting data, the system comprising:

at least one memory storing instructions; and

one or more processors configured to execute the instructions to perform operations comprising:

receiving data comprising a plurality of data value sequences;

generating conditional probabilities using a machine learning model based on the plurality of data value sequences;

determining a data format of at least a subset of the plurality of data value sequences based on the generated conditional probabilities;

generating reformatted data based on the received data by reformatting at least one of the plurality of data value sequences according to the determined data format; and

training a synthetic data model to generate synthetic data based on the reformatted data after the reformatting of the at least one of the plurality of data value sequences according to the determined data format.

2 . The system of claim 1 , wherein the machine learning model is configured to output conditional probabilities of subsequent data values based on preceding data values in the received data.

3 . The system of claim 1 , the operations further comprising displaying a probabilistic graph of the generated conditional probabilities.

4 . The system of claim 1 , the operations further comprising determining a frequency of the determined data format.

5 . The system of claim 4 , the operations further comprising displaying the frequency of the determined data format in a probabilistic graph of the generated conditional probabilities.

6 . The system of claim 1 , the operations further comprising at least one of storing or transmitting the reformatted data.

7 . The system of claim 1 , wherein determining the data format comprises using the machine learning model.

8 . The system of claim 1 , the operations further comprising:

generating a first probabilistic graph, the first probabilistic graph including a set of nodes corresponding to positions in received data value sequences, by iteratively:

determining conditional counts of occurrences of received data values at a subsequent node in the set of nodes, the conditional counts being based on counting instances of received data values at one or more preceding nodes in the set of nodes; and

determining conditional probabilities based on the conditional counts.

9 . The system of claim 8 , the operations further comprising determining a similarity metric of a second probabilistic graph and the first probabilistic graph, the second probabilistic graph being generated by the machine learning model or another machine learning model.

10 . The system of claim 8 , wherein:

generating the first probabilistic graph further includes determining total counts including a total of occurrences of the received data values; and

determining conditional probabilities is further based on the total counts.

11 . The system of claim 8 , the operations further comprising pruning the first probabilistic graph, the pruning comprising deleting a null value conditional probability of a node.

12 . The system of claim 1 , the operations further comprising generating synthetic data to replace sensitive information in the received data.

13 . The system of claim 1 , the operations further comprising determining at least one data schema of the received data value sequences.

14 . The system of claim 1 , the operations further comprising determining one or more foreign keys within the received data value sequences.

15 . The system of claim 1 , wherein the machine learning model is configured to learn relationships between sub-sequences in the received data value sequences.

16 . The system of claim 1 , wherein determining the conditional probabilities is further based on relationships between nonconsecutive data values in the received data value sequences.

17 . The system of claim 1 , wherein the machine learning model is a recurrent neural network.

18 . The system of claim 1 , the operations further comprising generating an expression for determining or reformatting data based on the conditional probabilities.

19 . A method for formatting data, the method comprising:

receiving data comprising a plurality of data value sequences;

generating conditional probabilities using a machine learning model based on the plurality of data value sequences, wherein the machine learning model is configured to output conditional probabilities of subsequent data values based on preceding data values in the received data;

determining a data format of at least a subset of the plurality of data value sequences based on the generated conditional probabilities;

generating reformatted data based on the determined data format; and

training a synthetic data model to generate synthetic data based on the reformatted data after generating the reformatted data based on the determined data format.

20 . A non-transitory computer-readable medium including instructions that are executable by one or more processors to perform operations comprising:

receiving data comprising a plurality of data value sequences;

generating conditional probabilities using a machine learning model based on the plurality of data value sequences, wherein the machine learning model is configured to output conditional probabilities of subsequent data values based on preceding data values in the received data;

generating reformatted data based on the received data by reformatting at least one of the plurality of data value sequences according to a standard format; and

training a synthetic data model to generate synthetic data based on the reformatted data after generating the reformatted data based on the determined data format.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2026
From: TRUONG, ANH; FARIVAR, REZA; WALTERS, AUSTIN; GOODSITT, JEREMY
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 074525/0249 →
Continuity (5)
Continuation 18340166 · Jun 23, 2023
Continuation 17833147 · Jun 6, 2022
Continuation 17078775 · Oct 23, 2020
Continuation 16810230 · Mar 5, 2020
Related Publication 20250045295A1 · Feb 6, 2025
References Cited (23)
US 9213944B1 · Do et al. · 2015 [cited by applicant]
US 9830315B1 · Xiao et al. · 2017 [cited by applicant]
US 10388272B1 · Thomson · 2019 [cited by examiner]
US 10713272B1 · Caldwell · 2020 [cited by examiner]
US 11341367B1 · Barbosa · 2022 [cited by examiner]
US 20040205474A1 · Eskin et al. · 2004 [cited by applicant]
US 20080281849A1 · Mineno · 2008 [cited by examiner]
US 20160342737A1 · Kaye · 2016 [cited by applicant]
US 20170124447A1 · Chang et al. · 2017 [cited by applicant]
US 20180082208A1 · Cormier et al. · 2018 [cited by applicant]
US 20180174671A1 · Cruz Huertas et al. · 2018 [cited by applicant]
US 20190260787A1 · Zou · 2019 [cited by examiner]
US 20200037962A1 · Shanbhag · 2020 [cited by examiner]
US 20200104361A1 · Zarrella · 2020 [cited by applicant]
US 20200118035A1 · Asawa et al. · 2020 [cited by applicant]
US 20200160178A1 · Kar et al. · 2020 [cited by applicant]
US 20210019309A1 · Yadav et al. · 2021 [cited by applicant]
US 20210074039A1 · Kholodkov et al. · 2021 [cited by applicant]
US 20210089375A1 · Ghafourifar et al. · 2021 [cited by applicant]
CN 113033240A · 2021 [cited by examiner]
WO WO2016145379A1 · 2016 [cited by applicant]
WO WO2019124724A1 · 2019 [cited by applicant]
English translation of CN 113033240 A (Year: 2021). [cited by examiner]