IP Library › Granted Patent US 11,934,486
Granted Patent B2
US 11,934,486 · App. 16/950,399 · Granted Mar 19, 2024

Systems and methods for data stream using synthetic data

Inventors: Anh Truong (Champaign, IL); Jeremy Goodsitt (Champaign, IL); Austin Walters (Savoy, IL)
Assignee: Capital One Services, LLC
G06F18/2148G06F3/0617G06F3/0644G06F3/065G06F3/0652G06F3/0653G06F3/0685G06F9/5016G06N3/049
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,934,486
App. No.
16/950,399
Granted
Mar 19, 2024
Kind
B2
Abstract

Systems and methods for synthetic data generation. A system includes at least one processor and a storage medium storing instructions that, when executed by the one or more processors, cause the at least one processor to perform operations including receiving a continuous data stream from an outside source, processing the continuous data stream in real-time, and using machine learning techniques to generating synthetic data to populate the dataset. The operations also include creating a plurality of bins, wherein the plurality of bins occupy a data range between the determined minimum and maximum values without overlapping; and determining a number of samples within each of the created bin, based on a bin edges, wherein the bin edges are bounds within the data range.

Claims (48)

1. A system for synthetic data generation, comprising:

at least one processor; and

at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:

determining a size of stored streaming data, stored in a first storage device, has reached a first threshold;

in response to the size determination, processing the stored streaming data, the processing comprising:

determining a total number of samples in the stored streaming data;

creating a plurality of bins, the bins having data ranges between a minimum and a maximum sample value;

assigning the samples to the bins, based on values of the samples and data ranges of the bins; and

determining a number of samples within the bins;

populating the bins with synthetic data, the populating comprising:

generating, by a synthetic data generator, a plurality of synthetic data points; and

assigning the synthetic data points to the bins based on values of the synthetic data points and data ranges of the bins;

determining a total number of the synthetic data points in the bins has reached a second threshold;

in response to the determination of whether the total number has reached a second threshold, pausing the synthetic data generator;

creating a processed dataset based on the bins; and

storing the processed dataset on a second storage device.

2. The system of claim 1 , wherein the first threshold is automatically adjusted based on changes to the system.

3. The system of claim 1 , wherein the processed dataset is a histogram.

4. The system of claim 1 , wherein the first storage device is a random access memory.

5. The system of claim 1 , wherein the second storage device is one of an electro-mechanical data storage device or a nonvolatile flash memory.

6. The system of claim 1 , wherein generating a plurality of synthetic data points comprises generating a plurality of synthetic data points using a recurrent neural network.

7. The system of claim 1 , wherein generating a plurality of synthetic data points comprises generating a plurality of synthetic data points using a generative adversarial network.

8. The system of claim 1 , wherein the operations further comprise adjusting the first threshold based on characteristics of the stored streaming data.

9. The system of claim 1 , wherein the operations further comprise adjusting the second threshold based on characteristics of the stored streaming data.

10. A system for synthetic data generation, comprising:

at least one processor; and

at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:

determining a size of stored streaming data has reached a first threshold, wherein the first threshold is dynamically adjusted based on changes to the stored streaming data or hardware of the system;

in response to the size determination, processing the stored streaming data, wherein the processing comprises determining a total number of samples in the stored streaming data and a range of values for the stored streaming data;

creating a synthetic dataset based on the processed stored streaming data, the creating comprising: generating, by a synthetic data generator, a plurality of synthetic data points, the plurality of synthetic data points being determined based on the total number of samples and the range of values for the stored streaming data; and

storing the synthetic dataset.

11. The system of claim 10 , wherein the first threshold specifies a maximum storage capacity.

12. The system of claim 10 , wherein the stored streaming data comprises numerical data.

13. The system of claim 10 , wherein the stored streaming data comprises image data.

14. The system of claim 10 , wherein the stored streaming data comprises video data.

15. The system of claim 10 , wherein the operations further comprise performing data profiling calculations on the synthetic dataset.

16. The system of claim 10 , wherein the stored streaming data contains sensitive data.

17. A method for synthetic data generation comprising:

determining a size of stored streaming data has reached a first threshold;

in response to the size determination, processing the stored streaming data, the processing comprising:

determining a size of the stored streaming data has reached a first threshold, wherein the first threshold is dynamically adjusted based on changes to the stored streaming data or hardware of the system; and

in response to the size determination, processing the stored streaming data, wherein the processing comprises determining a total number of samples in the stored data and a range of values for the stored data;

creating a synthetic dataset based on the processed data, the creating comprising:

generating, by a synthetic data generator, a plurality of synthetic data points, the plurality of synthetic data points being determined based on the total number of samples and the range of values for the stored data; and

storing the synthetic dataset.

18. The method of claim 17 , wherein the method further comprises performing data profiling calculations on the synthetic dataset.

19. The method of claim 17 , wherein the method further comprises adjusting the first threshold based on characteristics of the stored streaming data.

20. The method of claim 17 , wherein the method further comprises adjusting a second threshold based on characteristics of the stored streaming data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2020
From: TRUONG, ANH; GOODSITT, JEREMY; WALTERS, AUSTIN
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 054393/0070 →
Continuity (2)
Continuation 16596886 · Oct 9, 2019
Related Publication 20210133504A1 · May 6, 2021