IP Library › Granted Patent US 12,174,826
Granted Patent B2
US 12,174,826 · App. 18/364,704 · Granted Dec 24, 2024

Systems and methods for unified data validation

Inventor: Saisharath Kondakindi (Stamford, CT)
Assignee: Synchrony Bank
G06F16/2365G06F16/242G06F16/24532G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,174,826
App. No.
18/364,704
Granted
Dec 24, 2024
Kind
B2
Abstract

Examples described herein include implementations for big-data validation. One aspect includes generating a configuration file including dynamic matching data describing a first plurality of data entries and a second plurality of data entries, and generating a data action file. A plurality of data queries are generated based on the dynamic matching data indicated in the configuration file. The plurality of data queries are dynamically executed in parallel, including execution of a plurality of simultaneous data queries to the data source system. Fields of the first plurality of data entries and the second plurality of data entries are matched using the key type and the value structure, corresponding fields of the first data fields and the second data fields having a data mismatch are identified, and a mismatch database entry for the corresponding fields having the data mismatch is automatically generated.

Claims (100)

1. A computer implemented method comprising:

generating a configuration file, the configuration file including dynamic matching data describing a first plurality of data entries and a second plurality of data entries;

generating a data action file, the data action file identifying a first data set including the first plurality of data entries, a second data set including the second plurality of data entries, a key type, and a value structure associated with the key type, wherein the first data set and the second data set are stored in a data source system, wherein the first plurality of data entries includes first data fields, and wherein the second plurality of data entries includes second data fields;

dynamically generating a plurality of data queries including key type queries, value structure queries and dynamic value structure queries, wherein the dynamic value structure queries are generated in real-time based on the dynamic matching data indicated in the configuration file, and wherein the dynamic value structure queries are being generated based on a first data type, a second data type, and data field types present in the first data set and the second data set;

dynamically executing the plurality of data queries in parallel, wherein parallel execution of the plurality of data queries includes a plurality of simultaneous data queries to the data source system including the first data set and the second data set;

matching fields of the first plurality of data entries and the second plurality of data entries using the key type and the value structure associated with the key type;

identifying corresponding fields of the first data fields and the second data fields having a data mismatch; and

automatically generating a mismatch database entry for the corresponding fields having the data mismatch.

2. The computer implemented method of claim 1 , where the first data set includes at least one billion data fields, and wherein the second data set includes at least one billion data fields.

3. The computer implemented method of claim 1 , wherein the first data set comprises at least a terabyte of data, and wherein the second data set comprises at least a terabyte of data.

4. The computer implemented method of claim 1 , wherein the configuration file further identifies the first data type for the first data set and the second data type for the second data set.

5. The computer implemented method of claim 1 , further comprising:

generating the configuration file associated with the first data set, wherein the configuration file further indicates values used for generation of the dynamic value structure queries, and wherein the configuration file is generated using a machine learning algorithm trained from mismatch database entry training data;

generating a feedback value associated with the mismatch database entry, wherein the feedback value identifies a difference between an expected result and an actual result in the mismatch database entry, wherein the feedback value is associated with settings in the configuration file; and

updating the mismatch database entry training data and the machine learning algorithm using the feedback value and the mismatch database entry to generate update mismatch database entry training data.

6. The computer implemented method of claim 1 , further comprising using a machine learning algorithm to generate a dynamic query selection table using mismatch training data, wherein the dynamic value structure queries are selected from the dynamic query selection table.

7. The computer implemented method of claim 1 , further comprising processing the mismatch database entry to generate a mismatch report providing mismatch metrics for values of the corresponding fields of the first data set and the second data set.

8. The computer implemented method of claim 1 , further comprising processing the mismatch database entry using a machine learning algorithm to remove mismatch data where the corresponding fields for the first data set and the second data set include matching content with data mismatches.

9. The computer implemented method of claim 1 , further comprising identifying different tiers of matching standards for different corresponding fields of the first data set and the second data set; and

merging the first data set and the second data set;

wherein merging the corresponding fields of the first data fields and the second data fields having the data mismatch comprises a multi-tier machine learning algorithm for merging entries having identical key values with mismatched fields.

10. The computer implemented method of claim 1 , further comprising:

identifying corresponding chunks of data from the first data set and the second data set;

wherein dynamically executing the plurality of data queries in parallel includes generating separate data chunks for independent parallel processing using the corresponding chunks of data from the first data set and the second data set, and

wherein identifying the corresponding fields of the first data fields and the second data fields having the data mismatch is performed separately on the separate data chunks for the corresponding chunks of data from the first data set and the second data set.

11. The computer implemented method of claim 1 , further comprising:

wherein the plurality of data queries are queries for the first data set and the second data set,

wherein the key type queries are for fields of the first data fields and the second data fields associated with the key type,

wherein the dynamic value structure queries are for fields of the first data fields and the second data fields not indicated by the data action file.

12. The computer implemented method of claim 1 , wherein the configuration file is generated as the first plurality of data entries and the second plurality of data entries are being received.

13. A device comprising:

memory; and

one or more processors coupled to the memory and configured to perform operations including:

generating a configuration file, the configuration file including dynamic matching data describing a first plurality of data entries and a second plurality of data entries;

generating a data action file, the data action file identifying a first data set including the first plurality of data entries, a second data set including the second plurality of data entries, a key type, and a value structure associated with the key type, wherein the first data set and the second data set are stored in a data source system, wherein the first plurality of data entries includes first data fields, and wherein the second plurality of data entries includes second data fields;

dynamically generating a plurality of data queries including key type queries, value structure queries and dynamic value structure queries, wherein the dynamic value structure queries are generated in real-time based on the dynamic matching data indicated in the configuration file, and wherein the dynamic value structure queries are being generated based on a first data type, a second data type, and data field types present in the first data set and the second data set;

dynamically executing the plurality of data queries in parallel, wherein parallel execution of the plurality of data queries includes a plurality of simultaneous data queries to the data source system including the first data set and the second data set;

matching fields of the first plurality of data entries and the second plurality of data entries using the key type and the value structure associated with the key type;

identifying corresponding fields of the first data fields and the second data fields having a data mismatch; and

automatically generating a mismatch database entry for the corresponding fields having the data mismatch.

14. The device of claim 13 , where the first data set includes at least one billion data fields, and wherein the second data set includes at least one billion data fields.

15. The device of claim 13 , wherein the first data set comprises at least a terabyte of data, and wherein the second data set comprises at least a terabyte of data.

16. The device of claim 13 , wherein the configuration file further identifies the first data type for the first data set and the second data type for the second data set.

17. The device of claim 13 , wherein the one or more processors are configured for operations further comprising:

generating the configuration file associated with the first data set, wherein the configuration file further indicates values used for generation of the dynamic value structure queries, and wherein the configuration file is generated using a machine learning algorithm trained from mismatch database entry training data;

generating a feedback value associated with the mismatch database entry, wherein the feedback value identifies a difference between an expected result and an actual result in the mismatch database entry, wherein the feedback value is associated with settings in the configuration file; and

updating the mismatch database entry training data and the machine learning algorithm using the feedback value and the mismatch database entry to generate update mismatch database entry training data.

18. The device of claim 13 , wherein the one or more processors are configured for operations further comprising:

using a machine learning algorithm to generate a dynamic query selection table using mismatch training data, wherein the dynamic value structure queries are selected from the dynamic query selection table.

19. The device of claim 13 , wherein the one or more processors are configured for operations further comprising:

processing the mismatch database entry to generate a mismatch report providing mismatch metrics for values of the corresponding fields of the first data set and the second data set.

20. The device of claim 13 , wherein the one or more processors are configured for operations further comprising:

processing the mismatch database entry using a machine learning algorithm to remove mismatch data where the corresponding fields for the first data set and the second data set include matching content with data mismatches.

21. The device of claim 13 , wherein the one or more processors are configured for operations further comprising:

identifying different tiers of matching standards for different corresponding fields of the first data set and the second data set; and

merging the first data set and the second data set;

wherein merging the corresponding fields of the first data fields and the second data fields having the data mismatch comprises a multi-tier machine learning algorithm for merging entries having identical key values with mismatched fields.

22. The device of claim 13 , wherein the one or more processors are configured for operations further comprising:

identifying corresponding chunks of data from the first data set and the second data set;

wherein dynamically executing the plurality of data queries in parallel includes generating separate data chunks for independent parallel processing using the corresponding chunks of data from the first data set and the second data set, and

wherein identifying the corresponding fields of the first data fields and the second data fields having the data mismatch is performed separately on the separate data chunks for the corresponding chunks of data from the first data set and the second data set.

23. The device of claim 13 , wherein the one or more processors are configured for operations further comprising:

wherein the plurality of data queries are queries for the first data set and the second data set,

wherein the key type queries are for fields of the first data fields and the second data fields associated with the key type,

wherein the dynamic value structure queries are for fields of the first data fields and the second data fields not indicated by the data action file.

24. The device of claim 13 , wherein the configuration file is generated as the first plurality of data entries and the second plurality of data entries are being received.

25. A non-transitory computer readable storage medium comprising instructions that, when executed by one or more processors of a mobile device, cause the mobile device to perform operations including:

generating a configuration file, the configuration file including dynamic matching data describing a first plurality of data entries and a second plurality of data entries;

generating a data action file, the data action file identifying a first data set including the first plurality of data entries, a second data set including the second plurality of data entries, a key type, and a value structure associated with the key type, wherein the first data set and the second data set are stored in a data source system, wherein the first plurality of data entries includes first data fields, and wherein the second plurality of data entries includes second data fields;

dynamically generating a plurality of data queries including key type queries, value structure queries and dynamic value structure queries, wherein the dynamic value structure queries are generated in real-time based on the dynamic matching data indicated in the configuration file, and wherein the dynamic value structure queries are being generated based on a first data type, a second data type, and data field types present in the first data set and the second data set;

dynamically executing the plurality of data queries in parallel, wherein parallel execution of the plurality of data queries includes a plurality of simultaneous data queries to the data source system including the first data set and the second data set;

matching fields of the first plurality of data entries and the second plurality of data entries using the key type and the value structure associated with the key type;

identifying corresponding fields of the first data fields and the second data fields having a data mismatch; and

automatically generating a mismatch database entry for the corresponding fields having the data mismatch.

26. The non-transitory computer readable storage medium of claim 25 , where the first data set includes at least one billion data fields, and wherein the second data set includes at least one billion data fields.

27. The non-transitory computer readable storage medium of claim 25 , wherein the first data set comprises at least a terabyte of data, and wherein the second data set comprises at least a terabyte of data.

28. The non-transitory computer readable storage medium of claim 25 , wherein the configuration file further identifies the first data type for the first data set and the second data type for the second data set.

29. The non-transitory computer readable storage medium of claim 25 , wherein the one or more processors are configured for operations further comprising:

generating the configuration file associated with the first data set, wherein the configuration file further indicates values used for generation of the dynamic value structure queries, and wherein the configuration file is generated using a machine learning algorithm trained from mismatch database entry training data;

generating a feedback value associated with the mismatch database entry, wherein the feedback value identifies a difference between an expected result and an actual result in the mismatch database entry, wherein the feedback value is associated with settings in the configuration file; and

updating the mismatch database entry training data and the machine learning algorithm using the feedback value and the mismatch database entry to generate update mismatch database entry training data.

30. The non-transitory computer readable storage medium of claim 25 , wherein the one or more processors are configured for operations further comprising:

using a machine learning algorithm to generate a dynamic query selection table using mismatch training data, wherein the dynamic value structure queries are selected from the dynamic query selection table.

31. The non-transitory computer readable storage medium of claim 25 , wherein the one or more processors are configured for operations further comprising:

processing the mismatch database entry to generate a mismatch report providing mismatch metrics for values of the corresponding fields of the first data set and the second data set.

32. The non-transitory computer readable storage medium of claim 25 , wherein the one or more processors are configured for operations further comprising:

processing the mismatch database entry using a machine learning algorithm to remove mismatch data where the corresponding fields for the first data set and the second data set include matching content with data mismatches.

33. The non-transitory computer readable storage medium of claim 25 , wherein the one or more processors are configured for operations further comprising:

identifying different tiers of matching standards for different corresponding fields of the first data set and the second data set; and

merging the first data set and the second data set;

wherein merging the corresponding fields of the first data fields and the second data fields having the data mismatch comprises a multi-tier machine learning algorithm for merging entries having identical key values with mismatched fields.

34. The non-transitory computer readable storage medium of claim 25 , wherein the one or more processors are configured for operations further comprising:

identifying corresponding chunks of data from the first data set and the second data set;

wherein dynamically executing the plurality of data queries in parallel includes generating separate data chunks for independent parallel processing using the corresponding chunks of data from the first data set and the second data set, and

wherein identifying the corresponding fields of the first data fields and the second data fields having the data mismatch is performed separately on the separate data chunks for the corresponding chunks of data from the first data set and the second data set.

35. The non-transitory computer readable storage medium of claim 25 , wherein the one or more processors are configured for operations further comprising:

wherein the plurality of data queries are queries for the first data set and the second data set,

wherein the key type queries are for fields of the first data fields and the second data fields associated with the key type,

wherein the dynamic value structure queries are for fields of the first data fields and the second data fields not indicated by the data action file.

36. The non-transitory computer readable storage medium of claim 25 , wherein the configuration file is generated as the first plurality of data entries and the second plurality of data entries are being received.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 3, 2023
From: KONDAKINDI, SAISHARATH
To: SYNCHRONY BANK
Reel/Frame 064483/0372 →
Continuity (2)
Provisional Application 63370248 · Aug 3, 2022
Related Publication 20240045855A1 · Feb 8, 2024