IP Library Granted Patent US 11,086,906
Granted Patent B2
US 11,086,906 · App. 15/938,084 · Granted Aug 10, 2021

System and method for reconciliation of data in multiple systems using permutation matching

Inventors: Paul Burchard (Jersey City, NJ); Vladimir M. Zakharov (Bronx, NY)
Assignee: Goldman Sachs & Co. LLC
G06F16/285G06F16/245
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,086,906
App. No.
15/938,084
Granted
Aug 10, 2021
Kind
B2
Abstract

A method includes obtaining first and second data sets to be reconciled and, using matching rules, identifying discrepancies between the data sets. The matching rules include at least one permutation key, where each permutation key identifies a subset of data to be grouped together in one of the data sets. Identifying the discrepancies includes attempting to match one or more first characteristics associated with the grouped subset of data in one of the data sets to one or more second characteristics associated with another of the data sets. The matching rules could involve multiple matching characteristics, and the matching rules could be generated using a metric to select the matching characteristics of the matching rules. The metric could be based on a combination of a number of matched data items and a number of matched groups of data items.

Claims (48)

1. A method comprising:

obtaining, by at least one processing device, a first data set generated by a first computing system and a second data set generated by a second computing system, the first and second data sets configured to be reconciled;

enriching, by the at least one processing device, at least one of the first and second data sets by adding feature values to the at least one data set, the feature values selected to account for a difference in format between the first and second data sets;

generating, by the at least one processing device, matching rules based on the first and second data sets;

using the matching rules, identifying, by the at least one processing device, discrepancies between the first and second data sets; and

outputting, by the at least one processing device, the discrepancies between the first and second data sets to at least one of the first computing system, the second computing system, or a third computing system,

wherein the matching rules include at least one permutation key, each permutation key identifying a subset of data to be grouped together in one of the first and second data sets;

wherein identifying the discrepancies comprises attempting to match one or more first characteristics associated with the grouped subset of data in one of the first and second data sets to one or more second characteristics associated with another of the first and second data sets;

wherein generating the matching rules comprises (i) randomly adding or removing possible permutation keys in multiple simulation iterations, (ii) determining a metric for each simulation iteration, and (iii) selecting the at least one permutation key from the possible permutation keys based on the metrics of the simulation iterations; and

wherein randomly adding or removing the possible permutation keys is constrained to permutation keys with dissimilar frequency distributions in the first and second data sets.

2. The method of claim 1 , wherein the first and second characteristics comprise a key that needs to match exactly between data items in the first and second data sets in order to match the data items in the first and second data sets.

3. The method of claim 1 , wherein the first and second characteristics comprise a value that needs to match in aggregate between groups of data items in the first and second data sets in order to match the groups of data items in the first and second data sets.

4. The method of claim 1 , wherein the first and second characteristics comprise a value that needs to match in aggregate within a tolerance between groups of data items in the first and second data sets in order to match the groups of data items in the first and second data sets.

5. The method of claim 1 , wherein the metric for each simulation iteration is based on a combination of a number of matched data items and a number of matched groups of data items.

6. The method of claim 1 , wherein the matching rules include multiple permutation keys, different permutation keys associated with different subsets of data in different ones of the first and second data sets.

7. An apparatus comprising:

at least one memory configured to store a first data set generated by a first computing system and a second data set generated by a second computing system, the first and second data sets configured to be reconciled; and

at least one processing device configured to:

enrich at least one of the first and second data sets by adding feature values to the at least one data set, the feature values selected to account for a difference in format between the first and second data sets;

generate matching rules based on the first and second data sets;

identify discrepancies between the first and second data sets using the matching rules; and

control the apparatus to output the discrepancies between the first and second data sets to at least one of the first computing system, the second computing system, or a third computing system,

wherein the matching rules include at least one permutation key, each permutation key identifying a subset of data to be grouped together in one of the first and second data sets;

wherein, to identify the discrepancies, the at least one processing device is configured to attempt to match one or more first characteristics associated with the grouped subset of data in one of the first and second data sets to one or more second characteristics associated with another of the first and second data sets;

wherein, to generate the matching rules, the at least one processing device is configured to (i) randomly add or remove possible permutation keys in multiple simulation iterations, (ii) determine a metric for each simulation iteration, and (iii) select the at least one permutation key from the possible permutation keys based on the metrics of the simulation iterations; and

wherein, to randomly add or remove the possible permutation keys, the at least one processing device is constrained to permutation keys with dissimilar frequency distributions in the first and second data sets.

8. The apparatus of claim 7 , wherein the first and second characteristics comprise a key that needs to match exactly between data items in the first and second data sets in order to match the data items in the first and second data sets.

9. The apparatus of claim 7 , wherein the first and second characteristics comprise a value that needs to match in aggregate between groups of data items in the first and second data sets in order to match the groups of data items in the first and second data sets.

10. The apparatus of claim 7 , wherein the first and second characteristics comprise a value that needs to match in aggregate within a tolerance between groups of data items in the first and second data sets in order to match the groups of data items in the first and second data sets.

11. The apparatus of claim 7 , wherein the metric for each simulation iteration is based on a combination of a number of matched data items and a number of matched groups of data items.

12. The apparatus of claim 7 , wherein the matching rules include multiple permutation keys, different permutation keys associated with different subsets of data in different ones of the first and second data sets.

13. A non-transitory computer readable medium containing instructions that when executed cause at least one processor to:

obtain a first data set generated by a first computing system and a second data set generated by a second computing system, the first and second data sets configured to be reconciled;

enrich at least one of the first and second data sets by adding feature values to the at least one data set, the feature values selected to account for a difference in format between the first and second data sets;

generate matching rules based on the first and second data sets;

using the matching rules, identify discrepancies between the first and second data sets; and

output the discrepancies between the first and second data sets to at least one of the first computing system, the second computing system, or a third computing system,

wherein the matching rules include at least one permutation key, each permutation key identifying a subset of data to be grouped together in one of the first and second data sets;

wherein the instructions that when executed cause the at least one processor to identify the discrepancies comprise:

instructions that when executed cause the at least one processor to attempt to match one or more first characteristics associated with the grouped subset of data in one of the first and second data sets to one or more second characteristics associated with another of the first and second data sets;

wherein the instructions that when executed cause the at least one processor to generate the matching rules comprise:

instructions that when executed cause the at least one processor to (i) randomly add or remove possible permutation keys in multiple simulation iterations, (ii) determine a metric for each simulation iteration, and (iii) select the at least one permutation key from the possible permutation keys based on the metrics of the simulation iterations; and

wherein the randomly added or removed possible permutation keys are constrained to permutation keys with dissimilar frequency distributions in the first and second data sets.

14. The non-transitory computer readable medium of claim 13 , wherein the first and second characteristics comprise a key that needs to match exactly between data items in the first and second data sets in order to match the data items in the first and second data sets.

15. The non-transitory computer readable medium of claim 13 , wherein the first and second characteristics comprise a value that needs to match in aggregate between groups of data items in the first and second data sets in order to match the groups of data items in the first and second data sets.

16. The non-transitory computer readable medium of claim 13 , wherein the first and second characteristics comprise a value that needs to match in aggregate within a tolerance between groups of data items in the first and second data sets in order to match the groups of data items in the first and second data sets.

17. The non-transitory computer readable medium of claim 13 , wherein the metric for each simulation iteration is based on a combination of a number of matched data items and a number of matched groups of data items.

18. The non-transitory computer readable medium of claim 13 , wherein the matching rules include multiple permutation keys, different permutation keys associated with different subsets of data in different ones of the first and second data sets.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2022
From: PATHAN, SAHIR; BAXI, KUNAL; AMIR, BRANDON OREN
To: GOLDMAN SACHS & CO. LLC
Reel/Frame 058889/0269 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2018
From: BURCHARD, PAUL; ZAKHAROV, VLADIMIR M.
To: GOLDMAN SACHS & CO. LLC
Reel/Frame 045370/0474 →
Continuity (2)
Provisional Application 62485264 · Apr 13, 2017
Related Publication 20180300390A1 · Oct 18, 2018