IP Library Granted Patent US 11,693,821
Granted Patent B2
US 11,693,821 · App. 17/369,798 · Granted Jul 4, 2023

Systems and methods for performant data matching

Inventors: Curtiss W. Schuler (Glendale, AZ); Brett A. Norris (North Vancouver, CA); Satyender Goel (Chicago, IL)
Assignee: Collibra Belgium BV
G06F16/152G06F16/137
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,693,821
App. No.
17/369,798
Granted
Jul 4, 2023
Kind
B2
Abstract

The present disclosure is directed to systems and methods for performant data matching. Entities maintain large amounts of data and desire to reconcile duplicative records. One way to solve this problem is through data matching. However, standard data matching at the record level can be laborious and inefficient. To remedy these inefficiencies in data matching, the present disclosure describes a system where the token records are tokenized a second time into token sets based on the token records satisfying at least one token set rule. A token set rule may be based on the common presence of multiple tokens in a token record. If multiple token records have the required tokens from the set rule, then those token records can be hashed and rolled-up into the token set (i.e., tokenized a second time into the token set). The token set allows for more efficient data matching.

Claims (53)

1. A system for data matching, comprising:

a memory configured to store non-transitory computer readable instructions; and

a processor communicatively coupled to the memory, wherein the processor, when executing the non-transitory computer readable instructions, is configured to:

receive at least one token record from a first source;

receive at least one token record from a second source;

compare the at least one token record from the first source to the at least one token record from the second source;

based on the comparison of the at least one token record from the first source to the at least one token from the second source, identify at least one common token in the at least one token record from the first source and in the at least one token record from the second source, wherein the at least one common token is associated with at least one set rule;

based on the at least one set rule, match the at least one token record from the first source to the at least one token record from the second source; and

generate a token set by merging the at least one token record from the first source with the at least one token record from the second source using at least one hashing roll-up function, wherein the token set comprises the at least one token record from the first source and the at least one token record from the second source.

2. The system of claim 1 , the processor further configured to:

compare the token set to at least one universal reference token repository; and

based on a determination that the token set is not present in the at least one universal token repository, store the token set in the universal reference token repository.

3. The system of claim 1 , wherein the at least one token record from the first source is associated with at least two tokens.

4. The system of claim 1 , wherein the at least one token record from the second source is associated with at least two tokens.

5. The system of claim 1 , wherein the at least one universal reference token repository is comprised of at least one of: a customer token and a reference source token.

6. The system of claim 1 , wherein the first source and the second source share a common owner.

7. The system of claim 1 , wherein the set rule is comprised of at least two token requirements, wherein one of the at least two token requirements is the at least one common token.

8. The system of claim 1 , wherein the token set is comprised of a plurality of token records, wherein each of the plurality of token records comprise the at least one common token.

9. The system of claim 1 , wherein the first source or the second source is at least one of: a customer source and a reference source.

10. The system of claim 1 , wherein the processor is further configured to:

receive at least one token record from a third source, wherein the at least one token record from the third source comprises the at least one common token.

11. The system of claim 10 , wherein the processor is further configured to:

match the at least one token record from the third source with the token set, wherein the token set is based on the set rule, wherein the set rule requires the at least one common rule; and

merge the at least one token record from the third source with the token set.

12. A method for data matching, comprising:

receiving at least one token record from a first source;

receiving at least one token record from a second source;

receiving at least one token set rule, wherein the at least one token set rule is based on the presence of at least one common token;

comparing the at least one token record from the first source to the at least one token rule;

comparing the at least one token record from the second source to the at least one token rule; determining, by simple weighting, that the at least one token record from the first source and the at least one token record from the second source satisfy the at least one token rule;

merging, using at least one hashing roll-up function, the at least one token record from the first source with the at least one token record from the second source; and

generating a token set, wherein the token set comprises the merged at least one token record from the first source and the at least one token record from the second source.

13. The method of claim 12 , wherein the at least one token record from the first source and the at least one token record from the second source are encrypted with the same encryption algorithm.

14. The method of claim 12 , further comprising:

comparing the token set to at least one universal reference token repository; and

based on a determination that the token set is not present in the at least one universal token repository, storing the token set in the universal reference token repository.

15. The method of claim 14 , wherein the universal reference token repository is comprised of a plurality of token records from at least one of: a customer source and a reference source.

16. The method of claim 12 , further comprising:

comparing at least one token record from a third source to the token set; and

retokenizing, by at least one hashing roll-up function, the at least one token record from the third source into the token set.

17. The method of claim 16 , further comprising:

based on the at least one hashing roll-up function, identifying at least one duplicate token record in the same source.

18. The method of claim 16 , further comprising:

based on the at least one hashing roll-up function, identifying at least one overlapping token record in the third source.

19. A computer-readable media storing non-transitory computer executable instructions that when executed cause a computing system to perform a method for data matching, comprising:

receiving at least one token record from a first source, wherein the at least one record from the first source comprises at least one common token;

receiving at least one token record from a second source, wherein the at least one record from the second source comprises the at least one common token;

receiving at least one token set rule, wherein the at least one token set rule is based on the presence of the at least one common token;

comparing the at least one token record from the first source to the at least one token rule;

comparing the at least one token record from the second source to the at least one token rule;

determining that the at least one token record from the first source and the at least one token record from the second source satisfy the at least one token rule based on the presence of the at least one common token;

merging, using at least one hashing roll-up function, the at least one token record from the first source with the at least one token record from the second source; and

generating a token set, wherein the token set comprises the merged at least one token record from the first source and the at least one token record from the second source.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2023
From: COLLIBRA NV; CNV NEWCO B.V.; COLLIBRA B.V.
To: COLLIBRA BELGIUM BV
Reel/Frame 062989/0023 →
SECURITY INTEREST Recorded Jan 4, 2023
From: COLLIBRA BELGIUM BV
To: SILICON VALLEY BANK UK LIMITED
Reel/Frame 062269/0955 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2021
From: SCHULER, CURTISS W.; NORRIS, BRETT A.; GOEL, SATYENDER
To: COLLIBRA NV
Reel/Frame 056857/0907 →
Continuity (1)
Related Publication 20230014556A1 · Jan 19, 2023
Cited By (1)
US 12,619,575