IP Library › Granted Patent US 11,593,405
Granted Patent B2
US 11,593,405 · App. 14/692,543 · Granted Feb 28, 2023

Custodian disambiguation and data matching

Inventors: Lars Bremer (Boeblingen, DE); Thomas A. P. Hampp-Bahnmueller (Stuttgart, DE); Markus Lorch (Dettenhausen, DE); Pavlo Petrenko (Boeblingen, DE); Sebastian B. Schmid (Stuttgart, DE)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F16/285G06F16/215G06F16/22G06F16/2228G06F16/2365G06F16/25G06F16/258G06F21/604G06F21/62G06F21/6245G06Q10/10G06Q50/18G06Q50/265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,593,405
App. No.
14/692,543
Granted
Feb 28, 2023
Kind
B2
Abstract

Provided is a technique for matching different user representations of a person in a plurality of computer systems may be provided. The technique includes collecting information sets about user representations from a plurality of computer systems; normalizing the information sets to a unified format; grouping the information sets in the unified format into indexing buckets based on a user name using a non-phonetic algorithm; determining a similarity score for each pair of information sets in each of the indexing buckets; classifying each information set pair into a set of classes based on the similarity scores, wherein the set of classes comprise at least matches and non-matches; and using a data structure for merging information of information set pairs classified as matches.

Claims (32)

1. A system, the system comprising:

a processor; and

storage coupled to the processor, wherein the storage stores program instructions, and wherein the program instructions, when executed by the processor perform:

collecting information sets about user representations from a plurality of computer systems, wherein the user representations comprise heterogeneous formats, and wherein each of the user representations is used to access an application of a plurality of applications; and

matching the user representations by:

normalizing the information sets into normalized information sets, wherein each of the normalized information sets comprises one or more attributes, and wherein each of the one or more attributes has a weight;

grouping the normalized information sets into a plurality of buckets using user names;

for each pair of the normalized information sets in each of the plurality of buckets,

determining a sub-similarity score for each corresponding attribute by comparing each attribute in that pair; and

in response to determining the sub-similarity score,

determining a weighted similarity score for each corresponding attribute in that pair based on the sub-similarity score and the weight for that attribute; and

determining a similarity score for the pair by comparing each normalized information set in that pair based on each weighted similarity score for each corresponding attribute; and

merging each pair classified as a match based on the similarity score for that pair to integrate the user representations as matched user representations, wherein the merging uses a combination of a Friend-of-a-Friend ontology, a Resource Description Format, and a Web Ontology Language; and

performing a hold on data that a user has access to, wherein the user and the data are identified using the matched user representations for that user.

2. The system according to claim 1 , wherein a non-phonetic algorithm is any one of a q-gram algorithm and a suffix algorithm, and wherein the non-phonetic algorithm is used to group the normalized information sets.

3. The system according to claim 1 , wherein a longest common sub-string is determined using dynamic programming algorithms and is used to determine the similarity score.

4. A computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions being executable by at least one processor of a computer to cause the computer to perform:

collecting information sets about user representations from a plurality of computer systems, wherein the user representations comprise heterogeneous formats, and wherein each of the user representations is used to access an application of a plurality of applications; and

matching the user representations by:

normalizing the information sets into normalized information sets, wherein each of the normalized information sets comprises one or more attributes, and wherein each of the one or more attributes has a weight;

grouping the normalized information sets into a plurality of buckets using user names;

for each pair of the normalized information sets in each of the plurality of buckets,

determining a sub-similarity score for each corresponding attribute by comparing each attribute in that pair; and

in response to determining the sub-similarity score,

determining a weighted similarity score for each corresponding attribute in that pair based on the sub-similarity score and the weight for that attribute; and

determining a similarity score for the pair by comparing each normalized information set in that pair based on each weighted similarity score for each corresponding attribute; and

merging each pair classified as a match based on the similarity score for that pair to integrate the user representations as matched user representations, wherein the merging uses a combination of a Friend-of-a-Friend ontology, a Resource Description Format, and a Web Ontology Language; and

performing a hold on data that a user has access to, wherein the user and the data are identified using the matched user representations for that user.

5. The computer program product according to claim 4 , wherein a non-phonetic algorithm is any one of a q-gram algorithm and a suffix algorithm, and wherein the non-phonetic algorithm is used to group the normalized information sets.

6. The computer program product according to claim 4 , wherein a longest common sub-string is determined using dynamic programming algorithms and is used to determine the similarity score.

7. The system according to claim 1 , wherein determining the similarity score comprises using a longest common sub-string algorithm to generate an output and using any of an Overlap coefficient and a Dice coefficient on the output of the longest common sub-string algorithm.

8. The computer program product according to claim 4 , wherein determining the similarity score comprises using a longest common sub-string algorithm to generate an output and using any of an Overlap coefficient and a Dice coefficient on the output of the longest common sub-string algorithm.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2015
From: BREMER, LARS; HAMPP-BAHNMUELLER, THOMAS A.P.; LORCH, MARKUS; PETRENKO, PAVLO; SCHMID, SEBASTIAN B.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 035569/0183 →
Continuity (1)
Related Publication 20160314183A1 · Oct 27, 2016