IP Library Granted Patent US 10,657,287
Granted Patent B2
US 10,657,287 · App. 15/800,405 · Granted May 19, 2020

Identification of pseudonymized data within data sources

Inventors: Pedro Barbas (Dunboyne, IE); Austin Clifford (Glenageary, IE); Konrad Emanowicz (Maynooth, IE); Patrick G. O'Sullivan (Mulhuddart, IE)
Assignee: International Business Machines Corporation
G06F21/6254H04L9/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,657,287
App. No.
15/800,405
Granted
May 19, 2020
Kind
B2
Abstract

A computer-implemented method, computer program product and system for identifying pseudonymized data within data sources. One or more data repositories within one or more of the data sources are selected. One or more privacy data models are provided, where each of the privacy data models includes pattern(s) and/or parameter(s). One or more of the one or more privacy data models are selected. Data identification information is generated, where the data identification information indicates a presence or absence of pseudonymized data and of non-pseudonymized data within the one or more of the data sources. The data identification information is generated utilizing the pattern(s) and/or the parameter(s) to determine pseudonymized data.

Claims (34)

1. A computer program product for identifying pseudonymized data within data sources, the computer program product comprising a non-transitory computer readable storage medium having program code embodied therewith, the program code comprising the programming instructions for:

selecting one or more data repositories within one or more of said data sources;

providing one or more privacy data models, each of said privacy data models comprising one or both of one or more patterns and one or more parameters;

selecting one or more of said one or more privacy data models;

analyzing data stored in said one or more selected data repositories using one or more parameters from said selected one or more privacy data models, wherein said analysis using said one or more parameters from said selected one or more privacy data models comprises identifying a meaning of each individual data value within said one or more selected data repositories, wherein said meaning of each individual data value within said one or more selected data repositories is determined by using cognitive semantics to identify an absolute frequency of specific data values within said one or more selected data repositories;

calculating a frequency distribution table containing frequency distribution values based on said identified absolute frequency of specific data values within said one or more selected data repositories;

comparing said frequency distribution values of said frequency distribution table to a predefined threshold value to determine a presence or absence of pseudonymized data; and

determining a data value within a data source to be pseudonymized in response to its associated frequency distribution value exceeding said predefined threshold value.

2. The computer program product as recited claim 1 , wherein the program code further comprises the programming instructions for:

generating notifications for said pseudonymized data within said one or more of said data sources.

3. The computer program product as recited in claim 1 , wherein said one or more patterns correspond to an encoded cryptographic digest.

4. The computer program product as recited in claim 1 , wherein said one or more parameters correspond to absolute frequency information associated with specific data values.

5. The computer program product as recited in claim 1 , wherein said one or more parameters comprise a user parameter.

6. The computer program product as recited claim 1 , wherein the program code further comprises the programming instructions for:

analyzing data stored in said one or more selected data repositories using one or more patterns from said selected one or more privacy data models, wherein said analysis using said one or more patterns from said selected one or more privacy data models comprises using anonymization deconstruction techniques; and

identifying pseudonymized data by deconstructing cryptographic digests.

7. A system, comprising:

a memory unit for storing a computer program for identifying pseudonymized data within data sources; and

a processor coupled to the memory unit, wherein the processor is configured to execute the program instructions of the computer program comprising:

selecting one or more data repositories within one or more of said data sources;

providing one or more privacy data models, each of said privacy data models comprising one or both of one or more patterns and one or more parameters;

selecting one or more of said one or more privacy data models;

analyzing data stored in said one or more selected data repositories using one or more parameters from said selected one or more privacy data models, wherein said analysis using said one or more parameters from said selected one or more privacy data models comprises identifying a meaning of each individual data value within said one or more selected data repositories, wherein said meaning of each individual data value within said one or more selected data repositories is determined by using cognitive semantics to identify an absolute frequency of specific data values within said one or more selected data repositories;

calculating a frequency distribution table containing frequency distribution values based on said identified absolute frequency of specific data values within said one or more selected data repositories;

comparing said frequency distribution values of said frequency distribution table to a predefined threshold value to determine a presence or absence of pseudonymized data; and

determining a data value within a data source to be pseudonymized in response to its associated frequency distribution value exceeding said predefined threshold value.

8. The system as recited claim 7 , wherein the program instructions of the computer program further comprise:

generating notifications for said pseudonymized data within said one or more of said data sources.

9. The system as recited in claim 7 , wherein said one or more patterns correspond to an encoded cryptographic digest.

10. The system as recited in claim 7 , wherein said one or more parameters correspond to absolute frequency information associated with specific data values.

11. The system as recited in claim 7 , wherein said one or more parameters comprise a user parameter.

12. The system as recited claim 7 , wherein the program instructions of the computer program further comprise:

analyzing data stored in said one or more selected data repositories using one or more patterns from said selected one or more privacy data models, wherein said analysis using said one or more patterns from said selected one or more privacy data models comprises using anonymization deconstruction techniques; and

identifying pseudonymized data by deconstructing cryptographic digests.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2024
From: GREEN MARKET SQUARE LIMITED
To: WORKDAY, INC.
Reel/Frame 067801/0892 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2024
From: GREEN MARKET SQUARE LIMITED
To: WORKDAY, INC.
Reel/Frame 067556/0783 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2022
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: GREEN MARKET SQUARE LIMITED
Reel/Frame 058888/0675 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2017
From: BARBAS, PEDRO; CLIFFORD, AUSTIN; EMANOWICZ, KONRAD; O'SULLIVAN, PATRICK G.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 044004/0522 →
Continuity (1)
Related Publication 20190130132A1 · May 2, 2019