IP Library Granted Patent US 11,093,640
Granted Patent B2
US 11,093,640 · App. 15/951,251 · Granted Aug 17, 2021

Augmenting datasets with selected de-identified data records

Inventor: Aris Gkoulalas-Divanis (Waltham, MA)
Assignee: International Business Machines Corporation
G06F21/6254G06F21/604
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,093,640
App. No.
15/951,251
Granted
Aug 17, 2021
Kind
B2
Abstract

A computer system utilizes a dataset to support a research study. Regions of interestingness are determined within a model of data records of a first dataset that are authorized for the research study by associated entities. Data records from a second dataset are represented within the model, wherein the data records from the second dataset are relevant for supporting objectives of the research study. Data records from the second dataset that fail to satisfy de-identification requirements are removed. A resulting dataset is generated that including the first dataset records within a selected region of interestingness and selected records of the second dataset within the same region. The second dataset records within the resulting dataset are de-identified based on the de-identification requirements. Embodiments of the present invention further include a method and program product for utilizing a dataset to support a research study in substantially the same manner described above.

Claims (29)

1. A computer system for utilizing a dataset for supporting a research study, the computer system comprising:

one or more computer processors;

one or more computer readable storage media;

program instructions stored on the one or more computer readable storage media for execution by at least one of the one or more computer processors, the program instructions comprising instructions to:

determine one or more regions of interestingness within a model of data records of a first dataset, wherein the data records of the first dataset are authorized for the research study by associated entities and are not subject to de-identification requirements, wherein the model includes dimensions of quasi-identifiers, and wherein each data record of the first dataset is represented according to values of the quasi-identifiers of the data record;

represent within the model data records from a second dataset, wherein the data records from the second dataset are relevant for supporting objectives of the research study, correspond to entities other than those associated with the first dataset, and are authorized for the research study by associated entities after transformation to satisfy the de-identification requirements, and wherein each data record of the second dataset is represented in the model according to values of the quasi-identifiers of the data record;

remove data records of the second dataset from the one or more regions of interestingness based on those records failing to satisfy the de-identification requirements;

generate a resulting dataset for the research study including the data records of the first dataset within a selected region of interestingness and selected data records of the second dataset, wherein the selected region of interestingness corresponds to a cluster of data records of the first dataset that is identified based on the values of the quasi-identifiers of the data records of the first dataset, and wherein the selected data records of the second dataset are selected based on the values of the quasi-identifiers of the selected data records being within the selected region of interestingness; and

de-identify the data records of the second dataset within the resulting dataset based on the de-identification requirements.

2. The computer system of claim 1 , wherein generating a resulting dataset further comprises:

identifying data records of the second dataset within the selected region of interestingness that are similar to data records of the first dataset; and

generating the resulting dataset with the identified data records of the second dataset.

3. The computer system of claim 1 , wherein the data records from the second data set that are relevant for supporting objectives of the research comprise data records from the second data set that fit within a region of interestingness.

4. The computer system of claim 1 , wherein the data records of the second dataset are de-identified in separate groups corresponding to each of the one or more regions of interestingness.

5. The computer system of claim 1 , wherein each region of interestingness is identified by analyzing the data records from the first dataset to find clusters of records.

6. The computer system of claim 1 , wherein each data record of the first dataset includes one or more of: a direct identifier, and a quasi-identifier.

7. A computer program product for utilizing a dataset for supporting a research study, the computer program product comprising one or more computer readable storage media collectively having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to:

determine one or more regions of interestingness within a model of data records of a first dataset, wherein the data records of the first dataset are authorized for the research study by associated entities and are not subject to de-identification requirements, wherein the model includes dimensions of quasi-identifiers, and wherein each data record of the first dataset is represented according to values of the quasi-identifiers of the data record;

represent within the model data records from a second dataset, wherein the data records from the second dataset are relevant for supporting objectives of the research study, correspond to entities other than those associated with the first dataset, and are authorized for the research study by associated entities after transformation to satisfy the de-identification requirements, and wherein each data record of the second dataset is represented in the model according to values of the quasi-identifiers of the data record;

remove data records of the second dataset from the one or more regions of interestingness based on those records failing to satisfy the de-identification requirements;

generate a resulting dataset for the research study including the data records of the first dataset within a selected region of interestingness and selected data records of the second dataset, wherein the selected region of interestingness corresponds to a cluster of data records of the first dataset that is identified based on the values of the quasi-identifiers of the data records of the first dataset, and wherein the selected data records of the second dataset are selected based on the values of the quasi-identifiers of the selected data records being within the selected region of interestingness; and

de-identify the data records of the second dataset within the resulting dataset based on the de-identification requirements.

8. The computer program product of claim 7 , wherein generating a resulting dataset further comprises:

identifying data records of the second dataset within the selected region of interestingness that are similar to data records of the first dataset; and

generating the resulting dataset with the identified data records of the second dataset.

9. The computer program product of claim 7 , wherein the data records from the second data set that are relevant for supporting objectives of the research comprise data records from the second data set that fit within a region of interestingness.

10. The computer program product of claim 7 , wherein the data records of the second dataset are de-identified in separate groups corresponding to each of the one or more regions of interestingness.

11. The computer program product of claim 7 , wherein each region of interestingness is identified by analyzing the data records from the first dataset to find clusters of records.

12. The computer program product of claim 7 , wherein each data record of the first dataset includes one or more of: a direct identifier, and a quasi-identifier.

Assignments (3)
SECURITY INTEREST Recorded Oct 1, 2025
From: MERATIVE US L.P.; MERGE HEALTHCARE INCORPORATED
To: TCG SENIOR FUNDING L.L.C., AS COLLATERAL AGENT
Reel/Frame 072808/0442 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2022
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: MERATIVE US L.P.
Reel/Frame 061496/0752 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 12, 2018
From: GKOULALAS-DIVANIS, ARIS
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 045515/0159 →
Continuity (1)
Related Publication 20190318124A1 · Oct 17, 2019