IP Library Granted Patent US 11,093,646
Granted Patent B2
US 11,093,646 · App. 16/449,682 · Granted Aug 17, 2021

Augmenting datasets with selected de-identified data records

Inventor: Aris Gkoulalas-Divanis (Waltham, MA)
Assignee: International Business Machines Corporation
G06F21/6254G06F21/604
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,093,646
App. No.
16/449,682
Granted
Aug 17, 2021
Kind
B2
Abstract

A computer system utilizes a dataset to support a research study. Regions of interestingness are determined within a model of data records of a first dataset that are authorized for the research study by associated entities. Data records from a second dataset are represented within the model, wherein the data records from the second dataset are relevant for supporting objectives of the research study. Data records from the second dataset that fail to satisfy de-identification requirements are removed. A resulting dataset is generated that including the first dataset records within a selected region of interestingness and selected records of the second dataset within the same region. The second dataset records within the resulting dataset are de-identified based on the de-identification requirements. Embodiments of the present invention further include a method and program product for utilizing a dataset to support a research study in substantially the same manner described above.

Claims (13)

1. A method, in a data processing system comprising at least one processor and at least one memory, the at least one memory comprising instructions executed by the at least one processor to cause the at least one processor to utilize a dataset for supporting a research study, the method comprising:

determining one or more regions of interestingness within a model of data records of a first dataset, wherein the data records of the first dataset are authorized for the research study by associated entities and are not subject to de-identification requirements, wherein the model includes dimensions of quasi-identifiers, and wherein each data record of the first dataset is represented according to values of the quasi-identifiers of the data record;

representing within the model data records from a second dataset, wherein the data records from the second dataset are relevant for supporting objectives of the research study, correspond to entities other than those associated with the first dataset, and are authorized for the research study by associated entities after transformation to satisfy the de-identification requirements, and wherein each data record of the second dataset is represented in the model according to values of the quasi-identifiers of the data record;

removing data records of the second dataset from the one or more regions of interestingness based on those records failing to satisfy the de-identification requirements;

generating a resulting dataset for the research study including the data records of the first dataset within a selected region of interestingness and selected data records of the second dataset, wherein the selected region of interestingness corresponds to a cluster of data records of the first dataset that is identified based on the values of the quasi-identifiers of the data records of the first dataset, and wherein the selected data records of the second dataset are selected based on the values of the quasi-identifiers of the selected data records being within the selected region of interestingness; and

de-identifying the data records of the second dataset within the resulting dataset based on the de-identification requirements.

2. The method of claim 1 , wherein generating a resulting dataset further comprises:

identifying data records of the second dataset within the selected region of interestingness that are similar to data records of the first dataset; and

generating the resulting dataset with the identified data records of the second dataset.

3. The method of claim 1 , wherein the data records from the second data set that are relevant for supporting objectives of the research comprise data records from the second data set that fit within a region of interestingness.

4. The method of claim 1 , wherein the data records of the second dataset are de-identified in separate groups corresponding to each of the one or more regions of interestingness.

5. The method of claim 1 , wherein each region of interestingness is identified by analyzing the data records from the first dataset to find clusters of records.

6. The method of claim 1 , wherein each data record of the first dataset includes one or more of: a direct identifier, and a quasi-identifier.

Assignments (3)
SECURITY INTEREST Recorded Oct 1, 2025
From: MERATIVE US L.P.; MERGE HEALTHCARE INCORPORATED
To: TCG SENIOR FUNDING L.L.C., AS COLLATERAL AGENT
Reel/Frame 072808/0442 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2022
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: MERATIVE US L.P.
Reel/Frame 061496/0752 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2019
From: GKOULALAS-DIVANIS, ARIS
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 049564/0654 →