IP Library Granted Patent US 10,776,516
Granted Patent B2
US 10,776,516 · App. 16/051,881 · Granted Sep 15, 2020

Electronic medical record datasifter

Inventors: Ivaylo Dinov (Ann Arbor, MI); John Vandervest (Ann Arbor, MI); Simeone Marino (Ann Arbor, MI)
Assignee: THE REGENTS OF THE UNIVERSITY OF MICHIGAN
G06F21/6254G06F16/2272G06F17/11G06F17/18G06N7/08G16H10/60G16H50/70G06N5/003G06N7/005G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,776,516
App. No.
16/051,881
Granted
Sep 15, 2020
Kind
B2
Abstract

A method is presented for generating a data set from a database. The method involves iterative data manipulation that stochastically identifies candidate entries from the cases (subjects, participants) and variables (data elements) and subsequently selects, nullifies, and imputes the information. This process heavily relies on statistical multivariate imputation to preserve the joint distributions of the complex structured data archive. At each step, the algorithm generates a complete dataset that in aggregate closely resembles the intrinsic characteristics of the original data set, however, on an individual level the rows of data are substantially altered. This procedure drastically reduces the risk for subject reidentification by stratification, as meta-data for all subjects is repeatedly and lossily encoded.

Claims (55)

1. A computer-implemented method for generating a data set from a database, comprising:

receiving a set of records from a database, each record pertains to a particular entity and is comprised of a plurality of data fields;

identifying data fields in the set of records, wherein the data fields contain explicitly identifiable data for an entity;

obfuscating values in the data fields identified as containing explicitly identifiable data for the entity;

randomly selecting records in the set of records;

for each of the selected records, randomly selecting data fields in the selected records;

removing values in the selected data fields of the selected records in the set of records;

imputing values in data fields having missing values in the set of records using an imputation method; and

repeating the steps of removing values and imputing values for a number of iterations and thereby generate a new data set, where the number of iterations exceeds one.

2. The method of claim 1 further comprises obfuscating values in the data fields identified as containing explicitly identifiable data by removing the values from the identified data fields.

3. The method of claim 1 further comprises removing values in randomly selected data fields by randomly selecting data fields in the set of records using an alias method.

4. The method of claim 1 further comprises imputing values in data fields using multivariate imputation by chained equations.

5. The method of claim 1 further comprises, prior to the step of removing values in the data fields, identifying data fields in the set of records having unstructured data format and swapping values of at least one identified data field in a given pair of data records, where the given pair of data records is selected using a similarity measure between the data records in the set of records.

6. The method of claim 5 further comprises pairing a given data record in the set of records with each of the other data records in the set of records;

computing a similarity measure between values of data fields in each pair of records; and

identifying the given pair of records from amongst the pair of records, where the given pair of records has the highest similarity measure amongst the pair of records.

7. The method of claim 6 further comprises computing a similarity measures using a Bray-Curtis distance method.

8. The method of claim 5 further comprises

randomly selecting a given data field in the given data record;

pairing a given data record in the set of records with each of the other data records in the set of records;

computing a similarity measure between values of data fields in each pair of records;

forming a subset of data records from the set of records, where the records in the subset of data records have highest degree of similarity amongst the data records in the set of records;

randomly selecting a record from the subset of data records; and

swapping value of the given data field in the given data record with value in the corresponding data field on the randomly selected record from the subset of data records.

9. The method further comprises randomly selecting another data field in the given data record and, for the another data field in the given data record, repeating the steps of claim 8 .

10. The method further comprises retrieving another data record in the set of records and, for the another data record, repeating the steps of claim 8 .

11. A computer-implemented method for generating a data set from a database, comprising:

retrieving a set of records from a database, each record pertains to a particular person and is comprised of a plurality of data fields;

identifying data fields in the set of records, wherein the data fields that contain explicitly identifiable data for a person;

obfuscating values in the data fields identified as containing explicitly identifiable data for the person;

selecting at least one data field in a given data record in the set of records;

selecting a given pair of data records from the set of records, where the given pair of data records is selected using a similarity measure between the data records in the set of records;

swapping values of the at least one selected data field in the given pair of data records;

randomly selecting records in the set of records;

for each of the selected records, randomly selecting data fields in the selected records;

removing values in the selected data fields of the selected records in the set of records;

imputing values in data fields having missing values in the set of records using an imputation method; and

repeating the steps of removing values and imputing values for a number of iterations and thereby generate a new data set, where the number of iterations exceeds one.

12. The method of claim 11 further comprises obfuscating values in the data fields identified as containing explicitly identifiable data by removing the values from the identified data fields.

13. The method of claim 11 further comprises removing values in randomly selected data fields by randomly selecting data fields in the set of records using an alias method.

14. The method of claim 11 further comprises imputing values in data fields using multivariate imputation by chained equations.

15. The method of claim 11 wherein selecting a given pair of data records further comprises

pairing the given data record in the set of records with each of the other data records in the set of records;

computing a similarity measure between values of data fields in each pair of records; and

identifying the given pair of records from amongst the pair of records, where the given pair of records has the highest similarity measure amongst the pair of records.

16. The method of claim 15 further comprises computing a similarity measures using a Bray-Curtis distance method.

17. The method of claim 11 further comprises

randomly selecting a given data field in the given data record;

pairing a given data record in the set of records with each of the other data records in the set of records;

computing a similarity measure between values of data fields in each pair of records;

forming a subset of data records from the set of records, where the records in the subset of data records have highest degree of similarity amongst the data records in the set of records;

randomly selecting a record from the subset of data records; and

swapping value of the given data field in the given data record with value in the corresponding data field on the randomly selected record from the subset of data records.

18. The method further comprises randomly selecting another data field in the given data record and, for the another data field in the given data record, repeating the steps of claim 16 .

19. The method further comprises retrieving another data record in the set of records and, for the another data record, repeating the steps of claim 16 .

Assignments (2)
CONFIRMATORY LICENSE Recorded Sep 17, 2021
From: UNIVERSITY OF MICHIGAN
To: NATIONAL INSTITUTES OF HEALTH (NIH), U.S. DEPT. OF HEALTH AND HUMAN SERVICES (DHHS), U.S. GOVERNMENT
Reel/Frame 057516/0985 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2018
From: DINOV, IVAYLO; VANDERVEST, JOHN; MARINO, SIMEONE
To: THE REGENTS OF THE UNIVERSITY OF MICHIGAN
Reel/Frame 047363/0852 →
Continuity (2)
Provisional Application 62540184 · Aug 2, 2017
Related Publication 20190042791A1 · Feb 7, 2019