IP Library › Granted Patent US 12,189,820
Granted Patent B2
US 12,189,820 · App. 18/128,938 · Granted Jan 7, 2025

Systems and methods of data transformation for data pooling

Inventors: Lon Michel Luk Arbuckle (Ottawa, CA); Jordan Elijah Collins (Cambridge, CA); Khaldoun Zine El Abidine (Montreal, CA); Khaled El Emam (Ottawa, CA)
Assignee: Privacy Analytics Inc.
G06F21/6254G06N20/00G06F2221/2107
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,189,820
App. No.
18/128,938
Granted
Jan 7, 2025
Kind
B2
Abstract

A data anonymization pipeline system for managing holding and pooling data is disclosed. The data anonymization pipeline system transforms personal data at a source and then stores the transformed data in a safe environment. Furthermore, a re-identification risk assessment is performed before providing access to a user to fetch the de-identified data for secondary purposes.

Claims (49)

1. A method comprising:

determining a risk value representing a risk of re-identification for data received from a data source;

determining that the risk value does not satisfy a threshold risk value;

in response to determining that the risk value does not satisfy the threshold risk value, performing the following operations iteratively using a statistical model for a plurality of iterations, the operations comprising:

for each iteration of the plurality of iterations, updating the data with a respective transformation using the statistical model; and

determining a respective risk value representing the risk of re-identification for the updated data at the iteration; and

providing the updated data to a data pool.

2. The method of claim 1 , the data having been de-identified by:

generalizing information that identifies initial data stored in the data source using the statistical model; and

obtaining the data based on the generalized information.

3. The method of claim 1 , wherein the data is a cluster of a plurality of clusters of initial data that has been de-identified by generalizing information that identifies the initial data using the statistical model.

4. The method of claim 3 , wherein the plurality of clusters of initial data are generated using a machine learning model.

5. The method of claim 3 , wherein, for each iteration of the plurality of iterations, updating the data using the statistical model comprises: adjusting a cluster size of the data according to a comparison between the risk value and the threshold risk value.

6. The method of claim 1 , wherein determining that the risk value does not satisfy the threshold risk value comprises determining that the risk value is greater than the threshold risk value.

7. The method of claim 1 , wherein, for each iteration of the plurality of iterations, updating the data using the statistical model comprises one or more of: removing one or more individual data items from the data or adding one or more individual data items from the data source into the data.

8. The method of claim 1 , wherein determining the risk value comprises accounting for indirect identifiers present in the data.

9. The method of claim 8 , wherein accounting for the indirect identifiers present in the data comprises: accounting for a number of occurrences of the indirect identifiers and relationships of the indirect identifiers in the data.

10. The method of claim 1 , further comprising:

generating population statistics for the updated data according to the statistical model; and

tuning the updating process according to the generated population statistics.

11. The method of claim 1 , further comprising:

in response to determining that the risk value does not satisfy the threshold risk value, denying user access to the updated data.

12. The method of claim 1 , wherein the data comprises a stream of data or incremental data.

13. The method of claim 1 , wherein the data has been de-identified using a cryptographic key.

14. The method of claim 1 wherein the information represents one or more of demographic data and socio-economic data.

15. The method of claim 1 , further comprising:

upon providing the updated data from the data pool in response to a user request, applying one or more transformations to the updated data.

16. A system comprising:

one or more computers and one or more storage devices on which are stored instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

determining a risk value representing a risk of re-identification for data received from a data source;

determining that the risk value does not satisfy a threshold risk value;

in response to determining that the risk value does not satisfy the threshold risk value, performing the following operations iteratively using a statistical model for a plurality of iterations, the operations comprising:

for each iteration of the plurality of iterations, updating the data with a respective transformation using the statistical model; and

determining a respective risk value representing the risk of re-identification for the updated data at the iteration; and

providing the updated data to a data pool.

17. The system of claim 16 , the data having been de-identified by:

generalizing information that identifies initial data stored in the data source using the statistical model; and

obtaining the data based on the generalized information.

18. The system of claim 16 , wherein the data is a cluster of a plurality of clusters of initial data that has been de-identified by generalizing information that identifies the initial data using the statistical model.

19. One or more non-transitory computer storage media encoded with computer program instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

determining a risk value representing a risk of re-identification for data received from a data source;

determining that the risk value does not satisfy a threshold risk value;

in response to determining that the risk value does not satisfy the threshold risk value, performing the following operations iteratively using a statistical model for a plurality of iterations, the operations comprising:

for each iteration of the plurality of iterations, updating the data with a respective transformation using the statistical model; and

determining a respective risk value representing the risk of re-identification for the updated data at the iteration; and

providing the updated data to a data pool.

20. The one or more non-transitory computer storage media of claim 19 , the data having been de-identified by:

generalizing information that identifies initial data stored in the data source using the statistical model; and

obtaining the data based on the generalized information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2023
From: ARBUCKLE, LON MICHEL LUK; COLLINS, JORDAN ELIJAH; EL ABIDINE, KHALDOUN ZINE; EMAM, KHALED EL
To: PRIVACY ANALYTICS INC.
Reel/Frame 063201/0671 →
Continuity (3)
Continuation 16832212 · Mar 27, 2020
Provisional Application 62824696 · Mar 27, 2019
Related Publication 20230237196A1 · Jul 27, 2023
References Cited (18)
US 10803201B1 · Nicholls · 2020 [cited by examiner]
US 11620408B2 · Arbuckle et al. · 2023 [cited by applicant]
US 20150007249A1 · Bezzi et al. · 2015 [cited by applicant]
US 20170124351A1 · Scaiano et al. · 2017 [cited by applicant]
US 20180004978A1 · Hebert et al. · 2018 [cited by applicant]
US 20190026490A1 · Ahmed · 2019 [cited by examiner]
US 20190188292A1 · Gkoulalas-Divanis · 2019 [cited by applicant]
US 20190260784A1 · Stockdale · 2019 [cited by examiner]
US 20200082290A1 · Pascale · 2020 [cited by examiner]
US 20200311308A1 · Arbuckle et al. · 2020 [cited by applicant]
US 20200327252A1 · Derek et al. · 2020 [cited by applicant]
US 20220129584A1 · Blackport et al. · 2022 [cited by applicant]
WO WO2015047665 · 2015 [cited by applicant]
WO WO2018057479 · 2018 [cited by applicant]
[No Author Listed], “Privacy enhancing data de-identification terminology and classification of techniques,” ISO/IEC 20889, International Standard, Nov. 2019, 7 pages (preview only). [cited by applicant]
El Emam et al., “De-identification Methods for Open Health Data: The Case of the Heritage Health Prize Claims Dataset,” J Med Internet Res., Feb. 27, 2012, 14(1):e33. [cited by applicant]
El Emam, “Methods for the de-identification of electronic health records for genomic research,” Genome Med., 2011, 3(25):1-9. [cited by applicant]
Extended European Search Report in Application No. 20166365.5, dated Jul. 30, 2020, 8 pages. [cited by applicant]
Cited By (1)
US 12,651,090