IP Library Patent Application 13672318
Patent Application
App. No. 13/672,318

SYSTEM AND METHOD FOR EVALUATING MARKETER RE-IDENTIFICATION RISK

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
13/672,318
Abstract

Disclosures of databases for secondary purposes is increasing rapidly and any identification of personal data may from a dataset of database can be detrimental. A re-identification risk metric is determined for the scenario where an intruder wishes to re-identify as many records as possible in a disclosed database, known as a marketer risk. The dataset can be analyzed to determine equivalence classes for variables in the dataset and one or more equivalence class sizes. The re-identification risk metric associated with the dataset can be determined using a modified log-linear model by measuring a goodness of fit measure generalized for each of the one or more equivalence class sizes.

Claims (199)

1 . A method of assessing re-identification risk of a dataset containing personal information, the method executed by a processor comprising:

retrieving the dataset comprising a plurality of records from a storage device;

receiving variables selected from a plurality of variables present in the dataset, wherein the variables may be used as potential identifiers of personal information from the dataset; and

determining equivalence classes for each of the selected variables in the dataset and one or more equivalence class sizes;

determining a re-identification risk metric associated with the dataset using a modified log-linear model by measuring a goodness of fit measure generalized for each of the one or more equivalence class sizes.

2 . The method according to claim 1 , wherein determining a re-identification risk metric using a modified log-linear model comprises:

for the one or more equivalence classes:

determining the goodness of fit measure for the size of the equivalence class; and

determining a portion of a re-identification risk associated with the size of the equivalence class; and

determining the re-identification risk by summing all the determined portion of the re-identification risk.

3 . The method according to claim 2 , wherein determining the portion of the re-identification risk comprises:

calculating

h

k

(

γ

j

)

=

f

j

=

k

(

k

/

F

j

N

)

where h k is the portion of the re-identification risk associated with equivalence class size k, λ j is the actual re-identification risk, F j is the equivalence class sizes in an identification database, N is the set of records in the identification database.

4 . The method according to claim 2 , wherein the goodness of fit measures a bias arising from difference between an estimated re-identification risk and an actual re-identification risk.

5 . The method according to claim 4 , wherein measuring the bias comprises:

calculating

B

k

=

j

E

(

I

(

f

j

=

k

)

)

[

h

k

(

y

j

)

-

h

k

(

γ

j

)

]

where B k is the goodness of fit measure for equivalence class size k, f j is the equivalence sizes in the de-identified dataset, and {circumflex over (γ)} j is the estimated re-identification risk.

6 . The method according to claim 2 , wherein the risk threshold selected is less than

R

J

=

1

/

min

j

(

F

j

)

where R J is journalist risk.

7 . The method of claim 2 further comprising:

receiving a re-identification risk threshold value acceptable for the dataset; and

comparing the re-identification risk metric meets the risk threshold value.

8 . The method according to claim 7 , wherein if the re-identification metric is greater than the risk threshold the further comprising:

performing de-identification of the retrieved dataset based upon one or more equivalence classes to achieve the selected risk threshold.

9 . The method according to claim 8 wherein if the re-identification risk metric exceeds the selected risk threshold, the method repeats by performing de-identification of the retrieved dataset with increased suppression or generalization or both to meet the selected risk threshold.

10 . The method according to claim 1 , wherein a source database is equivalent to an identification database.

11 . The method according to claim 1 , wherein the de-identified dataset is a sample of the source database that has been de-identified.

12 . A system for assessing re-identification risk of a dataset containing personal information, the system comprising:

a memory;

a processor coupled to the memory, the processor performing:

retrieving the dataset comprising a plurality of records from the memory;

receiving variables selected from a plurality of variables present in the dataset, wherein the variables may be used as potential identifiers of personal information from the dataset; and

determining equivalence classes for each of the selected variables in the dataset and one or more equivalence class sizes;

determining a re-identification risk metric associated with the dataset using a modified log-linear model by measuring a goodness of fit measure generalized for each of the one or more equivalence class sizes.

13 . A computer readable memory containing instructions for assessing re-identification risk of a dataset containing personal information, the instructions when executed by a processor performing:

retrieving the dataset comprising a plurality of records from the memory;

receiving variables selected from a plurality of variables present in the dataset, wherein the variables may be used as potential identifiers of personal information from the dataset; and

determining equivalence classes for each of the selected variables in the dataset and one or more equivalence class sizes;

determining a re-identification risk metric associated with the dataset using a modified log-linear model by measuring a goodness of fit measure generalized for each of the one or more equivalence class sizes.

14 . The computer readable memory according to claim 13 , wherein determining a re-identification risk metric using a modified log-linear model comprises:

for the one or more equivalence classes:

determining the goodness of fit measure for the size of the equivalence class; and

determining a portion of a re-identification risk associated with the size of the equivalence class; and

determining the re-identification risk by summing all the determined portion of the re-identification risk.

15 . The computer readable memory according to claim 14 wherein determining the portion of the re-identification risk comprises:

calculating

h

k

(

γ

j

)

=

f

j

=

k

(

k

/

F

j

N

)

where h k is the portion of the re-identification risk associated with equivalence class size k, γ j is the actual re-identification risk, F j is the equivalence class sizes in an identification database, N is the set of records in the identification database.

16 . The computer readable memory according to claim 14 , wherein the goodness of fit measures a bias arising from difference between an estimated re-identification risk and an actual re-identification risk.

17 . The computer readable memory according to claim 16 , wherein measuring the bias comprises:

calculating

B

k

=

j

E

(

I

(

f

j

=

k

)

)

[

h

k

(

y

j

)

-

h

k

(

γ

j

)

]

where B k is the goodness of fit measure for equivalence class size k, f j is the equivalence sizes in the de-identified dataset, and {circumflex over (γ)} j is the estimated re-identification risk.

18 . The computer readable memory according to claim 14 , wherein the risk threshold selected is less than

R

J

=

1

/

min

j

(

F

j

)

where R J is journalist risk.

19 . The computer readable memory of claim 14 further comprising:

receiving a re-identification risk threshold value acceptable for the dataset; and

comparing the re-identification risk metric meets the risk threshold value.

20 . The computer readable memory according to claim 19 , wherein if the re-identification metric is greater than the risk threshold the further comprising:

performing de-identification of the retrieved dataset based upon one or more equivalence classes to achieve the selected risk threshold.

21 . The computer readable memory according to claim 20 wherein if the re-identification risk metric exceeds the selected risk threshold, the method repeats by performing de-identification of the retrieved dataset with increased suppression or generalization or both to meet the selected risk threshold.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2016
From: EL EMAM, KHALED; DANKAR, FIDA
To: UNIVERSITY OF OTTAWA
Reel/Frame 038407/0765 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2016
From: UNIVERSITY OF OTTAWA
To: PRIVACY ANALYTICS INC.
Reel/Frame 038054/0578 →