IP Library › Granted Patent US 12,406,088
Granted Patent B2
US 12,406,088 · App. 18/030,545 · Granted Sep 2, 2025

Method for evaluating the risk of re-identification of anonymised data

Inventors: Morgan Guillaudeux (Nantes, FR); Olivier Breillacq (Nantes, FR)
Assignee: BIG DATA SANTE
G06F21/6245G06F21/54G06F21/575G06F21/6254
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,088
App. No.
18/030,545
Granted
Sep 2, 2025
Kind
B2
Abstract

The method delivers a degree of protection (txP 3 ) representative of the risk of re-identification of data in the case of a correspondence search attack including a deterministic search based on an external information source and a correspondence search based on a distance. The method comprises steps of E) consolidating a set of original individuals (EDO) and a set of anonymous individuals (IA); F) identifying, in the set of original individuals, individuals at risk (IOrs) via the deterministic correspondence search; G) evaluating a degree of failure of re-identification (txP 1 ) for the sets of original individuals and of anonymous individuals, on the basis of the correspondence search based on distance; H) computing the degree of protection as a function of a total number of individuals in the original dataset, of a number (RS) of individuals at risk identified in step B) and of the degree of failure of re-identification (txP 1 ).

Claims (8)

1. A computer-implemented data processing method for evaluating a risk of re-identification of anonymized data, said method delivering a protection rate parameter (txP 3 ) representative of said risk of re-identification in case of a correspondence search attack including a deterministic search based on at least one external information source and a correspondence search based on a distance, said method comprising the steps of E) consolidating an original dataset (EDO) comprising a plurality of original individuals (IO) and an anonymized dataset (EDA) comprising a plurality of anonymous individuals (IA), said anonymous individuals (IA) being produced by a process of anonymizing said original individuals (IO); F) identifying, in said original dataset (EDO), original individuals at risk (IO rs ) as being original individuals (IO) having at least one noteworthy, or unique, value in at least one considered variable, or at least one combination of noteworthy, or unique, values in a set of considered variables, in a deterministic correspondence search and to which only one respective close anonymous individual (IA prs ) can be associated by said deterministic correspondence search; G) evaluating a re-identification failure rate parameter (txP 1 ) for said original datasets (EDO) and anonymized datasets (EDA), from said correspondence search based on a distance between each of said original individuals (IO) and one or more of the nearest of said anonymous individuals (IA) identified by a method called “k-NN” method; H) computing said protection rate parameter (txP 3 ) as a function of a total number (M) of original individuals (IO) in said original dataset (EDO), of a number (RS) of original individuals at risk (IO rs ) identified in step F) and of said re-identification failure rate parameter (txP 1 ) obtained in step G).

2. The method according to claim 1 , characterized in that, in step F), an anonymous individual (IA) is considered to be one of said nearest anonymous individuals (IA p , IA prs ) of one of said considered individuals at risk (IO rs ) when 1) said anonymous individual (IA) has a variable with the same modality as a considered variable of said original individual at risk (IO rs ) in said correspondence search in the case wherein said variable is a qualitative variable, or when 2) said anonymous individual has a value for said considered variable that is equal to a tolerance range close to the value of said same considered variable of said original individual at risk (IO rs ) in the case wherein said considered variable in said deterministic correspondence search is a continuous variable.

3. The method according to claim 1 , characterized in that step G) comprises the sub-steps of a) linking said original dataset (EDO) to said anonymized dataset (EDA); b) converting (PCA, MCA, FAMD) said original individuals (IO) and said anonymous individuals (IA) in a Euclidean space (A 1 , A 2 ), with said original individuals (IO) and anonymous individuals (IA) being represented by coordinates in said Euclidean space (A 1 , A 2 ); c) identifying, for each of said original individuals (IO), one or more of said nearest anonymous individuals (IA) based on said distance, using the “k-NN” method; and d) computing said re-identification failure rate parameter (txP 1 ) as being a percentage of cases where one of said nearest anonymous individuals (IA k ) identified in sub-step c) for one of said original individuals (IO i ) is not a valid anonymous individual (IA i ) corresponding to said original individual (IO i ).

4. The method according to claim 3 , characterized in that said distance is a Euclidean distance.

5. The method according to claim 3 , characterized in that the transformation of sub-step b) is carried out by a factor method (PCA, MCA, FAMD) and/or using an artificial neural network, called “autoencoder”.

6. The method according to claim 5 , characterized in that said factor method is a “Principal Component Analysis” (PCA) method when said individuals (IO, IA) comprise continuous type variables, a “Multiple Correspondence Analysis” (MCA) method when said individuals (IO, IA) comprise qualitative type variables, or a “Factor Analysis of Mixed Data” (FAMD) method when said individuals (IO, IA) comprise “continuous/qualitative” type variables.

7. A data anonymization computer system (SAD) including a data storage device (SD) storing program instructions (MET) for implementing the method according to claim 1 .

8. A computer program product including a medium in which program instructions (MET) are recorded that are readable by a processor for implementing the method according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2023
From: GUILLAUDEUX, MORGAN; BREILLACQ, OLIVIER
To: BIG DATA SANTE
Reel/Frame 063240/0681 →
Priority Claims (1)
FR 2010259 · Oct 7, 2020 · national
Continuity (1)
Related Publication 20240005035A1 · Jan 4, 2024
References Cited (13)
US 20190026490A1 · Ahmed · 2019 [cited by examiner]
EP 3567508A1 · 2019 [cited by applicant]
FR 3048101A1 · 2017 [cited by applicant]
Calvino Aida et al.: “Factor Analysis for Anonymization”, 2017 IEEE International Conference on Data Mining Workshops (ICDMW), IEEE, Nov. 18, 2017, pp. 984-991. [cited by applicant]
Josep Domingo-Ferrer et al.: “Disclosire risk assessment in statistical microdata protection via advanced record linkage”, Statistics and Computing, Kluwer Academic Publishers, BO, vol. 13, No. 4, Oct. 1, 2003, pp. 344-… [cited by applicant]
Pagliuca Daniela et al.: “Some Results of individual Ranking Method on the System of Enterprise Accounts Annual Survey”, In: “Statistical Disclosure Control”, Jul. 30, 1998, pp. 11-12, 15-16. [cited by applicant]
Robinson-Cox J.F., « A record-linkage approach to imputation of missing data : Analyzing tag retention in a tag-recapture experiment », Journal of Agricultural, Biological, and Environmental Statistics 3(1), 1998, pp. 4… [cited by applicant]
Winkler W.E., « Matching and record linkage », Cox B.G. (Ed.), Business Survey Methods, Wiley, New York, 1995, pp. 355-384. [cited by applicant]
Fellegi I.P. et al., Jaro M.A., et Winkler W.E., « A theory of record linkage », Journal of the American Statistical Association 64, 1969, pp. 1183-1210. [cited by applicant]
Jaro M. A., « Advances in record-linkage methodology as applied to matching the 1985 Census of Tampa, Florida », Journal of the American Statistical Association 84, 1989, pp. 414-420. [cited by applicant]
Domingo-Ferrer J. et al., « Disclosure risk assessment via record linkage by a maximum-knowledge attacker », 13th Annual Conference on Privacy, Security and Trust (PST), 2015. [cited by applicant]
Kounine A. et al., « Assessing Disclosure Risk in Anonymized Datasets », FloCon 2008 Conference. [cited by applicant]
International Search Report for corresponding International Application No. PCT/FR2021/000114, dated Feb. 23, 2022. [cited by applicant]