IP Library › Granted Patent US 12,632,579
Granted Patent B2
US 12,632,579 · App. 18/066,767 · Granted May 19, 2026

Dataset privacy management system

Inventors: Baya Dhouib (Valbonne, FR); Bini Samuel Yao (Cagnes-sur-Mer, FR)
Assignee: ACCENTURE GLOBAL SOLUTIONS LIMITED
G06F21/6218
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,579
App. No.
18/066,767
Granted
May 19, 2026
Kind
B2
Abstract

In some implementations, a dataset evaluation system may receive a target dataset. The dataset evaluation system may-processing the target dataset to generate a normalized target dataset. The dataset evaluation system may process the normalized target dataset with an intruder dataset to identify whether any quasi-identifiers are present in the normalized target dataset. The dataset evaluation system may determine a Cartesian product of the normalized target dataset and the intruder dataset. The dataset evaluation system may compute, using a distance linkage disclosure technique, an inference risk score for the target dataset with the intruder dataset based on the Cartesian product and whether any quasi-identifiers are present in the normalized target dataset. The dataset evaluation system may output information associated with the inference risk score.

Claims (91)

1 . A method, comprising:

receiving, by a dataset evaluation system comprising one or more processors, a target dataset;

pre-processing, by the dataset evaluation system, the target dataset to generate a normalized target dataset;

processing, by the dataset evaluation system, the normalized target dataset with an intruder dataset to identify whether any quasi-identifiers are present in the normalized target dataset;

generating, by the dataset evaluation system, an anonymized dataset based on the normalized target dataset, wherein the anonymized dataset includes an initial dataset of network traffic from user equipment's (UE), in proximity to a cell tower, for which user identifiers have been removed and replaced with anonymous unique identifiers, wherein the initial dataset is received by the dataset evaluation system via a wireless network;

joining, by the dataset evaluation system, the intruder dataset and the anonymized dataset to form a joined dataset;

determining, by the dataset evaluation system, a first Cartesian product of the anonymized dataset and the intruder dataset and a second Cartesian product of the joined dataset and the intruder dataset;

computing, by the dataset evaluation system and using a distance linkage disclosure technique based on whether any quasi-identifiers are present in the normalized target dataset, a first inference risk score for the target dataset with the intruder dataset with respect to the first Cartesian product and a second inference risk score for the target dataset and the intruder dataset with respect to the second Cartesian product;

outputting, by the dataset evaluation system, information associated with the first inference risk score and the second inference risk score,

wherein the one or more processors are configured to further cause the dataset evaluation system to:

apply group-level re-identification to the anonymized dataset and the intruder dataset to generate a set of equivalent clusters, wherein applying the group-level re-identification to the anonymized dataset and the intruder dataset comprises:

extracting a first feature set and a second feature set associated with the target dataset, by performing natural language processing on the received target dataset using a machine learning model, wherein the machine learning model is trained using datasets having different levels of de-identification techniques applied and wherein the machine learning model is trained to predict whether personal identifiable data is determined from the target dataset;

identifying a first set of clusters corresponding to the anonymized dataset with respect to the first feature set and the second feature set;

identifying a second set of clusters corresponding to the intruder dataset with respect to the first feature set and the second feature set;

determining intersections within the first set of clusters and the second set of clusters; and

applying the intersections of the first set of clusters from the anonymized dataset to the intersections of the second set of clusters from the intruder dataset to generate the set of equivalent clusters, and

wherein the one or more processors, when configured to cause the dataset evaluation system to output information, are configured to cause the dataset evaluation system to output the information regarding the generated set of equivalent clusters; and

performing automated actions based on the prediction by the trained machine learning model.

2 . The method of claim 1 , wherein pre-processing the target dataset comprises:

applying min-max scaling of numerical features of the target dataset to scale variables of the target dataset to generate the normalized target dataset.

3 . The method of claim 1 , wherein processing the normalized target dataset with the intruder dataset to identify whether any quasi-identifiers are present in the normalized target dataset comprises:

identifying whether one or more singleton records are present in the normalized target dataset.

4 . The method of claim 1 , wherein processing the normalized target dataset with the intruder dataset to identify whether any quasi-identifiers are present in the normalized target dataset comprises:

identifying whether a functional dependency is present with respect to a subset of records in the normalized target dataset.

5 . The method of claim 1 , wherein computing the first inference risk score and the second inference risk score comprises: computing a set of distances between sets of records using a Euclidean distance metric.

6 . The method of claim 1 , wherein the first inference risk score and the second inference risk score represent a disclosure risk of the target dataset relative to one or more external datasets.

7 . The method of claim 1 , wherein outputting information associated with the first inference risk score and the second inference risk score comprises:

outputting information identifying one or more compromised records in the target dataset.

8 . A dataset evaluation system, comprising:

one or more memories; and

one or more processors, coupled to the one or more memories, configured to:

receive a target dataset;

pre-process the target dataset to generate a normalized target dataset;

process the normalized target dataset with an intruder dataset to identify whether any quasi-identifiers are present in the normalized target dataset;

generate an anonymized dataset based on the normalized target dataset, wherein the anonymized dataset includes an initial dataset of network traffic from user equipment's (UE), in proximity to a cell tower, for which user identifiers have been removed and replaced with anonymous unique identifiers, wherein the initial dataset is received by the dataset evaluation system via a wireless network;

join the intruder dataset and the anonymized dataset to form a joined dataset;

determine a first Cartesian product of the anonymized dataset and the intruder dataset and a second Cartesian product of the joined dataset and the intruder dataset;

compute, using a distance linkage disclosure technique and based on whether any quasi-identifiers are present in the normalized target dataset, a first inference risk score for the target dataset with the intruder dataset with respect to the first Cartesian product and a second inference risk score for the target dataset and the intruder dataset with respect to the second Cartesian product;

compute a distance linkage variation value based on the first inference risk score and the second inference risk score;

output information associated with the first inference risk score and the second inference risk score,

wherein the one or more processors are further configured to cause the dataset evaluation system to:

apply group-level re-identification to the anonymized dataset and the intruder dataset to generate a set of equivalent clusters, wherein for applying the group-level re-identification to the anonymized dataset and the intruder dataset, the one or more processors are configured to:

extract a first feature set and a second feature set associated with the target dataset, by performing natural language processing on the received target dataset using a machine learning model, wherein the machine learning model is trained using datasets having different levels of de-identification techniques applied and wherein the machine learning model is trained to predict whether personal identifiable data is determined from the target dataset;

identify a first set of clusters corresponding to the anonymized dataset with respect to the first feature set and the second feature set;

identify a second set of clusters corresponding to the intruder dataset with respect to the first feature set and the second feature set;

determine intersections within the first set of clusters and the second set of clusters; and

apply the intersections of the first set of clusters from the anonymized dataset to the intersections of the second set of clusters from the intruder dataset to generate the set of equivalent clusters, and

wherein the one or more processors, when configured to cause the dataset evaluation system to output information, are configured to cause the dataset evaluation system to output the information regarding the generated set of equivalent clusters; and

performing automated actions based on the prediction by the trained machine learning model.

9 . The dataset evaluation system of claim 8 , wherein the one or more processors, when configured to cause the dataset evaluation system to output information associated with the first inference risk score and the second inference risk score, are configured to cause the dataset evaluation system to:

output information associated with anonymizing one or more records determined to be compromised based on the first inference risk score and the second inference risk score.

10 . The dataset evaluation system of claim 8 , wherein the target dataset is a de-identified version of a core dataset, and the first inference risk score or the second inference risk score represents a disclosure risk of the core dataset relative to the target dataset.

11 . The dataset evaluation system of claim 8 , wherein the first inference risk score or the second inference risk score is based on the set of clusters.

12 . A non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising:

one or more instructions that, when executed by one or more processors of a dataset evaluation system, cause the dataset evaluation system to:

receive an initial dataset via a wireless network;

anonymize the initial dataset to generate a target dataset;

pre-process the target dataset to generate a normalized target dataset;

process the normalized target dataset with an intruder dataset to identify whether any quasi-identifiers are present in the normalized target dataset;

generate an anonymized dataset based on the normalized target dataset, wherein the anonymized dataset includes the initial dataset of network traffic from user equipment's (UE), in proximity to a cell tower, for which user identifiers have been removed and replaced with anonymous unique identifiers;

join the intruder dataset and the anonymized dataset to form a joined dataset;

determine a first Cartesian product of the anonymized dataset and the intruder dataset and a second Cartesian product of the joined dataset and the intruder dataset;

compute, using a distance linkage disclosure technique and based on whether any quasi-identifiers are present in the normalized target dataset, a first inference risk score for the target dataset with the intruder dataset with respect to the first Cartesian product and a second inference risk score for the target dataset and the intruder dataset with respect to the second Cartesian product;

compute a distance linkage variation value based on the first inference risk score and the second inference risk score;

identify one or more compromised records based on the first inference risk score and the second inference risk score;

generate fuzzing data for the one or more compromised records;

update the initial dataset with the fuzzing data; and

output information associated with the initial dataset via the wireless network, based on updating the initial dataset, and the first inference risk score and the second inference risk score;

wherein the one or more instructions that, when executed by the one or more processors of a dataset evaluation system, further cause the dataset evaluation system to:

apply group-level re-identification to the anonymized dataset and the intruder dataset to generate a set of equivalent clusters, wherein for applying the group-level re-identification to the anonymized dataset and the intruder dataset, the one or more processors are configured to:

extract a first feature set and a second feature set associated with the target dataset, by performing natural language processing on the received target dataset using a machine learning model, wherein the machine learning model is trained using datasets having different levels of de-identification techniques applied and wherein the machine learning model is trained to predict whether personal identifiable data is determined from the target dataset;

identify a first set of clusters corresponding to the anonymized dataset with respect to the first feature set and the second feature set;

identify a second set of clusters corresponding to the intruder dataset with respect to the first feature set and the second feature set;

determine intersections within the first set of clusters and the second set of clusters; and

apply the intersections of the first set of clusters from the anonymized dataset to the intersections of the second set of clusters from the intruder dataset to generate the set of equivalent clusters, and

wherein the one or more processors, when configured to cause the dataset evaluation system to output information, are configured to cause the dataset evaluation system to output the information regarding the generated set of equivalent clusters; and

performing automated actions based on the prediction by the trained machine learning model.

13 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions, that cause the dataset evaluation system to generate the fuzzing data, cause the dataset evaluation system to:

generate one or more alternate records based on the one or more compromised records;

determine whether the one or more alternate records are compromised based on an updated inference risk score; and

replace the compromised one or more records with the one or more alternate records in the initial dataset based on determining that the one or more alternate records are not compromised.

14 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions, that cause the dataset evaluation system to process the normalized target dataset with the intruder dataset, cause the dataset evaluation system to:

identify whether one or more singleton records are present in the normalized target dataset.

15 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions, that cause the dataset evaluation system to output information associated with the initial dataset based on updating the initial dataset, cause the dataset evaluation system to:

output a subset of the initial dataset.

16 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions, that cause the dataset evaluation system to output information associated with the initial dataset based on updating the initial dataset, cause the dataset evaluation system to:

output a certification of the initial dataset.

17 . The non-transitory computer-readable medium of claim 12 , wherein the first inference risk score and the second inference risk score represent a disclosure risk of the target dataset relative to one or more external datasets.

18 . The non-transitory computer-readable medium of claim 12 , wherein the first inference risk score and the second inference risk score represent a disclosure risk of the target dataset relative to the initial dataset.

19 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions, that cause the dataset evaluation system to output information associated with the initial dataset, cause the dataset evaluation system to:

output information identifying the one or more compromised records in the target dataset.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2022
From: DHOUIB, BAYA; YAO, BINI SAMUEL
To: ACCENTURE GLOBAL SOLUTIONS LIMITED
Reel/Frame 062125/0330 →
Continuity (2)
Provisional Application 63376568 · Sep 21, 2022
Related Publication 20240095385A1 · Mar 21, 2024
References Cited (24)
US 9537880B1 · Jones · 2017 [cited by examiner]
US 9787638B1 · Adams · 2017 [cited by examiner]
US 12079814B2 · Harris · 2024 [cited by examiner]
US 12549603B2 · Ghosh · 2026 [cited by examiner]
US 20110258206A1 · El Emam · 2011 [cited by examiner]
US 20170286671A1 · Chari · 2017 [cited by examiner]
US 20180026944A1 · Phillips · 2018 [cited by examiner]
US 20190156028A1 · Hay · 2019 [cited by examiner]
US 20190220607A1 · Dodor · 2019 [cited by examiner]
US 20200057953A1 · Livny · 2020 [cited by examiner]
US 20200082290A1 · Pascale · 2020 [cited by examiner]
US 20200117833A1 · Pletea · 2020 [cited by examiner]
US 20200327252A1 · Mcfall · 2020 [cited by examiner]
US 20210279367A1 · Khan · 2021 [cited by examiner]
US 20220050917A1 · Jiang · 2022 [cited by examiner]
US 20220164471A1 · Braghin · 2022 [cited by examiner]
US 20220222374A1 · Antoniou · 2022 [cited by examiner]
US 20220300651A1 · Mondal · 2022 [cited by examiner]
US 20230052848A1 · Zhang · 2023 [cited by examiner]
US 20230237409A1 · Mallikarjun · 2023 [cited by examiner]
Y. Hua, Z. Li, B. Wang and J. Li, “A Method for Solving Quasi-Identifiers of Single Structured Relational Data,” in IEEE Access, vol. 9, pp. 166293-166302, 2021. (Year: 2021). [cited by examiner]
Di Vimercati, Sabrina De Capitani, et al. “Scalable distributed data anonymization for large datasets.” IEEE Transactions on Big Data 9.3 (2022): pp. 818-831. (Year: 2022). [cited by examiner]
Ghinita, Gabriel, et al. “A framework for efficient data anonymization under privacy and accuracy constraints.” ACM Transactions on Database Systems (TODS) 34.2 (2009): pp. 1-47. (Year: 2009). [cited by examiner]
Soria-Comas, J. et al., “Assessing Disclosure Risk via Record Linkage by a Maximum-Knowledge Intruder,” UNESCO Chair in Data Privacy, Dept. of Computer Engineering and Maths, Universitat Rovira i Virgili, Av. Pa [cited by applicant]