Dataset privacy management system
In some implementations, a dataset evaluation system may receive a target dataset. The dataset evaluation system may-processing the target dataset to generate a normalized target dataset. The dataset evaluation system may process the normalized target dataset with an intruder dataset to identify whether any quasi-identifiers are present in the normalized target dataset. The dataset evaluation system may determine a Cartesian product of the normalized target dataset and the intruder dataset. The dataset evaluation system may compute, using a distance linkage disclosure technique, an inference risk score for the target dataset with the intruder dataset based on the Cartesian product and whether any quasi-identifiers are present in the normalized target dataset. The dataset evaluation system may output information associated with the inference risk score.
1 . A method, comprising:
receiving, by a dataset evaluation system comprising one or more processors, a target dataset;
pre-processing, by the dataset evaluation system, the target dataset to generate a normalized target dataset;
processing, by the dataset evaluation system, the normalized target dataset with an intruder dataset to identify whether any quasi-identifiers are present in the normalized target dataset;
generating, by the dataset evaluation system, an anonymized dataset based on the normalized target dataset, wherein the anonymized dataset includes an initial dataset of network traffic from user equipment's (UE), in proximity to a cell tower, for which user identifiers have been removed and replaced with anonymous unique identifiers, wherein the initial dataset is received by the dataset evaluation system via a wireless network;
joining, by the dataset evaluation system, the intruder dataset and the anonymized dataset to form a joined dataset;
determining, by the dataset evaluation system, a first Cartesian product of the anonymized dataset and the intruder dataset and a second Cartesian product of the joined dataset and the intruder dataset;
computing, by the dataset evaluation system and using a distance linkage disclosure technique based on whether any quasi-identifiers are present in the normalized target dataset, a first inference risk score for the target dataset with the intruder dataset with respect to the first Cartesian product and a second inference risk score for the target dataset and the intruder dataset with respect to the second Cartesian product;
outputting, by the dataset evaluation system, information associated with the first inference risk score and the second inference risk score,
wherein the one or more processors are configured to further cause the dataset evaluation system to:
apply group-level re-identification to the anonymized dataset and the intruder dataset to generate a set of equivalent clusters, wherein applying the group-level re-identification to the anonymized dataset and the intruder dataset comprises:
extracting a first feature set and a second feature set associated with the target dataset, by performing natural language processing on the received target dataset using a machine learning model, wherein the machine learning model is trained using datasets having different levels of de-identification techniques applied and wherein the machine learning model is trained to predict whether personal identifiable data is determined from the target dataset;
identifying a first set of clusters corresponding to the anonymized dataset with respect to the first feature set and the second feature set;
identifying a second set of clusters corresponding to the intruder dataset with respect to the first feature set and the second feature set;
determining intersections within the first set of clusters and the second set of clusters; and
applying the intersections of the first set of clusters from the anonymized dataset to the intersections of the second set of clusters from the intruder dataset to generate the set of equivalent clusters, and
wherein the one or more processors, when configured to cause the dataset evaluation system to output information, are configured to cause the dataset evaluation system to output the information regarding the generated set of equivalent clusters; and
performing automated actions based on the prediction by the trained machine learning model.
2 . The method of claim 1 , wherein pre-processing the target dataset comprises:
applying min-max scaling of numerical features of the target dataset to scale variables of the target dataset to generate the normalized target dataset.
3 . The method of claim 1 , wherein processing the normalized target dataset with the intruder dataset to identify whether any quasi-identifiers are present in the normalized target dataset comprises:
identifying whether one or more singleton records are present in the normalized target dataset.
4 . The method of claim 1 , wherein processing the normalized target dataset with the intruder dataset to identify whether any quasi-identifiers are present in the normalized target dataset comprises:
identifying whether a functional dependency is present with respect to a subset of records in the normalized target dataset.
5 . The method of claim 1 , wherein computing the first inference risk score and the second inference risk score comprises: computing a set of distances between sets of records using a Euclidean distance metric.
6 . The method of claim 1 , wherein the first inference risk score and the second inference risk score represent a disclosure risk of the target dataset relative to one or more external datasets.
7 . The method of claim 1 , wherein outputting information associated with the first inference risk score and the second inference risk score comprises:
outputting information identifying one or more compromised records in the target dataset.
8 . A dataset evaluation system, comprising:
one or more memories; and
one or more processors, coupled to the one or more memories, configured to:
receive a target dataset;
pre-process the target dataset to generate a normalized target dataset;
process the normalized target dataset with an intruder dataset to identify whether any quasi-identifiers are present in the normalized target dataset;
generate an anonymized dataset based on the normalized target dataset, wherein the anonymized dataset includes an initial dataset of network traffic from user equipment's (UE), in proximity to a cell tower, for which user identifiers have been removed and replaced with anonymous unique identifiers, wherein the initial dataset is received by the dataset evaluation system via a wireless network;
join the intruder dataset and the anonymized dataset to form a joined dataset;
determine a first Cartesian product of the anonymized dataset and the intruder dataset and a second Cartesian product of the joined dataset and the intruder dataset;
compute, using a distance linkage disclosure technique and based on whether any quasi-identifiers are present in the normalized target dataset, a first inference risk score for the target dataset with the intruder dataset with respect to the first Cartesian product and a second inference risk score for the target dataset and the intruder dataset with respect to the second Cartesian product;
compute a distance linkage variation value based on the first inference risk score and the second inference risk score;
output information associated with the first inference risk score and the second inference risk score,
wherein the one or more processors are further configured to cause the dataset evaluation system to:
apply group-level re-identification to the anonymized dataset and the intruder dataset to generate a set of equivalent clusters, wherein for applying the group-level re-identification to the anonymized dataset and the intruder dataset, the one or more processors are configured to:
extract a first feature set and a second feature set associated with the target dataset, by performing natural language processing on the received target dataset using a machine learning model, wherein the machine learning model is trained using datasets having different levels of de-identification techniques applied and wherein the machine learning model is trained to predict whether personal identifiable data is determined from the target dataset;
identify a first set of clusters corresponding to the anonymized dataset with respect to the first feature set and the second feature set;
identify a second set of clusters corresponding to the intruder dataset with respect to the first feature set and the second feature set;
determine intersections within the first set of clusters and the second set of clusters; and
apply the intersections of the first set of clusters from the anonymized dataset to the intersections of the second set of clusters from the intruder dataset to generate the set of equivalent clusters, and
wherein the one or more processors, when configured to cause the dataset evaluation system to output information, are configured to cause the dataset evaluation system to output the information regarding the generated set of equivalent clusters; and
performing automated actions based on the prediction by the trained machine learning model.
9 . The dataset evaluation system of claim 8 , wherein the one or more processors, when configured to cause the dataset evaluation system to output information associated with the first inference risk score and the second inference risk score, are configured to cause the dataset evaluation system to:
output information associated with anonymizing one or more records determined to be compromised based on the first inference risk score and the second inference risk score.
10 . The dataset evaluation system of claim 8 , wherein the target dataset is a de-identified version of a core dataset, and the first inference risk score or the second inference risk score represents a disclosure risk of the core dataset relative to the target dataset.
11 . The dataset evaluation system of claim 8 , wherein the first inference risk score or the second inference risk score is based on the set of clusters.
12 . A non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising:
one or more instructions that, when executed by one or more processors of a dataset evaluation system, cause the dataset evaluation system to:
receive an initial dataset via a wireless network;
anonymize the initial dataset to generate a target dataset;
pre-process the target dataset to generate a normalized target dataset;
process the normalized target dataset with an intruder dataset to identify whether any quasi-identifiers are present in the normalized target dataset;
generate an anonymized dataset based on the normalized target dataset, wherein the anonymized dataset includes the initial dataset of network traffic from user equipment's (UE), in proximity to a cell tower, for which user identifiers have been removed and replaced with anonymous unique identifiers;
join the intruder dataset and the anonymized dataset to form a joined dataset;
determine a first Cartesian product of the anonymized dataset and the intruder dataset and a second Cartesian product of the joined dataset and the intruder dataset;
compute, using a distance linkage disclosure technique and based on whether any quasi-identifiers are present in the normalized target dataset, a first inference risk score for the target dataset with the intruder dataset with respect to the first Cartesian product and a second inference risk score for the target dataset and the intruder dataset with respect to the second Cartesian product;
compute a distance linkage variation value based on the first inference risk score and the second inference risk score;
identify one or more compromised records based on the first inference risk score and the second inference risk score;
generate fuzzing data for the one or more compromised records;
update the initial dataset with the fuzzing data; and
output information associated with the initial dataset via the wireless network, based on updating the initial dataset, and the first inference risk score and the second inference risk score;
wherein the one or more instructions that, when executed by the one or more processors of a dataset evaluation system, further cause the dataset evaluation system to:
apply group-level re-identification to the anonymized dataset and the intruder dataset to generate a set of equivalent clusters, wherein for applying the group-level re-identification to the anonymized dataset and the intruder dataset, the one or more processors are configured to:
extract a first feature set and a second feature set associated with the target dataset, by performing natural language processing on the received target dataset using a machine learning model, wherein the machine learning model is trained using datasets having different levels of de-identification techniques applied and wherein the machine learning model is trained to predict whether personal identifiable data is determined from the target dataset;
identify a first set of clusters corresponding to the anonymized dataset with respect to the first feature set and the second feature set;
identify a second set of clusters corresponding to the intruder dataset with respect to the first feature set and the second feature set;
determine intersections within the first set of clusters and the second set of clusters; and
apply the intersections of the first set of clusters from the anonymized dataset to the intersections of the second set of clusters from the intruder dataset to generate the set of equivalent clusters, and
wherein the one or more processors, when configured to cause the dataset evaluation system to output information, are configured to cause the dataset evaluation system to output the information regarding the generated set of equivalent clusters; and
performing automated actions based on the prediction by the trained machine learning model.
13 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions, that cause the dataset evaluation system to generate the fuzzing data, cause the dataset evaluation system to:
generate one or more alternate records based on the one or more compromised records;
determine whether the one or more alternate records are compromised based on an updated inference risk score; and
replace the compromised one or more records with the one or more alternate records in the initial dataset based on determining that the one or more alternate records are not compromised.
14 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions, that cause the dataset evaluation system to process the normalized target dataset with the intruder dataset, cause the dataset evaluation system to:
identify whether one or more singleton records are present in the normalized target dataset.
15 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions, that cause the dataset evaluation system to output information associated with the initial dataset based on updating the initial dataset, cause the dataset evaluation system to:
output a subset of the initial dataset.
16 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions, that cause the dataset evaluation system to output information associated with the initial dataset based on updating the initial dataset, cause the dataset evaluation system to:
output a certification of the initial dataset.
17 . The non-transitory computer-readable medium of claim 12 , wherein the first inference risk score and the second inference risk score represent a disclosure risk of the target dataset relative to one or more external datasets.
18 . The non-transitory computer-readable medium of claim 12 , wherein the first inference risk score and the second inference risk score represent a disclosure risk of the target dataset relative to the initial dataset.
19 . The non-transitory computer-readable medium of claim 12 , wherein the one or more instructions, that cause the dataset evaluation system to output information associated with the initial dataset, cause the dataset evaluation system to:
output information identifying the one or more compromised records in the target dataset.