IP Library Granted Patent US 11,615,209
Granted Patent B2
US 11,615,209 · App. 16/323,865 · Granted Mar 28, 2023

Big data k-anonymizing by parallel semantic micro-aggregation

Inventors: Andreas Hapfelmeier (Sauerlach, DE); Mike Imig (Cologne, DE); Michael Mock (Bonn, DE)
G06F21/6254G06F16/00G06F16/901G06F16/906
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,615,209
App. No.
16/323,865
Granted
Mar 28, 2023
Kind
B2
Abstract

Provided is a method for anonymizing datasets having sensitive information, including the steps of determining a dataset of records to be assigned to aggregation clusters; computing an average record of the dataset on the basis of a predefined repetition counter; finding a most distant first record to the average record using a distance measure; finding a most distant second record from the first record using the distance measure; forming a first aggregation cluster around the first record and a second aggregation cluster around the second record; and generating a new dataset by subtracting the first cluster and the second cluster from the previous dataset.

Claims (34)

1. A method for k-anonymizing datasets having sensitive information with improved processing speed, wherein k is a predetermined integer, the method comprising the steps of:

determining, by at least one hardware processor, a dataset of records from the datasets to be assigned to aggregation clusters; and iterating the steps of:

1) when an iteration number is a multiple of a predetermined integer parameter:

computing an average record of the dataset;

finding a most distant first record to the average record using a distance measure; and

finding a most distant second record from the first record using the distance measure; and

2) forming, by the at least one hardware processor, a first aggregation cluster around the first record and a second aggregation cluster around the second record, wherein the first aggregation cluster comprises k records closest to the first record and the second aggregation cluster comprises k records closest to the second record; and

generating, by the at least one hardware processor, a new dataset by subtracting the first cluster and the second cluster from the dataset;

wherein the predetermined integer parameter is a predefined repetition counter greater than 1 and the repetition counter is adapted to define after how many iterations a recalculation of the average record, the most distant first record, and the most distant second record is performed; and wherein the k-anonymizing is performed at least partially in parallel or distributed across a number of nodes using the predetermined repetition counter.

2. The method according to claim 1 , wherein the step of computing the average record is performed in parallel or distributed on the number of nodes.

3. The method according to claim 1 , wherein the step of finding the most distant first record is performed in parallel or distributed on the number of nodes.

4. The method according to claim 1 , wherein the step of finding the most distant second record is performed in parallel or distributed on the number of nodes.

5. The method according to claim 1 , wherein the average record is computed attribute-wise based on an aggregation function.

6. The method according to claim 1 , wherein a hierarchical generalization having a number of generalization stages is used for computing the average record of a nominal attribute of the dataset or a distance between nominal values.

7. A computer program product, comprising a computer readable hardware storage device having computer readable program code stored therein, said program code executable by a processor of a computer system to implement the method according to claim 1 .

8. A system for k-anonymizing datasets having sensitive information with improved processing speed, wherein k is a predetermined integer, the system comprising:

at least one hardware processor configured to perform the steps of:

determining a dataset of records from the datasets to be assigned to aggregation clusters; and

iterating the steps of:

1) When an iteration number is a multiple of a predetermined integer parameter:

computing an average record of the dataset;

finding a most distant first record to the average record using a distance measure; and

finding a most distant second record from the first record using the distance measure; and the steps of:

2) forming a first aggregation cluster around the first record and a second aggregation cluster around the second record, wherein the first aggregation cluster comprises k records closest to the first record and the second aggregation cluster comprises k records closest to the second record; and

generating a new dataset by subtracting the first cluster and the second cluster from the dataset;

wherein the predetermined integer parameter is a predefined repetition counter greater than 1 and the repetition counter is adapted to define after how many iterations a recalculation of the average record the most distant first record and most distant second record is performed; and wherein the k-anonymizing is performed at least partially in parallel or distributed across a number of nodes using the predetermined repetition counter.

9. A computer cluster having several nodes for k-anonymizing datasets having sensitive information with improved processing speed, wherein k is a predetermined integer, each node comprising:

at least one hardware processor configured for performing a method of k-anonymizing, by parallel or distributed processing, wherein the method of k-anonymizing comprises:

when an iteration number is a multiple of a predetermined integer parameter:

computing an average record of a dataset;

finding a most distant first record to the average record using a distance measure; and

finding a most distant second record from the first record using the distance measure; and

2) forming a first aggregation cluster around the first record and a second aggregation cluster around the second record, wherein the first aggregation cluster comprises k records closest to the first record and the second aggregation cluster comprises k records closest to the second record; and generating a new dataset by subtracting the first cluster and the second cluster from the dataset;

wherein the predetermined integer parameter is a predefined repetition counter greater than 1 and the repetition counter is adapted to define after how many iterations a recalculation of the average record the most distant first record and most distant second record is performed; and wherein the k-anonymizing is performed at least partially in parallel or distributed across a number of nodes using the predetermined repetition counter.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2019
From: HAPFELMEIER, ANDREAS
To: SIEMENS AKTIENGESELLSCHAFT
Reel/Frame 048844/0947 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2019
From: IMIG, MIKE; MOCK, MICHAEL
To: FRAUNHOFER GESELLSCHAFT ZUR FÖRDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 048845/0103 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2019
From: FRAUNHOFER GESELLSCHAFT ZUR FÖRDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
To: SIEMENS AKTIENGESELLSCHAFT
Reel/Frame 048845/0122 →
Continuity (1)
Related Publication 20190213357A1 · Jul 11, 2019