IP Library › Granted Patent US 10,528,761
Granted Patent B2
US 10,528,761 · App. 15/794,807 · Granted Jan 7, 2020

Data anonymization in an in-memory database

Inventor: Xinrong Huang (Shanghai, CN)
Assignee: SAP SE
G06F21/6254G06F16/285
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,528,761
App. No.
15/794,807
Granted
Jan 7, 2020
Kind
B2
Abstract

Disclosed herein are system, method, and computer program product embodiments for data anonymization in an in-memory database. An embodiment operates by identifying a textual attribute corresponding to data of an input table. A value corresponding to the textual attribute is determined for each of the plurality of records. A plurality of groups is generated based on the determined values. At least portion of the data as sorted into the plurality of groups is provided.

Claims (64)

1. A computer implemented method, comprising:

identifying a plurality of textual attributes, including both a first textual attribute and a second textual attribute, corresponding to personal identifying information stored across a plurality of records of an input table for anonymization based on at least a threshold for a number of values required in each of one or more groupings of the data;

determining a plurality of values for both the first and second textual attributes, wherein each value comprises one or more characters and corresponds to at least one of the plurality of records, and wherein a subset of the plurality of values comprises a plurality of unique values;

determining a width of both the first textual attribute and the second textual attribute based on the plurality of unique values, wherein the width corresponds to a range of the unique values associated with a respective attribute;

selecting the first textual attribute based on its width being greater than a width of the second textual attribute, wherein the greater width corresponds to a likelihood of reduced data loss through the anonymization;

generating a plurality of groups based on the determined plurality of values, wherein each group includes one or more of the determined plurality of values that share one or more common characters; and

providing at least portion of the personal identifying information as sorted into the plurality of groups, wherein a count of the values of each group satisfies the threshold.

2. The method of claim 1 , wherein the providing comprises:

determining that the count of values for a particular one of the plurality of groups is less than the threshold; and

suppressing the particular one of the plurality of groups that is less than the threshold, wherein the providing comprises providing the data sorted into the plurality of groups except for the particular group.

3. The method of claim 1 , wherein the personal information of the data includes:

an explicit identifier attribute from which a particular record of the data is distinguishable from one or more remaining records of the data, wherein based on the explicit identifier, an individual corresponding to the record is identifiable;

a first quasi-identifier attribute which when considered together with a second quasi-identifier attribute identify the individual corresponding to the record; and

a sensitive data attribute which includes personal information corresponding to the individual.

4. The method of claim 3 , wherein the textual attribute corresponds to the first quasi-identifier attribute.

5. The method of claim 1 , wherein the identifying comprises identifying a numerical attribute and a hierarchical attribute in addition to the plurality of textual attributes.

6. The method of claim 1 , wherein the selecting further comprises:

determining a weight corresponding to the first textual attribute;

determining a weight corresponding to the second textual attribute;

determining a weighted width for both the first textual attribute and the second textual attribute; and

selecting the first textual attribute based on its weighted width being greater than a width of the second textual attribute.

7. A system, comprising:

a memory; and

at least one processor coupled to the memory and configured to:

identify a plurality of textual attributes, including both a first textual attribute and a second textual attribute, corresponding to personal identifying information stored across a plurality of records of an input table for anonymization based on at least a threshold for a number of values required in each of one or more groupings of the data;

determine a plurality of values for both the first and second textual attributes, wherein each value comprises one or more characters and corresponds to at least one of the plurality of records, and wherein a subset of the plurality of values comprises a plurality of unique values;

determine a width of both the first textual attribute and the second textual attribute based on the plurality of unique values, wherein the width corresponds to a range of the unique values associated with a respective attribute;

select the first textual attribute based on its width being greater than a width of the second textual attribute, wherein the greater width corresponds to a likelihood of reduced data loss through the anonymization;

generate a plurality of groups based on the determined plurality of values, wherein each group includes one or more of the determined plurality of values that share one or more common characters; and

provide at least portion of the personal identifying information as sorted into the plurality of groups, wherein a count of the values of each group satisfies the threshold.

8. The system of claim 7 , wherein the processor that provides is configured to:

determine that the count of values for a particular one of the plurality of groups is less than the threshold; and

suppress the particular one of the plurality of groups that is less than the threshold, wherein the providing comprises providing the data sorted into the plurality of groups except for the particular group.

9. The system of claim 7 , wherein the personal information of the data includes:

an explicit identifier attribute from which a particular record of the data is distinguishable from one or more remaining records of the data, and wherein based on the explicit identifier, an individual corresponding to the record is identifiable;

a first quasi-identifier attribute which when considered together with a second quasi-identifier attribute identify the individual corresponding to the record; and

a sensitive data attribute which includes personal information corresponding to the individual.

10. The system of claim 9 , wherein the textual attribute corresponds to the first quasi-identifier attribute.

11. The system of claim 7 , wherein the processor that identifies is configured to:

identify a numerical attribute and a hierarchical attribute in addition to the plurality of textual attributes.

12. The system of claim 7 , wherein the processor that selects is further configured to:

determine a weight corresponding to the first textual attribute;

determine a weight corresponding to the second textual attribute;

determine a weighted width for both the first textual attribute and the second textual attribute; and

select the first textual attribute based on its weighted width being greater than a width of the second textual attribute.

13. A non-transitory computer-readable device having instructions stored thereon that, when executed by at least one computing device, causes the at least one computing device to perform operations comprising:

identifying a plurality of textual attributes, including both a first textual attribute and a second textual attribute, corresponding to personal identifying information stored across a plurality of records of an input table for anonymization based on at least a threshold for a number of values required in each of one or more groupings of the data;

determining a plurality of values for both the first and second textual attributes, wherein each value comprises one or more characters and corresponds to at least one of the plurality of records, and wherein a subset of the plurality of values comprises a plurality of unique values;

determining a width of both the first textual attribute and the second textual attribute based on the plurality of unique values, wherein the width corresponds to a range of the unique values associated with a respective attribute;

selecting the first textual attribute based on its width being greater than a width of the second textual attribute, wherein the greater width corresponds to a likelihood of reduced data loss through the anonymization;

generating a plurality of groups based on the determined plurality of values, wherein each group includes one or more of the determined plurality of values that share one or more common characters; and

providing at least portion of the personal identifying information as sorted into the plurality of groups, wherein a count of the values of each group satisfies the threshold.

14. The non-transitory computer-readable device of claim 13 , wherein the providing comprises:

determining that the count of values for a particular one of the plurality of groups is less than the threshold; and

suppressing the particular one of the plurality of groups that is less than the threshold, wherein the providing comprises providing the data sorted into the plurality of groups except for the particular group.

15. The non-transitory computer-readable device of claim 13 , wherein the personal information of the data includes:

an explicit identifier attribute from which a particular record of the data is distinguishable from one or more remaining records of the data, and wherein based on the explicit identifier, an individual corresponding to the record is identifiable;

a first quasi-identifier attribute which when considered together with a second quasi-identifier attribute identify the individual corresponding to the record; and

a sensitive data attribute which includes personal information corresponding to the individual.

16. The non-transitory computer-readable device of claim 15 , wherein the textual attribute corresponds to the first quasi-identifier attribute.

17. The non-transitory computer-readable device of claim 13 , wherein the identifying comprises:

identifying a numerical attribute and a hierarchical attribute in addition to the plurality of textual attributes.

18. The method of claim 1 , further comprising:

calculating a normalized certainty penalty corresponding to the information loss based on two or more different data types.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2017
From: HUANG, XINRONG
To: SAP SE
Reel/Frame 043978/0232 →
Continuity (1)
Related Publication 20190130131A1 · May 2, 2019