IP Library Granted Patent US 12,189,819
Granted Patent B2
US 12,189,819 · App. 17/744,630 · Granted Jan 7, 2025

Method and apparatus for de-identification of personal information

Inventors: Dae Woo Choi (Seoul, KR); Woo Seok Kwon (Seoul, KR); Myeong Sik Hwang (Seoul, KR); Sang Wook Kim (Seoul, KR); Gi Tae Kim (Seoul, KR)
Assignee: Fasoo
G06F21/6254G06F16/00G06F16/2379G06F21/62
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,189,819
App. No.
17/744,630
Granted
Jan 7, 2025
Kind
B2
Abstract

Disclosed are a method and an apparatus for de-identification of personal information. The method for de-identification of personal information comprises the steps of: obtaining, from a database, a raw table including records in which raw data indicating the personal information is recorded; generating generalized data by generalizing the raw data recorded in each of the records included in the raw table; setting a generalized hierarchical model consisting of the raw data and the generalized data; generating a raw lattice including a plurality of candidate nodes on the basis of the generalized hierarchical model; and setting, from among the plurality of candidate nodes included in the raw lattice, a final lattice including at least one candidate node satisfying a predetermined criterion. Thus, it is possible for the personal information to be efficiently de-identified.

Claims (53)

1. A personal information de-identification method performed by a personal information de-identification apparatus, the method comprising:

acquiring an original table including records in which original data indicating personal information is recorded from a database;

classifying respective records included in the original table based on attributes of the respective records, wherein the respective records are classified as one of classes of identifier (ID), quasi-identifier (QI), sensitive attribute (SA), and insensitive attribute (IA);

generalizing the original data recorded in the respective records included in the original table based on generalization levels;

setting up a generalization hierarchy model composed of the original data and the generalized data;

generating an original lattice including a plurality of candidate nodes indicating tables, which indicate generalization levels for types of personal information, based on a hierarchical structure indicated by the generalization hierarchy model; and

setting up a final lattice including one or more candidate nodes which satisfy a preset requirement among the plurality of candidate nodes included in the original lattice,

wherein a de-identified table generated in the generalizing of the original data is generated based on K-anonymity, generated based on K-anonymity and L-diversity, or generated based on K-anonymity and T-closeness, and

wherein the preset requirement includes a preset suppression requirement, which indicates a ratio of equivalence classes which do not satisfy a preset K-anonymity to equivalence classes constituting the de-identified table.

2. The personal information de-identification method of claim 1 , wherein the classifying respective records includes searching for personal information in the original table with regular expressions and setting up one of the classes for the respective records.

3. The personal information de-identification method of claim 1 , further comprising calculating a re-identification risk and a utility of a de-identified table corresponding to at least one final node included in the final lattice.

4. The personal information de-identification method of claim 1 , further comprising masking some or all of original data in records indicated by the identifier (ID) among the records included in the original table, or deleting original data in records indicated by the identifier (ID).

5. The personal information de-identification method of claim 1 , wherein the setting up of the final lattice comprises:

selecting one or more candidate nodes from among the plurality of candidate nodes included in the original lattice;

generating de-identified tables by de-identifying the original table based on generalization levels indicated by the one or more candidate nodes;

setting a candidate node corresponding to a de-identified table satisfying a preset suppression requirement to a final node; and

setting up the final lattice including the final node corresponding to the candidate node satisfying the preset requirement.

6. The personal information de-identification method of claim 1 , wherein at least one of:

the identifier (ID) indicates an equivalence class including a record in which original data indicating personal information whereby a specific individual is explicitly identified is recorded;

the quasi-identifier (QI) indicates an equivalence class including a record in which original data indicating personal information whereby a specific individual is inexplicitly identified is recorded;

the sensitive attribute (SA) indicates an equivalence class including a record in which original data indicating personal information having a sensitivity of a preset reference value or higher is recorded; and

the insensitive attribute (IA) indicates an equivalence class including a record in which original data indicating personal information having a lower sensitivity than sensitive attribute (SA) is recorded.

7. A personal information de-identification method performed by a personal information de-identification apparatus, the method comprising:

acquiring an original table including records in which original data indicating personal information is recorded from a database;

searching for personal information in the original table with regular expressions;

setting up respective records included in the original table as one of classes of identifier (ID), quasi-identifier (QI), sensitive attribute (SA), and insensitive attribute (IA) based on attributes of the respective records according to results of the searching; and

generalizing the original data recorded in the respective records included in the original table based on generalization levels,

wherein a de-identified table generated in the generalizing of the original data is generated based on K-anonymity, generated based on K-anonymity and L-diversity, or generated based on K-anonymity and T-closeness, and

wherein the preset requirement includes a preset suppression requirement, which indicates a ratio of equivalence classes which do not satisfy a preset K-anonymity to equivalence classes constituting the de-identified table.

8. The personal information de-identification method of claim 7 , further comprising setting up a generalization hierarchy model composed of the original data and the generalized data.

9. The personal information de-identification method of claim 8 , further comprising generating an original lattice including a plurality of candidate nodes indicating tables, which indicate generalization levels for types of personal information, based on a hierarchical structure indicated by the generalization hierarchy model.

10. The personal information de-identification method of claim 9 , further comprising setting up a final lattice including one or more candidate nodes which satisfy a preset requirement among the plurality of candidate nodes included in the original lattice.

11. The personal information de-identification method of claim 7 , further comprising calculating a re-identification risk and a utility of a de-identified table corresponding to at least one final node included in the final lattice.

12. The personal information de-identification method of claim 7 , further comprising masking some or all of original data in records indicated by the identifier (ID) among the records included in the original table, or deleting original data in records indicated by the identifier (ID).

13. The personal information de-identification method of claim 7 , wherein the setting up of the final lattice comprises:

selecting one or more candidate nodes from among the plurality of candidate nodes included in the original lattice;

generating de-identified tables by de-identifying the original table based on generalization levels indicated by the one or more candidate nodes;

setting a candidate node corresponding to a de-identified table satisfying a preset suppression requirement to a final node; and

setting up the final lattice including the final node corresponding to the candidate node satisfying the preset requirement.

14. A personal information de-identification apparatus comprising:

a processor; and

a memory configured to store at least one command executed by the processor,

wherein the at least one command is executable to:

acquire an original table including records in which original data indicating personal information is recorded from a database;

search for personal information in the original table on the basis of regular expressions;

set up respective records included in the original table as one of classes of identifier (ID), quasi-identifier (QI), sensitive attribute (SA), and insensitive attribute (IA) based on attributes of the respective records according to results of the search; and

generalize the original data recorded in the respective records included in the original table based on generalization levels,

wherein a de-identified table generated in the generalizing of the original data is generated based on K-anonymity, generated based on K-anonymity and L-diversity, or generated based on K-anonymity and T-closeness, and

wherein the preset requirement includes a preset suppression requirement, which indicates a ratio of equivalence classes which do not satisfy a preset K-anonymity to equivalence classes constituting the de-identified table.

15. The personal information de-identification apparatus of claim 14 , wherein at least one command is further executable to:

set up a generalization hierarchy model composed of the original data and the generalized data;

generate an original lattice including a plurality of candidate nodes indicating tables, which indicate generalization levels for types of personal information, based on a hierarchical structure indicated by the generalization hierarchy model; and

set up a final lattice including one or more candidate nodes which satisfy a preset requirement among the plurality of candidate nodes included in the original lattice.

Assignments (2)
CHANGE OF NAME Recorded May 17, 2022
From: FASOO.COM
To: FASOO
Reel/Frame 060073/0899 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2022
From: CHOI, DAE WOO; KWON, WOO SEOK; HWANG, MYEONG SIK; KIM, SANG WOOK; KIM, GI TAE
To: FASOO.COM CO., LTD.
Reel/Frame 060072/0240 →
Priority Claims (3)
KR 10-2016-0082839 · Jun 30, 2016 · national
KR 10-2016-0082860 · Jun 30, 2016 · national
KR 10-2016-0082878 · Jun 30, 2016 · national
Continuity (2)
Continuation 16314202
Related Publication 20220277106A1 · Sep 1, 2022
References Cited (17)
US 9600673B2 · Chen · 2017 [cited by applicant]
US 20020169793A1 · Sweeney · 2002 [cited by applicant]
US 20070255704A1 · Baek · 2007 [cited by applicant]
US 20100077006A1 · El Emam · 2010 [cited by applicant]
US 20100332537A1 · El Emam · 2010 [cited by applicant]
US 20130138698A1 · Harada · 2013 [cited by applicant]
US 20150007249A1 · Bezzi · 2015 [cited by examiner]
US 20160154978A1 · Baker · 2016 [cited by examiner]
US 20180114037A1 · Scaiano · 2018 [cited by applicant]
JP 2011113285 · 2011 [cited by applicant]
JP 2013080375 · 2013 [cited by applicant]
JP 2013161428 · 2013 [cited by applicant]
WO 2008069011 · 2008 [cited by applicant]
WO 2011145401 · 2011 [cited by applicant]
Khaled El Eman et al., A Globally Optimal k-Anonymity Method for the De-Identification of Health Data, Journal of the American Medical Informatics Association, Sep./Oct. 2009, pp. 670-682, vol. 16, No. 5, United States. [cited by applicant]
Koji Sedna, Revision of the Personal Information Protection Law and new trends in data science Anonymization technology that reduces the risk of individual identification, Communications of the Operations Research Socie… [cited by applicant]
Florian Kohlmayer et al., Flash: Efficient, Stable and Optimal K-Anonymity, 2012 ASE/IEEE International Conference on Social Computing and 2012 ASE/IEEE International Conference on Privacy, Security, Risk and Trust, Sep… [cited by applicant]