IP Library Granted Patent US 12,619,783
Granted Patent B2
US 12,619,783 · App. 18/812,227 · Granted May 5, 2026

Automatic pseudonymization technique recommendation method using artificial intelligence

Inventors: Gi Chang Shim (Seoul, KR); Yong Kyu Park (Seoul, KR)
Assignee: EASYCERTI INC.
G06F21/6254
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,783
App. No.
18/812,227
Granted
May 5, 2026
Kind
B2
Abstract

Provided is an automatic pseudonymization technique recommendation method using artificial intelligence, including receiving a training dataset including a plurality of pieces of data including a column name, a data type, a pseudonymization technique recommendation, adding, a data type vector obtained from the data type to a word vector obtained corresponding to the column name to obtain a numeric vector, and labeling the numeric vector with the pseudonymization technique recommendation to generate a plurality of pieces of training data, training a first learning model using the plurality of pieces of training data so that the first learning model outputs a pseudonymization technique recommendation in response to an input of the numeric vector, obtaining a numeric vector corresponding to each column of the dataset to be pseudonymized, and inputting the numeric vector obtained into the trained first learning model to obtain a pseudonymization technique recommendation for each column of the dataset.

Claims (28)

1 . An automatic pseudonymization technique recommendation method using artificial intelligence, wherein the method is implemented on a computing device comprising a processor, and a memory storing instructions or programs executable by the processor, and comprises:

receiving, at the computer device, a training dataset including a plurality of pieces of data including a column name, a data type, and a pseudonymization technique recommendation;

connecting, at the computer device, for each of the plurality of pieces of data, a data type vector obtained from the data type with a word vector obtained corresponding to the column name to obtain a numeric vector, and labeling the numeric vector with the pseudonymization technique recommendation to generate a plurality of pieces of training data;

training, at the computer device, a first learning model using the plurality of pieces of training data so that the first learning model outputs a pseudonymization technique recommendation in response to an input of the numeric vector;

obtaining, at the computer device, a numeric vector corresponding to each column of the dataset to be pseudonymized;

inputting, at the computer device, the numeric vector obtained corresponding to each column of the dataset to be pseudonymized into the trained first learning model to obtain a pseudonymization technique recommendation for each column of the dataset to be pseudonymized; and

performing pseudonymization on the dataset to be pseudonymized based on the pseudonymization technique recommendation obtained for each column of the dataset to be pseudonymized,

wherein

the plurality of pieces of data included in the training dataset further includes industry classification information;

the numeric vector obtained when generating the training data further includes an industry classification vector obtained from industry classification information; and

the obtaining the numeric vector corresponding to each column of the dataset to be pseudonymized includes:

inputting a column name corresponding to each column of the dataset to be pseudonymized into a second learning model to obtain a word vector of each column of the dataset to be pseudonymized;

obtaining a data type vector corresponding to the data type of each column of the dataset to be pseudonymized;

obtaining an industry classification vector corresponding to industry classification information previously set for the dataset to be pseudonymized; and

connecting the data type vector and the industry classification vector with the word vector obtained for each column of the dataset to be pseudonymized to obtain a numeric vector corresponding to each column of the dataset to be pseudonymized,

wherein the data type includes at least a numeric type and a character type.

2 . The method of claim 1 , further comprising training or retraining the second learning model using a plurality of column names obtained from the training dataset.

3 . The method of claim 2 , wherein

the first learning model is a decision tree model, and

the second learning model is a word embedding model.

4 . The method of claim 3 , wherein

the word embedding model is a FastText model, and

the decision tree model is a Gradient Boosting Trees model.

5 . A non-transitory computer-readable recording medium recording a program for executing the automatic pseudonymization technique recommendation method using artificial intelligence as set forth in claim 1 .

6 . A computing device, comprising:

a processor; and

a memory storing instructions or programs executable by the processor, wherein,

when the instructions or programs are executed by the processor, the automatic pseudonymization technique recommendation method using artificial intelligence as set forth in claim 1 is executed.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2024
From: SHIM, GI CHANG; PARK, YONG KYU
To: EASYCERTI INC.
Reel/Frame 068750/0967 →
Priority Claims (1)
KR 10-2023-0111703 · Aug 25, 2023 · national
Continuity (1)
Related Publication 20250068769A1 · Feb 27, 2025
References Cited (14)
US 11526508B1 · McCallie, Jr. · 2022 [cited by examiner]
US 11599749B1 · Ziaeefard · 2023 [cited by examiner]
US 20160254911A1 · Manchepalli et al. · 2016 [cited by applicant]
US 20190354720A1 · Tucker · 2019 [cited by examiner]
US 20210133557A1 · Iyoob · 2021 [cited by examiner]
US 20220188699A1 · Matlick · 2022 [cited by examiner]
US 20220382906A1 · Wu et al. · 2022 [cited by applicant]
US 20230195933A1 · Satish Padmanabhan · 2023 [cited by examiner]
US 20240062011A1 · Kanuga · 2024 [cited by examiner]
US 20240062044A1 · Subramanian · 2024 [cited by examiner]
US 20240095394A1 · Madhavan · 2024 [cited by examiner]
US 20240119175A1 · Middleton · 2024 [cited by examiner]
JP 2021503648A · 2021 [cited by applicant]
JP 2023070942A · 2023 [cited by applicant]