IP Library › Granted Patent US 11,556,514
Granted Patent B2
US 11,556,514 · App. 17/184,122 · Granted Jan 17, 2023

Semantic data type classification in rectangular datasets

Inventors: Roger C. Raphael (San Jose, CA); Mu Qiao (Belmont, CA); Scott Schumacher (Porter Ranch, CA); Angineh Aghakiant (San Jose, CA)
Assignee: International Business Machines Corporation
G06F16/2282G06F16/213G06F16/221G06F16/258G06F40/30G06K9/6256G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,556,514
App. No.
17/184,122
Filed
Feb 24, 2021
Granted
Jan 17, 2023
Kind
B2
Art Unit
2168
USPC
707/803
Abstract

Provided is a method, computer program product, and system for automatically predicting unknown semantic data types in a rectangular dataset using a holistic knowledge of said dataset. A processor may receive one or more rectangular datasets, the one or more rectangular datasets comprising a plurality of columns having a set of known semantic data types. The processor may extract a set of features from the plurality of columns, where the set of features is used to determine a relationship among each column of the plurality of columns. The processor may construct a set of training data based on the extracted set of features. Using the training data, the processor may train a machine learning model to predict a semantic data type of a target column in a rectangular dataset having an unknown semantic data type.

Claims (75)

1. A computer-implemented method comprising:

receiving one or more rectangular datasets, wherein the one or more rectangular datasets comprise a plurality of columns having a set of known semantic data types;

extracting a set of features from the plurality of columns, wherein the set of features is used to determine a relationship among each column of the plurality of columns;

constructing a set of training data based on the extracted set of features;

training, using the set of training data, a machine learning model to predict a semantic data type of a target column in a rectangular dataset having an unknown semantic data type;

receiving a first rectangular dataset, wherein the first rectangular dataset contains a first target column having a first unknown semantic data type and a first plurality of other columns containing known semantic data types;

analyzing, using the machine learning model, a first set of local features associated with the first target column;

correlating the first set of local features of the first target column with a set of known features of the first plurality of other columns of the first rectangular dataset, wherein at least one known feature includes a name of each column of the first plurality of columns; and

predicting, based on the correlating, a likelihood that the first unknown semantic data type is a specific semantic data type.

2. The computer-implemented method of claim 1 , wherein the first set of local features associated with the first target column is selected from the group of local features consisting of:

primitive data type of the first target column;

name of the first target column;

length of the first target column name; and

entropy of the first target column.

3. The computer-implemented method of claim 1 , wherein the set of known features associated with the first plurality of other columns of the first rectangular dataset is further selected from the group of known features consisting of:

a total number of columns in the first rectangular dataset;

entropy of each column of the plurality of other columns; and

a correlation coefficient between the first target column and each of the plurality of other columns.

4. The computer-implemented method of claim 1 , further comprising:

receiving a user verification indicating that the likelihood of the first unknown semantic data type being the specific semantic data type is accurate; and

retraining, in response to the user verification, the machine learning model using the first rectangular dataset, wherein the first target column is classified as the specific semantic data type.

5. The computer-implemented method of claim 4 , further comprising:

analyzing, using the retrained machine learning model, a second rectangular dataset to determine a second unknown semantic data type of a second target column.

6. The computer-implemented method of claim 1 , wherein the likelihood includes a likelihood value and the method further comprises:

comparing the likelihood value to a predetermined threshold; and

retraining automatically, in response the likelihood value meeting the predetermined threshold, the machine learning model using the specific semantic data type of the first target column of the first rectangular dataset.

7. The computer-implemented method of claim 1 , wherein the set of training data comprises one or more feature vectors.

8. The computer-implemented method of claim 1 , wherein the at least one known feature further includes an entropy value for a set of data of each column of the plurality of other columns.

9. The computer-implemented method of claim 8 , wherein the entropy value for the set of data of each column of the plurality of other columns is compared to an entropy threshold.

10. The computer-implemented method of claim 9 , wherein the set of data for a first other column is ignored if the entropy value exceeds the entropy threshold.

11. A system comprising:

a processor; and

a computer-readable storage medium communicatively coupled to the processor and storing program instructions which, when executed by the processor, cause the processor to perform a method comprising:

receiving one or more rectangular datasets, wherein the one or more rectangular datasets comprise a plurality of columns having a set of known semantic data types;

extracting a set of features from the plurality of columns, wherein the set of features is used to determine a relationship among each column of the plurality of columns;

constructing a set of training data based on the extracted set of features;

training, using the set of training data, a machine learning model to predict a semantic data type of a target column in a rectangular dataset having an unknown semantic data type;

receiving a first rectangular dataset, wherein the first rectangular dataset contains a first target column having a first unknown semantic data type and a first plurality of other columns containing known semantic data types;

analyzing, using the machine learning model, a first set of local features associated with the first target column;

correlating the first set of local features of the first target column with a set of known features of the first plurality of other columns of the first rectangular dataset, wherein at least one known feature includes a name of each column of the first plurality of columns; and

predicting, based on the correlating, a likelihood that the first unknown semantic data type is a specific semantic data type.

12. The system of claim 11 , wherein the first set of local features associated with the first target column is selected from the group of local features consisting of:

primitive data type of the first target column;

name of the first target column;

length of the first target column name; and

entropy of the first target column.

13. The system of claim 11 , wherein the set of known features associated with the first plurality of other columns of the first rectangular dataset is further selected from the group of known features consisting of:

a total number of columns in the first rectangular dataset;

entropy of each column of the plurality of other columns; and

a correlation coefficient between the first target column and each of the plurality of other columns.

14. The system of claim 11 , wherein the method performed by the processor further comprises:

receiving a user verification indicating that the likelihood of the first unknown semantic data type being the specific semantic data type is accurate;

retraining, in response to the user verification, the machine learning model using the first rectangular dataset, wherein the first target column is classified as the specific semantic data type; and

analyzing, using the retrained machine learning model, a second rectangular dataset to determine a second unknown semantic data type of a second target column.

15. The system of claim 11 , wherein the likelihood includes a likelihood value and the method further comprises:

comparing the likelihood value to a predetermined threshold; and

retraining automatically, in response the likelihood value meeting the predetermined threshold, the machine learning model using the specific semantic data type of the first target column of the first rectangular dataset.

16. A computer program product comprising a computer-readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to perform a method comprising:

receiving one or more rectangular datasets, wherein the one or more rectangular datasets comprise a plurality of columns having a set of known semantic data types;

extracting a set of features from the plurality of columns, wherein the set of features is used to determine a relationship among each column of the plurality of columns;

constructing a set of training data based on the extracted set of features;

training, using the set of training data, a machine learning model to predict a semantic data type of a target column in a rectangular dataset having an unknown semantic data type;

receiving a first rectangular dataset, wherein the first rectangular dataset contains a first target column having a first unknown semantic data type and a first plurality of other columns containing known semantic data types;

analyzing, using the machine learning model, a first set of local features associated with the first target column;

correlating the first set of local features of the first target column with a set of known features of the first plurality of other columns of the first rectangular dataset, wherein at least one known feature includes a name of each column of the first plurality of columns; and

predicting, based on the correlating, a likelihood that the first unknown semantic data type is a specific semantic data type.

17. The computer program product of claim 16 , wherein the first set of local features associated with the first target column is selected from the group of local features consisting of:

primitive data type of the first target column;

name of the first target column;

length of the first target column name; and

entropy of the first target column.

18. The computer program product of claim 16 , wherein the set of known features associated with the first plurality of other columns of the first rectangular dataset is further selected from the group of known features consisting of:

a total number of columns in the first rectangular dataset;

entropy of each column of the plurality of other columns; and

a correlation coefficient between the first target column and each of the plurality of other columns.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 24, 2021
From: RAPHAEL, ROGER C.; QIAO, MU; SCHUMACHER, SCOTT; AGHAKIANT, ANGINEH
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 055394/0520 →
Continuity (1)
Related Publication 20220269663A1 · Aug 25, 2022
Cited By (3)
US 12,339,858 US 12,554,932 US 12,646,002