IP Library Granted Patent US 11,226,970
Granted Patent B2
US 11,226,970 · App. 16/145,825 · Granted Jan 18, 2022

System and method for tagging database properties

Inventors: Tomoya Wada (New York, NY); Winnie Cheng (West New York, NJ); Rohit Mahajan (Iselin, NJ); Alex Mylnikov (Edison, NJ)
Assignee: HITACHI VANTARA LLC
G06F16/24578G06F16/221G06F16/9024G06F16/9535
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,226,970
App. No.
16/145,825
Granted
Jan 18, 2022
Kind
B2
Abstract

A method and system for tagging database columns are presented. The method includes receiving an input column name of at least one column in a database; performing signature matching of the input column name to contents of a seed table; determining a first confidence score for the signature matching; and tagging a matching value in the seed table as a tag for the input column name, when a first confidence score exceeds a first threshold value.

Claims (51)

1. A method for tagging database columns, comprising:

receiving an input column name of at least one column in a database;

performing signature matching of the input column name to contents of a seed table, wherein the seed table contains previously generated tags associated with their respective column names, and wherein the signature matching is based on a natural language processing;

determining a first confidence score for the signature matching;

tagging a matching value in the seed table as a tag for the input column name, when the first confidence score exceeds a first threshold value, wherein the first confidence score is determined as a probability of the column-pair matching computed based on similarity of the column names; and

performing graph signature matching of the input column name to discovery-assistance data (DAD);

determining a third confidence score for the graph signature matching; and

tagging a matching value in a DAD table as a tag for the input column name when the third confidence score exceeds a third threshold value.

2. The method of claim 1 , further comprising:

performing probabilistic signature matching of the input column name to contents of a data corpus table;

determining a second confidence score for the probabilistic signature matching; and

tagging a matching value in the seed table as a tag for the input column name when the second confidence score exceeds a second threshold value.

3. The method of claim 1 , wherein performing signature matching of the input column name to contents of a seed table further comprising:

performing a natural language process to match a key to a closest value, wherein the key is the input column name and the value is an entry in the seed table.

4. The method of claim 3 , wherein further comprising:

performing a dictionary search and phonetic n-gram search to identify a matching key in the seed table.

5. The method of claim 1 , wherein performing graph signature matching of the input column name to contents of the seed table further comprising:

performing a natural language process to match a key to metadata representing relationship as maintained by the DAD table.

6. The method of claim 5 , wherein the DAD table maintains data flows, wherein each data flow represents similarities of two columns based on their contents.

7. The method of claim 6 , wherein relationship metadata is based on a graph of the data flows.

8. A non-transitory computer readable medium having stored thereon instructions for causing a processing circuitry to execute a process, the process comprising:

receiving an input column name of at least one column in a database;

performing a signature matching of the input column name to contents of a seed table, wherein the seed table contains previously generated tags associated with their respective column names, and wherein the signature matching is based on a natural language processing;

determining a first confidence score for the signature matching;

tagging a matching value in the seed table as a tag for the input column name, when the first confidence score exceeds a first threshold value, wherein the first confidence score is determined as a probability of the column-pair matching computed based on similarity of the column names; and

performing graph signature matching of the input column name to discovery-assistance data (DAD);

determining a third confidence score for the graph signature matching; and

tagging a matching value in a DAD table as a tag for the input column name when the third confidence score exceeds a third threshold value.

9. A system for tagging database columns, comprising:

a processing circuitry; and

a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to:

receive an input column name of at least one column in a database;

perform a signature matching of the input column name to contents of a seed table, wherein the seed table contains previously generated tags associated with their respective column names, and wherein the signature matching is based on a natural language processing;

determine a first confidence score for the signature matching;

tag a matching value in the seed table as a tag for the input column name, when the first confidence score exceeds a first threshold value, wherein the first confidence score is determined as a probability of the column-pair matching computed based on similarity of the column names; and

perform graph signature matching of the input column name to a DAD;

determine a third confidence score for the graph signature matching; and

tag a matching value in a DAD table as a tag for the input column name when the third confidence score exceeds a third threshold value.

10. The system of claim 9 , wherein the system is further configured to:

perform probabilistic signature matching of the input column name to contents of a data corpus table;

determine a second confidence score for the probabilistic signature matching; and

tag a matching value in the seed table as a tag for the input column name when the second confidence score exceeds a second threshold value.

11. The system of claim 10 , wherein the system is further configured to:

perform a natural language process to match a key to a closest value, wherein the key is the input column name and the value is an entry in the seed table.

12. The system of claim 11 , wherein the system is further configured to:

perform a dictionary search and phonetic n-gram search to identify a matching key in the seed table.

13. The system of claim 11 , wherein the seed table includes previously discovered tags associated with their respective column names.

14. The system of claim 13 , wherein the system is further configured to:

perform a natural language process to match a key to metadata representing relationship as maintained by the DAD table.

15. The system of claim 14 , wherein the DAD table maintains data flows, wherein each data flow represents similarities of two columns based on their contents.

16. The system of claim 15 , wherein relationship metadata is based on a graph of the data flows.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2021
From: IO-TAHOE LLC
To: HITACHI VANTARA LLC
Reel/Frame 057323/0314 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2018
From: WADA, TOMOYA; CHENG, WINNIE; MAHAJAN, ROHIT; MYLNIKOV, ALEX
To: IO-TAHOE LLC.
Reel/Frame 047006/0530 →
Continuity (1)
Related Publication 20200104379A1 · Apr 2, 2020