IP Library Granted Patent US 12,326,871
Granted Patent B2
US 12,326,871 · App. 18/507,769 · Granted Jun 10, 2025

Detecting duplicate tables in data lake databases

Inventors: Karina Elayne Kervin (Sacramento, CA); Jian Wu (Round Rock, TX); Sibasis Das (Kolkata, IN); Radha Mohan De (Howrah, IN); Swaminathan Balasubramanian (Troy, MI); Cheranellore Vasudevan (Bastrop, TX)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F16/254G06F16/215G06F16/2246G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,326,871
App. No.
18/507,769
Granted
Jun 10, 2025
Kind
B2
Abstract

Detecting duplicate tables by converting relational tables to knowledge graphs; and mapping nodes in the knowledge graphs to sources in the relational tables. The method may further include applying graph matching to the knowledge graphs; and assessing degree of matching between matched knowledge graphs.

Claims (35)

1. A computer implemented method of detecting duplicate tables comprising:

converting syntactical data from relational tables from a schema of a database to knowledge graphs having semantic data;

mapping nodes in the knowledge graphs to sources in the relational tables;

applying graph matching to the knowledge graphs;

detecting semantically equivalent tables from the relational tables by assessing a degree of matching between matched knowledge graphs based on equivalent data from common linked-pair nodes, defined from relationships, from the knowledge graphs; and

generating a corrective flag to eliminate the semantically equivalent tables from the relational tables from the schema of the database based on the degree of matching.

2. The computer implemented method of claim 1 , wherein the relational tables comprise data arranged in columns.

3. The computer implemented method of claim 1 , wherein the duplicate tables are present in a cloud based data lake.

4. The computer implemented method of claim 1 , wherein the graph matching comprises using at least one of domain specific ontologies, metadata of the nodes, metadata of entities, metadata of fields, and graph manipulation algorithms to determine matches in graphs.

5. The computer implemented method of claim 4 , wherein graph matching by domain specific ontologies includes separating the matched knowledge graphs into table-based sub-graphs and look for similarity between sub-graphs.

6. The computer implemented method of claim 1 , wherein graph matching is a full match between at least two knowledge graphs.

7. The computer implemented method of claim 1 , wherein graph matching comprises a partial match between at least two knowledge graphs.

8. The computer implemented method of claim 1 , wherein graph matching comprises an algorithm selected from the group consisting of graph edit distance, graph kernel, graph embedded for general vector representation, graph neural networks, and combinations thereof.

9. A system for detecting duplicate tables including a hardware processor; and a memory that stores a computer program product, the computer program product of the system includes instructions comprising:

convert, using a hardware processor, syntactical data from relational tables from a schema of a database to knowledge graphs having semantic data;

map, using a hardware processor, nodes in the knowledge graphs to sources in the relational tables;

apply, using a hardware processor, graph matching to the knowledge graphs; and

detect, using the hardware processor, semantically equivalent tables from the relational tables by assessing; a degree of matching between matched knowledge graphs based on equivalent data from common linked-pair nodes, defined from relationships, from the knowledge graphs; and

generate, using the hardware processor, a corrective flag to eliminate the semantically equivalent tables from the relational tables from the schema of the database based on the degree of matching.

10. The system of claim 9 , wherein the relational tables comprise data arranged in columns.

11. The system of claim 9 , wherein the graph matching comprises using at least one of domain specific ontologies, metadata of the nodes, metadata of entities, metadata of fields, and graph manipulation algorithms to determine matches in graphs.

12. The system of claim 11 , wherein graph matching by domain specific ontologies includes separating the matched knowledge graphs into table-based sub-graphs and look for similarity between sub-graphs.

13. The system of claim 9 , wherein graph matching is a full match between at least two knowledge graphs.

14. The system of claim 9 wherein graph matching comprises a partial match between at least two knowledge graphs.

15. The system of claim 9 , wherein graph matching comprises an algorithm selected from the group consisting of graph edit distance, graph kernel, graph embedded for general vector representation, graph neural networks, and combinations thereof.

16. The system of claim 9 , wherein the duplicate tables are present in a cloud based data lake.

17. A computer program product for detecting duplicate tables, the computer program product including a computer readable storage medium having computer readable program code embodied therewith, the program instructions executable by a processor to cause the processor to:

convert syntactical data from relational tables from a schema of a database to knowledge graphs having semantic data;

map nodes in the knowledge graphs to sources in the relational tables;

apply graph matching to the knowledge graphs;

detect semantically equivalent tables from the relational tables by assessing a degree of matching between matched knowledge graphs based on equivalent data from common linked-pair nodes, defined from relationships, from the knowledge graphs; and

generate a corrective flag to eliminate the semantically equivalent tables from the relational tables from the schema of the database based on the degree of matching.

18. The computer program product of claim 17 , wherein the duplicate tables are present in a cloud based data lake.

19. The computer program product of claim 17 , wherein the relational tables comprise data arranged in columns.

20. The computer program product of claim 17 , wherein graph matching is a full match between at least two knowledge graphs, or a partial match between at least two knowledge graphs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: KERVIN, KARINA ELAYNE; WU, JIAN; DAS, SIBASIS; DE, RADHA MOHAN; BALASUBRAMANIAN, SWAMINATHAN; VASUDEVAN, CHERANELLORE
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 065544/0393 →
Continuity (1)
Related Publication 20250156435A1 · May 15, 2025
References Cited (36)
US 8380681B2 · Oltean et al. · 2013 [cited by applicant]
US 9223794B2 · Therrien et al. · 2015 [cited by applicant]
US 9830383B2 · Balasubramanian et al. · 2017 [cited by applicant]
US 10198460B2 · Gorelik et al. · 2019 [cited by applicant]
US 10242016B2 · Gorelik et al. · 2019 [cited by applicant]
US 10445062B2 · Oberbreckling et al. · 2019 [cited by applicant]
US 11119980B2 · Szczepanik et al. · 2021 [cited by applicant]
US 11204907B2 · VanderSpek et al. · 2021 [cited by applicant]
US 11379506B2 · Stojanovic et al. · 2022 [cited by applicant]
US 11449499B1 · Alsaadi · 2022 [cited by examiner]
US 11544566B2 · Gupta et al. · 2023 [cited by applicant]
US 20100063973A1 · Cao · 2010 [cited by examiner]
US 20100306412A1 · Therrien et al. · 2010 [cited by applicant]
US 20120158672A1 · Oltean et al. · 2012 [cited by applicant]
US 20150281292A1 · Murayama · 2015 [cited by examiner]
US 20210019232A1 · Murti et al. · 2021 [cited by applicant]
US 20210397738A1 · Talreja et al. · 2021 [cited by applicant]
US 20220021652A1 · Moghe · 2022 [cited by examiner]
US 20220222543A1 · Khatibi · 2022 [cited by examiner]
US 20220269659A1 · Wang et al. · 2022 [cited by applicant]
US 20230004347A1 · Slager · 2023 [cited by examiner]
US 20230229644A1 · Bremer · 2023 [cited by examiner]
US 20240095241A1 · Zheng · 2024 [cited by examiner]
CN 11436418A · 2022 [cited by applicant]
WO 2020135048A1 · 2020 [cited by applicant]
Article entitled “Transforming Table to Knowledge Graph Using a Rule-Based Pipeline”, by Yulianti et al., dated 2021. (Year: 2021). [cited by examiner]
Article entitled “Construction and Application of a Knowledge Graph”, by Hao et al., dated Jun. 26, 2021 (Year: 2021). [cited by examiner]
Article entitled “Interpreting Language Models Through Knowledge Graph Extraction”, by Swamy et al., dated Nov. 16, 2021. ( Year: 2021). [cited by examiner]
Article entitled Demonstrating MATE and COCOA for Data Discovery, by Becktepe et al., dated Jun. 23, 2023. (Year: 2023). [cited by examiner]
Article entitled “Models of Similarity in Complex Networks”, by Shvydun dated May 2, 2023. (Year: 2023). [cited by examiner]
Article entitled “Structural-Semantic Approach for Approximate Frequent Subgraph Mining”, by Moussaoui et al., dated 2015. (Year: 2015). [cited by examiner]
Ma, G., Ahmed, N. K., Willke, T. L., & Yu, P. S. (Oct. 4, 2020). Deep graph similarity learning: A survey. Data Mining and Knowledge Discovery, 35, 688-725. [cited by applicant]
Dobroshinksy, S. (Feb. 5, 2020). Integrate and deduplicate datasets using AWS Lake Formation FindMatches. AWS Big Data Blog, Retrieved from https://aws. amazon. com/blogs/big-data/integrate-and-deduplicate-datasets-usin… [cited by applicant]
Livi, L., & Rizzi, A. (Aug. 21, 2012). The graph matching problem. Pattern Analysis and Applications, 16, 253-283. [cited by applicant]
Dibowski, et al., Using Knowledge Graphs to Manage a Data Lake, GI-INFORMATIK, Jan. 2021, 11 pages. [cited by applicant]
Lee et al. “Table2Graph: A Scalable Graph Construction from Relational Tables using Map-Reduce”, IEEE First International Conference on Big Data Computing Service and Applications, 2015, 8 pages. [cited by applicant]