IP Library › Granted Patent US 12,333,253
Granted Patent B2
US 12,333,253 · App. 17/529,899 · Granted Jun 17, 2025

Automatic data domain identification

Inventors: Malolan Chetlur (Jakkur, IN); Arvind Agarwal (New Delhi, IN); Subhendu Dey (Kolkata, IN); Sameep Mehta (Bangalore, IN); Sandipan Sarkar (Kolkata, IN)
Assignee: International Business Machines Corporation
G06F40/30G06F40/242
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,253
App. No.
17/529,899
Granted
Jun 17, 2025
Kind
B2
Abstract

An apparatus is disclosed which includes at least one processing device comprising a processor coupled to a memory. The at least one processing device, when executing program code, is configured to: extract one or more entities identified in a plurality of data artifacts based at least in part on one or more datasets, extract one or more entities identified in a plurality of code artifacts based at least in part on the one or more datasets, extract one or more entities identified in a plurality of user interface artifacts based at least in part on the one or more datasets, generate a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities, and perform one or more of a lexical analysis and a semantic analysis on the set of dependency graphs to identify a data domain of the one or more datasets.

Claims (41)

1. An apparatus, comprising:

at least one processing device comprising a processor coupled to a memory, the at least one processing device, when executing program code, is configured to:

extract one or more entities identified in a plurality of data artifacts based at least in part on one or more datasets;

extract one or more entities identified in a plurality of code artifacts based at least in part on the one or more datasets;

extract one or more entities identified in a plurality of user interface artifacts based at least in part on the one or more datasets;

generate a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities; and

perform one or more of a lexical analysis and a semantic analysis on the set of dependency graphs to identify a data domain of the one or more datasets.

2. The apparatus of claim 1 , wherein the plurality of data artifacts comprises one or more of (a) one or more schemas with their table names and associated column names, (b) index, trigger, and stored procedures associated with the schemas, (c) relationships between the different tables and databases, (d) table data, and (e) documentation, logs, performance and operational profile of datasets.

3. The apparatus of claim 1 , wherein the plurality of code artifacts comprises source code and associated libraries.

4. The apparatus of claim 1 , wherein the plurality of user interface artifacts comprises one or more of (a) user interface screens with natural language text, (b) user interface form objects and formatting, and (c) user interface modalities.

5. The apparatus of claim 1 , wherein generating a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities comprises generating a first dependency graph of the set of dependency graphs based at least in part on one or more relationships between the extracted one or more entities of the data artifacts and the extracted one or more entities of the code artifacts.

6. The apparatus of claim 5 , wherein generating a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities further comprises generating a second dependency graph of the set of dependency graphs based at least in part on one or more relationships between the extracted one or more entities of the data artifacts and the extracted one or more entities of the user interface artifacts.

7. The apparatus of claim 6 , wherein generating a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities further comprises generating a third dependency graph of the set of dependency graphs based at least in part on one or more relationships between the extracted one or more entities of the code artifacts and the extracted one or more entities of the user interface artifacts.

8. The apparatus of claim 1 , wherein the at least one processing device, when executing program code, is further configured to:

retrieve the data artifacts from at least one of a database and a file system; and

apply one of a data definition language operation and a data manipulation language operation to identify the one or more entities in each data artifact and to determine one or more relationships between the one or more entities.

9. A computer-implemented method, comprising:

extracting one or more entities identified in a plurality of data artifacts based at least in part on one or more datasets;

extracting one or more entities identified in a plurality of code artifacts based at least in part on the one or more datasets;

extracting one or more entities identified in a plurality of user interface artifacts based at least in part on the one or more datasets;

generating a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities; and

performing one or more of a lexical analysis and a semantic analysis on the set of dependency graphs to identify a data domain of the one or more datasets;

wherein the method is carried out by at least one computing device.

10. The computer-implemented method of claim 9 , wherein the plurality of data artifacts comprises one or more of (a) one or more schemas with their table names and associated column names, (b) index, trigger, and stored procedures associated with the schemas, (c) relationships between the different tables and databases, (d) table data, and (e) documentation, logs, performance and operational profile of datasets.

11. The computer-implemented method of claim 9 , wherein the plurality of code artifacts comprises source code and associated libraries.

12. The computer-implemented method of claim 9 , wherein the plurality of user interface artifacts comprises one or more of (a) user interface screens with natural language text, (b) user interface form objects and formatting, and (c) user interface modalities.

13. The computer-implemented method of claim 9 , wherein generating a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities comprises generating a first dependency graph of the set of dependency graphs based at least in part on one or more relationships between the extracted one or more entities of the data artifacts and the extracted one or more entities of the code artifacts.

14. The computer-implemented method of claim 13 , wherein generating a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities further comprises generating a second dependency graph of the set of dependency graphs based at least in part on one or more relationships between the extracted one or more entities of the data artifacts and the extracted one or more entities of the user interface artifacts.

15. The computer-implemented method of claim 14 , wherein generating a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities further comprises generating a third dependency graph of the set of dependency graphs based at least in part on one or more relationships between the extracted one or more entities of the code artifacts and the extracted one or more entities of the user interface artifacts.

16. The computer-implemented method of claim 9 , further comprising:

retrieving the data artifacts from at least one of a database and a file system; and

applying one of a data definition language operation and a data manipulation language operation to identify the one or more entities in each data artifact and determine one or more relationships between the one or more entities.

17. A computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computing device to cause the computing device to:

extract one or more entities identified in a plurality of data artifacts based at least in part on one or more datasets;

extract one or more entities identified in a plurality of code artifacts based at least in part on the one or more datasets;

extract one or more entities identified in a plurality of user interface artifacts based at least in part on the one or more datasets;

generate a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities; and

perform one or more of a lexical analysis and a semantic analysis on the set of dependency graphs to identify a data domain of the one or more datasets.

18. The computer program product of claim 17 , wherein generating a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities comprises generating a first dependency graph of the set of dependency graphs based at least in part on one or more relationships between the extracted one or more entities of the data artifacts and the extracted one or more entities of the code artifacts.

19. The computer program product of claim 18 , wherein generating a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities further comprises generating a second dependency graph of the set of dependency graphs based at least in part on one or more relationships between the extracted one or more entities of the data artifacts and the extracted one or more entities of the user interface artifacts.

20. The computer program product of claim 19 , wherein generating a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities further comprises generating a third dependency graph of the set of dependency graphs based at least in part on one or more relationships between the extracted one or more entities of the code artifacts and the extracted one or more entities of the user interface artifacts.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2021
From: CHETLUR, MALOLAN; AGARWAL, ARVIND; DEY, SUBHENDU; MEHTA, SAMEEP; SARKAR, SANDIPAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058153/0774 →
Continuity (1)
Related Publication 20230153537A1 · May 18, 2023
References Cited (16)
US 6026413A · Challenger et al. · 2000 [cited by applicant]
US 7996413B2 · Cotichini et al. · 2011 [cited by applicant]
US 9589037B2 · Hill · 2017 [cited by examiner]
US 12192230B2 · Glazier · 2025 [cited by examiner]
US 20060136467A1 · Avinash et al. · 2006 [cited by applicant]
US 20220350745A1 · Xia et al. · 2022 [cited by applicant]
CN 107665252A · 2018 [cited by applicant]
EP 3428813A1 · 2019 [cited by applicant]
WO 2018155816A1 · 2018 [cited by applicant]
WO 2021217502A1 · 2021 [cited by applicant]
WO 2004006118A1 · 2022 [cited by applicant]
S. Lalithsena et al., “Automatic Domain Identification for Linked Open Data,” 2013 IEEE/WIC/ACM International Joint Conferences on Web Intelligence (WI) and Intelligent Agent Technologies (IAT), Nov. 2013, 8 pages. [cited by applicant]
A. Abele et al., “Linked Data Profiling: Identifying the Domain of Datasets Based on Data Content and Metadata,” WWWW'16 Companion, Apr. 15, 2016, pp. 287-291. [cited by applicant]
M. Ota et al., “Data-Driven Domain Discovery for Structured Datasets,” Proceedings of the VLDB Endowment, Mar. 2020, pp. 953-965. [cited by applicant]
P. Mell et al., “The NIST Definition of Cloud Computing,” Recommendations of the National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]
International Search Report and Written Opinion of PCT/CN2022/130287, Jan. 18, 2023, 9 pages. [cited by applicant]