IP Library Granted Patent US 12,547,837
Granted Patent B2
US 12,547,837 · App. 18/409,341 · Granted Feb 10, 2026

Artificial intelligence based metadata semantic enrichment

Inventors: Md Faisal Mahbub Chowdhury (Ardsley, NY); Alfio Massimiliano Gliozzo (Brooklyn, NY); Nandana Sampath Mihindukulasooriya (Sunnyside, NY); Michael Robert Glass (Bayonne, NJ); Sarthak Dash (Jersey City, NJ); Sugato Bagchi (White Plains, NY); Gaetano Rossiello (Brooklyn, NY)
Assignee: International Business Machines Corporation
G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,547,837
App. No.
18/409,341
Granted
Feb 10, 2026
Kind
B2
Abstract

Mechanisms are provided for automatically generating semantical enhanced metadata for a structured data structure. Multi-task machine learning training is performed, based on data comprising separate sets of training data samples for each of a plurality of semantic metadata enhancement tasks, of a base artificial intelligence (AI) computer model to thereby generate a fine-tuned AI computer model trained to specifically generate semantically enhanced metadata for structured data structures. A prompt is received that specifies a structure of an input structured data structure and requests a semantic metadata enhancement task from the plurality of semantic metadata enhancement tasks. The fine-tuned AI computer model processes the prompt to generate semantically enhanced metadata for the structure of the input structured data structure and provide it to a downstream computing system for performing a downstream computing operation based on the semantically enhanced metadata.

Claims (37)

1 . A method, in a data processing system, for automatically generating semantical enhanced metadata for a structured data structure, the method comprising:

generating multi-task machine learning training data comprising separate sets of training data samples for each of a plurality of semantic metadata enhancement tasks;

executing multi-task machine learning training of a base artificial intelligence (AI) computer model based on the multi-task machine learning training data to thereby generate a fine-tuned AI computer model trained to specifically generate semantically enhanced metadata for structured data structures, wherein the fine-tuned AI computer model is trained to perform the plurality of semantic metadata enhancement tasks;

receiving a prompt specifying a structure of an input structured data structure and requesting at least one of the semantic metadata enhancement tasks from the plurality of semantic metadata enhancement tasks;

processing, by the fine-tuned AI computer model, the prompt to generate semantically enhanced metadata for the structure of the input structured data structure; and

providing the semantically enhanced metadata to a downstream computing system for performing a downstream computing operation based on the semantically enhanced metadata.

2 . The method of claim 1 , wherein the input structured data structure is a table data structure and the structure of the input structured data structure comprises column names of the table data structure.

3 . The method of claim 2 , wherein the semantic metadata enhancement tasks comprise at least one of generating table captions, generating column descriptions, generating column datatypes, or generating tags for portions of the table data structure.

4 . The method of claim 1 , wherein the base AI computer model is a previously trained foundation model that is trained on a large and varied set of data for general applicability.

5 . The method of claim 4 , wherein the previously trained foundation model is one of a large language model (LLM), a generative adversarial network (GAN), or a variational autoencoder (VAE).

6 . The method of claim 1 , wherein executing multi-task machine learning training of a base artificial intelligence (AI) computer model based on the multi-task machine learning training data comprises generating a plurality of batches of training data samples from the multi-task machine learning training data, wherein each batch comprises a proportional number of training data samples for each of the semantic metadata enhancement tasks.

7 . The method of claim 6 , wherein each batch of training data samples comprises at least two training data samples for opposite semantic metadata enhancement tasks.

8 . The method of claim 1 , wherein the multi-task machine learning training data is generated by curating a larger size training data to a selected subset of training data at least by filtering the larger size training data based on filtering criteria that eliminates structured data structures, or portions of structured data structures, that are determined to be not informative with regard to the structure of the structured data structure.

9 . The method of claim 2 , wherein the downstream computing operation comprises at least one of correlating the table data structure, or portions of the table data structure, with one of dictionaries or ontologies, or performing a table data structure similarity evaluation with one or more other table data structures.

10 . The method of claim 1 , wherein processing the prompt is performed without accessing the data within the structured input structured data structure.

11 . A computer program product comprising a computer readable storage medium having a computer readable program stored therein, wherein the computer readable program, when executed on a computing device, causes the computing device to:

generate multi-task machine learning training data comprising separate sets of training data samples for each of a plurality of semantic metadata enhancement tasks;

execute multi-task machine learning training of a base artificial intelligence (AI) computer model based on the multi-task machine learning training data to thereby generate a fine-tuned AI computer model trained to specifically generate semantically enhanced metadata for structured data structures, wherein the fine-tuned AI computer model is trained to perform the plurality of semantic metadata enhancement tasks;

receive a prompt specifying a structure of an input structured data structure and requesting at least one of the semantic metadata enhancement tasks from the plurality of semantic metadata enhancement tasks;

process, by the fine-tuned AI computer model, the prompt to generate semantically enhanced metadata for the structure of the input structured data structure; and

provide the semantically enhanced metadata to a downstream computing system for performing a downstream computing operation based on the semantically enhanced metadata.

12 . The computer program product of claim 11 , wherein the input structured data structure is a table data structure and the structure of the input structured data structure comprises column names of the table data structure.

13 . The computer program product of claim 12 , wherein the semantic metadata enhancement tasks comprise at least one of generating table captions, generating column descriptions, generating column datatypes, or generating tags for portions of the table data structure.

14 . The computer program product of claim 11 , wherein the base AI computer model is a previously trained foundation model that is trained on a large and varied set of data for general applicability.

15 . The computer program product of claim 14 , wherein the previously trained foundation model is one of a large language model (LLM), a generative adversarial network (GAN), or a variational autoencoder (VAE).

16 . The computer program product of claim 11 , wherein executing multi-task machine learning training of a base artificial intelligence (AI) computer model based on the multi-task machine learning training data comprises generating a plurality of batches of training data samples from the multi-task machine learning training data, wherein each batch comprises a proportional number of training data samples for each of the semantic metadata enhancement tasks.

17 . The computer program product of claim 16 , wherein each batch of training data samples comprises at least two training data samples for opposite semantic metadata enhancement tasks.

18 . The computer program product of claim 11 , wherein the multi-task machine learning training data is generated by curating a larger size training data to a selected subset of training data at least by filtering the larger size training data based on filtering criteria that eliminates structured data structures, or portions of structured data structures, that are determined to be not informative with regard to the structure of the structured data structure.

19 . The computer program product of claim 11 , wherein processing the prompt is performed without accessing the data within the structured input structured data structure.

20 . An apparatus comprising:

at least one processor; and

at least one memory coupled to the at least one processor, wherein the at least one memory comprises instructions which, when executed by the at least one processor, cause the at least one processor to:

generate multi-task machine learning training data comprising separate sets of training data samples for each of a plurality of semantic metadata enhancement tasks;

execute multi-task machine learning training of a base artificial intelligence (AI) computer model based on the multi-task machine learning training data to thereby generate a fine-tuned AI computer model trained to specifically generate semantically enhanced metadata for structured data structures, wherein the fine-tuned AI computer model is trained to perform the plurality of semantic metadata enhancement tasks;

receive a prompt specifying a structure of an input structured data structure and requesting at least one of the semantic metadata enhancement tasks from the plurality of semantic metadata enhancement tasks;

process, by the fine-tuned AI computer model, the prompt to generate semantically enhanced metadata for the structure of the input structured data structure; and

provide the semantically enhanced metadata to a downstream computing system for performing a downstream computing operation based on the semantically enhanced metadata.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2024
From: CHOWDHURY, MD FAISAL MAHBUB; GLIOZZO, ALFIO MASSIMILIANO; MIHINDUKULASOORIYA, NANDANA SAMPATH; GLASS, MICHAEL ROBERT; DASH, SARTHAK; BAGCHI, SUGATO; ROSSIELLO, GAETANO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 066086/0444 →
Continuity (1)
Related Publication 20250225328A1 · Jul 10, 2025
References Cited (27)
US 5752253A · Geymond et al. · 1998 [cited by applicant]
US 7209923B1 · Cooper · 2007 [cited by applicant]
US 10692484B1 · Merritt · 2020 [cited by examiner]
US 11250872B2 · Thomas et al. · 2022 [cited by applicant]
US 11763832B2 · Nesta · 2023 [cited by examiner]
US 11989519B2 · Platt · 2024 [cited by examiner]
US 20200226327A1 · Matusov · 2020 [cited by examiner]
US 20200410981A1 · Merritt · 2020 [cited by examiner]
US 20230325725A1 · Lester · 2023 [cited by examiner]
US 20240378196A1 · Lester · 2024 [cited by examiner]
CN 112559556B · 2021 [cited by applicant]
CN 112988785B · 2021 [cited by applicant]
CN 113986958A · 2022 [cited by applicant]
CN 114547329A · 2022 [cited by applicant]
WO WO2023022727A1 · 2023 [cited by examiner]
WO WO2023026166A1 · 2023 [cited by applicant]
Aghajanyan, Armen et al., “Muppet: Massive Multi-task Representations with Pre-Finetuning”, arXiv:2101.11038v1 [cs.CL], Jan. 26, 2021, 12 pages. [cited by applicant]
Bao, Junwei et al., “Text Generation From Tables”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 27, Issue 2, Feb. 2019, published Oct. 26, 2018, 19 pages. [cited by applicant]
Caruana, Rich, “Multitask Learning”, Machine Learning, vol. 28, No. 1, Jul. 1997, 35 pages. [cited by applicant]
Deng, Xiang et al., “TURL: Table Understanding through Representation Learning”, arXiv:2006.14806v2 [cs.IR], Dec. 3, 2020, 14 pages. [cited by applicant]
Gottumukkala, Ananth et al., “Dynamic Sampling Strategies for Multi-Task Reading Comprehension”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 5-10, 2020, 5 pages. [cited by applicant]
Liu, Shikun et al., “End-to-End Multi-Task Learning with Attention”, IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, Jun. 16-20, 2019, 10 pages. [cited by applicant]
Muennighoff, Niklas et al., “Crosslingual Generalization through Multitask Finetuning”, arXiv:2211.01786v2 [cs.CL], May 29, 2023, 119 pages. [cited by applicant]
Raffel, Colin et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”, Journal of Machine Learning Research, vol. 21, Issue 1, Article 140, Jan. 1, 2020, 67 pages. [cited by applicant]
Ruder, Sebastian, “An Overview of Multi-Task Learning in Deep Neural Networks”, arXiv:1706.05098v1 [cs.LG], Jun. 15, 2017, 14 pages. [cited by applicant]
Weller, Orion et al., “When to Use Multi-Task Learning vs Intermediate Fine-Tuning for Pre-Trained Encoder Transfer Learning”, arXiv:2205.08124v1 [cs.CL], May 17, 2022, 11 pages. [cited by applicant]
Xu, Junjie H. et al., “Table Caption Generation in Scholarly Documents Leveraging Pre-trained Language Models”, 2021 IEEE 10th Global Conference on Consumer Electronics (GCCE), Oct. 12-15, 2021, 5 pages. [cited by applicant]