IP Library › Granted Patent US 12,411,814
Granted Patent B2
US 12,411,814 · App. 17/110,435 · Granted Sep 9, 2025

Metadata based mapping assist

Inventors: Ramkumar Ramalingam (Theni, IN); Subhojeet Pramanik (Kolkata, IN); Jothiponsundar Radhakrishnan (Bangalore, IN); Saptarshi Misra (Kolkata, IN); Nagarjuna Surabathina (Prakasam district, IN); Matu Agarwal (Bangalore, IN)
Assignee: International Business Machines Corporation
G06F16/211G06F16/221G06F16/2237G06F16/24573G06F18/214G06F18/22G06N3/044G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,411,814
App. No.
17/110,435
Granted
Sep 9, 2025
Kind
B2
Abstract

Methods, computer program products, and/or systems are provided that perform the following operations: obtaining source schema metadata, wherein the source schema metadata is associated with fields of a source schema; obtaining target schema metadata with a target schema, wherein the target schema metadata is associated with fields of a target schema; determining, for each field of the source schema and each field of the target schema, a representation for each field based, at least in part, on the source schema metadata or the target schema metadata associated with each field; and providing the representation for each field of the source schema and each field of the target schema for use in generating data mappings between the source schema and the target schema.

Claims (60)

1. A computer-implemented method for data or schema mapping, the computer-implemented method comprising:

obtaining, by a server computer, source schema metadata, wherein the source schema metadata is associated with fields of a source schema;

obtaining, by the server computer, target schema metadata with a target schema, wherein the target schema metadata is associated with fields of a target schema;

determining, by the server computer, for each field of the source schema and each field of the target schema, a representation for each field based, at least in part, on the source schema metadata or the target schema metadata associated with each field, and generating schema field representations for a dynamic object schema, wherein the dynamic object schema enable extending schemas of an existing object, and wherein the determining of the representation for each field of the source schema and each field of the target schema comprises:

training a machine learning model, wherein training the machine learning model comprises:

obtaining a first training dataset;

generating a first stage machine learning model based on the first training dataset, wherein the first stage machine learning model is trained to encode a sentence included in a dataset into a vector representation;

obtaining a second training dataset, wherein the second training dataset includes schema metadata from various schema objects;

generating a trained machine learning model based on the second training dataset and the first stage machine learning model, wherein the trained machine learning model is trained to encode sentences associated with metadata columns of a schema field into a single vector representation, wherein an input layer of the machine learning model utilizes a same encoder to encode a premise and a hypothesis; and

utilizing triplets of related data in an unsupervised fashion to learn sentence representations directly from the metadata columns;

generating, through a sequence neural network, sequence aware embeddings based on the representation for each field by combining the representations from the source schema metadata and the target schema metadata; and

providing, by the server computer, the representation for each field of the source schema and the representation for each field of the target schema for use in generating data mappings between the source schema and the target schema.

2. The computer-implemented method of claim 1 , wherein the representation for each field of the source schema and each field of the target schema comprises a fixed size vector embedding that describes a combination of metadata associated with a particular field of the source schema or the target schema.

3. The computer-implemented method of claim 1 , further comprising:

calculating field-to-field matches between the source schema and the target schema using the representations for the fields of the source schema and the target schema, wherein the representations are generated based on multiple metadata columns associated with a schema field; and

generating data mapping suggestions between fields of the source schema and the target schema based on the field-to-field matches.

4. The computer-implemented method of claim 1 , wherein the representations are determined for fields of a dynamic object schema.

5. The computer-implemented method of claim 1 , wherein the determining the representation for each field of the source schema and each field of the target schema comprises:

providing the source schema metadata or the target schema metadata associated with each field individually as input to a machine learning model, the machine learning model having been trained to generate the representation for each field based on one or more metadata columns associated with each field; and

receiving, for each field as output from the machine learning model, a vector representation that describes a combination of the one or more metadata columns associated with each field as the representation for each field and be used as the representation for that field in determining matches between fields of the source and target schemas for use in data mapping.

6. The computer-implemented method of claim 5 , wherein the machine learning model is trained to identify relevant information from multiple metadata columns associated with a schema field to generate a single vector representation for the schema field.

7. The computer-implemented method of claim 1 , further comprising:

generating a confidence score for each field of the target schema as compared to each field of the source schema, wherein the confidence score for each field of the target schema is based on comparing the representations for each field of the source schema to the representation for the field of the target schema, and an overall calculated score between two fields using cosine similarity; and

generating data mapping suggestions between the source schema and the target schema for the fields of the target schema based, at least in part, on the confidence scores.

8. The computer-implemented method of claim 7 , wherein the generating the confidence score for each field of the target schema comprises calculating an overall confidence score between two fields using cosine similarity.

9. A computer-implemented method comprising:

determining, by a server computer, for each field of a source schema and each field of a target schema, a representation for each field based, at least in part, on source schema metadata or target schema metadata associated with each field, and generating schema field representations for a dynamic object schema, wherein the dynamic object schema enable extending schemas of an existing object, and wherein the determining of the representation for each field of the source schema and each field of the target schema comprises:

training a machine learning model, wherein training the machine learning model comprises:

obtaining a first training dataset, wherein the first training dataset is a sentence related dataset;

generating a first stage machine learning model based on the first training dataset, wherein the first stage machine learning model is trained to encode a sentence included in a dataset into a vector representation, wherein the machine learning model generates fixed length sentence representations by utilizing a natural language inference dataset;

obtaining a second training dataset, wherein the second training dataset includes schema metadata from various schema objects;

generating a trained machine learning model based on the second training dataset and the first stage machine learning model, wherein the trained machine learning model is trained to encode metadata columns associated with a schema field into a single vector representation and to learn domain specific semantic representations relating to a schema domain, wherein an input layer of the machine learning model utilizes a same encoder to encode a premise and a hypothesis; and

utilizing triplets of related data in an unsupervised fashion to learn sentence representations directly from the metadata columns

providing the trained machine learning model for use in generating data mapping suggestions between the source schema and the target schema.

10. The computer-implemented method of claim 9 , wherein the generating the trained machine learning model for use in generating data mapping does not require any prior schema mapping data.

11. The computer-implemented method of claim 9 , wherein the second training dataset includes schema metadata that describes fields included in a schema.

12. A computer program product for data or schema mapping comprising a computer readable storage medium having stored thereon:

program instructions programmed to obtain source schema metadata, wherein the source schema metadata is associated with fields of a source schema;

program instructions programmed to obtain, by a server computer, target schema metadata with a target schema, wherein the target schema metadata is associated with fields of a target schema;

program instructions programmed to determine, by the server computer, for each field of the source schema and each field of the target schema, a representation for each field based, at least in part, on the source schema metadata or the target schema metadata associated with each field, and generating schema field representations for a dynamic object schema, wherein the dynamic object schema enable extending schemas of an existing object, and wherein the determining of the representation for each field of the source schema and each field of the target schema comprises:

program instructions programmed to train a machine learning model to learn domain specific semantic representations relating to a schema domain, wherein training the machine learning model comprises:

program instructions programmed to obtain a first training dataset;

program instructions programmed to generate a first stage machine learning model based on the first training dataset, wherein the first stage machine learning model is trained to encode a sentence included in a dataset into a vector representation;

program instructions programmed to obtain a second training dataset, wherein the second training dataset includes schema metadata from various schema objects;

program instructions programmed to generate a trained machine learning model based on the second training dataset and the first stage machine learning model, wherein the trained machine learning model is trained to encode sentences associated with metadata columns of a schema field into a single vector representation, wherein an input layer of the machine learning model utilizes a same encoder to encode a premise and a hypothesis, and wherein the trained machine learning model does not require use of prior schema mapping data; and

utilizing triplets of related data in an unsupervised fashion to learn sentence representations directly from the metadata columns;

program instructions programmed to generating, through a sequence neural network, sequence aware embeddings based on the representation for each field by combining the representations from the source schema metadata and the target schema metadata; and

program instructions programmed to provide, by the server computer, the representation for each field of the source schema and the representation for each field of the target schema for use in generating data mappings between the source schema and the target schema.

13. The computer program product of claim 12 , wherein the representation for each field of the source schema and each field of the target schema comprises a fixed size vector embedding that describes a combination of metadata associated with a particular field of the source schema or the target schema.

14. The computer program product of claim 12 , the computer readable storage medium having further stored thereon:

program instructions programmed to calculate field-to-field matches between the source schema and the target schema using the representations for the fields of the source schema and the target schema, wherein the representations are generated based on multiple metadata columns associated with a schema field; and

program instructions programmed to generate data mapping suggestions between fields of the source schema and the target schema based on the field-to-field matches.

15. The computer program product of claim 12 , wherein the program instructions programmed to determine, for each field of the source schema and each field of the target schema, the representation for each field, further comprise:

program instructions programmed to provide the source schema metadata or the target schema metadata associated with each field individually as input to a machine learning model, the machine learning model having been trained to generate the representation for each field based on one or more metadata columns associated with each field; and

program instructions programmed to receive, for each field as output from the machine learning model, a vector representation that describes a combination of the one or more metadata columns associated with each field as the representation for each field.

16. The computer program product of claim 15 , wherein the machine learning model is trained to identify relevant information from multiple metadata columns associated with a schema field to generate a single vector representation for the schema field.

17. The computer program product of claim 12 , the computer readable storage medium having further stored thereon:

program instructions programmed to generate a confidence score for each field of the target schema as compared to each field of the source schema, wherein the confidence score for each field of the target schema is based on comparing the representations for each field of the source schema to the representation for the field of the target schema; and

program instructions programmed to generate data mapping suggestions between the source schema and the target schema for the fields of the target schema based, at least in part, on the confidence scores.

18. The computer program product of claim 17 , wherein the generating the confidence score for each field of the target schema comprises calculating an overall confidence score between two fields using cosine similarity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2020
From: RAMALINGAM, RAMKUMAR; PRAMANIK, SUBHOJEET; RADHAKRISHNAN, JOTHIPONSUNDAR; MISRA, SAPTARSHI; SURABATHINA, NAGARJUNA; AGARWAL, MATU
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054527/0257 →
Continuity (1)
Related Publication 20220179833A1 · Jun 9, 2022
References Cited (28)
US 8442999B2 · Gorelik · 2013 [cited by applicant]
US 8954375B2 · Kehoe · 2015 [cited by applicant]
US 20040205573A1 · Carlson · 2004 [cited by examiner]
US 20190286978A1 · Aggarwal · 2019 [cited by applicant]
US 20190318272A1 · Sassin · 2019 [cited by examiner]
US 20200051550A1 · Baker · 2020 [cited by examiner]
US 20200081899A1 · Shapur · 2020 [cited by applicant]
US 20200104746A1 · Strope · 2020 [cited by examiner]
US 20200334520A1 · Chen · 2020 [cited by examiner]
US 20210126881A1 · Ball · 2021 [cited by examiner]
US 20210232908A1 · Xian · 2021 [cited by examiner]
US 20220100772A1 · Kadarundalagi Raghura · 2022 [cited by examiner]
CN 109858032A · 2019 [cited by examiner]
CN 110532399A · 2019 [cited by examiner]
“Anypoint Platform—One platform for APIs and integrations”, MuleSoft, retrieved from the internet on Aug. 17, 2020, 7 pages, <https://www.mulesoft.com/platform/enterprise-integration>. [cited by applicant]
“Boomi Suggest and Your Privacy”, boomi, A Dell Technology Business, retrieved from the Internet on Aug. 17, 2020, 2 pages, <https://boomi.com/privacy/suggest/>. [cited by applicant]
“Facebookresearch / InferSent”, GitHub, retrieved from the Internet on Aug. 19, 2020, 3 pages, <https://github.com/facebookresearch/InferSent>. [cited by applicant]
“IBM Cloud Pak for Integration 2020.2.1 adds support for IBM Cloud Transformation Advisor and IBM App Connect Enterprise Mapping Assist”, IBM United States Software Announcement 220-300 Jun. 23, 2020, 10 pages, Grace Pe… [cited by applicant]
“Schema Matching using Machine Learning”, Author Identity hidden due to Anonymous Submission, 7 pages, provided by the inventors on Jun. 1, 2020, <https://pdfs.semanticscholar.org/0786/c8586bc4f3491a394b07975b7d7780a1ce… [cited by applicant]
“The Stanford Natural Language Inference (SNLI) Corpus”, The Stanford Natural Language Processing Group, retrieved from the Internet on Aug. 17, 2020, 7 pages, <https://nlp.stanford.edu/projects/snli/>. [cited by applicant]
Agarwal, Matu, “IBM App Connect ‘Mapping Assist’—Putting AI to Work for Your Business”, Published on Jun. 24, 2020 / Updated on Jun. 26, 2020, 6 pages, <https://developer.ibm.com/integration/blog/2020/06/24/ibm-app-conn… [cited by applicant]
Berlin et al., “Database Schema Matching Using Machine Learning with Feature Selection”, CAISE 2002, LNCS 2348, pp. 452-466, 2002, A. Banks Pidduck et al. (Eds.), <https://link.springer.com/content/pdf/10.1007%2F3-540-4… [cited by applicant]
Conneau et al., “Supervised Learning of Universal Sentence Representations from Natural Language Inference Data”, arXiv:1705.02364v5 [cs.CL] Jul. 8, 2018, 12 pages. [cited by applicant]
Fatima, Nida, “Understanding Data Mapping and its Techniques”, Astera, Aug. 7, 2020, 6 pages. [cited by applicant]
Heinzerling et al., “BPEmb: Subword Embeddings in 275 Languages”, 2018, 5 pages, <https://nlp.h-its.org/bpemb/>. [cited by applicant]
Lu et al., “A Deep Architecture for Matching Short Texts”, Advances in Neural Information Processing Systems 26 (NIPS 2013), 10 pages, <http://www.hangli-hl.com/uploads/3/4/4/6/34465961/nips_2013.pdf>. [cited by applicant]
Rahm et al., “A survey of approaches to automatic schema matching”, The VLDB Journal 10: pp. 334-350 (2001), Digital Object Identifier (DOI) 10.1007/s007780100057, Published online: Nov. 21, 2001—copyright Springer-Verl… [cited by applicant]
Rahm et al., “On Matching Schemas Automatically”, Microsoft Research, Microsoft Corporation, Feb. 2001, Technical Report MSR-TR-2001-17, 22 pages. [cited by applicant]