IP Library › Granted Patent US 12,619,683
Granted Patent B2
US 12,619,683 · App. 17/411,410 · Granted May 5, 2026

Artificial intelligence (AI) based data matching and alignment

Inventors: Neda Abolhasssani (San Mateo, CA); Maziyar Baran Pouyan (Emeryville, CA); Teresa Sheausan Tung (Tustin, CA); Andrew Fano (Lincolnshire, IL); Sayantan Mitra (Bangalore, IN)
Assignee: ACCENTURE GLOBAL SOLUTIONS LIMITED
G06F18/24147G06F18/24143G06F40/30G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,683
App. No.
17/411,410
Granted
May 5, 2026
Kind
B2
Abstract

An Artificial Intelligence (AI)-based data matching and alignment system identifies similar data sources for a target data source from a data corpus and generates a knowledge graph that enables downstream applications seamless access to data in the data corpus. The system extracts column features at different levels for the target data source and a plurality of data sources from the data corpus. Feature matrices are built from the features of the target data source and the plurality of data sources. Candidate data sources similar to the target data source are filtered from the plurality of data sources using the feature matrices. The tree-based similarity is estimated and K Nearest Neighbor (KNN) graphs are built to identify columns from the candidate data sources that are similar to columns of the target data source to build the knowledge graph.

Claims (63)

1 . An Artificial Intelligence (AI) based data matching and alignment system, comprising:

at least one processor;

a non-transitory processor-readable medium storing machine-readable instructions that cause the processor to:

receive a request for identifying similar data for a target data source from a plurality of data sources;

generate feature matrices for the target data source and the plurality of data sources, wherein each feature matrix of the feature matrices includes respective features of the target data source and the plurality of data sources;

identify at least one candidate data source that is similar to the target data source from the plurality of data sources, wherein the at least one candidate data source is identified based on corresponding feature matrices of the plurality of data sources and the feature matrix of the target data source, wherein the identification of the at least one candidate data source provides for preliminary filtering to select a subset of the plurality of data sources;

identify columns of the at least one candidate data source that are similar to columns of the target data source by matching columns of the target data source with columns of the subset of the plurality of data sources and by applying tree-based similarity calculations to a feature matrix of the at least one candidate data source and the feature matrix of the target data source, wherein the similarities are determined based at least on feature matrices of the target data source and the features of the at least one candidate data sources;

enable a display of one or more of the columns of the at least one candidate data source that are similar columns to the columns of the target data source;

generate a knowledge graph that represents the similar columns of the at least one candidate data source and the target data source; and

enable functioning of a downstream application to access the knowledge graph for information extraction,

wherein enabling the downstream application to access the knowledge graph comprising:

obtaining output that includes at least one of:

a portion of the knowledge graph, the knowledge graph representing the columns as nodes, wherein the similar columns are connected by edges of the knowledge graph, and a distance between the nodes signifies a column similarity; and

a ranked list of similarity mappings between the columns of the target data source and the candidate data sources along with respective similarity percentages for the similarity mappings; and

executing the functions of the downstream application based on the obtained output.

2 . The data matching and alignment system of claim 1 , wherein the features include character level features, semantic level features, and data dependency features.

3 . The data matching and alignment system of claim 1 , wherein the character level features include at least column data type features and character distribution features.

4 . The data matching and alignment system of claim 1 , wherein the semantic level features include at least semantic text features and numeric distribution comparison features.

5 . The data matching and alignment system of claim 1 , wherein to identify the at least one candidate data source similar to the target data source, the processor is to:

generate a K Nearest Neighbor (KNN) graph on implementing a distance metric to the corresponding feature matrices of the plurality of data sources and the feature matrix of the target data source.

6 . The data matching and alignment system of claim 5 , wherein the distance metric includes Mahanalobis distance.

7 . The data matching and alignment system of claim 1 , wherein to identify the at least one candidate data source similar to the target data source, the processor is to:

select nearest N neighbors from the KNN graph as a plurality of candidate data sources, wherein N is a natural number and the at least one candidate data source includes the plurality of candidate data sources.

8 . The data matching and alignment system of claim 1 , wherein to identify the columns of the at least one candidate data source that are similar to the columns of the target data source, the processor is to:

generate K Nearest Neighbor (KNN) graphs from the tree-based similarity calculations; and

identify the columns of the at least one candidate data source that are similar to the columns of the target data source from the KNN graphs.

9 . A method of generating similarity mappings between data sources comprising:

receiving a request for identifying matching data for a target data source of a plurality of data sources from a data corpus;

extracting column features of the target data source and the plurality of data sources, wherein the column features are stored as corresponding feature matrices;

identifying one or more candidate data sources from the plurality of data sources that are similar to the target data source, wherein the candidate data sources are identified based on a distance measure obtained for the feature matrix of the target data source and the corresponding feature matrices of the plurality of data sources, wherein the identification of the one or more candidate data source provides for preliminary filtering to select a subset of the plurality of data sources;

identifying columns of the one or more candidate data sources that are similar to columns of the target data source by matching columns of the target data source with columns of the subset of the plurality of data sources and by applying tree-based similarity calculations to a feature matrix of the one or more candidate data source and the feature matrix of the target data source, wherein the similarities between the columns are determined based at least on feature matrices of the target data source and the features of one of the one or more candidate data sources;

generating a knowledge graph representing the similarities of the columns of the one or more candidate data sources and the columns of the target data source; and

enabling functioning of a downstream application by enabling the downstream application to access the knowledge graph,

wherein enabling the downstream application to access the knowledge graph comprising:

obtaining output that includes at least one of:

a portion of the knowledge graph, the knowledge graph representing the columns as nodes, wherein the similar columns are connected by edges of the knowledge graph, and a distance between the nodes signifies a column similarity; and

a ranked list of similarity mappings between the columns of the target data source and the candidate data sources along with respective similarity percentages for the similarity mappings; and

executing the functions of the downstream application based on the obtained output.

10 . The method of claim 9 , further comprising:

providing a display of the columns of the one or more candidate data sources that are similar to the columns of the target data source.

11 . The method of claim 10 , further comprising:

providing via the display, a percentage of similarity between each of the similar columns of the one or more candidate data sources and the columns of the target data source.

12 . The method of claim 10 , further comprising:

receiving user input selecting one or more of the similar columns for generating the knowledge graph, wherein the user input is received via the display.

13 . The method of claim 9 , wherein the column features include at least character level features, semantic level features, and dependency level features.

14 . The method of claim 13 , further comprising:

generating the feature matrices for the plurality of data sources including the target data source, by stacking the character level features, semantic level features, and dependency level features of each of the columns adjacent to each other.

15 . The method of claim 9 , further comprising:

outputting reasons for identifying the columns of the one or more candidate data sources as being similar to the columns of the target data source.

16 . A non-transitory processor-readable storage medium comprising machine-readable instructions that cause a processor to:

receiving a request for identifying matching data for a target data source of a plurality of data sources from a data corpus;

extracting column features of the target data source and the plurality of data sources, wherein the column features are stored as corresponding feature matrices;

identifying one or more candidate data sources from the plurality of data sources that are similar to the target data source, wherein the candidate data sources are identified based on a distance measure obtained for the feature matrix of the target data source and the corresponding feature matrices of the plurality of data sources, wherein the identification of the one or more candidate data source provides for preliminary filtering to select a subset of the plurality of data sources;

identifying columns of the one or more candidate data sources that are similar to columns of the target data source by matching columns of the target data source with columns of the subset of the plurality of data sources and by applying tree-based similarity calculations to a feature matrix of the one or more candidate data source and the feature matrix of the target data source, wherein the similarities between the columns are determined based at least on feature matrices of the target data source and the one or more candidate data sources;

generating a knowledge graph representing the similarities of the columns of the one or more candidate data sources and the columns of the target data source; and

enabling functioning of a downstream application by enabling the downstream application to access the knowledge graph,

wherein enabling the downstream application to access the knowledge graph comprising:

obtaining output that includes at least one of:

a portion of the knowledge graph, the knowledge graph representing the columns as nodes, wherein the similar columns are connected by edges of the knowledge graph, and a distance between the nodes signifies a column similarity; and

a ranked list of similarity mappings between the columns of a target data source and the candidate data sources along with respective similarity percentages for the similarity mappings; and

executing the functions of the downstream application based on the obtained output.

17 . The non-transitory processor-readable storage medium of claim 16 , wherein the instructions to identify the at least one candidate data source as similar to the target data source, further cause the processor to:

apply K Nearest Neighbor (KNN) methodology on implementing the distance measure to the corresponding feature matrices of the plurality of data sources and the feature matrix of the target data source.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2021
From: ABOLHASSSANI, NEDA; POUYAN, MAZIYAR BARAN; TUNG, TERESA SHEAUSAN; FANO, ANDREW; MITRA, SAYANTAN
To: ACCENTURE GLOBAL SOLUTIONS LIMITED
Reel/Frame 058072/0480 →
Priority Claims (1)
IN 202111020371 · May 4, 2021 · national
Continuity (1)
Related Publication 20220358336A1 · Nov 10, 2022
References Cited (18)
US 9436760B1 · Tacchi · 2016 [cited by examiner]
US 20060058624A1 · Kimura · 2006 [cited by examiner]
US 20070203908A1 · Wang · 2007 [cited by examiner]
US 20110246466A1 · Ekin · 2011 [cited by examiner]
US 20120254143A1 · Varma · 2012 [cited by examiner]
US 20130013291A1 · Bullock · 2013 [cited by examiner]
US 20150169758A1 · Assom et al. · 2015 [cited by applicant]
US 20150178272A1 · Geigel · 2015 [cited by examiner]
US 20170221240A1 · Stetson et al. · 2017 [cited by applicant]
US 20180101800A1 · Lecue · 2018 [cited by examiner]
US 20190278777A1 · Malik · 2019 [cited by examiner]
US 20190278850A1 · Atasu · 2019 [cited by examiner]
US 20190384571A1 · Oberbreckling · 2019 [cited by examiner]
US 20210192364A1 · Wang · 2021 [cited by examiner]
US 20220269936A1 · Zhu · 2022 [cited by examiner]
Kayva Chandra, “Classify your data using Azure Purview”, Dec. 9, 2020, (5 pages). [cited by applicant]
“Entity Matching”, https://docs.cognite.com/cdf/integration/guides/contextualization/matching.html, (6 pages). [cited by applicant]
Christos Koutras et al., “REMA: Graph Embeddings-based Relational Schema Matching”, Delft University of Technology, 2020, (4 page). [cited by applicant]