IP Library › Granted Patent US 12,530,333
Granted Patent B2
US 12,530,333 · App. 18/045,030 · Granted Jan 20, 2026

Structural data matching using neural network encoders

Inventors: Rajalingappaa Shanmugamani (Singapore, SG); Jiaxuan Zhang (Singapore, SG)
Assignee: SAP SE
G06F16/2237G06F16/221G06F16/2264G06F16/248G06F16/283G06N3/045G06N3/08H03M7/3082G06N3/082G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,333
App. No.
18/045,030
Granted
Jan 20, 2026
Kind
B2
Abstract

Implementations of the present disclosure include methods, systems, and computer-readable storage mediums for receiving first and second data sets, both the first and second data sets including structured data in a plurality of columns, for each of the first data set and the second data set, inputting each column into an encoder specific to a column type of a respective column, the encoder providing encoded data for the first data set, and the second data set, respectively, providing a first multi-dimensional vector based on encoded data of the first data set, providing a second multi-dimensional vector based on encoded data of the second data set, and outputting the first multi-dimensional vector and the second multi-dimensional vector to a loss-function, the loss-function processing the first multi-dimensional vector and the second multi-dimensional vector to provide an output, the output representing matched data points between the first and second data sets.

Claims (40)

1 . A computer-implemented method executed by one or more processors, the method comprising:

receiving a first data set and a second data set, both the first data set and the second data set comprising structured data in a plurality of columns;

pre-processing data values of each column of each of the first data set and the second data set, such that data values within individual columns are of a same length in terms of number of characters;

for each of the first data set and the second data set, inputting each column into an encoder specific to a column type of a respective column, the encoder providing encoded data for the first data set, and the second data set, respectively;

providing a first multi-dimensional vector based on encoded data of the first data set by mapping a first output of first fully connected layers to a latent space independently of a second output of second fully connected layers;

providing a second multi-dimensional vector based on encoded data of the second data set by mapping the second output of the second fully connected layers to the latent space independently of the first output of the first fully connected layers; and

outputting the first multi-dimensional vector and the second multi-dimensional vector to a loss-function, the loss function being computed during a supervised training process using triplet mining comprising anchor points, positive points that match respective anchor points, and negative points that are non-matching to respective anchor points, the loss-function processing the first multi-dimensional vector and the second multi-dimensional vector to provide an output, the output representing an exact match between a data point of the first data set and a data point of the second data set.

2 . The method of claim 1 , wherein a same encoder is used to provide the encoded data of the first data set, and the encoded data of the second data set.

3 . The method of claim 1 , wherein, prior to the encoder providing encoded data, data values of one or more of the first data set, and the second data set are pre-processed to provide revised data values.

4 . The method of claim 3 , wherein pre-processing comprises pre-appending one or more zeros to a numerical data value.

5 . The method of claim 3 , wherein pre-processing comprises pre-appending one or more spaces to a string data value.

6 . The method of claim 1 , further comprising filtering at least one column from each of the first data set, and the second data set prior to providing encoded data.

7 . The method of claim 1 , further comprising determining a column type for each column of the plurality of columns.

8 . A non-transitory computer-readable storage medium coupled to one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations, the operations comprising:

receiving a first data set and a second data set, both the first data set and the second data set comprising structured data in a plurality of columns;

pre-processing data values of each column of each of the first data set and the second data set, such that data values within individual columns are of a same length in terms of number of characters;

for each of the first data set and the second data set, inputting each column into an encoder specific to a column type of a respective column, the encoder providing encoded data for the first data set, and the second data set, respectively;

providing a first multi-dimensional vector based on encoded data of the first data set by mapping a first output of first fully connected layers to a latent space independently of a second output of second fully connected layers;

providing a second multi-dimensional vector based on encoded data of the second data set by mapping the second output of the second fully connected layers to the latent space independently of the first output of the first fully connected layers; and

outputting the first multi-dimensional vector and the second multi-dimensional vector to a loss-function, the loss function being computed during a supervised training process using triplet mining comprising anchor points, positive points that match respective anchor points, and negative points that are non-matching to respective anchor points, the loss-function processing the first multi-dimensional vector and the second multi-dimensional vector to provide an output, the output representing an exact match between a data point of the first data set and a data point of the second data set.

9 . The computer-readable storage medium of claim 8 , wherein a same encoder is used to provide the encoded data of the first data set, and the encoded data of the second data set.

10 . The computer-readable storage medium of claim 8 , wherein, prior to the encoder providing encoded data, data values of one or more of the first data set, and the second data set are pre-processed to provide revised data values.

11 . The computer-readable storage medium of claim 10 , wherein pre-processing comprises pre-appending one or more zeros to a numerical data value.

12 . The computer-readable storage medium of claim 10 , wherein pre-processing comprises pre-appending one or more spaces to a string data value.

13 . The computer-readable storage medium of claim 8 , wherein operations further comprise filtering at least one column from each of the first data set, and the second data set prior to providing encoded data.

14 . The computer-readable storage medium of claim 8 , wherein operations further comprise determining a column type for each column of the plurality of columns.

15 . A system, comprising:

a computing device; and

a computer-readable storage device coupled to the computing device and having instructions stored thereon which, when executed by the computing device, cause the computing device to perform operations, the operations comprising:

receiving a first data set and a second data set, both the first data set and the second data set comprising structured data in a plurality of columns;

pre-processing data values of each column of each of the first data set and the second data set, such that data values within individual columns are of a same length in terms of number of characters;

for each of the first data set and the second data set, inputting each column into an encoder specific to a column type of a respective column, the encoder providing encoded data for the first data set, and the second data set, respectively;

providing a first multi-dimensional vector based on encoded data of the first data set by mapping a first output of first fully connected layers to a latent space independently of a second output of second fully connected layers;

providing a second multi-dimensional vector based on encoded data of the second data set by mapping the second output of the second fully connected layers to the latent space independently of the first output of the first fully connected layers; and

outputting the first multi-dimensional vector and the second multi-dimensional vector to a loss-function, the loss function being computed during a supervised training process using triplet mining comprising anchor points, positive points that match respective anchor points, and negative points that are non-matching to respective anchor points, the loss-function processing the first multi-dimensional vector and the second multi-dimensional vector to provide an output, the output representing an exact match between a data point of the first data set and a data point of the second data set.

16 . The system of claim 15 , wherein a same encoder is used to provide the encoded data of the first data set, and the encoded data of the second data set.

17 . The system of claim 15 , wherein, prior to the encoder providing encoded data, data values of one or more of the first data set, and the second data set are pre-processed to provide revised data values.

18 . The system of claim 17 , wherein pre-processing comprises pre-appending one or more zeros to a numerical data value.

19 . The system of claim 17 , wherein pre-processing comprises pre-appending one or more spaces to a string data value.

20 . The system of claim 15 , wherein operations further comprise filtering at least one column from each of the first data set, and the second data set prior to providing encoded data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2022
From: SHANMUGAMANI, RAJALINGAPPAA; ZHANG, JIAXUAN
To: SAP SE
Reel/Frame 061408/0038 →
Continuity (2)
Continuation 15937216 · Mar 27, 2018
Related Publication 20230059579A1 · Feb 23, 2023
References Cited (38)
US 10127495B1 · Bopardikar · 2018 [cited by examiner]
US 10726016B2 · Chavan · 2020 [cited by examiner]
US 10970629B1 · Dirac et al. · 2021 [cited by applicant]
US 20060149692A1 · Hercus · 2006 [cited by examiner]
US 20100030796A1 · Netz · 2010 [cited by examiner]
US 20120150531A1 · Bangalore · 2012 [cited by examiner]
US 20160267397A1 · Carlsson · 2016 [cited by applicant]
US 20170032035A1 · Gao · 2017 [cited by examiner]
US 20170262491A1 · Brewster · 2017 [cited by examiner]
US 20180096000A1 · Harrison · 2018 [cited by examiner]
US 20180240243A1 · Kim et al. · 2018 [cited by applicant]
US 20180314938A1 · Andoni · 2018 [cited by examiner]
US 20190005313A1 · Vemulapalli · 2019 [cited by examiner]
US 20190155904A1 · Santos et al. · 2019 [cited by applicant]
US 20190179896A1 · Anisimovich · 2019 [cited by examiner]
US 20190228312A1 · Andoni · 2019 [cited by examiner]
US 20190303465A1 · Shanmugamani et al. · 2019 [cited by applicant]
CN 105210064 · 2015 [cited by applicant]
CN 105719001 · 2016 [cited by applicant]
EP 1995878 · 2008 [cited by applicant]
Baxter [online], “How to Match Similar Data Tables in Excel with Fuzzy Lookup,” Builtvisible, Nov. 16, 2017, [retrieved on Mar. 26, 2019], retrieved from: URL<https://builtvisible.com/match- similar-not-exact-data-point… [cited by applicant]
Bellet et al. [online], “A Survey on Metric Learning for Feature Vectors and Structured Data,” Arxiv.org: arXiv preprint arXiv: 1306.6709, Jun. 27, 2013, [retrieved on: Mar. 25, 2019], retrieved from: URL<https://arxiv.… [cited by applicant]
Chatterjee et al. [online], “Similarity Learning with (or without) Convolutional Neural Network,” CS 598 LAZ: Cutting-Edge Trends in Deep Learning and Recognition, Feb. 16, 2017, [retrieved on Mar. 25, 2019, retrieved f… [cited by applicant]
Communication Pursuant to Article 94 (3) EPC issued in European Application No. 18196739.9 on Apr. 13, 2021, 9 pages. [cited by applicant]
Costa, “Probabilistic Interpretation of Feedforward Network Outputs, with Relationships to Statistical Prediction of Ordinal Quantities,” International Journal of Neural Systems, vol. 7, No. 5, Nov. 1996, 14 pages. [cited by applicant]
Extended European Search Report issued in European Application No. 18196739.9 on Apr. 10, 2019, 13 pages. [cited by applicant]
Final Office Action in U.S. Appl. No. 15/937,216, dated Aug. 27, 2021, 59 pages. [cited by applicant]
Final Office Action in U.S. Appl. No. 15/937,216, dated Oct. 19, 2020, 33 pages. [cited by applicant]
Li et al. [online], “Generative Moment Matching Networks,” Arxiv.org: arXiv:1502.02761v1g, Feb. 10, 2015, [retrieved on: Mar. 25, 2019], retrieved from: URL<https://arxiv.org/abs/1502.02761>, 9 pages. [cited by applicant]
Non-Final Office Action in U.S. Appl. No. 15/937,216, dated Apr. 16, 2020, 51 pages. [cited by applicant]
Non-Final Office Action in U.S. Appl. No. 15/937,216, dated Mar. 9, 2021, 39 pages. [cited by applicant]
Schroff et al., “FaceNet: A Unified Embedding for Face Recognition and Clustering,” Proceedings of the IEEE conference on computer vision and pattern recognition, Jun. 2015, 9 pages. [cited by applicant]
U.S. Appl. No. 16/208,681, Saito et al., “Representing Sets of Entitites for Matching Problems,” filed on Dec. 4, 2018. [cited by applicant]
U.S. Appl. No. 16/210,070, Le et al., “Graphical Approach To Multi-Matching,” filed Dec. 5, 2018. [cited by applicant]
U.S. Appl. No. 16/217,148, Saito et al., “Utilizing Embeddings for Efficient Matching of Entities, ” filed Dec. 12, 2018. [cited by applicant]
Wikipedia.org [online], “Microsoft Excel—Wikipedia” Mar. 2018, [retrieved on Apr. 7, 2021], retrieved from: URL <https://en.wikipedia.org/w/index.php?title=Microsoft_Excel&oldid=830205165>, 23 pages. [cited by applicant]
Zhang et al., “Character-level Confultional Networks for Text Classification,” Advances in neural information processing systems, Sep. 2015, 9 pages. [cited by applicant]
Office Action in Chinese Appln. No. 201811148169.6, dated Nov. 10, 2023, 8 pages (with English translation). [cited by applicant]