IP Library Granted Patent US 12,625,870
Granted Patent B2
US 12,625,870 · App. 18/955,028 · Granted May 12, 2026

Systems and methods for associating data entries

Inventors: Lu Zhang (Bellevue, WA); Nichole Haas (Monroe, WA); Joshua Manoj (Bellevue, WA); Sri Raja Harshini Koka (Bellevue, WA)
Assignee: SAP SE
G06F16/24558G06F16/2379
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,625,870
App. No.
18/955,028
Granted
May 12, 2026
Kind
B2
Abstract

In one embodiment, a first entry in a first database is modified to include data from a highest-ranked one of one or more available data tables that correspond to the first entry. Each of one or more characters fields of the modified first entry are converted into a respective one or more first-entry tokens, and each of one or more character fields of each of a plurality of second entries in a second database is converted into a respective one or more second-entry tokens. The first-entry tokens are compared to the second-entry tokens, and, in response to the comparison, it is determined whether the first entry matches one of the second entries. In response to determining that the first entry matches one of the second entries, the first entry and the matching second entry are associated with one another in one or both the first and second databases.

Claims (58)

1 . A method, comprising:

loading first data entries from an expense report database and an itinerary database into a data entry store;

converting, by a tokenizer module of a computing device, the first data entries in the data entry store into one or more first-entry tokens;

assigning, by a weighting engine executed by the computing device, a weight to each of the one or more first-entry tokens based on a frequency of the of the first-entry tokens;

filtering, by a filtering component of the computing device, the first data entries to remove the first data entries in expense reports that are not associated with a travel segment expense;

converting, by the tokenizer, each of one or more character fields of each of a plurality of second entries in a second database into a respective one or more second-entry tokens and weighting the second-entry tokens based on a frequency of the second-entry tokens;

comparing, by a comparison engine of the computing device, the first-entry tokens to the second-entry tokens;

determining, by the computing device, whether the first data entries matches one of the second entries based on the comparing and weights of the first-entry tokens and the second-entry tokens, wherein determining that the first data entries matches one of the plurality of second entries is further in response to:

each of a first plurality of character fields of the first data entries exactly matching a respective character field of the one of the plurality of second entries; and

each of at least two of a second plurality of character fields of the first data entries at least partially matching a respective character field of the one of the plurality of second entries;

associating, by the computing device, a first data entry with one of the second entries in response to determining that the first data entry matches the one of the second entries; and

determining that the first data entry matches one of the plurality of second entries.

2 . The method of claim 1 wherein the filtering removes first data entries that have been submitted but have not been approved.

3 . The method of claim 1 wherein the filtering removes first data entries that have not been submitted for approval.

4 . The method of claim 1 wherein the filtering removes first data entries that are already associated with a travel segment entry in the itinerary database.

5 . The method of claim 1 wherein converting each of one or more character fields of each of the plurality of second entries in the second database into a respective one or more second-entry tokens includes converting each of the one or more character fields into a respective plurality of overlapping second-entry tokens.

6 . The method of claim 1 , wherein determining that the first data entries match one of the plurality of second entries is further in response to:

each of a first plurality of character fields of the first data entries exactly matching a respective character field of the one of the plurality of second entries;

each of at least one of a second plurality of character fields of the first data entries at least partially matching a respective character field of the one of the plurality of second entries; and

each of at least one of a third plurality of character fields of the first data entries at least partially matching a respective character field of the one of the plurality of second entries.

7 . A system, comprising:

one or more processors; and

a non-transitory machine-readable medium storing a program executable by the one or more processors, the program comprising sets of instructions for:

loading first data entries from an expense report database and an itinerary database into a data entry store;

converting, by a tokenizer module, the first data entries in the data entry store into one or more first-entry tokens;

assigning, by a weighting engine, a weight to each of the one or more first-entry tokens based on a frequency of the of the first-entry tokens;

filtering, by a filtering component, the first data entries to remove the first data entries in expense reports that are not associated with a travel segment expense;

converting, by the tokenizer, each of one or more character fields of each of a plurality of second entries in a second database into a respective one or more second-entry tokens and weighting the second-entry tokens based on a frequency of the second-entry tokens;

comparing, by a comparison engine using the one or more processors, the first-entry tokens to the second-entry tokens;

determining, by the one or more processors, whether the first data entries matches one of the second entries based on the comparing and weights of the first-entry tokens and the second-entry tokens wherein determining that the first data entries matches one of the plurality of second entries is further in response to:

each of a first plurality of character fields of the first data entries exactly matching a respective character field of the one of the plurality of second entries; and

each of at least two of a second plurality of character fields of the first data entries at least partially matching a respective character field of the one of the plurality of second entries;

associating, by the one or more processors, a first data entry with one of the second entries in response to determining that the first data entry matches the one of the second entries; and

determining, by the one or more processors, that the first data entry matches one of the plurality of second entries.

8 . The system of claim 7 wherein the filtering removes first data entries that have been submitted but have not been approved.

9 . The system of claim 7 wherein the filtering removes first data entries that have not been submitted for approval.

10 . The system of claim 7 wherein the filtering removes first data entries that are already associated with a travel segment entry in the itinerary database.

11 . The system of claim 7 wherein converting each of one or more character fields of each of the plurality of second entries in the second database into a respective one or more second-entry tokens includes converting each of the one or more character fields into a respective plurality of overlapping second-entry tokens.

12 . The system of claim 7 , wherein determining that the first data matches entries match one of the plurality of second entries is further in response to:

each of a first plurality of character fields of the first data entries exactly matching a respective character field of the one of the plurality of second entries;

each of at least one of a second plurality of character fields of the first data entries at least partially matching a respective character field of the one of the plurality of second entries; and

each of at least one of a third plurality of character fields of the first data entries at least partially matching a respective character field of the one of the plurality of second entries.

13 . A non-transitory machine-readable medium storing a program executable by at least one processor of a computer, the program comprising sets of instructions for:

loading first data entries from an expense report database and an itinerary database into a data entry store;

converting, by a tokenizer module of the computer, the first data entries in the data entry store into one or more first-entry tokens;

assigning, by a weighting engine, a weight to each of the one or more first-entry tokens based on a frequency of the of the first-entry tokens;

filtering, by a filtering component the first data entries to remove the first data entries in expense reports that are not associated with a travel segment expense;

converting, by the tokenizer of the computer, each of one or more character fields of each of a plurality of second entries in a second database into a respective one or more second-entry tokens and weighting the second-entry tokens based on a frequency of the second-entry tokens;

comparing, by the computer, the first-entry tokens to the second-entry tokens;

determining, by the computer, whether the first data entries matches one of the second entries based on the comparing and weights of the first-entry tokens and the second-entry tokens, wherein determining that the first data entries matches one of the plurality of second entries is further in response to:

each of a first plurality of character fields of the first data entries exactly matching a respective character field of the one of the plurality of second entries; and

each of at least two of a second plurality of character fields of the first data entries at least partially matching a respective character field of the one of the plurality of second entries;

associating, by the computer, a first data entry with one of the second entries in response to determining that the first data entry matches the one of the second entries; and

determining, but the computer, that the first data entry matches one of the plurality of second entries.

14 . The non-transitory machine-readable medium of claim 13 wherein the filtering removes first data entries that have been submitted but have not been approved.

15 . The non-transitory machine-readable medium of claim 13 wherein the filtering removes first data entries that have not been submitted for approval.

16 . The non-transitory machine-readable medium of claim 13 wherein the filtering removes first data entries that are already associated with a travel segment entry in the itinerary database.

17 . The non-transitory machine-readable medium of claim 13 wherein converting each of one or more character fields of each of the plurality of second entries in the second database into a respective one or more second-entry tokens includes converting each of the one or more character fields into a respective plurality of overlapping second-entry tokens.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2024
From: ZHANG, LU; HAAS, NICHOLE; MANOJ, JOSHUA; HARSHINI KOKA, SRI RAJA
To: SAP SE
Reel/Frame 069377/0645 →
Continuity (3)
Continuation 17531546 · Nov 19, 2021
Continuation 16201807 · Nov 27, 2018
Related Publication 20250086185A1 · Mar 13, 2025
References Cited (22)
US 5850507A · Ngai et al. · 1998 [cited by applicant]
US 8495077B2 · Bayliss · 2013 [cited by applicant]
US 8639596B2 · Chew · 2014 [cited by applicant]
US 8825502B2 · Bormann et al. · 2014 [cited by applicant]
US 9336256B2 · Boukobza · 2016 [cited by applicant]
US 9514123B2 · Goel et al. · 2016 [cited by applicant]
US 9922290B2 · Thomas · 2018 [cited by examiner]
US 10108688B2 · Busch et al. · 2018 [cited by applicant]
US 10115128B2 · Depasquale · 2018 [cited by examiner]
US 20010042087A1 · Kephart et al. · 2001 [cited by applicant]
US 20020016764A1 · Hoffman · 2002 [cited by applicant]
US 20050273452A1 · Molloy et al. · 2005 [cited by applicant]
US 20080103822A1 · Arora et al. · 2008 [cited by applicant]
US 20080222111A1 · Hoang et al. · 2008 [cited by applicant]
US 20090287600A1 · Amorosa et al. · 2009 [cited by applicant]
US 20100318640A1 · Mehta et al. · 2010 [cited by applicant]
US 20110295864A1 · Betz et al. · 2011 [cited by applicant]
US 20130085905A1 · Menon · 2013 [cited by applicant]
US 20140337331A1 · Hassanzadeh et al. · 2014 [cited by applicant]
US 20160078566A1 · Farrell et al. · 2016 [cited by applicant]
US 20170116206A1 · Gumerato et al. · 2017 [cited by applicant]
US 20190295158A1 · Wu · 2019 [cited by applicant]