IP Library › Granted Patent US 12,596,963
Granted Patent B2
US 12,596,963 · App. 17/858,581 · Granted Apr 7, 2026

Machine-learning based record processing systems

Inventors: Gunther Havel (Portsmouth, NH); John Wilson (Richmond, VA); Ashwin Assysh Sharma (Minneapolis, MN)
Assignee: Capital One Services, LLC
G06N20/20G06F18/2113G06F18/22G06Q20/389
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,963
App. No.
17/858,581
Granted
Apr 7, 2026
Kind
B2
Abstract

Aspects described herein may allow unmatched records in databases be matched automatically. For example, a computing device may receive source records and target records to be matched together. The computing device may determine, based on a machine learning model and for each of the plurality of target records, a distance value between the respective target record, and a subset of the plurality of source records. A matched record that identifies the subset of the plurality of source records and a selected one or more target records may be generated. The matched record may be configured to update records in a database.

Claims (75)

1 . A method comprising:

receiving, by a computing device from a cloud storage device, a trigger command indicating a threshold is satisfied by a quantity of records that have been uploaded from a first database;

receiving, by the computing device based on the trigger command, the uploaded records comprising a plurality of source records, from a first database, and a plurality of target records from a second database, wherein the first database is configured to store records of a different format from records in the second database;

inputting, based on the trigger command and into a machine learning model, the plurality of source records and the plurality of target records;

receiving, as output from the machine learning model, a selection of a subset of the plurality of source records and a subset of the plurality of target records, wherein the selection is based on:

determining, for each of the plurality of target records, a distance value between:

the respective target record, and

the subset of the plurality of source records;

ranking, based on the distance value of each of the plurality of target records, the plurality of target records;

selecting, based on the ranking, the subset of target records of the plurality of target records;

generating a matched record that indicates a matching between:

the selected s subset of the plurality of target records, and

the selected subset of the plurality of source records;

sending, to the first database, the matched record; and

updating, based on the matched record, the first database by changing a status of the subset of the plurality of source records from unmatched to matched.

2 . The method of claim 1 , wherein each of the plurality of source records and the plurality of target records is a transaction record that indicates a balance amount, and wherein generating the matched record is further based on determining that a total balance amount of the subset of the plurality of source records equals to a total balance amount of the selected one or more target records.

3 . The method of claim 1 , wherein the receiving the plurality of source records comprises receiving, from the first database via a cloud storage bucket, the plurality of source records.

4 . The method of claim 3 , wherein the plurality of source records and the plurality of target records are unmatched before being input into the machine learning model.

5 . The method of claim 1 , wherein each of the plurality of target records comprises a plurality of data fields, and wherein determining the distance value comprises:

identifying one or more data fields of the plurality of data fields; and

calculating the distance value based on data within the one or more data fields.

6 . The method of claim 1 , wherein the distance value is a gower distance value.

7 . The method of claim 1 , further comprising:

dividing, based on a second machine learning model comprising a clustering algorithm, the plurality of source records into one or more subsets.

8 . The method of claim 1 , wherein selecting the subset of the plurality of source records comprises determining a distance value between each two source records of the subset does not exceed a threshold.

9 . A system comprising:

a computing device, and

a first database;

wherein the computing device comprising:

one or more processors;

instructions, when executed by the one or more processors, cause the computing device to:

receive, from a cloud storage device, a trigger command indicating a threshold is satisfied by a quantity of records that have been uploaded from a first database;

receive, based on the trigger command, the uploaded records comprising, a plurality of source records, from a first database, and a plurality of target records from a second database, wherein the first database is configured to store records of a different format from records in the second database;

input, based on the trigger command and into a machine learning model, the plurality of source records and the plurality of target records;

receive, as output from the machine learning model, a selection of a subset of the plurality of source records and a subset of the plurality of target records, wherein the selection is based on:

determine, for each of the plurality of target records, a distance value between:

the respective target record, and

the subset of the plurality of source records;

rank, based on the distance value of each of the plurality of target records, the plurality of target records;

select, based on the ranking, the subset of the plurality of target records;

generate a matched record that indicates a matching between:

the selected subset of the plurality of target records, and

the selected subset of the plurality of source records;

send, to the first database, the matched record; and

update, based on the matched record, the first database by changing a status of the subset of the plurality of source records from unmatched to matched; and

wherein the first database is configured to receive, from the computing device, the matched record.

10 . The system of claim 9 , wherein each of the plurality of source records and the plurality of target records is a transaction record that indicates a balance amount, and wherein the computing device is configured to generate the matched record further based on determining that a total balance amount of the subset of the plurality of source records equals to a total balance amount of the selected one or more target records.

11 . The system of claim 9 , wherein the instructions, when executed by the one or more processors, cause the computing device to receive the plurality of source records from the first database via a cloud bucket.

12 . The system of claim 11 , wherein the plurality of source records and the plurality of target records are unmatched before being input into the machine learning model.

13 . The system of claim 9 , wherein each of the plurality of target records comprises a plurality of data fields, and wherein the instructions, when executed by the one or more processors, cause the computing device to determine the distance value by:

identifying one or more data fields of the plurality of data fields; and

calculating the distance value based on data within the one or more data fields.

14 . The system of claim 9 , wherein the distance value is a gower distance value.

15 . The system of claim 9 , wherein the instructions, when executed by the one or more processors, further cause the computing device to:

divide, based on a second machine learning model comprising a clustering algorithm, the plurality of source records into one or more subsets.

16 . The system of claim 9 , wherein the instructions, when executed by the one or more processors, cause the computing device to select the subset of the plurality of source records by determining a distance value between each two source records of the subset does not exceed a threshold.

17 . A non-transitory computer-readable medium storing computer instruction that, when executed by one or more processors, cause performance of actions comprising:

receiving, by a computing device, from a cloud storage device, a trigger command indicating a threshold is satisfied by a quantity of records that have been uploaded from a first database;

receiving, by the computing device based on the trigger command, the uploaded records comprising a plurality of source records, from a first database, and a plurality of target records from a second database, wherein the first database is configured to store records of a different format from records in the second database;

inputting, based on the trigger command and into a machine learning model, the plurality of source records and the plurality of target records;

receiving, as output from the machine learning model and based on a clustering algorithm, a selection of a subset of the plurality of source records and a subset of the plurality of target records, wherein the selection is based on:

determining a distance value between each two source records of the subset does not exceed a threshold;

determining, based on a second machine learning model and for each of the plurality of target records, a distance value between:

the respective target record, and

the subset of the plurality of source records;

ranking, based on the distance value of each of the plurality of target records, the plurality of target records;

selecting, based on the ranking, the subset of the plurality of target records;

generating a matched record that indicates a matching between:

the selected subset of the plurality of target records, and

the selected subset of the plurality of source records;

sending, to the first database, the matched record; and

updating, based on the matched record, the first database by changing a status of the subset of the plurality of source records from unmatched to matched.

18 . The non-transitory computer-readable medium of claim 17 , wherein each of the plurality of source records and the plurality of target records is a transaction record that indicates a balance amount, and wherein generating the matched record is further based on determining that a total balance amount of the subset of the plurality of source records equals to a total balance amount of the selected one or more target records.

19 . The non-transitory computer-readable medium of claim 17 , wherein the instructions, when executed by the one or more processor, cause receiving the plurality of source records from the first database via a cloud storage bucket.

20 . The non-transitory computer-readable medium of claim 19 , wherein the plurality of source records and the plurality of target records are unmatched before being input into the machine learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2022
From: HAVEL, GUNTHER; WILSON, JOHN; SHARMA, ASHWIN ASSYSH
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 060430/0993 →
Continuity (1)
Related Publication 20240013100A1 · Jan 11, 2024
References Cited (17)
US 11687575B1 · Nguyen · 2023 [cited by examiner]
US 11887079B2 · Godshall · 2024 [cited by examiner]
US 12346306B2 · Katz · 2025 [cited by examiner]
US 12400247B2 · Saito · 2025 [cited by examiner]
US 12412221B2 · Jeske · 2025 [cited by examiner]
US 20160364794A1 · Chari et al. · 2016 [cited by applicant]
US 20190318347A1 · Aguiar · 2019 [cited by examiner]
US 20200012980A1 · Li · 2020 [cited by examiner]
US 20200126037A1 · Tatituri · 2020 [cited by examiner]
US 20200184281A1 · Le · 2020 [cited by examiner]
US 20220076231A1 · Farrell · 2022 [cited by examiner]
US 20220351210A1 · Ramani · 2022 [cited by examiner]
US 20220398583A1 · Crudele · 2022 [cited by examiner]
US 20220405836A1 · Gorman · 2022 [cited by examiner]
US 20230067073A1 · McCormick · 2023 [cited by examiner]
US 20230073140A1 · Richter · 2023 [cited by examiner]
Conceptual-Framework-for-Improving-Bank-Reconciliation-Accuracy-Using-Intelligent-Audit Controls, Journal of Frontiers in Multidisciplinary Research, Author: Sandra Orobosa Ikponmwoba, E-ISSN: 3050-9726, P-ISSN: 3050-97… [cited by examiner]