IP Library Granted Patent US 12694004
Granted Patent B2
US 12694004 · App. 18/673,776 · Granted Jul 28, 2026

Machine-learning based data entry duplication detection and mitigation and methods thereof

Inventors: Srinivasarao Daruna (Ashburn, VA); Vijay Sahebgouda Bantanur (Gaithersburg, MD); Marisa Lee (Washington, DC)
Assignee: Capital One Services, LLC
G06F16/215G06F3/0481G06F16/2322G06F18/22G06F18/24323G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12694004
App. No.
18/673,776
Granted
Jul 28, 2026
Kind
B2
Abstract

Systems and methods of the present disclosure enable a processor to automatically detect duplicate data entries by receiving data entries associated with a user, where each data entry includes a value, a time, an entity identifier, and a location. Pairs of similar data entries are determined by matching the entity identifier and the location pairs data entries. Candidate duplicate data entries are determined based on a proximity in time between data entries of the similar data entries. For each candidate duplicate data entry, a feature vector is generated including the entity identifier, location, value and time, and each feature vector is submitted to a duplicate classification model to automatically determine duplicate data entries from the candidate duplicate data entries, the duplicate classification model being trained according to a historical dispute entries.

Claims (51)

1 . A method comprising:

receiving, by at least one processor, an electronic request from at least one terminal device, the electronic request requesting performance of at least one activity associated with a user;

accessing, by the at least one processor, a rolling data log of a plurality of data entries associated with a user;

wherein each data entry of the plurality of data entries comprises at least one attribute;

wherein the rolling data log comprises an automatically updating set of data entries associated with a rolling window of time;

determining, by the at least one processor, based at least in part on at least one similarity metric produced for a plurality of possible data entry pairs of the plurality data entries to measure a similarity between the plurality of possible data entry pairs of the plurality of data entries, a plurality of pairs of candidate duplicate data entries, each pair comprising at least one data entry in the rolling data log and the electronic request;

inputting, by the at least one processor, each pair of candidate duplicate data entries of the plurality of pairs of candidate duplicate data entries into a duplicate classification model to automatically generate a duplicate classification indicative of at least one pair of duplicate data entries from the plurality of pairs of candidate duplicate data entries based at least in part on the at least one attribute of each candidate duplicate data entry;

wherein the duplicate classification model comprises model parameters trained according to a plurality of historical data entries and a plurality of historical dispute entries disputing past data entries;

causing to display, by the at least one processor, a duplicate graphical user interface (GUI) on at least one display of at least computing device associated with the user, the duplicate GUI comprising:

at least one visualization indicative of the duplicate classification indicative of the at least one pair of duplicate data entries, and

an interface element associated with the at least one visualization, wherein selection of the interface element causes an electronic request to automatically remove at least one duplicate data entry of the at least one pair of duplicate data entries prior to the performance of the at least one activity; and

training, by the at least one processor, upon selection of the interface element, the model parameters of the duplicate classification model based at least in part on:

the duplicate classification of the at least one pair of duplicate data entries and

the at least one indication that the user has selected the user selectable element.

2 . The method as recited in claim 1 , wherein the alert is communicated prior to a posting of the at least one duplicate data entry.

3 . The method as recited in claim 1 , wherein the duplicate classification comprises a binary classification for each candidate duplicate data entry of the plurality of candidate duplicate data entries.

4 . The method as recited in claim 1 , wherein the duplicate classification model comprises a random forest model.

5 . The method as recited in claim 1 , wherein the plurality of historical data entries comprises historical dispute entries from a rolling time period preceding the data entry.

6 . The method as recited in claim 5 , wherein the rolling time period comprises three months preceding the data entry.

7 . The method as recited in claim 5 , wherein the duplicate classification model is retrained according to a predetermined schedule.

8 . The method as recited in claim 7 , wherein the predetermined schedule comprises once per week.

9 . The method as recited in claim 1 , wherein the plurality of data entries comprises a plurality of financial transactions.

10 . The method as recited in claim 1 , further comprising:

utilizing, by the at least one processor, at least one similarity measurement to determine a similarity between each pairing of a plurality of pairings of the plurality of data entries; and

determining, by the at least one processor, the plurality of pairs of candidate duplicate data entries based at least in part on the at least one similarity measurement and at least one threshold similarity value.

11 . A system comprising:

at least one processor in communication with at least one non-transitory computer readable medium having software instructions stored thereon, wherein, upon execution of the software instructions, the at least one processor is configured to perform a method comprising:

receiving an electronic request from at least one terminal device, the electronic request requesting performance of at least one activity associated with a user;

accessing a rolling data log of a plurality of data entries associated with a user;

wherein each data entry of the plurality of data entries comprises at least one attribute;

wherein the rolling data log comprises an automatically updating set of data entries associated with a rolling window of time;

determining, based at least in part on at least one similarity metric produced for a plurality of possible data entry pairs of the plurality data entries to measure a similarity between the plurality of possible data entry pairs of the plurality of data entries, a plurality of pairs of candidate duplicate data entries, each pair comprising at least one data entry in the rolling data log and the electronic request;

inputting each pair of candidate duplicate data entries of the plurality of pairs of candidate duplicate data entries into a duplicate classification model to automatically generate a duplicate classification indicative of at least one pair of duplicate data entries from the plurality of pairs of candidate duplicate data entries based at least in part on the at least one attribute of each candidate duplicate data entry;

wherein the duplicate classification model comprises model parameters trained according to a plurality of historical data entries and a plurality of historical dispute entries disputing past data entries;

causing to display a duplicate graphical user interface (GUI) on at least one display of at least computing device associated with the user, the duplicate GUI comprising:

at least one visualization indicative of the duplicate classification indicative of the at least one pair of duplicate data entries, and

an interface element associated with the at least one visualization, wherein selection of the interface element causes an electronic request to automatically remove at least one duplicate data entry of the at least one pair of duplicate data entries prior to the performance of the at least one activity; and

training, by the at least one processor, upon selection of the interface element, the model parameters of the duplicate classification model based at least in part on:

the duplicate classification of the at least one pair of duplicate data entries and

the at least one indication that the user has selected the user selectable element.

12 . The system as recited in claim 11 , wherein the alert is communicated prior to a posting of the at least one duplicate data entry.

13 . The system as recited in claim 11 , wherein the duplicate classification comprises a binary classification for each candidate duplicate data entry of the plurality of candidate duplicate data entries.

14 . The system as recited in claim 11 , wherein the duplicate classification model comprises a random forest model.

15 . The system as recited in claim 11 , wherein the plurality of historical data entries comprises historical dispute entries from a rolling time period preceding the data entry.

16 . The system as recited in claim 15 , wherein the rolling time period comprises three months preceding the data entry.

17 . The system as recited in claim 15 , wherein the duplicate classification model is retrained according to a predetermined schedule.

18 . The system as recited in claim 17 , wherein the predetermined schedule comprises once per week.

19 . The system as recited in claim 11 , wherein the plurality of data entries comprises a plurality of financial transactions.

20 . The system as recited in claim 11 , wherein the method further comprises:

utilizing at least one similarity measurement to determine a similarity between each pairing of a plurality of pairings of the plurality of data entries; and

determining the plurality of pairs of candidate duplicate data entries based at least in part on the at least one similarity measurement and at least one threshold similarity value.