IP Library Granted Patent US 11,995,054
Granted Patent B2
US 11,995,054 · App. 17/902,569 · Granted May 28, 2024

Machine-learning based data entry duplication detection and mitigation and methods thereof

Inventors: Srinivasarao Daruna (Ashburn, VA); Vijay Sahebgouda Bantanur (Gaithersburg, MD); Marisa Lee (Washington, DC)
Assignee: Capital One Services, LLC
G06F16/215G06F3/0481G06F16/2322G06F18/22G06F18/24323G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,995,054
App. No.
17/902,569
Granted
May 28, 2024
Kind
B2
Abstract

Systems and methods of the present disclosure enable a processor to automatically detect duplicate data entries by receiving data entries associated with a user, where each data entry includes a value, a time, an entity identifier, and a location. Pairs of similar data entries are determined by matching the entity identifier and the location pairs data entries. Candidate duplicate data entries are determined based on a proximity in time between data entries of the similar data entries. For each candidate duplicate data entry, a feature vector is generated including the entity identifier, location, value and time, and each feature vector is submitted to a duplicate classification model to automatically determine duplicate data entries from the candidate duplicate data entries, the duplicate classification model being trained according to a historical dispute entries.

Claims (49)

1. A method comprising:

receiving, by at least one processor, a plurality of data entries associated with a user;

wherein each data entry of the plurality of data entries comprises at least one attribute;

determining, by the at least one processor, a plurality of pairs of candidate duplicate data entries based at least in part on:

a proximity between each data entry of each pair of candidate duplicate data entries, and

the at least one attribute between each candidate duplicate data entry of each pair of candidate duplicate data entries;

wherein the proximity comprises a predetermined interval of time;

inputting, by the at least one processor, each pair of candidate duplicate data entries of the plurality of pairs of candidate duplicate data entries into a duplicate classification model to automatically generate a duplicate classification indicative of at least one pair of duplicate data entries from the plurality of pairs of candidate duplicate data entries based at least in part on the at least one attribute of each candidate duplicate data entry;

wherein the duplicate classification model comprises model parameters trained according to a plurality of historical data entries and a plurality of historical dispute entries disputing past data entries;

generating, by the at least one processor, a duplicate graphical user interface (GUI) comprising an alert message and a one-click resolution interface element;

wherein the alert message represents the duplicate classification of an incorrect duplicate data entry;

wherein the one-click resolution interface element comprises a user selectable element that upon selection causes an electronic request to automatically remove at least one duplicate data entry of the at least one pair of duplicate data entries;

causing to display, by the at least one processor, the duplicate GUI on a user computing device associated with the user of the at least one pair of duplicate data entries;

receiving, by the at least one processor, at least one indication that the user has selected the user selectable element; and

training, by the at least one processor, the model parameters of the duplicate classification model based at least in part on:

the duplicate classification of the at least one pair of duplicate data entries and

the at least one indication that the user has selected the user selectable element.

2. The method as recited in claim 1 , wherein the alert is communicated prior to a posting of the at least one duplicate data entry.

3. The method as recited in claim 1 , wherein the duplicate classification comprises a binary classification for each candidate duplicate data entry of the plurality of candidate duplicate data entries.

4. The method as recited in claim 1 , wherein the duplicate classification model comprises a random forest model.

5. The method as recited in claim 1 , wherein the plurality of historical data entries comprises historical dispute entries from a rolling time period preceding the data entry.

6. The method as recited in claim 5 , wherein the rolling time period comprises three months preceding the data entry.

7. The method as recited in claim 5 , wherein the duplicate classification model is retrained according to a predetermined schedule.

8. The method as recited in claim 7 , wherein the predetermined schedule comprises once per week.

9. A system comprising:

at least one processor in communication with at least one non-transitory computer readable medium having software instructions stored thereon, wherein, upon execution of the software instructions, the at least one processor is configured to:

receive a plurality of data entries associated with a user;

wherein each data entry of the plurality of data entries comprises at least one attribute;

determine a plurality of pairs of candidate duplicate data entries based at least in part on:

a proximity between each data entry of each pair of candidate duplicate data entries, and

the at least one attribute between each candidate duplicate data entry of each pair of candidate duplicate data entries;

wherein the proximity comprises a predetermined interval of time;

input each pair of candidate duplicate data entries of the plurality of pairs of candidate duplicate data entries into a duplicate classification model to automatically generate a duplicate classification indicative of at least one pair of duplicate data entries from the plurality of pairs of candidate duplicate data entries based at least in part on the at least one attribute of each candidate duplicate data entry;

wherein the duplicate classification model comprises model parameters trained according to a plurality of historical data entries and a plurality of historical dispute entries disputing past data entries;

generate a duplicate graphical user interface (GUI) comprising an alert message and a one-click resolution interface element;

wherein the alert message represents the duplicate classification of an incorrect duplicate data entry;

wherein the one-click resolution interface element comprises a user selectable element that upon selection causes an electronic request to automatically remove at least one duplicate data entry of the at least one pair of duplicate data entries;

cause to display the duplicate GUI on a user computing device associated with the user of the at least one pair of duplicate data entries;

receive at least one indication that the user has selected the user selectable element; and

train the model parameters of the duplicate classification model based at least in part on:

the duplicate classification of the at least one pair of duplicate data entries and

the at least one indication that the user has selected the user selectable element.

10. The system as recited in claim 9 , wherein the alert is communicated prior to a posting of the at least one duplicate data entry.

11. The system as recited in claim 9 , wherein the duplicate classification comprises a binary classification for each candidate duplicate data entry of the plurality of candidate duplicate data entries.

12. The system as recited in claim 9 , wherein the duplicate classification model comprises a random forest model.

13. The system as recited in claim 9 , wherein the plurality of historical data entries comprises historical dispute entries from a rolling time period preceding the data entry.

14. The system as recited in claim 13 , wherein the rolling time period comprises three months preceding the data entry.

15. The system as recited in claim 13 , wherein the duplicate classification model is retrained according to a predetermined schedule.

16. The system as recited in claim 15 , wherein the predetermined schedule comprises once per week.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 2, 2022
From: DARUNA, SRINIVASARAO; BANTANUR, VIJAY SAHEBGOUDA; LEE, MARISA
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 060984/0182 →
Continuity (2)
Continuation 17139128 · Dec 31, 2020
Related Publication 20220414074A1 · Dec 29, 2022
Cited By (1)
US 12,614,196