IP Library Granted Patent US 10,152,497
Granted Patent B2
US 10,152,497 · App. 15/052,382 · Granted Dec 11, 2018

Bulk deduplication detection

Inventors: Dai Duong Doan (Alameda, CA); Arun Kumar Jagota (Sunnyvale, CA); Chenghung Ker (Burlingame, CA); Parth Vaishnav (Cupertino, CA); Danil Dvinov (San Francisco, CA); Dmytro Kudriavtsev (Belmont, CA)
Assignee: salesforce.com, inc.
G06F17/30303G06F17/30489G06F7/32G06F17/30598
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,152,497
App. No.
15/052,382
Granted
Dec 11, 2018
Kind
B2
Abstract

Some embodiments of the present invention include a system and method for removing duplicate records from a group of records in a database system. The method includes generating a first cluster of records from the group of records, generating a second cluster of records from the group of records, identifying sets of duplicate records in the first cluster of records, and identifying sets of duplicate records in the second cluster of records. The method also includes merging at least two sets of duplicate records associated with both the first cluster and the second cluster of records to form a merged set of duplicate records. The merging is performed based on the at least two sets of duplicate records having a common record. Duplicate records in the group of records may then be removed by removing duplicate records from the merged set of duplicate records.

Claims (49)

1. A computer-implemented method comprising:

generating, by a database system, a first cluster of records from a group of records;

generating, by the database system, a second cluster of records from the group of records;

causing, by the database system, sets of duplicate records in the first cluster of records to be identified;

causing, by the database system, sets of duplicate records in the second cluster of records to be identified;

merging, by the database system, at least two sets of duplicate records associated with both the first cluster and the second cluster of records to form a merged set of duplicate records, wherein a set of duplicate records is implemented using a linked list having a head node and a body node for each record in the set of duplicate records and wherein the merging is performed based on the at least two sets of duplicate records having a common record and comprises merging a linked list associated with each set of duplicate records; and

removing, by the database system, one or more duplicate records from the merged set of duplicate records.

2. The method of claim 1 , wherein removing the one or more duplicate records from the merged set of duplicate records comprises removing the one or more duplicate records from the group of records.

3. The method of claim 2 , wherein the first cluster of records and the second cluster of records are generated based on one or more keys.

4. The method of claim 3 , wherein the first cluster of records and the second cluster of records are subsets of the group of records, and wherein the records in the first cluster and the records in the second cluster are not mutually exclusive.

5. The method of claim 4 , wherein the merging of the at least two sets of duplicate records comprises:

selecting, by the database system, a record from a first set of duplicate records;

comparing, by the database system, the selected record with records in a second set of duplicate records; and

merging, by the database system, the first set of duplicate records with the second set of duplicate records based on matching the selected record with any one record in the second set of duplicate records.

6. The method of claim 5 , wherein the merging the first set of duplicate records with the second set of duplicate records comprises merging a set of duplicate records with few records to a set of duplicate records with more records.

7. The method of claim 6 , wherein size information of the set of duplicate records and an identification information of the set of duplicate records are stored in the head node, and wherein the merging the first set of duplicate records with the second set of duplicate records comprises merging a linked list associated with the first set of duplicate records with a linked list associated with the second set of duplicate records.

8. An apparatus for identifying duplicate records in a database object, the apparatus comprising:

one or more processors; and

a non-transitory computer readable medium storing a plurality of instructions, which when executed, cause the one or more processors to:

generate a first cluster of records from a group of records;

generate a second cluster of records from the group of records;

cause sets of duplicate records in the first cluster of records to be identified;

cause sets of duplicate records in the second cluster of records to be identified;

merge at least two sets of duplicate records associated with both the first cluster and the second cluster of records to form a merged set of duplicate records, wherein a set of duplicate records is implemented using a linked list having a head node and a body node for each record in the set of duplicate records and wherein the merging is performed based on the at least two sets of duplicate records having a common record and comprises merging a linked list associated with each set of duplicate records; and

remove one or more duplicate records from the merged set of duplicate records.

9. The apparatus of claim 8 , wherein removing the one or more duplicate records from the merged set of duplicate records comprises removing the one or more duplicate records from the group of records.

10. The apparatus of claim 9 , wherein the first cluster of records and the second cluster of records are generated based on one or more keys.

11. The apparatus of claim 10 , wherein the first cluster of records and the second cluster of records are subsets of the group of records, and wherein the records in the first cluster and the records in the second cluster are not mutually exclusive.

12. The apparatus of claim 11 , wherein the merging of the at least two sets of duplicate records comprises:

selecting a record from a first set of duplicate records;

comparing the selected record with records in a second set of duplicate records; and

merging the first set of duplicate records with the second set of duplicate records based on matching the selected record with any one record in the second set of duplicate records.

13. The apparatus of claim 12 , wherein the merging the first set of duplicate records with the second set of duplicate records comprises merging a set of duplicate records with few records to a set of duplicate records with more records.

14. The apparatus of claim 13 , wherein size information of the set of duplicate records and an identification information of the set of duplicate records are stored in the head node, and wherein the merging the first set of duplicate records with the second set of duplicate records comprises merging a linked list associated with the first set of duplicate records with a linked list associated with the second set of duplicate records.

15. A computer program product comprising computer-readable program code to be executed by one or more processors when retrieved from a non-transitory computer-readable medium, the program code including instructions to:

generate a first cluster of records from the group of records;

generate a second cluster of records from the group of records;

cause sets of duplicate records in the first cluster of records to be identified;

cause sets of duplicate records in the second cluster of records to be identified;

merge at least two sets of duplicate records associated with both the first cluster and the second cluster of records to form a merged set of duplicate records, wherein a set of duplicate records is implemented using a linked list having a head node and a body node for each record in the set of duplicate records and wherein the merging is performed based on the at least two sets of duplicate records having a common record and comprises merging a linked list associated with each set of duplicate records; and

remove one or more duplicate records from the merged set of duplicate records.

16. The computer program product of claim 15 , wherein removing the one or more duplicate records from the merged set of duplicate records comprises removing the one or more duplicate records from the group of records.

17. The computer program product of claim 16 , wherein the first cluster of records and the second cluster of records are generated based on one or more keys.

18. The computer program product of claim 17 , wherein the first cluster of records and the second cluster of records are subsets of the group of records, and wherein the records in the first cluster and the records in the second cluster are not mutually exclusive.

19. The computer program product of claim 18 , wherein the merging of the at least two sets of duplicate records comprises:

selecting a record from a first set of duplicate records;

comparing the selected record with records in a second set of duplicate records; and

merging the first set of duplicate records with the second set of duplicate records based on matching the selected record with any one record in the second set of duplicate records.

20. The computer program product of claim 19 , wherein the merging the first set of duplicate records with the second set of duplicate records comprises merging a set of duplicate records with few records to a set of duplicate records with more records, wherein size information of the set of duplicate records and an identification information of the set of duplicate records are stored in the head node, and wherein the merging the first set of duplicate records with the second set of duplicate records comprises merging a linked list associated with the first set of duplicate records with a linked list associated with the second set of duplicate records.

Assignments (2)
CHANGE OF NAME Recorded Oct 3, 2023
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 065114/0983 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 24, 2016
From: DOAN, DAI DUONG; JAGOTA, ARUN KUMAR; KER, CHENGHUNG; VAISHNAV, PARTH; DVINOV, DANIL; KUDRIAVTSEV, DMYTRO
To: SALESFORCE.COM, INC.
Reel/Frame 037817/0275 →
Continuity (1)
Related Publication 20170242868A1 · Aug 24, 2017
Cited By (6)
US 12,210,579 US 12,235,849 US 12,450,234 US 12,572,826 US 12,681,918 US 12,688,205