IP Library Patent Application 14221682
Patent Application
App. No. 14/221,682

SYSTEM AND METHOD FOR ORGANIZING DATA

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
14/221,682
Abstract

A system and method for organizing raw data from one or more sources uses an improved mechanism for identifying duplicate data between fields (e.g., columns) in the databases. The fields may be similar fields within a single database or similar or identical fields within a pair of databases and as organized as arrays or field vectors. The present invention sorts each of the field vectors and if necessary, partitions them by common value. A number of comparisons required to identify the duplicate data between the field vectors is reduced by feeding back a difference between the compared values. This difference is used to adjust indices into the field vectors for subsequent comparison.

Claims (31)

1 . A computer-implemented method for identifying duplicate data between a first vector and a second vector comprising:

sorting, by at least one computing processor, values in the first vector in a decreasing order;

partitioning, by the at least one computing processor, the sorted values in the first vector into first sets, wherein at least one of the first sets includes a plurality of sorted values that have a common value;

sorting, by the at least one computing processor, values in the second vector in said decreasing order;

partitioning, by the at least one computing processor, the sorted values in the second vector into second sets, wherein at least one of the second sets includes a plurality of members that have a common value;

comparing, by the at least one computing processor, a first sorted value at a first index in a first one of the first sets of the partitioned first vector with a second sorted value at a second index in a first one of the second sets of the partitioned second vector;

adjusting, by the at least one computing processor, said first index to a next one of the first sets of the partitioned first vector if said first sorted value is greater than said second sorted value;

adjusting, by the at least one computing processor, said second index to a next one of the second sets of the partitioned second vector if said second sorted value is greater than said first sorted value; and

identifying, by the at least one computing processor, said first and second sorted values as duplicate data if said first sorted value is equal to said second sorted value.

2 . A computer-implemented method for identifying duplicate data between a first vector and a second vector comprising:

sorting, by at least one computing processor, values in the first vector in an increasing order;

partitioning, by the at least one computing processor, the sorted values in the first vector into first sets, each of the first sets having at least one sorted value, all sorted values in each of the first sets having a common value, wherein at least one of the first sets includes a plurality of sorted values;

sorting, by the at least one computing processor, values in the second vector in said increasing order;

partitioning, by the at least one computing processor, the sorted values in the second vector into second sets, each of the second sets having at least one sorted value, all sorted values in each of the second sets having a common value, wherein at least one of the second sets includes a plurality of sorted values;

comparing, by the at least one computing processor, a first sorted value at a first index in a first one of the first sets of the partitioned first vector with a second sorted value at a second index in a first one of the second sets of the partitioned second vector;

adjusting, by the at least one computing processor, said first index to a next one of the first sets of the partitioned first vector if said first sorted value is less than said second sorted value;

adjusting, by the at least one computing processor, said second index to a next one of the second sets of the partitioned second vector if said second sorted value is less than said first sorted value; and

identifying, by the at least one computing processor, said first and second sorted values as duplicate data if said first sorted value is equal to said second sorted value.

3 . The method of claim 2 , wherein at least one of the second sets includes a plurality of sorted values.

4 . The method of claim 1 , further comprising storing said duplicate data.

5 . The method of claim 2 , further comprising storing said duplicate data.

6 . A computer-implemented method comprising:

sorting, by at least one computing processor, values in a first vector in an increasing order;

partitioning, by the at least one computing processor, the sorted values in the first vector into first sets, each of the first sets having at least one sorted value, all sorted values in each of the first sets sharing a common value, wherein at least one of the first sets includes a plurality of sorted values;

sorting, by the at least one computing processor, values in a second vector in said increasing order;

partitioning, by the at least one computing processor, the sorted values in the second vector into second sets, each of the second sets having at least one sorted value, all sorted values in each of the second sets having a common value;

comparing, by the at least one computing processor, a first sorted value at a first index in a first one of the first sets of the partitioned first vector with a second sorted value at a second index in a first one of the second sets of the partitioned second vector, wherein the first one of the first sets of the partitioned first vector includes a plurality of sorted values;

when said first sorted value is less than said second sorted value, adjusting said first index to a next one of the first sets of the partitioned first vector;

when said second sorted value is less than said first sorted value, adjusting said second index to a next one of the second sets of the partitioned second vector; and

when said first sorted value is equal to said second sorted value, adjusting said first index to a next one of the first sets of the partitioned first vector and adjusting said second index to a next one of the second sets of the partitioned second vector.

7 . The method of claim 6 , wherein the first sorted value is one of the at least one sorted value in the first one of the first sets of the partitioned first vector, and wherein the second sorted value is one of the at least one sorted value in the second one of the second sets of the partitioned second vector.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2014
From: GRUENWALD, BJORN J.
To: INMENTIA, INC.
Reel/Frame 032550/0285 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2014
From: INMENTIA, INC.
To: INMENTIA IPH, INC.
Reel/Frame 032550/0358 →
CHANGE OF NAME Recorded Mar 28, 2014
From: INMENTIA IPH, INC.
To: PRIMENTIA IPH, INC.
Reel/Frame 032555/0262 →