IP Library Patent Application 14466231
Patent Application
App. No. 14/466,231

AUTOMATIC JOINING OF DATA SETS BASED ON STATISTICS OF FIELD VALUES IN THE DATA SETS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
14/466,231
Abstract

A computer system processes arbitrary data sets to identify fields of data that can be the basis of a join operation. Each data set has a plurality of entries, with each entry having a plurality of fields. For each pair of data sets, the computer system compares the values of fields in a first data set in the pair of data sets to the values of fields in a second data set in the pair of data sets, to identify fields having substantially similar sets of values. Given pairs of fields that have similar sets of values, the computer system measures entropy with respect to an intersection of the sets of values of the pair of fields. The computer system can recommend fields for a join operation between any pair of data sets in the plurality of data sets based on such statistical measures.

Claims (53)

1 . A computer-implemented process comprising:

receiving a plurality of data sets, each data set having a plurality of entries, each entry having a plurality of fields, wherein a field in the plurality of fields has at least one value;

for each pair of data sets in the plurality of data sets:

comparing the values of fields in a first data set in the pair of data sets to the values of fields in a second data set in the pair of data sets to identify fields having substantially similar sets of values, and

measuring entropy with respect to an intersection of the sets of values of the identified fields from the pair of data sets; and

suggesting fields for a join operation between any pair of data sets in the plurality of data sets, based at least on the measured entropy with respect to the intersection of the sets of values of the identified fields from the pair of data sets.

2 . The computer-implemented process of claim 1 , wherein, for each pair of data sets in the plurality of data sets, the process further comprises:

measuring density of at least one of the identified fields in the pair of data sets; and

wherein suggesting fields is further based at least on the measured density.

3 . The computer-implemented process of claim 2 , wherein, for each pair of data sets in the plurality of data sets, the process further comprises:

measuring a likelihood that a value in the identified field in the first data set matches a value in the identified field in the second data set; and

wherein suggesting fields is further based at least on the measured likelihood.

4 . The computer-implemented process of claim 1 , wherein, for each pair of data sets in the plurality of data sets, the process further comprises:

measuring a likelihood that a value in the identified field in the first data set matches a value in the identified field in the second data set; and

wherein suggesting fields is further based at least on the measured likelihood.

5 . The computer-implemented process of claim 1 , wherein suggesting comprises:

generating a ranked list of identified fields.

6 . The computer-implemented process of claim 5 , wherein suggesting comprises:

presenting the ranked list on a display; and

receiving an input indicating a selection of identified fields from the ranked list.

7 . The computer-implemented process of claim 5 , wherein suggesting comprises:

the processor selecting identified fields from the ranked list.

8 . The computer-implemented process of claim 7 , further comprising:

presenting the selected identified fields on a display.

9 . The computer-implemented process of claim 1 , wherein the plurality of data sets includes N data sets, where N is a positive integer greater than 2.

10 . The computer-implemented process of claim 7 , further comprising:

receiving a query results for a query applied to the plurality of data sets;

for each data set in the results, performing a join operation using the selected identified fields in the ranked list.

11 . The computer-implemented process of claim 10 , further comprising:

presenting the joined results on a display.

12 . The computer-implemented process of claim 1 , wherein the plurality of data sets includes data from different tables in a relational database management system.

13 . The computer-implemented process of claim 1 , wherein the plurality of data sets includes data from different tables in an object oriented database system.

14 . The computer-implemented process of claim 1 , wherein the plurality of data sets includes data from different tables in an index of documents.

15 . A computer system comprising:

memory in which a plurality of data sets are stored, each data set having a plurality of entries, each entry having a plurality of fields, wherein a field in the plurality of fields has at least one value;

one or more processing units programmed by a computer program to be instructed to, for each pair of data sets in the plurality of data sets:

compare the values of fields in a first data set in the pair of data sets to the values of fields in a second data set in the pair of data sets to identify fields having substantially similar sets of values, and

measure entropy with respect to an intersection of the sets of values of the identified fields from the pair of data sets; and

suggest fields for a join operation between any pair of data sets in the plurality of data sets, based at least on the measured entropy with respect to the intersection of the sets of values of the identified fields from the pair of data sets.

16 . The computer system of claim 15 , wherein, for each pair of data sets in the plurality of data sets, the one or more processing units are further programmed to be instructed to:

measure density of at least one of the identified fields in the pair of data sets; and

wherein suggesting fields is further based at least on the measured densities.

17 . The computer system of claim 16 , wherein, for each pair of data sets in the plurality of data sets, the one or more processing units are further programmed to be instructed to:

measure a likelihood that a value in the identified field in the first data set matches a value in the identified field in the second data set; and

wherein suggesting fields is further based at least on the measured likelihood.

18 . The computer system of claim 15 , wherein, for each pair of data sets in the plurality of data sets, the one or more processing units are further programmed to be instructed to:

measure a likelihood that a value in the identified field in the first data set matches a value in the identified field in the second data set; and

wherein suggesting fields is further based at least on the measured likelihood.

19 . The computer system of claim 15 , wherein suggesting comprises:

generating a ranked list of identified fields.

20 . The computer system of claim 19 , wherein suggesting comprises:

presenting the ranked list on a display; and

receiving an input indicating a selection of identified fields from the ranked list.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2019
From: ATTIVIO, INC.
To: SERVICENOW, INC.
Reel/Frame 051166/0096 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2014
From: YOUNG, JONATHAN; O'NEIL, JOHN; JOHNSON, WILLIAM K., III; SERRANO, MARTIN; GEORGE, GREGORY
To: ATTIVIO, INC.
Reel/Frame 033593/0345 →