IP Library › Granted Patent US 8,856,085
Granted Patent B2
US 8,856,085 · App. 13/185,601 · Granted Oct 7, 2014

Automatic consistent sampling for data analysis

Inventor: Alexander Gorelik (Palo Alto, CA)
Assignee: International Business Machines Corporation
G06F17/30371
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,856,085
App. No.
13/185,601
Granted
Oct 7, 2014
Kind
B2
Abstract

A method, computer program product, and system for analyzing data within one or more databases, comprising selecting one or more databases for analysis, each database comprising one or more database objects comprising one or more data values, applying a function to each data value in each database object within the one or more databases, where the function produces function values limited to a predetermined range, identifying for analysis the data values producing a certain function value within the predetermined range to form a sampled data set, and analyzing the sampled data set to determine relationships between the database objects within and across the one or more databases.

Claims (27)

1. A computer program product for analyzing data within one or more databases, comprising:

a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code comprising computer readable program code configured to:

select one or more databases for analysis, each database comprising one or more database objects comprising one or more data values, wherein the data values in each database object are arranged in columns;

apply a function to each data value in each database object within the one or more databases, wherein the function produces function values limited to a predetermined range;

identify for analysis the data values producing a certain function value within the predetermined range to form a sampled data set;

identify for analysis the data values that produce function values other than the certain function value and reside in one or more columns lacking high cardinality to form an unsampled data set, wherein a column has a high cardinality when data values in the column satisfy one or more from a group of a predetermined cardinality threshold and a predetermined selectivity threshold; and

analyze the sampled data set with the unsampled data set by matching data values within these data sets to determine relationships between the database objects within and across the one or more databases.

2. The computer program product of claim 1 , wherein the computer readable program code is further configured to:

determine one or more primary key-foreign key relationships between the database objects within and across the one or more databases.

3. The computer program product of claim 1 , wherein a column is a high cardinality column if a number of data values in the column that produce the certain function value exceeds the predetermined cardinality threshold.

4. The computer program product of claim 1 , wherein a column is a high cardinality column if a number of unique data values in the column divided by the number of data values in the column exceeds the predetermined selectivity threshold.

5. The computer program product of claim 1 , wherein the function is a hash function.

6. The computer program product of claim 1 , wherein the database objects are tables.

7. A system for analyzing data within one or more databases, comprising:

one or more databases, each database comprising one or more database objects comprising one or more data values stored in memory, wherein the data values in each database object are arranged in columns; and

a processor configured with logic to:

select one or more databases from the one or more databases for analysis;

apply a function to each data value in each database object within the selected one or more databases, wherein the function produces function values limited to a predetermined range;

identify for analysis the data values producing a certain function value within the predetermined range to form a sampled data set;

identify for analysis the data values that produce function values other than the certain function value and reside in one or more columns lacking high cardinality to form an unsampled data set, wherein a column has a high cardinality when data values in the column satisfy one or more from a group of a predetermined cardinality threshold and a predetermined selectivity threshold; and

analyze the sampled data set with the unsampled data set by matching data values within these data sets to determine relationships between the database objects within and across the one or more databases.

8. The system of claim 7 , wherein the processor being further configured with logic to:

determine one or more primary key-foreign key relationships between the database objects within and across the one or more databases.

9. The system of claim 7 , wherein a column is a high cardinality column if a number of data values in the column that produce the certain function value exceeds the predetermined cardinality threshold.

10. The system of claim 7 , wherein the function is a hash function.

11. The system of claim 7 , wherein the database objects are tables.

12. The system of claim 7 , wherein a column is a high cardinality column if a number of unique data values in the column divided by the number of data values in the column exceeds the predetermined selectivity threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 19, 2011
From: GORELIK, ALEXANDER
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 026612/0086 →
Continuity (1)
Related Publication 20130024430A1 · Jan 24, 2013