IP Library › Granted Patent US 10,445,062
Granted Patent B2
US 10,445,062 · App. 15/706,082 · Granted Oct 15, 2019

Techniques for dataset similarity discovery

Inventors: Robert James Oberbreckling (Boulder, CO); Luis E. Rivas (Denver, CO)
Assignee: Oracle International Corporation
G06F7/02G06F16/2255G06F16/25G06F16/9535
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,445,062
App. No.
15/706,082
Granted
Oct 15, 2019
Kind
B2
Abstract

The present disclosure relates to techniques for analysis of data from multiple different data sources to determine similarity amongst the datasets. Determining a similarity between datasets may be useful for downstream processing of those datasets for different uses. A graphical interface may be provided to display detailed results including: a similarity prediction, data similarity prediction, column order similarity prediction, document type similarity prediction, prediction of overlapping or related columns, orphaned column prediction (e.g., a left orphaned column or a right orphaned column). Detecting similarities may be useful for leveraging prior data transformations generated for the datasets that are analyzed.

Claims (49)

1. A method comprising, at a computer system:

generating a first plurality of hash signatures, wherein each hash signature of the first plurality of hash signatures are generated based on first profile metadata for a different column in a first plurality of columns in a first dataset stored at a first data source;

generating a second plurality of hash signatures, wherein each hash signature of the second plurality of hash signatures are generated based on second profile metadata for a different column in a second plurality of columns in a second dataset stored at a second data source;

determining a set of column pairs based on comparison of the first plurality of hash signatures with the second plurality of hash signatures, wherein each column pair in the set of column pairs includes a different first column of the first plurality of columns and a different second column of the second plurality of columns based on similarity between a first hash signature for the different first column and a second hash signature for the different second column;

determining a first measure of similarity between each column pair in the set of column pairs;

determining a second measure of similarity between each column pair in the set of column pairs;

determining a third measure of similarity between each column pair in the set of column pairs; and

determining a measure of similarity between the first data set and the second dataset based on the third measure of similarity.

2. The method of claim 1 , further comprising:

generating information about a comparison of similarity of the first dataset to the second dataset; and

generating a graphical interface to display the information about the comparison of similarity of the first dataset to the second dataset.

3. The method of claim 1 , further comprising:

receiving input corresponding to a selection for combining the first dataset with the second dataset based on the measure of similarity; and

generating a transform script for generating a third dataset based on the input for combining the first dataset and the second dataset.

4. The method of claim 1 , wherein a hash signature based on profile metadata for a column is generated using a scalar hashing function.

5. The method of claim 4 , wherein profile metadata includes a type profile of the column, a subtype profile of the column, a compounding attribute of the column, a pattern of data in the column, one or more delimiters of the column, or a combination thereof.

6. The method of claim 1 , wherein the comparison to determine set of column pairs includes analyzing the first plurality of hash signatures with the second plurality of hash signatures.

7. The method of claim 1 , wherein determining the first measure of similarity between each column pair in the set of column pairs includes determining a summation of predictions for the column pair based on a normalized scale, and determining the first measure of similarity based on dividing the summation by a count of columns.

8. The method of claim 1 , wherein the third measure of similarity is a weighted summation of the first measure of similarity and the second measure of similarity.

9. The method of claim 1 , wherein the second measure of similarity is based on a count of column pairs predicted for overlap in the set of column pairs, based on a count of column pairs that have equal column indices in column pairs predicted for overlap, and based on a count of column pairs that have unequal column indices in column pairs predicted for overlap.

10. A system comprising:

one or more processors; and

a memory accessible to the one or more processors, the memory storing instructions that, upon execution by the one or more processors, causes the one or more processors to:

generate a first plurality of hash signatures, wherein each hash signature of the first plurality of hash signatures are generated based on first profile metadata for a different column in a first plurality of columns in a first dataset stored at a first data source;

generate a second plurality of hash signatures, wherein each hash signature of the second plurality of hash signatures are generated based on second profile metadata for a different column in a second plurality of columns in a second dataset stored at a second data source;

determine a set of column pairs based on comparison of the first plurality of hash signatures with the second plurality of hash signatures, wherein each column pair in the set of column pairs includes a different first column of the first plurality of columns and a different second column of the second plurality of columns based on similarity between a first hash signature for the different first column and a second hash signature for the different second column;

determine a first measure of similarity between each column pair in the set of column pairs;

determine a second measure of similarity between each column pair in the set of column pairs;

determine a third measure of similarity between each column pair in the set of column pairs; and

determine a measure of similarity between the first data set and the second dataset based on the third measure of similarity.

11. The system of claim 10 , wherein the instructions, upon execution by the one or more processors, further causes the one or more processors to:

generate information about a comparison of similarity of the first dataset to the second dataset; and

generate a graphical interface to display the information about the comparison of similarity of the first dataset to the second dataset.

12. The system of claim 10 , wherein a hash signature based on profile metadata for a column is generated using a scalar hashing function.

13. The system of claim 12 , wherein profile metadata includes a type profile of the column, a subtype profile of the column, a compounding attribute of the column, a pattern of data in the column, one or more delimiters of the column, or a combination thereof.

14. The system of claim 10 , wherein the comparison to determine set of column pairs includes analyzing the first plurality of hash signatures with the second plurality of hash signatures.

15. The system of claim 10 , wherein determining the first measure of similarity between each column pair in the set of column pairs includes determining a summation of predictions for the column pair based on a normalized scale, and determining the first measure of similarity based on dividing the summation by a count of columns.

16. The system of claim 10 , wherein the third measure of similarity is a weighted summation of the first measure of similarity and the second measure of similarity.

17. The system of claim 10 , wherein the second measure of similarity is based on a count of column pairs predicted for overlap in the set of column pairs, based on a count of column pairs that have equal column indices in column pairs predicted for overlap, and based on a count of column pairs that have unequal column indices in column pairs predicted for overlap.

18. A non-transitory computer readable medium storing instructions that are executable by one or more processors to cause the one or more processors to:

generate a first plurality of hash signatures, wherein each hash signature of the first plurality of hash signatures are generated based on first profile metadata for a different column in a first plurality of columns in a first dataset stored at a first data source;

generate a second plurality of hash signatures, wherein each hash signature of the second plurality of hash signatures are generated based on second profile metadata for a different column in a second plurality of columns in a second dataset stored at a second data source;

determine a set of column pairs based on comparison of the first plurality of hash signatures with the second plurality of hash signatures, wherein each column pair in the set of column pairs includes a different first column of the first plurality of columns and a different second column of the second plurality of columns based on similarity between a first hash signature for the different first column and a second hash signature for the different second column;

determine a first measure of similarity between each column pair in the set of column pairs;

determine a second measure of similarity between each column pair in the set of column pairs;

determine a third measure of similarity between each column pair in the set of column pairs; and

determine a measure of similarity between the first data set and the second dataset based on the third measure of similarity.

19. The non-transitory computer readable medium of claim 18 , wherein a hash signature based on profile metadata for a column is generated using a scalar hashing function.

20. The non-transitory computer readable medium of claim 18 , wherein the comparison to determine set of column pairs includes analyzing the first plurality of hash signatures with the second plurality of hash signatures, wherein determining the first measure of similarity between each column pair in the set of column pairs includes determining a summation of predictions for the column pair based on a normalized scale, and determining the first measure of similarity based on dividing the summation by a count of columns, wherein the third measure of similarity is a weighted summation of the first measure of similarity and the second measure of similarity, and wherein the second measure of similarity is based on a count of column pairs predicted for overlap in the set of column pairs, based on a count of column pairs that have equal column indices in column pairs predicted for overlap, and based on a count of column pairs that have unequal column indices in column pairs predicted for overlap.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2018
From: OBERBRECKLING, ROBERT JAMES; RIVAS, LUIS E.
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 044715/0846 →
Continuity (2)
Provisional Application 62395351 · Sep 15, 2016
Related Publication 20180074786A1 · Mar 15, 2018
Cited By (6)
US 12,197,927 US 12,205,163 US 12,326,871 US 12,339,864 US 12,450,234 US 12,596,694