IP Library Granted Patent US 11,163,769
Granted Patent B2
US 11,163,769 · App. 16/443,958 · Granted Nov 2, 2021

Joining two data tables on a join attribute

Inventors: Michal Bodziony (Tegoborze, PL); Konrad K. Skibski (Zielonki, PL); Tomasz Kazalski (Balice, PL); Artur M. Gruszecki (Cracow, PL); Lukasz Gaza (Jankowice, PL)
Assignee: International Business Machines Corporation
G06F16/24544G06F16/24532G06F16/24537
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,163,769
App. No.
16/443,958
Granted
Nov 2, 2021
Kind
B2
Abstract

A computer-implemented method for joining two data tables on a join attribute, where the data tables have at least a first and a second attribute and the second attribute is the join attribute. The method provides a function for associating a computing node to a given record. The function may be used to determine the associated computing node. The records of the two data tables may be distributed to the respective determined computing nodes. The relationship between the values of the first and second attributes may be modelled using a predefined dataset. For each record of the two data tables the values of the first attribute may be re-determined using the corresponding values of the second attribute. The function may be used to re-determine the associated computing node.

Claims (30)

1. A computer implemented method for joining two data tables on a join attribute, the data tables having at least a first and a second attribute, the second attribute being the join attribute, the method comprising:

providing a function for associating a computing node to a given record based on a value of the first attribute of the given record;

determining, using the function for each record of the two data tables, the associated computing node;

distributing the records of the two data tables to the respective determined computing nodes;

modeling a relationship between the values of the first and second attributes using a predefined dataset;

re-determining, for the each record of the two data tables, the values of the first attribute using the corresponding values of the second attribute based on the model of the relationship;

re-determining, using the function for the each record of the two data tables, the associated computing node using the re-determined values of the first attribute;

for the each record of the two data tables, in response to determining that the re-determined computing node of the each record is different from the node on which the each record is currently distributed, re-distributing the each record to the re-determined computing node; and

performing collocated joins of the distributed records and the redistributed record.

2. The method of claim 1 further comprising:

wherein the distributing results in a pair of partitions of the two data tables in the each of the computing nodes; and

in response to determining that the number of pairs of partitions in the sample, on which a collocated join can be performed is higher than a predefined threshold, determining a sample of the resulting pairs.

3. The method of claim 2 , wherein the computing nodes corresponding to the sample are in idle mode.

4. The method of claim 1 ,

wherein the modeling comprising determining a regression function describing the relationship between the values of the first and second attributes; and

wherein the re-determining the values of the first attribute is performed using the regression function.

5. The method of claim 1 , wherein the predefined dataset comprising the two data tables different from the two data tables.

6. The method of claim 1 , wherein the values of the first and second attributes in each data table of the two data tables have a correlation higher than a predetermined minimum correlation threshold.

7. The method of claim 1 , wherein the re-distributing comprises re-distributing the whole data tables using the re-determined values of the first attribute in response to determining that a number of records whose re-determined computing node is different from the nodes on which the records are currently stored is higher than a predefined number of records.

8. The method of claim 1 , wherein the collocated joins of the two data tables are done in parallel.

9. The method of claim 1 , providing, to the respective computing nodes in response to receiving a request for redistributing, the results of the collocated joins and redistributing the remaining non-joined records in response to the performing is done for all possible collocated joins of portions of the distributed records separately on the respective computing nodes.

10. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to perform a method comprising:

providing a function for associating a computing node to a given record based on a value of the first attribute of the given record;

determining, using the function for each record of the two data tables, the associated computing node;

distributing the records of the two data tables to the respective determined computing nodes;

modeling a relationship between the values of the first and second attributes using a predefined dataset;

re-determining, for the each record of the two data tables, the values of the first attribute using the corresponding values of the second attribute based on the model of the relationship;

re-determining, using the function for the each record of the two data tables, the associated computing node using the re-determined values of the first attribute;

for the each record of the two data tables, in response to determining that the re-determined computing node of the each record is different from the node on which the each record is currently distributed, re-distributing the each record to the re-determined computing node; and

performing collocated joins of the distributed records and the redistributed record.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2019
From: BODZIONY, MICHAL; SKIBSKI, KONRAD K.; KAZALSKI, TOMASZ; GRUSZECKI, ARTUR M.; GAZA, LUKASZ
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 049495/0926 →
Continuity (2)
Continuation 15663896 · Jul 31, 2017
Related Publication 20190303370A1 · Oct 3, 2019