IP Library › Granted Patent US 12,259,905
Granted Patent B2
US 12,259,905 · App. 17/448,713 · Granted Mar 25, 2025

Data distribution in data analysis systems

Inventors: Felix Beier (Haigerloch, DE); Dennis Butterstein (Stuttgart, DE); Einar Lueck (Filderstadt, DE); Sabine Perathoner-Tschaffler (Nufringen, DE)
Assignee: International Business Machines Corporation
G06F16/27G06F11/1471G06F16/2255G06F16/2358
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,259,905
App. No.
17/448,713
Granted
Mar 25, 2025
Kind
B2
Abstract

The present disclosure relates to a computer implemented method for data synchronization in a data analysis system. The data analysis system comprises a source and target database system. The method comprises: receiving a change record describing an operation performed on a data record in the source database system. The change record may be read for determining a value of a distribution key of the data record. The value of the distribution key may be used for selecting a target database node of the target database system where the operation is to be performed. A direct connection may be established to the selected target database node and the change record may be provided to the selected target database node through the direct connection.

Claims (40)

1. A computer implemented method for data synchronization in a data analysis system, the data analysis system comprising a target database system and a source database system, the method comprising

retrieving a change record describing an operation performed on a data record in the source database system of the data analysis system based on reading a transaction log associated with the source database system with a frequency higher than a defined minimum frequency;

determining a distribution key that is configured to be used by the target database system to distribute records over target database nodes of the target database system, wherein the target database system comprises a metadata catalog that includes cluster metadata and table metadata, the table meta comprising information on a total number of the target database nodes and storage properties of the target database nodes and the cluster metadata comprising information on the distribution key;

reading the change record for determining a value of the distribution key of the data record;

using the value of the distribution key for selecting a target database node of the target database nodes where the operation is to be performed, wherein selecting the target database node for storing data record comprises providing a distribution map of hash values to connection numbers, wherein each connection number indicates a connection between the source database system and a respective target database node, computing a hash value of the determined value of the distribution key of the data record, and using the distribution map for assigning the computed hash value to a connection number, wherein the connection is established according to the connection number;

establishing a direct TCP/IP connection between the source database system and the selected target database node; and

providing the change record to the selected target database node through the direct connection.

2. The method of claim 1 , repeating the method for further received change records, thereby distributing the change records to respective target database nodes through respective direct connections.

3. The method of claim 2 , the method being concurrently performed for the change records.

4. The method of claim 1 , further comprising

determining a distribution rule of the target database system, the distribution rule assigning values of the distribution key to respective target database nodes;

wherein selecting the target database node comprises applying by the source database system the distribution rule on the determined value of the distribution key.

5. The method of claim 1 , wherein receiving the change record comprises reading a transaction recovery log indicating transactions to be replicated to the target database system.

6. The method of claim 1 , wherein the operation includes at least one of inserting, deleting or updating a data record.

7. The method of claim 1 , the distribution key comprising one or more attributes of the data record.

8. A computer program product for data synchronization in a data analysis system, the data analysis system comprising a target database system and a source database system, the computer program product comprising a non-transitory computer readable hardware storage device, and program instructions stored on the computer readable hardware storage device, to:

retrieve a change record describing an operation performed on a data record in the source database system of the data analysis system based on reading a transaction log associated with the source database system with a frequency higher than a defined minimum frequency;

determine a distribution key that is configured to be used by the target database system to distribute records over target database nodes of the target database system, wherein the target database system comprises a metadata catalog that includes cluster metadata and table metadata, the table meta comprising information on a total number of the target database nodes and storage properties of the target database nodes and the cluster metadata comprising information on the distribution key;

read the change record for determining a value of the distribution key of the data record;

use the value of the distribution key for selecting a target database node of the target database nodes where the operation is to be performed, wherein selecting the target database node for storing data record comprises providing a distribution map of hash values to connection numbers, wherein each connection number indicates a connection between the source database system and a respective target database node, computing a hash value of the determined value of the distribution key of the data record, and using the distribution map for assigning the computed hash value to a connection number, wherein the connection is established according to the connection number;

establish a direct TCP/IP connection between the source database system and the selected target database node; and

provide the change record to the selected target database node through the direct connection.

9. The computer program product of claim 8 , wherein the computer readable storage device further comprises instructions to repeatedly receive change records, thereby distributing the change records to respective target database nodes through respective direct connections.

10. The computer program product of claim 9 , wherein the computer readable storage device further comprises instructions that are concurrently performed for the change records.

11. The computer program product of claim 8 , wherein the computer readable storage device further comprising instructions to:

determine a distribution rule of the target database system, the distribution rule assigning values of the distribution key to respective target database nodes; and

wherein selecting the target database node comprises applying by the source database system the distribution rule on the determined value of the distribution key.

12. The computer program product of claim 8 , wherein the computer readable storage device instructions to receive the change record further comprise reading a transaction recovery log indicating transactions to be replicated to the target database system.

13. The computer program product of claim 8 , wherein the operation includes at least one of inserting, deleting or updating a data record.

14. The computer program product of claim 8 , wherein the distribution key comprising one or more attributes of the data record.

15. A computer system for a data analysis system, the data analysis system comprising a source database system and a target database system, the computer system including one or more non-transitory computer-readable storage media configured to store program instructions and one or more computer processors configured to execute said program instructions store on the one or more non-transitory computer-readable storage media, the computer system being configured for:

receiving a change record describing an operation performed on a data record in the source database system;

determining a distribution key that is configured to be used by the target database system to distribute records over target database nodes of the target database system, wherein the target database system comprises a metadata catalog that includes cluster metadata and table metadata, the table meta comprising information on a total number of the target database nodes and storage properties of the target database nodes and the cluster metadata comprising information on the distribution key;

reading the change record for determining a value of the distribution key of the data record;

using the value of the distribution key for selecting a target database node of the target database nodes where the operation is to be performed, wherein selecting the target database node for storing data record comprises providing a distribution map of hash values to connection numbers, wherein each connection number indicates a connection between the source database system and a respective target database node, computing a hash value of the determined value of the distribution key of the data record, and using the distribution map for assigning the computed hash value to a connection number, wherein the connection is established according to the connection number;

establishing a direct connection to the selected target database node; and

providing the change record to the selected target database node through the direct connection.

16. The computer system of claim 15 , wherein the computer system is further configured to repeatedly receive change records, thereby distributing the change records to respective target database nodes through respective direct connections.

17. The computer system of claim 15 , wherein the computer system is comprised in the source database system.

18. The computer system of claim 15 , wherein the computer system is remotely connected to the source database system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2021
From: BEIER, FELIX; BUTTERSTEIN, DENNIS; LUECK, EINAR; PERATHONER-TSCHAFFLER, SABINE
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 057585/0788 →
Continuity (1)
Related Publication 20230101740A1 · Mar 30, 2023
References Cited (27)
US 7206861B1 · Callon · 2007 [cited by examiner]
US RE48243E · Pareek et al. · 2020 [cited by applicant]
US 20080104133A1 · Chellappa · 2008 [cited by applicant]
US 20100332448A1 · Holenstein · 2010 [cited by applicant]
US 20120254175A1 · Horowitz · 2012 [cited by applicant]
US 20140108421A1 · Isaacson · 2014 [cited by applicant]
US 20160232208A1 · Duan et al. · 2016 [cited by applicant]
US 20170316026A1 · Kanthak · 2017 [cited by applicant]
US 20180253483A1 · Lee · 2018 [cited by applicant]
US 20180357264A1 · Rice · 2018 [cited by applicant]
US 20190332582A1 · Kumar · 2019 [cited by applicant]
US 20200026714A1 · Brodt · 2020 [cited by applicant]
US 20200034365A1 · Martin · 2020 [cited by applicant]
US 20200320059A1 · Kumar · 2020 [cited by applicant]
US 20200364185A1 · Beier · 2020 [cited by applicant]
WO WO0207006A1 · 2002 [cited by examiner]
WO WO2022082406A1 · 2022 [cited by examiner]
“About CDC Replication”, IBM, Mar. 19, 2021, 5 pages, <https://www.ibm.com/support/knowledgecenter/SSTRGZ_11.3.3/com.ibm.cdcdoc.sysreq.doc/concepts/aboutcdc.html>. [cited by applicant]
“IBM InfoSphere Data Replication (Q and SQL Replication) and related PIDs, considerations for GDPR readiness”, Mar. 22, 2021 ,9 pages, <https://www.ibm.com/support/knowledgecenter/SSTRGZ_11.4.0/com.ibm.swg.im.iis.db.rep… [cited by applicant]
“Synopsis tables”, IBM Documentation, May 9, 2021, 3 pages, < https://www.ibm.com/docs/en/db2/11.5?topic=organization-synopsis-tables>. [cited by applicant]
Beier et al., “Dynamic Selection of Data Apply Strategy During Database Replication”, U.S. Appl. No. 16/815,415, filed Mar. 11, 2020, 53 pages. [cited by applicant]
Db2 tech talk using info sphere information server with db2, May 9, 21, 46 pages, <https://www.slideshare.net/albertspijkers/db2-tech-talk-using-info-sphere-information-server-with-db2>. [cited by applicant]
IBM, “Configuring the Db2 connector to use direct connections”, IBM Documentation, Search in InfoSphere Information Server 11.7.0, printed on Jul. 27, 2021, 4 pages, <https://www.ibm.com/support/knowledgecenter/en/SSZJP… [cited by applicant]
Sait et al., “Strategies for Migrating Oracle Databases to AWS”, Last Update: Aug. 2018, Amazon Web Services, 38 pages, <https://d1.awsstatic.com/whitepapers/strategies-for-migrating-oracle-database-to-aws.pdf?did=wp_ca… [cited by applicant]
Stolze et al., “IDAA Cluster Load”, printed on Jul. 27, 2021, 3 pages. [cited by applicant]
Beier et al., “Data Distribution in Target Database Systems”, U.S. Appl. No. 17/448,715, filed Sep. 24, 2021, 48 pages. [cited by applicant]
IBM Appendix P, list of patents and patent applications treated as related, Filed Herewith, 2 pages. [cited by applicant]