IP Library Granted Patent US 11,544,290
Granted Patent B2
US 11,544,290 · App. 16/741,479 · Granted Jan 3, 2023

Intelligent data distribution and replication using observed data access patterns

Inventors: Stefano Braghin (Dublin, IE); Srikumar Venugopal (Dublin, IE)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F16/278G06F16/215G06F16/285G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,544,290
App. No.
16/741,479
Granted
Jan 3, 2023
Kind
B2
Abstract

Embodiments for providing intelligent data replication and distribution in a computing environment. Data access patterns of one or more queries issued to a plurality of data partitions may be forecasted. Data may be dynamically distributed and replicated to one or more existing data partitions or additional of the plurality of data partitions according to the forecasting.

Claims (29)

1. A method for providing intelligent data replication and distribution by one or more processors, comprising:

forecasting data access patterns of one or more queries issued to a plurality of data partitions using pattern data identified from historical queries specifically directed to join operations and issued to at least two of the plurality of data partitions in conjunction with one another, wherein the pattern data is used to correlate those of the at least two of the plurality of data partitions to a particular type of the join operations specified in a respective query of the one or more queries; and

dynamically distributing and replicating data to one or more existing data partitions or additional of the plurality of data partitions according to the forecasting.

2. The method of claim 1 , further including using a garbage collection operation to remove excess partitions to facilitate future replications of different partition groups.

3. The method of claim 1 , further including identifying one or more clusters of the plurality of data partitions being simultaneously queried with the one or more queries.

4. The method of claim 1 , further including replicating one or more new copies of those of the plurality of data partitions having a greater frequency of access or use as compared to other data partitions of the plurality of data partitions.

5. The method of claim 1 , further including replicating one or more new copies of those of the plurality of data partitions according to a query type of the one or more queries and identifying those of the plurality of data partitions ready for distribution.

6. The method of claim 1 , further including replicating one or more new copies of those of the plurality of data partitions from a plurality of different nodes together as a cluster on a single node.

7. The method of claim 1 , further including replicating one or more new copies of those of the plurality of data partitions from a plurality of different nodes together as co-located cluster across a plurality of nodes.

8. A system for providing intelligent data replication and distribution in a computing environment, comprising:

one or more computers with executable instructions that when executed cause the system to:

forecast data access patterns of one or more queries issued to a plurality of data partitions using pattern data identified from historical queries specifically directed to join operations and issued to at least two of the plurality of data partitions in conjunction with one another, wherein the pattern data is used to correlate those of the at least two of the plurality of data partitions to a particular type of the join operations specified in a respective query of the one or more queries; and

dynamically distribute and replicate data to one or more existing data partitions or additional of the plurality of data partitions according to the forecasting.

9. The system of claim 8 , wherein the executable instructions use a garbage collection operation to remove excess partitions to facilitate future replications of different partition groups.

10. The system of claim 8 , wherein the executable instructions identify one or more clusters of the plurality of data partitions being simultaneously queried with the one or more queries.

11. The system of claim 8 , wherein the executable instructions replicate one or more new copies of those of the plurality of data partitions having a greater frequency of access or use as compared to other data partitions of the plurality of data partitions.

12. The system of claim 8 , wherein the executable instructions replicate one or more new copies of those of the plurality of data partitions according to a query type of the one or more queries and identifying those of the plurality of data partitions ready for distribution.

13. The system of claim 8 , wherein the executable instructions replicate one or more new copies of those of the plurality of data partitions from a plurality of different nodes together as a cluster on a single node.

14. The system of claim 8 , wherein the executable instructions replicate one or more new copies of those of the plurality of data partitions from a plurality of different nodes together as co-located cluster across a plurality of nodes.

15. A computer program product for providing intelligent data replication and distribution in a computing environment by a processor, the computer program product comprising a non-transitory computer-readable storage medium having computer-readable program code portions stored therein, the computer-readable program code portions comprising:

an executable portion that forecasts data access patterns of one or more queries issued to a plurality of data partitions using pattern data identified from historical queries specifically directed to join operations and issued to at least two of the plurality of data partitions in conjunction with one another, wherein the pattern data is used to correlate those of the at least two of the plurality of data partitions to a particular type of the join operations specified in a respective query of the one or more queries; and

an executable portion that dynamically distributes and replicates data to one or more existing data partitions or additional of the plurality of data partitions according to the forecasting.

16. The computer program product of claim 15 , wherein the executable portion uses a garbage collection operation to remove excess partitions to facilitate future replications of different partition groups.

17. The computer program product of claim 15 , wherein the executable portion identifies one or more clusters of the plurality of data partitions being simultaneously queried with the one or more queries.

18. The computer program product of claim 15 , wherein the executable portion replicates one or more new copies of those of the plurality of data partitions having a greater frequency of access or use as compared to other data partitions of the plurality of data partitions.

19. The computer program product of claim 15 , wherein the executable portion:

replicates one or more new copies of those of the plurality of data partitions according to a query type of the one or more queries and identifying those of the plurality of data partitions ready for distribution; or

replicates one or more new copies of those of the plurality of data partitions from a plurality of different nodes together as a cluster on a single node.

20. The computer program product of claim 15 , wherein the executable portion replicates one or more new copies of those of the plurality of data partitions from a plurality of different nodes together as co-located cluster across a plurality of nodes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2020
From: BRAGHIN, STEFANO; VENUGOPAL, SPIKUMAR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 051500/0208 →
Continuity (1)
Related Publication 20210216572A1 · Jul 15, 2021