IP Library › Granted Patent US 12,271,396
Granted Patent B2
US 12,271,396 · App. 18/225,827 · Granted Apr 8, 2025

Discovery of discrete partitioning information

Inventors: Rohit Jaykumar Gattani (Pleasanton, CA); Rahul Gupta (Dublin, CA)
Assignee: Oracle International Corporation
G06F16/278G06F16/254
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,271,396
App. No.
18/225,827
Granted
Apr 8, 2025
Kind
B2
Abstract

A system for data partitioning based on discovery of discrete partitioning information. The system can receive data sets in table format from source system. The data can be stored in the source system to be partitioned and transmitted from the source system to a target system. The system can determine a respective partitioning column for each data set. The system can determine a number of partitions. The system can determine, for each data set, a respective set of discrete values from the plurality of discrete values of the respective partitioning column. The number of the discrete values of the set of discrete values can be based at least in part on the number of partitions. The system can the discrete value sets with each other. The system can determine a final set of discrete values based at least in part on the comparison.

Claims (63)

1. A method, comprising:

receiving, by a computing system, a plurality of data sets sampled from data stored in a source system, wherein partitioning information is not available or not provided at the source system;

extracting, by the computing system and for each data set, a respective partitioning column, each respective partitioning column comprising a plurality of discrete values;

extracting, by the computing system and for each data set, a respective set of discrete values from the plurality of discrete values of the respective partitioning column;

generating, by the computing system, a respective cumulative deviation score for each respective set of discrete values;

comparing, by the computing system, the generated cumulative deviation scores for the respective sets of discrete values;

extracting, by the computing system, a final set of discrete values as the partitioning information to be used for partitioning the data based at least in part on the compared cumulative deviation scores; and

partitioning, by the computing system, the data stored in the source system based at least in part on the partitioning information.

2. The method of claim 1 , wherein the extracting of the respective set of discrete values from the plurality of discrete values of the respective partitioning column comprises:

determining a frequency of each discrete value of the plurality of discrete values in the respective partitioning column;

comparing each frequency of each discrete value to a threshold frequency; and

determining the respective set of discrete values based at least in part on the comparison of each frequency of each discrete value to the threshold frequency.

3. The method of claim 1 , wherein the method further comprises:

determining a number of partitions to partition the data stored in the source system.

4. The method of claim 3 , wherein the method further comprises:

comparing the number of partitions to a number of discrete values of the respective set of discrete values;

determining that the number of partitions is different than the number of discrete values of the respective set of discrete values; and

grouping two or more discrete values of the respective set of discrete values such that the number of partitions corresponds to the number of discrete values.

5. The method of claim 3 , wherein the method further comprises:

determining the number of the discrete values of the set of discrete values being based at least in part on the number of partitions.

6. The method of claim 5 , wherein the determining of the number of the discrete values of the set of discrete values comprises determining a distribution of the discrete values.

7. The method of claim 3 , wherein the determining of the number of partitions comprises:

determining a capability of a downstream process to receive the partitioned data, wherein the number of partitions is based at least in part on the capability of the downstream process.

8. The method of claim 3 , wherein the method further comprises:

initializing a set of virtual machines based at least in part on the number partitions; and

transmitting, using the set of virtual machines, the partitioned data in parallel from the source system to a target system.

9. The method of claim 1 , wherein the generating of the respective cumulative deviation score for each respective set of discrete values comprises:

comparing each discrete value of each set of discrete values with each discrete value of each other respective set of discrete values to determine a deviation of discrete values, wherein the respective cumulative deviation score for each respective set of discrete values is based at least in a part on the deviation of discrete values.

10. A computing system, comprising:

one or more processors; and

a computer-readable medium including instructions that, when executed by the one or more processors, cause the one or more processors to:

receive a plurality of data sets sampled from data stored in a source system, wherein partitioning information is not available or not provided at the source system;

extract, for each data set, a respective partitioning column, each respective partitioning column comprising a plurality of discrete values;

extract, for each data set, a respective set of discrete values from the plurality of discrete values of the respective partitioning column;

generate a respective cumulative deviation score for each respective set of discrete values;

compare the generated cumulative deviation scores for each the respective sets of discrete values;

extract a final set of discrete values as the partitioning information to be used for partitioning the data based at least in part on the compared cumulative deviation scores; and

partition the data stored in the source system based at least in part on the partitioning information.

11. The computing system of claim 10 , wherein the extracting of the respective set of discrete values from the plurality of discrete values of the respective partitioning column comprises:

determining a frequency of each discrete value of the plurality of discrete values in the respective partitioning column;

comparing each frequency of each discrete value to a threshold frequency; and

determining the respective set of discrete values based at least in part on the comparison of each frequency of each discrete value to the threshold frequency.

12. The computing system of claim 10 , wherein the instructions that, when executed by the one or more processors, further cause the one or more processors to:

determine a number of partitions to partition the data stored in the source system.

13. The computing system of claim 10 , wherein

determining the generating of the respective cumulative deviation score for each respective set of discrete values comprises:

comparing each discrete value of each set of discrete values with each discrete value of each other respective set of discrete values to determine a deviation of discrete values, wherein the respective cumulative deviation score for each respective set of discrete values is based at least in a part on the deviation of discrete values.

14. A non-transitory computer-readable medium including stored thereon a sequence of instructions that, when executed by one or more processors, causes the one or more processors to:

receive a plurality of data sets sampled from data stored in a source system, wherein partitioning information is not available or not provided at the source system;

extract, for each data set, a respective partitioning column, each respective partitioning column comprising a plurality of discrete values;

extract, for each data set, a respective set of discrete values from the plurality of discrete values of the respective partitioning column;

generate a respective cumulative deviation score for each respective set of discrete values;

compare the generated cumulative deviation scores for the respective sets of discrete values;

extract a final set of discrete values as the partitioning information to be used for partitioning the data based at least in part on the compared cumulative deviation scores; and

partition the data stored in the source system based at least in part on the partitioning information.

15. The non-transitory computer-readable medium of claim 14 , wherein the instructions that, when executed by the one or more processors, further cause the one or more processors to:

determine a number of partitions to partition the data stored in the source system.

16. The non-transitory computer-readable medium of claim 15 , wherein the instructions that, when executed by the one or more processors, further cause the one or more processors to:

compare the number of partitions to a number of discrete values of the respective set of discrete values;

determine that the number of partitions is different than the number of discrete values of the respective set of discrete values; and

group two or more discrete values of the respective set of discrete values such that the number of partitions corresponds to the number of discrete values.

17. The non-transitory computer-readable medium of claim 14 , wherein the generating of the respective cumulative deviation score for each respective set of discrete values comprises:

comparing each discrete value of each set of discrete values with each discrete value of each other respective set of discrete values to determine a deviation of discrete values, wherein the respective cumulative deviation score for each respective set of discrete values is based at least in a part on the deviation of discrete values.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNED PROPERTY/APPLICATION NUMBER FROM 18225825 TO 18225827 PREVIOUSLY RECORDED AT REEL: 064373 FRAME: 0469. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 28, 2023
From: GATTANI, ROHIT JAYKUMAR; GUPTA, RAHUL
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 066253/0789 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2023
From: GATTANI, ROHIT JAYKUMAR; GUPTA, RAHUL
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 064373/0469 →
Continuity (1)
Related Publication 20250036652A1 · Jan 30, 2025
References Cited (19)
US 7024401B2 · Harper et al. · 2006 [cited by applicant]
US 7814142B2 · Mamou et al. · 2010 [cited by applicant]
US 10356150B1 · Meyers · 2019 [cited by applicant]
US 11621966B1 · Yu · 2023 [cited by examiner]
US 20050160055A1 · Boulle · 2005 [cited by examiner]
US 20080313246A1 · Shankar · 2008 [cited by examiner]
US 20110066593A1 · Ahluwalia et al. · 2011 [cited by applicant]
US 20120143090A1 · Hay · 2012 [cited by examiner]
US 20150356149A1 · Dagli et al. · 2015 [cited by applicant]
US 20200125666A1 · Eadon · 2020 [cited by examiner]
US 20200311062A1 · Mihm et al. · 2020 [cited by applicant]
US 20220261390A1 · Creasey · 2022 [cited by examiner]
US 20240202210A1 · Gattani · 2024 [cited by examiner]
CN 110737683 · 2020 [cited by applicant]
KR 20220096049 · 2022 [cited by applicant]
“Source Partitioning”, Amazon Redshift Connectors, Cloud Data Integration Connectors, Nov. 29, 2022, 1 page. [cited by applicant]
Ives et al., “Adapting to Source Properties in Processing Data Integration Queries”, Available Online at: https://homes.cs.washington.edu/˜alon/files/aqp04.pdf, Jun. 13-18, 2004, 12 pages. [cited by applicant]
Vieira, “Use PK Chunking to Extract Large Data Sets from Salesforce”, Available Online at: https://developer.salesforce.com/blogs/engineering/2015/03/use-pk-chunking-extract-large-data-sets-salesforce, Mar. 23, 2015, 4 … [cited by applicant]
U.S. Appl. No. 18/084,421 , “Non-Final Office Action”, Feb. 27, 2024, 27 pages. [cited by applicant]