IP Library › Granted Patent US 12,737,371
Granted Patent B2
US 12,737,371 · App. 19/073,997 · Granted Sep 15, 2026

Pipeline for efficient processing of structured data

Inventors: Shucheng Liang (Shanghai, CN); Tian Liu (Shanghai, CN); Amy Chen (Shanghai, CN)
Assignee: Capital One Financial Corporation
G06F16/248G06F16/211G06F16/221G06F16/2255
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,371
App. No.
19/073,997
Granted
Sep 15, 2026
Kind
B2
Abstract

A computing platform is configured to (i) identify one or more data sets to be used as input for one or more data processing operations, (ii) determine a computing resource allocation, (iii) receive an indication of a user-defined function, the user-defined function corresponding to a desired output of the one or more data processing operations, (iv) based on (a) the computing resource allocation and (b) the user-defined function, divide the data set into a plurality of data subsets each to be processed in a respective sub-task by one of a plurality of parallel worker processes, (v) execute, in each worker process simultaneously, a respective batch of sub-tasks by processing corresponding batch of respective data subsets using the user-defined function to generate a respective batch result, (vi) combine the respective results from each worker process to generate a combined result, and (vii) output the combined result.

Claims (78)

1 . A computing platform comprising:

at least one processor;

a non-transitory computer-readable medium; and

program instructions stored on the non-transitory computer-readable medium that, when executed by the at least one processor, cause the computing platform to:

receive, from a client station, an indication of one or more data sets to be used as input for one or more data processing operations;

receive, from the client station, an indication of a computing resource allocation to be used for the one or more data processing operations;

receive, from the client station, an indication of at least one user-defined function to be applied to the one or more data sets during the one or more data processing operations, the at least one user-defined function comprising at least one operator to be applied to the one or more data sets;

determine that a size of the one or more data sets exceeds a physical memory capacity of the computing resource allocation;

based on (i) determining that the size of the one or more data sets exceeds a physical memory capacity of the computing resource allocation and (ii) the at least one operator:

decompose the at least one operator into a plurality of sub-tasks; and

divide the one or more data sets into a plurality of data subsets, each data subset comprising an array of values from the one or more data sets to be processed in a respective sub-task of the plurality of sub-tasks by one of a plurality of parallel worker processes;

using the computing resource allocation, execute, in each worker process simultaneously, a respective batch of sub-tasks by processing a corresponding batch of respective data subsets using the at least one user-defined function;

based on processing a corresponding batch of respective data subsets using the at least one user-defined function, generate a respective batch result associated with the worker process;

combine the respective batch results from each worker process of the plurality of parallel worker processes to generate a combined result; and

output, to the client station, the combined result.

2 . The computing platform of claim 1 , wherein:

the one or more data sets comprises a plurality of rows of data and one or more columns of data; and

wherein the program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets comprise program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets each comprising one or more rows of the plurality of rows of data.

3 . The computing platform of claim 2 , wherein the program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets comprise program instructions that, when executed by the at least one processor, cause the computing platform to divide the one or more data sets into a plurality of data subsets each comprising one or more rows of the plurality of rows of data by randomly assigning one or more rows of the plurality of rows of data to a data subset of the plurality of data subsets.

4 . The computing platform of claim 2 , wherein the program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets comprise program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets each comprising one or more rows of the plurality of rows of data by:

assigning each row of the plurality of rows of data a hash key from a plurality of hash keys; and

grouping rows of the plurality of rows of data based on assigned hash keys.

5 . The computing platform of claim 1 , wherein:

the one or more data sets comprises a plurality of columns of data and one or more rows of data; and

wherein the program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets comprise program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets, wherein each data subset of the plurality of data subsets comprises one or more columns of the plurality of columns of data.

6 . The computing platform of claim 5 , wherein the program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets comprise program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets, wherein each data subset of the plurality of data subsets comprises one or more columns of the plurality of columns of data, by one or both of (i) randomly assigning one or more columns of the plurality of columns of data to a data subset of the plurality of data subsets and (ii) assigning one or more columns of the plurality of columns of data to a data subset of the plurality of data subsets based on a user-defined schema.

7 . The computing platform of claim 1 , wherein the program instructions that cause the computing platform to receive, from a client station, an indication of one or more data sets to be used as input for one or more data processing operations comprise program instructions that cause the computing platform to receive, from the client station, a first data set and at least a second data set to be used as input for the one or more data processing operations, wherein the one or more data processing operations comprise joining the first data set with at least the second data set to form a merged data set; and

wherein the program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets comprise program instructions that cause the computing platform to divide each data set into a plurality of data subsets each comprising one or more rows of a plurality of rows of data by:

assigning each row of the plurality of rows of data a hash key from a plurality of hash keys; and

grouping rows of the plurality of rows of data based on assigned hash keys; and

wherein the computing platform further comprises program instructions stored on the non-transitory computer-readable medium that, when executed by the at least one processor, cause the computing platform to join each respective subset of the first data set with corresponding subsets of at least the second data set having the same hash keys based on a user-defined merge key, thereby forming the merged data set.

8 . The computing platform of claim 1 , wherein the computing resource allocation comprises a plurality of processor cores, wherein each processor core in the plurality of processor cores executes the respective batch of sub-tasks in a respective one of the plurality of parallel worker processes.

9 . The computing platform of claim 1 , wherein the program instructions that cause the computing platform to receive, from the client station, an indication of a computing resource allocation to be used for the one or more data processing operations comprise program instructions that cause the computing platform to receive, from the client station, an indication of the computing resource allocation based on total available computing resources.

10 . A non-transitory computer-readable medium, wherein the non-transitory computer-readable medium is provisioned with program instructions that, when executed by at least one processor, cause a computing platform to:

receive, from a client station, an indication of one or more data sets to be used as input for one or more data processing operations;

receive, from the client station, an indication of a computing resource allocation to be used for the one or more data processing operations;

receive, from the client station, an indication of at least one user-defined function to be applied to the one or more data sets during the one or more data processing operations, the at least one user-defined function comprising at least one operator to be applied to the one or more data sets;

determine that a size of the one or more data sets exceeds a physical memory capacity of the computing resource allocation:

based on (i) determining that the size of the one or more data sets exceeds a physical memory capacity of the computing resource allocation and (ii) the at least one operator:

decompose the at least one operator into a plurality of sub-tasks; and

divide the one or more data sets into a plurality of data subsets, each data subset comprising an array of values from the one or more data sets to be processed in a respective sub-task of the plurality of sub-tasks by one of a plurality of parallel worker processes;

using the computing resource allocation, execute, in each worker process simultaneously, a respective batch of sub-tasks by processing a corresponding batch of respective data subsets using the at least one user-defined function;

based on processing a corresponding batch of respective data subsets using the at least one user-defined function, generate a respective batch result associated with the worker process;

combine the respective batch results from each worker process of the plurality of parallel worker processes to generate a combined result; and

output, to the client station, the combined result.

11 . The non-transitory computer-readable medium of claim 10 , wherein:

the one or more data sets comprises a plurality of rows of data and one or more columns of data; and

wherein the program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets comprise program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets each comprising one or more rows of the plurality of rows of data.

12 . The non-transitory computer-readable medium of claim 11 , wherein the program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets comprise program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets each comprising one or more rows of the plurality of rows of data by randomly assigning one or more rows of the plurality of rows of data to a data subset of the plurality of data subsets.

13 . The non-transitory computer-readable medium of claim 11 , wherein the program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets comprise program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets each comprising one or more rows of the plurality of rows of data by:

assigning each row of the plurality of rows of data a hash key from a plurality of hash keys; and

grouping rows of the plurality of rows of data based on assigned hash keys.

14 . The non-transitory computer-readable medium of claim 10 , wherein:

the one or more data sets comprises a plurality of columns of data and one or more rows of data; and

wherein the program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets comprise program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets, wherein each data subset of the plurality of data subsets comprises one or more columns of the plurality of columns of data.

15 . The non-transitory computer-readable medium of claim 14 , wherein the program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets comprise program instructions that cause the computing platform to divide the one or more data sets into a plurality of data subsets, wherein each data subset of the plurality of data subsets comprises one or more columns of the plurality of columns of data, by one or both of (i) randomly assigning one or more columns of the plurality of columns of data to a data subset of the plurality of data subsets and (ii) assigning one or more columns of the plurality of columns of data to a data subset of the plurality of data subsets based on a user-defined schema.

16 . A method carried out by a computing platform, the method comprising:

receiving, from a client station, one or more data sets to be used as input for one or more data processing operations;

receiving, from the client station, a computing resource allocation to be used for the one or more data processing operations;

receiving, from the client station, an indication of at least one user-defined function to be applied to the one or more data sets during the one or more data processing operations, the at least one user-defined function comprising at least one operator to be applied to the one or more data sets;

determining that a size of the one or more data sets exceeds a physical memory capacity of the computing resource allocation:

based on (i) determining that the size of the one or more data sets exceeds a physical memory capacity of the computing resource allocation and (ii) the at least one operator:

decomposing the at least one operator into a plurality of sub-tasks; and

dividing the one or more data sets into a plurality of data subsets, each data subset comprising an array of values from the one or more data sets to be processed in a respective sub-task of the plurality of sub-tasks by one of a plurality of parallel worker processes;

using the computing resource allocation, executing, in each worker process simultaneously, a respective batch of sub-tasks by processing a corresponding batch of respective data subsets using the at least one user-defined function;

based on processing a corresponding batch of respective data subsets using the at least one user-defined function, generating a respective batch result associated with the worker process;

combining the respective batch results from each worker process of the plurality of parallel worker processes to generate a combined result; and

outputting, to the client station, the combined result.

17 . The method of claim 16 , wherein:

the one or more data sets comprises a plurality of rows of data and one or more columns of data; and

wherein dividing the one or more data sets into a plurality of data subsets comprises dividing the one or more data sets into a plurality of data subsets each comprising one or more rows of the plurality of rows of data.

18 . The method of claim 17 , wherein dividing the one or more data sets into a plurality of data subsets comprises:

assigning each row of the plurality of rows of data a hash key from a plurality of hash keys; and

grouping rows of the plurality of rows of data based on assigned hash keys.

19 . The method of claim 16 , wherein:

the one or more data sets comprises a plurality of columns of data and one or more rows of data; and

wherein dividing the one or more data sets into a plurality of data subsets comprises dividing the one or more data sets into a plurality of data subsets, wherein each data subset of the plurality of data subsets comprises one or more columns of the plurality of columns of data.

20 . The method of claim 19 , wherein dividing the one or more data sets into a plurality of data subsets comprises dividing the one or more data sets into a plurality of data subsets, wherein each data subset of the plurality of data subsets comprises one or more columns of the plurality of columns of data, by one or both of (i) randomly assigning one or more columns of the plurality of columns of data to a data subset of the plurality of data subsets and (ii) assigning one or more columns of the plurality of columns of data to a data subset of the plurality of data subsets based on a user-defined schema.

Assignments (2)
MERGER Recorded Jul 2, 2025
From: DISCOVER FINANCIAL SERVICES
To: CAPITAL ONE FINANCIAL CORPORATION
Reel/Frame 071784/0903 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 12, 2025
From: LIANG, SHUCHENG; LIU, TIAN; CHEN, AMY
To: DISCOVER FINANCIAL SERVICES
Reel/Frame 070483/0964 →
Continuity (2)
Continuation PCTCN2025076149 · Feb 7, 2025
Related Publication 20260236482A1 · Aug 13, 2026
References Cited (9)
US 9170848B1 · Goldman · 2015 [cited by examiner]
US 10437848B2 · Agarwal et al. · 2019 [cited by applicant]
US 11216454B1 · Cole · 2022 [cited by examiner]
US 20150379072A1 · Dirac et al. · 2015 [cited by applicant]
US 20210248143A1 · Khillar · 2021 [cited by examiner]
US 20240028591A1 · Balakrishnan · 2024 [cited by examiner]
US 20240289306A1 · Alickaj · 2024 [cited by examiner]
US 20250390276A1 · Liao · 2025 [cited by examiner]
CN 105677486B · 2019 [cited by applicant]