IP Library Granted Patent US 11,061,734
Granted Patent B2
US 11,061,734 · App. 16/264,399 · Granted Jul 13, 2021

Performing customized data compaction for efficient parallel data processing amongst a set of computing resources

Inventors: Zhidong Ke (Milpitias, CA); Kevin Terusaki (Oakland, CA); Praveen Innamuri (Sunnyvale, CA); Narek Asadorian (San Francisco, CA)
Assignee: salesforce.com, inc.
G06F9/5055G06F9/52G06F16/1744
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,061,734
App. No.
16/264,399
Granted
Jul 13, 2021
Kind
B2
Abstract

Described is a system and method for compacting data into customized (e.g. optimal) file sizes for processing by computing resources. The mechanism may leverage various computing resources such as a cluster computing frameworks combined with a stream processing platform to efficiently process the activity data. For example, activity data of an organization may be processed by a set of jobs (or sub-jobs) as part of a data stream by a set of distributed computing resources. In order to efficiently process such data, the mechanism may compact the data into customized (e.g. optimal) file sizes. For example, the customized file sizes may provide an optimal (or near optimal) amount of data to be processed by each job, for example, to improve performance.

Claims (50)

1. A system comprising:

one or more processors; and

a non-transitory computer readable medium storing a plurality of instructions, which when executed, cause the one or more processors to:

store metadata created from activity data associated with an organization, the activity data stored in one or more activity data files as part of a data repository;

estimate, using the metadata, an amount of activity data created for the organization for a specified time period comprising determining a total number and average file size of the activity data files and identifying a time span of activity data stored in the activity data files;

determine an optimal customized activity data file size for processing the activity data by a set of computing resources;

determine a number of specified time periods required to create a compacted activity data file having the optimal customized activity data file size based on the estimated amount of activity data;

assign a set of jobs to compact the activity data for the determined number of the specified time periods comprising determining a number of jobs required to create the set of the compacted activity data files for the time span of activity; and

initiate the computing resources to execute the set of jobs to create a set of the compacted activity data files, one or more of the jobs executing concurrently.

2. The system of claim 1 , wherein the specified time period includes one or more days, and each of the jobs compacts activity data for the number of days required to create the compacted activity file having the customized activity data file size.

3. The system of claim 1 , wherein storing the metadata created from activity data associated with an organization includes:

collecting metadata from the data repository; and

updating, by scheduling a real-time job, to update a metadata repository.

4. The system of claim 1 , the plurality of instructions when executed further causing the one or more processors to:

track the set of jobs by storing a current state for each job; and

update each of the jobs to a completed state once the compacted activity file is created.

5. The system of claim 4 , the plurality of instructions when executed further causing the one or more processors to:

resume the creation of the compacted activity data file in response to a failure of a job by referencing the stored current state for the job.

6. A computer program product, comprising a non-transitory computer readable medium having a computer-readable program code embodied therein to be executed by one or more processors, the program code including instructions to:

store metadata created from activity data associated with an organization, the activity data stored in one or more activity data files as part of a data repository;

estimate, using the metadata, an amount of activity data created for the organization for a specified time period comprising determining a total number and average file size of the activity data files and identifying a time span of activity data stored in the activity data files;

determine an optimal customized activity data file size for processing the activity data by a set of computing resources;

determine a number of specified time periods required to create a compacted activity data file having the optimal customized activity data file size based on the estimated amount of activity data;

assign a set of jobs to compact the activity data for the determined number of the specified time periods comprising determining a number of jobs required to create the set of the compacted activity data files for the time span of activity; and

initiate the computing resources to execute the set of jobs to create a set of the compacted activity data files, one or more of the jobs executing concurrently.

7. The computer program product of claim 6 , wherein the specified time period includes one or more days, and each of the jobs of the set of jobs compacts activity data for the number of days required to create the compacted activity file having the customized activity data file size.

8. The computer program product of claim 6 , wherein storing the metadata created from activity data associated with an organization includes:

collecting metadata from the data repository; and

updating, by scheduling a real-time job, to update a metadata repository.

9. The computer program product of claim 6 , the program code including further instructions to:

track the set of jobs by storing a current state for each job; and

update each of the jobs to a completed state once the compacted activity file is created.

10. The computer program product of claim 9 , the program code including further instructions to:

resume the creation of the compacted activity data file in response to a failure of a job by referencing the stored current state for the job.

11. A method comprising:

storing, by a database system, metadata created from activity data associated with an organization, the activity data stored in one or more activity data files as part of a data repository;

estimating, by the database system using the metadata, an amount of activity data created for the organization for a specified time period comprising determining a total number and average file size of the activity data files and identifying a time span of activity data stored in the activity data files;

determining, by the database system, an optimal customized activity data file size for processing the activity data by a set of computing resources;

determining, by the database system, a number of specified time periods required to create a compacted activity data file having the customized activity data file size based on the estimated amount of activity data;

assigning, by the database system, a set of jobs to compact the activity data for the determined number of the specified time periods comprising determining a number of jobs required to create the set of the compacted activity data files for the time span of activity; and

initiating, by the database system, the computing resources to execute the set of jobs to create a set of the compacted activity data files, one or more of the jobs executing concurrently.

12. The method of claim 11 , wherein assigning the set of jobs to compact the activity data for the number of the specified time periods includes:

determining, by the database system, a number of jobs required for the set of jobs to create the set of the compacted activity data files for the time span of activity.

13. The method of claim 12 , wherein the specified time period includes one or more days, and each of the jobs compacts activity data for the number of days required to create the compacted activity file having the customized activity data file size.

14. The method of claim 11 , wherein storing the metadata created from activity data associated with an organization includes:

collecting metadata from the data repository; and

updating, by scheduling a real-time job, to update a metadata repository.

15. The method of claim 11 , further comprising:

tracking the set of jobs by storing a current state for each job; and

updating each of the jobs to a completed state once the compacted activity file is created.

Assignments (3)
CHANGE OF NAME Recorded Sep 20, 2023
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 064975/0956 →
CORRECTIVE ASSIGNMENT TO CORRECT THE GIVEN NAME OF INVENTOR INNAMURI TO RECITE-- PRAVEEN-- PREVIOUSLY RECORDED AT REEL: 048214 FRAME: 0056. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 21, 2021
From: KE, ZHIDONG; TERUSAKI, KEVIN; INNAMURI, PRAVEEN; ASADORIAN, NAREK
To: SALESFORCE.COM, INC.
Reel/Frame 056325/0620 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2019
From: KE, ZHIDONG; TERUSAKI, KEVIN; INNAMURI, TERUSAKI; ASADORIAN, NAREK
To: SALESFORCE.COM, INC.
Reel/Frame 048214/0056 →
Continuity (1)
Related Publication 20200250007A1 · Aug 6, 2020