IP Library Granted Patent US 9,898,504
Granted Patent B1
US 9,898,504 · App. 14/520,249 · Granted Feb 20, 2018

System, method, and computer program for accessing data on a big data platform

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,898,504
App. No.
14/520,249
Granted
Feb 20, 2018
Kind
B1
Abstract

A system, method, and computer program product are provided for accessing data on a big data platform. In use, a request associated with a data processing job to process data stored in a big data store is identified, the data being stored in a plurality of rows with each row being associated with a unique key. Additionally, a data processing job input associated with the request is received, the data processing job input including a set of keys required to be read for processing. Further, the set of keys is translated into one or more queries, the one or more queries including at least one of a request to read an individual key or a request to read a range of keys. Moreover, the data is loaded from the big data store based on the one or more queries.

Claims (28)

1. A method, comprising:

identifying, by a computer processor, a request associated with a data processing job to process data stored in a big data store, the data being stored in a plurality of rows with each row being associated with a unique key, and the big data store supporting both key based random access to individual rows and range based retrieval of rows given a start row key and an end row key;

receiving, by the computer processor, a data processing job input associated with the request, the data processing job input including a set of keys required to be read for processing;

translating, by the computer processor using an algorithm, the set of keys into one or more queries to the big data store, the one or more queries including a request to read an individual key and a request to read a range of keys;

loading, by the computer processor to the data processing job, the data from the big data store by executing the one or more queries; and

responsive to loading the data to the data processing job, processing, by the computer processor, the data by the data processing job.

2. The method of claim 1 , further comprising splitting the request into multiple queries based on prior knowledge of the data processing job input.

3. The method of claim 1 , further comprising splitting the request into multiple queries by algorithmically determining an optimal set of queries to be performed by heuristically approximating an amount of redundant data to be loaded, balancing sequential disk reads and potential reading of redundant data and fine grained random access reads.

4. The method of claim 1 , wherein the big data store includes a plurality of hard drives directly connected to a plurality of host systems.

5. The method of claim 1 , wherein the plurality of rows are stored on one or more disks sorted by an associated unique key with adjacent rows placed on common physical disk blocks.

6. A computer program product embodied on a non-transitory computer readable medium, comprising:

computer code for identifying, by a computer processor, a request associated with a data processing job to process data stored in a big data store, the data being stored in a plurality of rows with each row being associated with a unique key, and the big data store supporting both key based random access to individual rows and range based retrieval of rows given a start row key and an end row key;

computer code for receiving, by a computer processor, a data processing job input associated with the request, the data processing job input including a set of keys required to be read for processing;

computer code for translating, by a computer processor using an algorithm, the set of keys into one or more queries to the big data store, the one or more queries including a request to read an individual key and a request to read a range of keys;

computer code for loading, by a computer processor to the data processing job, the data from the big data store by executing the one or more queries; and

responsive to loading the data to the data processing job, processing, by the computer processor, the data by the data processing job.

7. The computer program product of claim 6 , further comprising computer code for splitting the request into multiple queries based on prior knowledge of the data processing job input.

8. The computer program product of claim 6 , further comprising computer code for splitting the request into multiple queries by algorithmically determining an optimal set of queries to be performed by heuristically approximating an amount of redundant data to be loaded balancing sequential disk reads and potential reading of redundant data and fine grained random access reads.

9. The computer program product of claim 6 , wherein the computer program product is operable such that the big data store includes a plurality of hard drives directly connected to a plurality of host systems.

10. The computer program product of claim 6 , wherein the computer program product is operable such that the plurality of rows are stored on one or more disks sorted by an associated unique key with adjacent rows placed on common physical disk blocks.

11. A system comprising:

a memory system; and

one or more processing cores coupled to the memory system and that are each configured to:

identify a request associated with a data processing job to process data stored in a big data store, the data being stored in a plurality of rows with each row being associated with a unique key, and the big data store supporting both key based random access to individual rows and range based retrieval of rows given a start row key and an end row key;

receive a data processing job input associated with the request, the data processing job input including a set of keys required to be read for processing;

translate, using an algorithm, the set of keys into one or more queries to the big data store, the one or more queries including a request to read an individual key and a request to read a range of keys;

load, to the data processing job, the data from the big data store by executing the one or more queries; and

responsive to loading the data to the data processing job, process the data by the data processing job.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2016
From: AMDOCS SOFTWARE SYSTEMS LIMITED
To: AMDOCS DEVELOPMENT LIMITED; AMDOCS SOFTWARE SYSTEMS LIMITED
Reel/Frame 039695/0965 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2015
From: PEDHAZUR, NIR; ROTEM-GAL-OZ, ARNON WALTER; KAFKA, OREN; GOFER, ZOHAR
To: AMDOCS SOFTWARE SYSTEMS LIMITED
Reel/Frame 035135/0672 →