IP Library › Granted Patent US 8,775,425
Granted Patent B2
US 8,775,425 · App. 12/862,273 · Granted Jul 8, 2014

Systems and methods for massive structured data management over cloud aware distributed file system

Inventors: Himanshu Gupta (New Delhi, IN); Rajeev Gupta (Noida, IN); Mukesh Kumar Mohania (Rajpur Chung, IN); Ullas Balan Nambiar (Haryana, IN)
Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,775,425
App. No.
12/862,273
Granted
Jul 8, 2014
Kind
B2
Abstract

Methods and arrangements for accommodating a query, directing the query to datasets, creating partitions and partitioning the datasets, and returning a response to the query, the response being structured in accordance with the created partitions.

Claims (45)

1. A method comprising:

accommodating a query;

directing the query to datasets which include data from files in a distributed file system;

creating partitions of the datasets, wherein said creating comprises creating smart replicas, and wherein the smart replicas comprise reordered content of the datasets;

indexing the partitions via mapping partition keys to a database atop the distributed file system;

parsing the query;

looking up indexed values based on the parsed query;

identifying partitions based on said looking up of indexed values;

rewriting the query, with partition information, into the database atop the distributed file system; and

returning a response to the query, wherein the response is structured based on the created partitions.

2. The method according to claim 1 , wherein said creating of partitions comprises creating horizontal partitions.

3. The method according to claim 2 , wherein the horizontal partitions comprise tuples of high similarity.

4. The method according to claim 1 , wherein the datasets include data from a data warehouse.

5. The method according to claim 1 , comprising repartitioning each partition locally.

6. The method according to claim 1 , wherein said indexing comprises employing inverted index maps identifying pairs where attributes are equivalent to predetermined values.

7. The method according to claim 6 , wherein said employing of inverted index maps comprises indicating occurrences of attribute-value equivalencies among partitions.

8. The method according to claim 1 , wherein said creating of smart replicas comprises keeping replicas sorted over different attribute sets.

9. An apparatus comprising:

one or more processors; and

a computer readable storage medium having computer readable program code embodied therewith and executable by the one or more processors, the computer readable program code being configured to:

accommodate a query;

direct the query to datasets which include data from files in a distributed file system;

create partitions of the datasets, wherein to create partitions comprises creating smart replicas, and wherein the smart replicas comprise reordered content of the datasets;

index the partitions via mapping partition keys to a database atop the distributed file system;

parse the query;

look up indexed values based on the parsed query;

identify partitions based on the looking up of indexed values;

rewrite the query, with partition information, into the database atop distributed file system; and

return a response to the query, wherein the response is structured based on the created partitions.

10. A computer program product embedded in a computer readable storage medium embodied with computer readable program code which, when executed, causes a computing device to perform operations, the computer readable program code being configured to:

accommodate a query;

direct the query to datasets which include data from files in a distributed file system;

create partitions of the datasets, wherein to create partitions comprises creating smart replicas, and wherein the smart replicas comprise reordered content of the datasets;

index the partitions via mapping partition keys to a database atop the distributed file system;

parse the query;

look up indexed values based on the parsed query;

identify partitions based on the looking up of indexed values;

rewrite the query, with partition information, into the database atop distributed file system; and

return a response to the query, wherein the response is structured based on the created partitions.

11. The computer program product according to claim 10 , wherein said computer readable program code is configured to create horizontal partitions.

12. The computer program product according to claim 10 , wherein the datasets include data from a data warehouse.

13. The computer program product according to claim 10 , wherein said computer readable program code is configured to re-partition each partition locally.

14. The computer program product according to claim 10 , wherein said computer readable program code is configured to employ inverted index maps identifying pairs where attributes are equivalent to predetermined values.

15. The computer program product according to claim 14 , wherein said computer readable program code is configured to indicate occurrences of attribute-value equivalencies among partitions.

16. The computer program product according to claim 10 , wherein said computer readable program code is configured to keep replicas sorted over different attribute sets.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 1, 2010
From: GUPTA, HIMANSHU; GUPTA, RAJEEV; MOHANIA, MUKESH KUMAR; NAMBIAR, ULLAS BALAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 024924/0620 →
Continuity (1)
Related Publication 20120054182A1 · Mar 1, 2012