IP Library Patent Application 13671896
Patent Application
App. No. 13/671,896

SYSTEM AND METHOD FOR OPERATING A BIG-DATA PLATFORM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
13/671,896
Abstract

A system and method for operating a big-data platform that includes at a data analysis platform, receiving discrete client data; storing the client data in a network accessible distributed storage system that includes: storing the client data in a real-time storage system; and merging the client data into a columnar-based distributed archive storage system; receiving a data query request through a query interface; and selectively interfacing with the client data from the real-time storage system and archive storage system according to the query.

Claims (25)

1 . A method for operating a big-data platform comprising:

at a data analysis platform, receiving discrete client data;

storing the client data in a network accessible distributed storage system that includes:

storing the client data in a real-time storage system;

merging the client data into a columnar-based distributed archive storage system;

receiving a data query request through a query interface; and

selectively interfacing with the client data from the real-time storage system and archive storage system according to the query.

2 . The method of claim 1 , wherein selectively interfacing with the client data includes at a data-intensive processing cluster, processing the data query according to a data mapping and reduction processes in querying data from the real-time storage system and archive storage system.

3 . The method of claim 2 , wherein the data-intensive processing cluster is a Hadoop cluster and the data mapping and reduction process is a MapReduce implementation.

4 . The method of claim 1 , wherein discrete client data is received and stored with dynamic schema.

5 . The method of claim 4 , wherein the data query request includes a schema definition and wherein selectively interfacing with the client data includes applying the schema definition to the dynamic schema.

6 . The method of claim 1 , further comprising at a client data agent collecting client data and transmitting the client data to the data analysis platform.

7 . The method of claim 6 , wherein the client data agent is integrated into an event channel from which client data is collected.

8 . The method of claim 7 , wherein the event channel is selected from the list consisting of syslog, a relational database, cloud data, and sensor data.

9 . The method of claim 6 , further comprising at the client data agent serializing data into a binary serialization data-interchange that is transmitted to the data analysis platform.

10 . The method of claim 6 , wherein collecting client data is collected through a client agent data-input plugin.

11 . The method of claim 1 , wherein the columnar-based distributed archive storage system stores client data in time series order, and wherein selectively interfacing with client data includes querying data from distributed storage system.

12 . The method of claim 11 , wherein querying client data from distributed storage system includes cooperatively querying the real-time storage system and the archive storage system for a cohesive query result.

13 . The method of claim 1 , wherein receiving a data query includes converting relational database styled query to data-intensive cluster query process.

14 . The method of claim 1 , wherein the data query request is received through a infographics interface and further comprising returning an infographic from the selectively interfaced client data.

15 . The method of claim 1 , wherein receiving a data query includes receiving the data query through a business intelligence tool driver and further comprising returning data analytics results to the business intelligence tool driver.

16 . The method of claim 1 , wherein client data is associated with a user account through unique identifier.

17 . The method of claim 16 , [scalable querying and isolated data storage], wherein client data merged into the archive data storage is isolated according to the user account associated with the client data and a query processing cluster interfaces with the distributed storage system, and the query processing cluster is shared between by a plurality of user accounts.

18 . The method of claim 1 , further comprising at a client data agent collecting client data and transmitting the client data to the data analysis platform; wherein the columnar-based distributed archive storage system stores client data in time series order with a dynamic schema, and wherein selectively interfacing with client data includes cooperatively querying data from the real-time storage system and the archive storage system for a cohesive query result.

19 . The method of claim 18 , wherein distributed storage system includes over one petabyte of data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2012
From: FURUHASHI, SADAYUKI; YOSHIKAWA, HIRONOBU; OTA, KAZUKI
To: TREASURE DATA, INC.
Reel/Frame 029319/0645 →