IP Library Patent Application 13300523
Patent Application
App. No. 13/300,523

REAL-TIME ANALYTICS OF STREAMING DATA

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
13/300,523
Abstract

Storage media, systems and methods are disclosed herein for analyzing data streams in real time. More particularly, storage media, systems and methods are presented for processing data streams to calculate results for prospective queries. The results may be advantageously computed prior to the formulation of the specific query, for example, based on a pre-established framework of potential query parameters. More particularly, a universe of potential queries may be extrapolated from the pre-established framework of potential query parameters. Results for each of the potential queries may them be tracked in real time. For example, results for each of the potential queries may be continuously updated based on real-time processing of events in a data stream.

Claims (43)

1 . A method for performing real time analytics on streaming data, the method comprising:

processing events in a data stream to extract from each event a set of attribute-value pairs for one or more dimension attributes and one or more value attributes;

identifying one or more tuples in a multidimensional data structure implicated by the extracted attribute-value pairs for the one or more dimension attributes;

for each implicated tuple, updating one or more stored aggregates associated therewith, based on the extracted attribute-value pairs for the one or more value attributes.

2 . The method of claim 1 further comprising tracking over an interval a tuple frequency for each of the implicated tuples, wherein the updating the one or more stored aggregates includes discarding each implicated tuple with a low tuple frequency over the interval.

3 . The method of claim 1 further comprising tracking over an interval, for each value attribute, aggregates for a first plurality of implicated tuples having same attribute-value pairs for zero or more ordinary dimension attributes and different attribute-value pairs for one or more leaderboard dimension attributes, and determining a top-N values for the one or more leadership dimension attributes over the interval.

4 . The method of claim 3 , wherein the Top-N values are characterized as resulting in one of (i) the highest aggregates (ii) the lowest aggregates, and (iii) the aggregates closest to a selected value, over the interval.

5 . The method of claim 3 further comprising determining a set of top-N values for each of a plurality of intervals in a time window and determining a top-N values for the one or more leadership dimensions attributes over the time window based on the plurality of sets of top-N values.

6 . The method of claim 1 , wherein the one or more dimension attributes include a K-Gram for identifying topics of interest.

7 . The method of claim 6 , further comprising tracking over an interval a tuple frequency for each of the implicated tuples including a K-Gram, wherein the updating the one or more stored aggregates includes discarding each implicated tuple with a low tuple frequency over the interval, whereby statistics for trending K-Gram-value pairs are tracked.

8 . A method for implementing a real time analytics platform, the method comprising:

establishing an analytics platform framework characterized by one or more time windows, one or more dimension attributes, and one or more value attribute;

generating a first multi-dimensional data structure for maintaining, for each tuple of the one or more dimension attributes, an aggregate of each of the one or more value attributes over each of the one or more time windows.

9 . A system for performing real time analytics on streaming data, the system comprising:

a processor for processing an event in a data stream to extract a set of attribute-value pairs for one or more dimension attributes and one or more value attributes;

a mapper for identifying one or more tuples in a multidimensional data structure implicated by the extracted attribute-value pairs for the one or more dimension attributes; and

one or more updaters for updating, for each implicated tuple, one or more stored aggregates associated therewith, based on the extracted attribute-value pairs for the one or more value attributes.

10 . The system of claim 9 , wherein the system is configured to track over an interval a tuple frequency for each of the implicated tuples, wherein the updating the one or more stored aggregates includes discarding each implicated tuple with a low tuple frequency over the interval.

11 . The system of claim 9 wherein the system is configured to: (i) update, over an interval, for each value attribute, aggregates for a first plurality of implicated tuples having same attribute-value pairs for zero or more ordinary dimension attributes and different attribute-value pairs for one or more leaderboard dimension attributes, and (ii) determine a top-N values for the one or more leadership dimension attributes over the interval.

12 . The system of claim 9 , wherein the one or more dimension attributes include a K-Gram for identifying topics of interest, wherein the system is configured to track over an interval a tuple frequency for each of the implicated tuples including a K-Gram, wherein the updating the one or more stored aggregates includes discarding each implicated tuple with a low tuple frequency over the interval, whereby statistics for trending K-Gram-value pairs are tracked.

13 . A multi-dimensional data structure for implementing a real-time analytics platform characterized by one or more time windows, one or more dimension attributes, and one or more value attributes, the data structure comprising:

a plurality of tuples associated with the one or more dimension attributes; and

a slate associated with each tuple for maintaining an aggregate for each of the one or more value attributes over each of the one or more time windows.

14 . A method for performing real-time analytics on a data stream, the methods comprising:

processing a data stream to maintain a plurality of stored aggregates for a universe of prospective queries extrapolated from a pre-established framework of possible query parameters;

returning one of the stored aggregates in response to a query.

15 . A system for performing real-time analytics on a data stream the system comprising:

a processor for processing a data stream to maintain a plurality of stored aggregates for a universe of prospective queries extrapolated from a pre-established framework of possible query parameters; and

memory for storing the plurality of stored aggregates.

16 . The system of claim 15 , further comprising an interface for providing a query, wherein the processor is configured to return one of the stored aggregates in response to the query.

17 . A multi-dimensional data structure for implementing a real-time analytics platform, the data structure comprising:

a plurality of stored tuples each representing a set of search query parameters for prospective queries extrapolated from a pre-established framework of possible query parameters; and

one or more stored aggregates associated with each of the stored tuples, wherein each aggregate represents a result for a prospective query characterized by the set of search query parameters represented in the tuple associated with that aggregate.

18 . A non-transitory computer readable medium storing processor executable instructions for performing real time analytics on streaming data, including instructions for:

processing events in a data stream to extract from each event a set of attribute-value pairs for one or more dimension attributes and one or more value attributes;

identifying one or more tuples in a multidimensional data structure implicated by the extracted attribute-value pairs for the one or more dimension attributes;

for each implicated tuple, updating one or more stored aggregates associated therewith, based on the extracted attribute-value pairs for the one or more value attributes.

19 . A non-transitory computer readable medium storing processor executable instructions for performing real time analytics on streaming data, including instructions for:

establishing an analytics platform framework characterized by one or more time windows, one or more dimension attributes, and one or more value attribute;

generating a first multi-dimensional data structure for maintaining, for each tuple of the one or more dimension attributes, an aggregate of each of the one or more value attributes over each of the one or more time windows.

20 . A non-transitory computer readable medium storing processor executable instructions for performing real time analytics on streaming data, including instructions for:

processing a data stream to maintain a plurality of stored aggregates for a universe of prospective queries extrapolated from a pre-established framework of possible query parameters; and

returning one of the stored aggregates in response to a query.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2018
From: WAL-MART STORES, INC.
To: WALMART APOLLO, LLC
Reel/Frame 045817/0115 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2012
From: GATTANI, ABHISHEK; RAJARAMAN, ANAND
To: WAL-MART STORES, INC.
Reel/Frame 029248/0065 →