IP Library Granted Patent US 9,323,809
Granted Patent B2
US 9,323,809 · App. 14/644,081 · Granted Apr 26, 2016

System and methods for rapid data analysis

Inventors: Robert Johnson (Palo Alto, CA); Lior Abraham (San Francisco, CA); Ann Johnson (Palo Alto, CA); Boris Dimitrov (Portola Valley, CA); Don Fossgreen (Scott's Valley, CA)
Assignee: Interana, Inc.
G06F17/30469G06F17/30584
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,323,809
App. No.
14/644,081
Granted
Apr 26, 2016
Kind
B2
Abstract

A method for rapid data analysis comprising receiving and interpreting a query, collecting a first data sample from the first set of data shards, calculating an intermediate result to the query based on analysis of the first data sample, identifying a second set of data shards based on the intermediate result, collecting a second data sample from the second set of data shards, and calculating a final result to the query based on analysis of the second data sample.

Claims (40)

1. A method for rapid data analysis comprising:

receiving and interpreting a query, wherein interpreting the query comprises translating strings of the query to integers using a string translator, wherein interpreting the query further comprises identifying a first set of data shards containing data relevant to the query, wherein the first set of data shards are partitioned according to a set of shard partitioning rules, wherein identifying the first set of data shards comprises identifying the first set of data shards using the set of shard partitioning rules;

collecting a first data sample from the first set of data shards, wherein collecting the first data sample comprises collecting data from each of the first set of data shards, wherein collecting data from each of the first set of data shards comprises collecting a strict subset of data contained within each of the first set of data shards;

calculating an intermediate result to the query based on analysis of the first data sample;

identifying a second set of data shards based on the intermediate result,

wherein the second set of data shards contains data not contained in the first set of data shards;

collecting a second data sample from the second set of data shards, wherein collecting the second data sample comprises collecting data from each of the second set of data shards, wherein collecting data from each of the second set of data shards comprises collecting a complete set of data contained within each of the second set of data shards; and

calculating a final result to the query based on analysis of the second data sample.

2. The method of claim 1 , wherein collecting the first data sample from the first set of data shards comprises collecting data from columnar datasets of the first set of data shards.

3. The method of claim 2 , wherein the first set of data shards comprises event data organized by time.

4. The method of claim 1 , wherein receiving and interpreting the query further comprises interpreting references to implicit data.

5. The method of Claim 4 , wherein receiving and interpreting the query further comprises selecting at least one of an ordering function and a grouping function.

6. The method of claim 1 , wherein identifying a first set of data shards comprises identifying node locations of the first set of data shards using a configuration database.

7. The method of claim 1 , wherein translating strings of the query to integers using a string translator comprises, for each of the strings of the query:

identifying prefix of the string,

splitting the string into the prefix and a suffix,

assigning a first code to the string based on the prefix,

assigning a second code to the string based on the suffix and

concatenating the first code and the second code.

8. The method of claim 1 , wherein the query includes at least one time range and at least one event data source.

9. The method of claim 8 , wherein calculating the final result to the query further comprises calculating confidence bands for estimated result accuracy based on analysis of a statistical distribution of sampled data.

10. The method of Claim 9 , wherein calculating the final result to the query further comprises returning both a cohort and aggregate data associated with the cohort as a query result.

11. The method of Claim 9 , wherein calculating the intermediate result to the query further comprises calculating confidence bands for estimated result accuracy based on analysis of a statistical distribution of sampled data.

12. A method for rapid data analysis comprising:

receiving and interpreting a query, wherein interpreting the query comprises translating strings of the query to integers using a string translator, wherein interpreting the query further comprises identifying a first set of data shards containing data relevant to the query;

collecting a first data sample from the first set of data shards, wherein collecting the first data sample comprises collecting data from each of the first set of data shards, wherein collecting data from each of the first set of data shards comprises collecting a strict subset of data contained within each of the first set of data shards;

calculating a first intermediate result to the query based on analysis of the first data sample;

performing a non-zero number of intermediate searches, each intermediate search comprising:

identifying an additional set of data shards based on at least one of the first intermediate result and additional intermediate results, wherein the additional set of data shards contains data not contained in the first set of data shards,

collecting additional data samples from the additional set of data shards, and

calculating additional intermediate results based on analysis of the additional data samples; and

calculating a final result to the query.

13. The method of claim 12 , wherein the number of intermediate searches is a fixed number.

14. The method of claim 12 , further comprising calculating confidence bands for each additional intermediate result based on analysis of a statistical distribution of sampled data.

15. The method of claim 14 , wherein performing a number of intermediate searches comprises performing intermediate searches until a confidence band of an additional intermediate result passes a confidence threshold.

16. The method of claim 15 , wherein the confidence threshold is automatically set in response to a speed/accuracy variable.

17. The method of claim 16 , wherein the speed/accuracy variable is passed as part of the query.

18. The method of claim 17 , wherein the query includes at least one time range and at least one event data source.

19. The method of claim 15 , wherein receiving and interpreting the query further comprises parsing SQL-like query strings into a query tree.

20. The method of claim 14 , further comprising notifying a user that the confidence bands are below a confidence threshold.

Assignments (4)
CHANGE OF NAME Recorded Feb 3, 2021
From: INTERANA, INC.
To: SCUBA ANALYTICS, INC.
Reel/Frame 055210/0293 →
RELEASE OF SECURITY INTEREST Recorded Mar 17, 2020
From: VENTURE LENDING & LEASING VIII, INC.; VENTURE LENDING & LEASING IX, INC.
To: INTERANA, INC.
Reel/Frame 052143/0398 →
SECURITY INTEREST Recorded Jun 25, 2018
From: INTERANA, INC.
To: VENTURE LENDING & LEASING IX, INC.; VENTURE LENDING & LEASING VIII, INC.
Reel/Frame 047262/0292 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2015
From: JOHNSON, ROBERT; ABRAHAM, LIOR; JOHNSON, ANN; DIMITROV, BORIS; FOSSGREEN, DON
To: INTERANA, INC.
Reel/Frame 036192/0705 →
Continuity (2)
Provisional Application 61950827 · Mar 10, 2014
Related Publication 20150254307A1 · Sep 10, 2015