IP Library Granted Patent US 10,713,240
Granted Patent B2
US 10,713,240 · App. 15/645,698 · Granted Jul 14, 2020

Systems and methods for rapid data analysis

Inventors: Robert Johnson (Palo Alto, CA); Lior Abraham (San Francisco, CA); Ann Johnson (Palo Alto, CA); Boris Dimitrov (Portola Valley, CA); Don Fossgreen (Scotts Valley, CA)
Assignee: Interana, Inc.
G06F16/2425G06F16/2462G06F16/2471G06F16/24545G06F16/24554G06F16/278
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,713,240
App. No.
15/645,698
Granted
Jul 14, 2020
Kind
B2
Abstract

A method for rapid data analysis includes receiving and interpreting a first query operating on a first dataset partitioned into shards by a first field; collecting a first data sample from a first set of data shards; calculating a first result to the first query based on analysis of the first data sample; and partitioning a second dataset into shards by a second field based on the first result.

Claims (40)

1. A method for rapid data analysis comprising:

receiving and interpreting a first query, wherein interpreting the first query comprises identifying a first set of data shards of a first dataset containing data relevant to the first query; wherein the first dataset is partitioned by a first field;

for a first query pass of the first query, collecting a first data sample from the first set of data shards, wherein collecting the first data sample comprises collecting data from each of the first set of data shards;

for the first query pass, calculating a first result to the first query based on analysis of the first data sample; and

for a second query pass of the first query that uses the first result as input, partitioning a second dataset based on a second field, wherein the second data set contains data identical to the first dataset, wherein the second field is identified by a set of shard partitioning rules, based on the first field; wherein the second field is non-identical to the first field.

2. The method of claim 1 , wherein the second dataset is the first dataset; wherein partitioning the second dataset comprises re-partitioning the first dataset.

3. The method of claim 1 , wherein the second dataset is distinct from the first dataset.

4. The method of claim 1 , wherein partitioning the second dataset comprises automatically partitioning the second dataset to improve query performance for queries similar to the first query.

5. The method of claim 4 , wherein automatically partitioning the second dataset to improve query performance comprises automatically partitioning the second dataset only after queries similar to the first query have been identified as common queries.

6. The method of claim 4 , further comprising generating a data aggregate of the first dataset to improve query performance for the queries similar to the first query.

7. The method of claim 1 , wherein partitioning the second dataset comprises identifying the second field as containing data relevant to the first query.

8. The method of claim 1 , further comprising detecting that the first dataset is used less than a use threshold and, in response, removing the first dataset.

9. The method of claim 1 , further comprising:

analyzing the first result to identify a set of query-relevant data sources;

identifying a second set of data shards from the set of query-relevant data sources;

collecting a second data sample from the second set of data shards, wherein collecting the second data sample comprises collecting data from each of the second set of data shards; and

calculating a final result to the first query based on analysis of the second data sample.

10. The method of claim 9 , wherein the second set of data shards contains data not contained in the first set of data shards.

11. A system for rapid data analysis comprising:

an event database, comprising first and second datasets; wherein the first and second datasets contain identical data; wherein the first dataset is partitioned by a first field;

a string lookup database that stores information linking strings to integers that uniquely identify the strings;

a string translator that converts strings in incoming data to integer identifiers using the string lookup database;

a query engine that processes queries on the event database and returns at least a first query result for a first query pass of a first query, and a second query result for a second query pass of the first query, wherein the second query pass uses the first query result as input; and

a data manager that, based on a second field, partitions the second dataset; wherein the second field is identified by a set of shard partitioning rules, based on the first field; wherein the a second field is non-identical to the first field.

12. The system of claim 11 , wherein the data manager also repartitions the first data set.

13. The system of claim 12 , wherein the data manager repartitions the first data set by the first field.

14. The system of claim 11 , wherein the data manager automatically partitions the second dataset to improve query performance for future queries similar to past queries.

15. The system of claim 11 , wherein the data manager automatically partitions the second dataset to improve query performance for queries identified as common queries.

16. The system of claim 11 , wherein the data manager further generates a data aggregate of the first dataset to improve query performance for future queries similar to past queries.

17. The system of claim 11 , wherein the data manager identifies the second field as containing data relevant to the first query prior to partitioning the second dataset by the second field.

18. The system of claim 11 , wherein the data manager identifies and removes datasets used less than a use threshold.

19. The system of claim 11 , wherein the query engine processes an incoming query by:

identifying a first set of data shards of the first dataset containing data relevant to the incoming query;

collecting a first data sample from the first set of data shards;

calculating a first result to the incoming query based on analysis of the first data sample;

analyzing the first result to identify a set of query-relevant data sources;

identifying a second set of data shards from the set of query-relevant data sources;

collecting a second data sample from the second set of data shards, wherein collecting the second data sample comprises collecting data from each of the second set of data shards; and

calculating a final result to the incoming query based on analysis of the second data sample.

20. The system of claim 19 , wherein the second set of data shards contains data not contained in the first set of data shards.

Assignments (4)
CHANGE OF NAME Recorded Feb 3, 2021
From: INTERANA, INC.
To: SCUBA ANALYTICS, INC.
Reel/Frame 055210/0293 →
RELEASE OF SECURITY INTEREST Recorded Mar 17, 2020
From: VENTURE LENDING & LEASING VIII, INC.; VENTURE LENDING & LEASING IX, INC.
To: INTERANA, INC.
Reel/Frame 052143/0398 →
SECURITY INTEREST Recorded Jun 25, 2018
From: INTERANA, INC.
To: VENTURE LENDING & LEASING IX, INC.; VENTURE LENDING & LEASING VIII, INC.
Reel/Frame 047262/0292 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2017
From: JOHNSON, ROBERT; ABRAHAM, LIOR; JOHNSON, ANN; DIMITROV, BORIS; FOSSGREEN, DON
To: INTERANA, INC.
Reel/Frame 042955/0489 →
Continuity (4)
Continuation 15077800 · Mar 22, 2016
Continuation 14644081 · Mar 10, 2015
Provisional Application 61950827 · Mar 10, 2014
Related Publication 20170308570A1 · Oct 26, 2017