IP Library Granted Patent US 11,263,215
Granted Patent B2
US 11,263,215 · App. 16/924,613 · Granted Mar 1, 2022

Methods for enhancing rapid data analysis

Inventors: Robert Johnson (Palo Alto, CA); Oleksandr Barykin (Sunnyvale, CA); Alex Suhan (Menlo Park, CA); Lior Abraham (San Francisco, CA); Don Fossgreen (Scotts Valley, CA)
Assignee: SCUBA ANALYTICS, INC.
G06F16/24554G06F16/278
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,263,215
App. No.
16/924,613
Granted
Mar 1, 2022
Kind
B2
Abstract

A method for enhancing rapid data analysis includes receiving a set of data; storing the set of data in a first set of data shards sharded by a first field; and identifying anomalous data from the set of data by monitoring a range of shard indices associated with a first shard of the first set of data shards, detecting that the range of shard indices is smaller than an expected range by a threshold value, and identifying data of the first shard as anomalous data.

Claims (41)

1. A method for identifying anomalous data in a computer database, comprising:

determining a set of data;

sharding the set of data into a plurality of database shards;

storing the set of data in the plurality of database shards;

determining that a first database shard of the plurality contains anomalous data based on a structure of the first database shard;

identifying data of the first database shard as the anomalous data; and

in response to identification of the anomalous data, compressing the anomalous data in a different manner than the non-anomalous data;

wherein the set of data is stored in persistent memory; and wherein the persistent memory comprises a plurality of contiguous blocks, wherein the plurality of database shards is stored in the plurality of contiguous blocks, wherein a block size of each block is defined by a threshold that represents a comparison between a cost of scanning a current block compared to a cost of scanning a next block of the plurality of contiguous blocks.

2. The method of claim 1 , wherein the structure is a shard density.

3. The method of claim 1 , wherein the structure is a shard range.

4. The method of claim 3 , wherein determining that a first database shard of the plurality contains anomalous data based on the shard range, and comprises detecting that the shard range is outside of an expected range by a threshold value.

5. The method of claim 1 , wherein compressing the set of anomalous data in a different manner than the set of non-anomalous data comprises using a first string dictionary to compress the set of anomalous data and a second string dictionary to compress the set of non-anomalous data, wherein the first and second string dictionaries are distinct.

6. The method of claim 1 , further comprising:

receiving a query that comprises custom query code;

converting the custom query code from a foreign language to a native query language; and

interpreting the query in the native query language.

7. The method of claim 6 , wherein the foreign language is a natural language; wherein interpreting the query comprises interpreting the natural language.

8. The method of claim 1 , further comprising:

receiving and interpreting a query;

collecting a first data sample comprising performing data sampling to ignore the identified anomalous data;

calculating a query result using the first data sample;

determining that a performance time associated with calculating the query result is below a threshold value; and

in response to the performance time below the threshold value, restructuring the set of data.

9. The method of claim 1 , wherein storing the set of data comprises partitioning the set of data into the plurality of database shards using a sampling function.

10. The method of claim 1 , wherein a respective shard size of each of the plurality of shards is set based on the block size.

11. A method for identifying anomalous data in a computer database, comprising:

receiving event data;

structuring the event data, comprising compressing the event data;

calculating a statistical distribution of a first field of the event data;

identifying anomalous data from non-anomalous data of the event data based on the statistical distribution; and

in response to identification of the anomalous data, restructuring the anomalous data;

wherein a set of data comprises the event data; where in the set of data is stored in a plurality of database shards in persistent memory; and wherein the persistent memory comprises a plurality of contiguous blocks, wherein the plurality of database shards is stored in the plurality of contiguous blocks, wherein a block size of each block is defined by a threshold that represents a comparison between a cost of scanning a current block compared to a cost of scanning a next block of the plurality of contiguous blocks.

12. The method of claim 11 , wherein the event data is associated with an object attribute field, and wherein structuring the event data comprises structuring the event data explicitly and structuring the object attribute field implicitly.

13. The method of claim 12 , wherein the object attribute field is directly accessible by a user query.

14. The method of claim 11 , further comprising:

receiving and interpreting a query, wherein the query comprises a grouping function; and

calculating multiple query results using the event data, wherein the query results are grouped using the grouping function.

15. The method of claim 14 , wherein the grouping function comprises a cohort function that divides the query results into a set of cohorts, wherein each query result appears in exactly one cohort.

16. The method of claim 11 , wherein the event data is represented by integers.

17. The method of claim 11 , wherein structuring the event data comprises organizing the event data into a columnar database.

18. The method of claim 11 , wherein structuring the event data comprises duplicating data elements of the event data.

Assignments (2)
CHANGE OF NAME Recorded Feb 3, 2021
From: INTERANA, INC.
To: SCUBA ANALYTICS, INC.
Reel/Frame 055210/0293 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2020
From: JOHNSON, ROBERT; BARYKIN, OLEKSANDR; SUHAN, ALEX; ABRAHAM, LIOR; FOSSGREEN, DON
To: INTERANA, INC.
Reel/Frame 053163/0855 →
Continuity (4)
Continuation 16384603 · Apr 15, 2019
Continuation 15043333 · Feb 12, 2016
Provisional Application 62115404 · Feb 12, 2015
Related Publication 20200341985A1 · Oct 29, 2020