IP Library Granted Patent US 11,995,086
Granted Patent B2
US 11,995,086 · App. 17/581,835 · Granted May 28, 2024

Methods for enhancing rapid data analysis

Inventors: Robert Johnson (Palo Alto, CA); Oleksandr Barykin (Sunnyvale, CA); Alex Suhan (Menlo Park, CA); Lior Abraham (San Francisco, CA); Don Fossgreen (Scotts Valley, CA)
Assignee: Scuba Analytics, Inc.
G06F16/24554G06F16/278H04L63/1425
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,995,086
App. No.
17/581,835
Granted
May 28, 2024
Kind
B2
Abstract

A method for enhancing rapid data analysis includes receiving a set of data; storing the set of data in a first set of data shards sharded by a first field; and identifying anomalous data from the set of data by monitoring a range of shard indices associated with a first shard of the first set of data shards, detecting that the range of shard indices is smaller than an expected range by a threshold value, and identifying data of the first shard as anomalous data.

Claims (43)

1. An anomalous data identification system, comprising:

a non-transitory computer readable medium storing instructions that, when executed by a processing system, cause the processing system to:

receive a dataset;

identify anomalous data comprising a subset of the dataset;

shard the dataset into database shards according to a set of partitioning rules while accounting for the anomalous data;

compress the anomalous data in a different manner than non-anomalous data;

receive a query;

identify a set of shards containing data relevant to the query;

collect a data sample from the set of shards; and

calculate a result to the query based on an analysis of the data sample.

2. The system of claim 1 , wherein the analysis of the data sample excludes the anomalous data.

3. The system of claim 1 , wherein the analysis of the data sample is based on weights assigned to the anomalous data and non-anomalous data, wherein the weights are different for the anomalous data and the non-anomalous data.

4. The system of claim 1 , wherein identifying the anomalous data comprises flagging the anomalous data with metadata.

5. The system of claim 4 , wherein the metadata comprises handling information.

6. The system of claim 1 , wherein the processing system is further configured to physically structure the anomalous data on disk or in memory differently from non-anomalous data.

7. The system of claim 1 , wherein the anomalous data is evenly distributed across the database shards.

8. The system of claim 1 , wherein compressing the anomalous data comprises using a different string dictionary than the non-anomalous data.

9. The system of claim 1 , wherein the processing system is further configured to store the anomalous data in different manner than non-anomalous data.

10. The system of claim 1 , wherein the anomalous data is identified using a manually configured ruleset.

11. The system of claim 1 , wherein the dataset is stored in persistent memory; and wherein the persistent memory comprises a plurality of contiguous blocks, wherein the database shards are stored in the plurality of contiguous blocks.

12. The system of claim 11 , wherein a block size of each block is defined by a threshold that represents a comparison between a cost of scanning a current block compared to a cost of scanning a next block of the plurality of contiguous blocks.

13. The system of claim 1 , wherein at least one of the non-anomalous data or the anomalous data is compressed lossless.

14. A system for anomalous data identification, comprising:

a database comprising a set of database shards collectively storing a dataset, wherein the dataset is partitioned into the database shards according to a set of partitioning rules; and

a processing system configured to identify anomalous data comprising a subset of the dataset by:

analyzing the dataset to determine a statistical distribution of data element counts;

identifying the anomalous data based on the statistical distribution;

structuring the anomalous data; and

compressing the anomalous data in a different manner than non-anomalous data.

15. The system of claim 14 , wherein identifying the anomalous data comprises tagging the anomalous data with metadata.

16. The system of claim 14 , wherein structuring the anomalous data comprises placing the anomalous data on slower disks compared to non-anomalous data.

17. The system of claim 14 , wherein the anomalous data is structured by partitioning the anomalous data based on a set of data fields of the anomalous data.

18. The system of claim 14 , wherein the dataset is stored in persistent memory; and wherein the persistent memory comprises a plurality of contiguous blocks, wherein the database shards are stored in the plurality of contiguous blocks.

19. The system of claim 14 , wherein the anomalous data is identified before the dataset is partitioned.

20. An anomalous data identification system, comprising:

a non-transitory computer readable medium storing instructions that, when executed by a processing system, cause the processing system to:

receive a dataset;

store the dataset as database shards, wherein the database shards are partitioned according to a set of partitioning rules;

receive a query;

identify a set of shards containing data relevant to the query;

collecting a data sample from the set of shards;

identify anomalous data comprising a subset of the data sample; and

calculate a result to the query based on an analysis of the data sample that treats the subset of the data sample differently from a remainder of the data sample, wherein the subset of the data sample is compressed in a different manner than the remainder of the data sample.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2022
From: JOHNSON, ROBERT; BARYKIN, OLEKSANDR; SUHAN, ALEX; ABRAHAM, LIOR; FOSSGREEN, DON
To: INTERANA, INC.
Reel/Frame 059183/0291 →
CHANGE OF NAME Recorded Mar 7, 2022
From: INTERANA, INC.
To: SCUBA ANALYTICS, INC.
Reel/Frame 059332/0848 →
Continuity (5)
Continuation 16924613 · Jul 9, 2020
Continuation 16384603 · Apr 15, 2019
Continuation 15043333 · Feb 12, 2016
Provisional Application 62115404 · Feb 12, 2015
Related Publication 20220147530A1 · May 12, 2022