IP Library Granted Patent US 9,842,152
Granted Patent B2
US 9,842,152 · App. 14/518,931 · Granted Dec 12, 2017

Transparent discovery of semi-structured data schema

Inventors: Benoit Dageville (Foster City, CA); Vadim Antonov (Belmont, CA)
Assignee: Snowflake Computing, Inc.
G06F17/30575G06F9/4881G06F9/5016G06F9/5088G06F17/302G06F17/3048G06F17/30292G06F17/30315G06F17/30371G06F17/30463G06F17/30466G06F17/30498G06F17/30545G06F17/30598G06F17/30864G06F17/30867G06F17/30914H04L67/1095H04L67/2842
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,842,152
App. No.
14/518,931
Granted
Dec 12, 2017
Kind
B2
Abstract

A system, apparatus, and method for managing data storage and data access for semi-structured data systems.

Claims (46)

1. A method for managing semi-structured data comprising:

receiving semi-structured data elements from a data source that is connected over a computer network;

performing statistical analysis on collections of the semi-structured data elements as they are added to the database via a computer processor, wherein separate collections comprising portions of the semi-structured data are stored in separate files having different subsets of the semi-structured data elements that have been extracted;

identifying common data elements from within the semi-structured data;

combining common data elements from the data source into separate pseudo-columns;

storing non-common semi-structured data elements in an overflow serialized column in computer memory; and

deriving metadata corresponding to the pseudo-columns of the common data elements from the statistical analysis.

2. The method of claim 1 , wherein the common data elements are extracted from the semi-structured data and stored separately in a columnar format.

3. The method of claim 2 , wherein the pseudo-columns are not made visible to users.

4. The method of claim 1 , further comprising storing the non-common semi-structured data in a serialized format in a main column.

5. The method of claim 1 , further comprising filtering with a bloom filter comprising identifiers of data elements contained within the separate collections.

6. The method of claim 1 , wherein metadata for each pseudo column comprises at least one of:

minimum and maximum values within a corresponding pseudo column,

a number representing the number of times distinct values appear in the column.

7. The method of claim 1 , further comprising extracting data elements from overflow serialized data if the data element requested is not in a pseudo-column.

8. The method of claim 1 , further comprising reconstructing semi-structured data to an original form by extracting data elements from pseudo-columns and from the overflow serialized data.

9. The method of claim 1 , further comprising updating the metadata corresponding to the pseudo-columns as additional semi-structured data is received.

10. A system for aggregating semi-structured data comprising:

one or more processors;

memory operably connected to the one or more processors; and

the memory storing one or more modules programmed to:

receive semi-structured data elements from a data source;

perform statistical analysis on collections of the semi-structured data elements as they are added to the database, wherein separate collections comprising portions of the semi-structured data are stored in separate files having different subsets of the semi-structured data elements that have been extracted;

identify common data elements from within the semi-structured data and combine the common data elements from the data source into separate pseudo-columns;

store non-common semi-structured data elements in an overflow serialized column; and

derive metadata corresponding to the pseudo-columns of the common data elements from the statistical analysis.

11. The system of claim 10 , wherein the combining the common data elements from the data source into separate pseudo-columns comprises extracting common data elements from the semi-structured data and storing the common data elements separately in a columnar format.

12. The system of claim 10 , wherein the columnar format is invisible to users.

13. The system of claim 10 , wherein metadata for each pseudo-column comprises at least one of:

minimum and maximum values within a corresponding pseudo column, and

a number representing the number of times distinct values appear in the column.

14. An apparatus for aggregating semi-structured data comprising:

one or more processors;

memory operably connected to the one or more processors; and

the memory storing:

a receiving module configured to receive semi-structured data elements from a data source;

a statistical module configured to perform statistical analysis on collections of the semi-structured data elements as they are added to the database, wherein separate collections comprising portions of the semi-structured data are stored in separate files having different subsets of the semi-structured data elements that have been extracted;

an aggregation means for identifying common data elements from within the semi-structured data and combining the common data elements from the data source into separate pseudo-columns;

the aggregation means further for serializing and storing non-common semi-structured data elements in an overflow serialized column; and

the aggregation means further for deriving metadata corresponding to the pseudo-columns of the common data elements from the statistical analysis.

15. The apparatus of claim 14 , wherein the combining the common data elements from the data source into separate pseudo-columns comprises extracting common data elements from the semi-structured data and storing the common data elements separately in a columnar format.

16. The apparatus of claim 14 , wherein the columnar format is invisible to users.

17. The apparatus of claim 14 , wherein metadata for each pseudo column comprises at least one of:

minimum and maximum values within a corresponding pseudo column, and

a number representing the number of times distinct values appear in the column.

18. The apparatus of claim 14 , wherein the aggregation means is further for updating the metadata corresponding to the pseudo-columns as additional semi-structured data is received.

Assignments (2)
CHANGE OF NAME Recorded Apr 11, 2019
From: SNOWFLAKE COMPUTING, INC.
To: SNOWFLAKE INC.
Reel/Frame 049127/0027 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2014
From: DAGEVILLE, BENOIT; ANTONOV, VADIM
To: SNOWFLAKE COMPUTING INC.
Reel/Frame 034008/0852 →
Continuity (2)
Provisional Application 61941986 · Feb 19, 2014
Related Publication 20150234931A1 · Aug 20, 2015