IP Library Granted Patent US 11,334,625
Granted Patent B2
US 11,334,625 · App. 16/899,542 · Granted May 17, 2022

Loading collaborative datasets into data stores for queries via distributed computer networks

Inventors: Bryon Kristen Jacob (Austin, TX); David Lee Griffith (Austin, TX); Triet Minh Le (Austin, TX); Jon Loyens (Austin, TX); Brett A Hurt (Austin, TX); Arthur Albert Keen (Austin, TX)
Assignee: data.world, Inc.
G06F16/9024G06F9/3877G06F16/24535G06F16/258G06F16/906
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,334,625
App. No.
16/899,542
Granted
May 17, 2022
Kind
B2
Abstract

Various embodiments relate generally to data science and data analysis, computer software and systems, and wired and wireless network communications to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby a collaborative data layer and associated logic facilitate, for example, efficient access to, and implementation of, collaborative datasets. In some examples, a system may include an atomized workflow loader configured to receive an atomized dataset to load into a data store, and to determine resource requirements data to describe at least one resource requirement. The atomized workflow loader may be further configured to select a data store type based on a resource requirement, and perform a load operation of the atomized dataset as a function of the data store type.

Claims (49)

1. A method comprising:

receiving an atomized dataset to load into a graph-based data store, the atomized dataset including a data arrangement in which data is stored as an atomized data point with one or more other atomized data points of one or more data types as a consolidated dataset, the atomized data point being implemented as a triple, the data arrangement representing at least a portion of a graph, the atomized data point being a representation for a relationship between two data units, and the consolidated dataset having a plurality of atomized data points of the one or more data types also having links that, when parsed, identify one or more relationships between the plurality of atomized data points and the one or more data types including a resource associated with each of the atomized and the other atomized data points and a data type associated with the resource;

converting the atomized dataset, after being received, from a first data format to a second data format, the second data format being a collaborative data format configured to be used to form a portion of the graph;

determining resource requirements data to describe a capability to operate a database configured to access graph-based data to identify at least one resource requirement;

selecting a data store type based on the at least one resource requirement;

performing a load operation of the atomized dataset as a function of the data store;

managing versioning of the atomized dataset to include a hierarchy of files, comprising tracking each version as one of an immutable collection of data files;

receiving a query to access the atomized dataset;

classifying at least a portion of the query directed to the dataset to determine a classification type, whereby the classification type is associated with a type of query for a query portion associated with a specific entity; and

applying the portion of the query as a sub-query to at least one of a number of data stores, a subset of which includes the one or more types of triplestore-based graph databases.

2. The method of claim 1 wherein tracking each version comprises:

tracking at least one pointer to at least one of the data files.

3. The method of claim 1 wherein performing the load operation of the atomized dataset comprises:

loading the dataset into a graph database.

4. The method of claim 1 wherein determining the resource requirements comprises:

identifying operating characteristics of a data store to load the atomized dataset as graph data.

5. The method of claim 4 wherein determining the resource requirements comprises:

determining an operating characteristic of the data store related to a text search; and

identifying the data store for selection as a function of the classification type.

6. The method of claim 4 wherein determining the resource requirements comprises:

determining an operating characteristic of the data store related to geo-spatial information; and

identifying the data store for selection as a function of the classification type.

7. The method of claim 4 wherein determining the resource requirements comprises:

determining an operating characteristic of the data store related to graphic processing unit (“GPU”)-optimized data; and

identifying the data store for selection as a function of the classification type.

8. The method of claim 1 wherein selecting the data store type comprises:

selecting a product having a proprietary storage architecture.

9. The method of claim 8 wherein selecting the product comprises:

selecting a triple having a specific storage architecture.

10. A system comprising:

a processor and a memory to store one or more executable instructions, the processor configured to execute instructions to implement an atomized workflow loader configured to receive an atomized dataset to load into a graph-based data store, the atomized dataset including a data arrangement in which data is stored as an atomized data point with one or more other atomized data points of one or more data types as a consolidated dataset, the atomized data point being implemented as a triple, the data arrangement representing at least a portion of a graph, the atomized data point being a representation for a relationship between two data units, and the consolidated dataset having a plurality of atomized data points of the one or more data types also having links that, when parsed, identify one or more relationships between the plurality of atomized data points and the one or more data types including a resource associated with each of the atomized and the other atomized data points and a data type associated with the resource, to convert the atomized dataset, after being received, from a first data format to a second data format, the second data format being a collaborative data format configured to be used to form a portion of the graph, to determine resource requirements data to describe a capability to operate a database configured to access graph-based data to identify at least one resource requirement, the atomized workflow loader further configured to select a data store type based on the at least one resource requirement, perform a load operation of the atomized dataset as a function of the data store type, manage versioning of the atomized dataset to include a hierarchy of files, comprising a version controller configured to track each version as one of an immutable collection of data files to manage versioning of the atomized dataset, receive a query to access the atomized dataset, classify at least a portion of the query directed to the dataset to determine a classification type, whereby the classification type is associated with a type of query for a query portion associated with a specific entity, and apply the portion of the query as a sub-query to at least one of a number of data stores, a subset of which includes the one or more types of triplestore-based graph databases.

11. The system of claim 10 where in the version controller is further configured to:

track at least one pointer to at least one of the data files.

12. The system of claim 10 further comprising:

a dataset requirement determinator configured to identify operating characteristics of a data store to load the atomized dataset as graph data.

13. The system of claim 12 wherein the data store is a triple store.

14. The system of claim 12 wherein the dataset requirement determinator is configured to:

determine an operating characteristic of the data store related to a text search; and

identify the data store for selection.

15. The system of claim 12 wherein the dataset requirement determinator is configured to:

determine an operating characteristic of the data store related to geo-spatial information; and

identify the data store for selection.

16. The system of claim 12 wherein the dataset requirement determinator is configured to:

determine an operating characteristic of the data store related to graphic processing unit (“GPU”)-optimized data; and

identify the data store for selection.

17. The system of claim 10 further comprises:

a product selector configured to select a product having a proprietary storage architecture, the product selector is further configured to select a triple having a specific storage architecture.

18. The system of claim 10 further comprising:

a data representation including one or more models configured to direct how to load datasets into the graph-based data store.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE APPLICATION NUMBERS PREVIOUSLY RECORDED AT REEL: 73004 FRAME: 844. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST . Recorded Dec 2, 2025
From: DATA.WORLD, INC.
To: SERVICENOW, INC.
Reel/Frame 073834/0037 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2020
From: JACOB, BRYON KRISTEN; LOYENS, JON; GRIFFITH, DAVID LEE; HURT, BRETT A.; LE, TRIET MINH; KEEN, ARTHUR ALBERT
To: DATA.WORLD, INC.
Reel/Frame 053064/0416 →
Continuity (2)
Continuation 15186519 · Jun 19, 2016
Related Publication 20210390141A1 · Dec 16, 2021
Cited By (1)
US 12,681,950