IP Library Granted Patent US 11,163,755
Granted Patent B2
US 11,163,755 · App. 16/395,043 · Granted Nov 2, 2021

Query generation for collaborative datasets

Inventors: Bryon Kristen Jacob (Austin, TX); David Lee Griffith (Austin, TX); Triet Minh Le (Austin, TX); Jon Loyens (Austin, TX); Brett A. Hurt (Austin, TX); Arthur Albert Keen (Austin, TX)
Assignee: data.world, Inc.
G06F16/242G06F16/248G06F16/2455G06F16/2471G06F16/24535G06F16/24542G06F21/6227
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,163,755
App. No.
16/395,043
Granted
Nov 2, 2021
Kind
B2
Abstract

Various embodiments relate generally to data science and data analysis, computer software and systems, and wired and wireless network communications to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby a collaborative data layer and associated logic facilitate, for example, efficient access to, and implementation of, collaborative datasets. In some examples, a method may include receiving data representing a query of a consolidated dataset that may include datasets formatted atomized datasets, analyzing the query to classify portions of the query to form classified query portions, partitioning the query into sub-queries as a function of a classification type for each of the classified query portions, and retrieving data representing a query result from distributed data repositories.

Claims (61)

1. A method comprising:

receiving, into a collaborative dataset consolidation system, data representing a query of a consolidated dataset comprising a plurality of datasets formatted as plurality of atomized datasets;

analyzing the query to classify portions of the query to form classified query portions;

partitioning the query into a plurality of sub-queries as a function of a classification type associated with each of the classified query portions, at least one classified query portion being classified by a type of dataset;

determining a type of triple store loaded based on the type of dataset;

executing the query and the sub-queries as a federated query to one or more of the plurality of datasets using one or more triple stores stored in one or more distributed data repositories configured to be stored and accessed by the collaborative data consolidation system, each of the one or more triple stores being further formatted to query using a format associated with at least one of the one or more of the plurality of data sets, at least one of the one or more triple stores being the type of triple store;

retrieving responsive data representing a query result returned in response to the query or sub-queries from at least one of the one or more distributed data repositories;

identifying the query results generated by the query as a data-related activity associated with a subset of datasets in the plurality of datasets;

identifying one or more user account identifiers associated with the subsets of datasets associated with the query; and

disseminating an update as the data-related activity to a community of networked users associated with the one or more user account identifiers.

2. The method of claim 1 wherein disseminating the update to the community of networked users comprises:

generating data representing a notification to a user associated with the one or more other account identifiers.

3. The method of claim 2 wherein generating data representing the notification comprises:

causing presentation of the notification in an activity feed portion of a user interface of a computing device.

4. The method of claim 1 wherein analyzing the query comprises:

detecting discovery of modified data associated with the data-related activity.

5. The method of claim 4 wherein detecting discovery of the modified data associated with the data-related activity comprises:

determining an activity of querying the dataset.

6. The method of claim 4 wherein detecting discovery of the modified data associated with the data-related activity comprises:

determining an activity of modifying the dataset.

7. The method of claim 4 further comprising:

generating data representing a recommendation.

8. The method of claim 1 further comprising:

identifying a number of atomized datasets associated with the query.

9. The method of claim 8 further comprising:

implementing per-dataset permissions to determine whether a user account is authorized to query each of the atomized datasets.

10. The method of claim 8 further comprising:

identifying an identifier associated with the query;

determining authorization data between the identifier and each of atomized datasets in the number of atomized datasets; and

granting authorization to apply the query to the number of atomized datasets based on authorization to each of the atomized datasets.

11. A collaborative dataset consolidation system, comprising:

a data store configured to receive, into the collaborative dataset consolidation system, data representing a query of a consolidated dataset comprising a plurality of datasets formatted as plurality of atomized datasets; and

a processor configured to execute instructions to implement a dataset query engine configured to:

analyze the query to classify portions of the query to form classified query portions;

partition the query into a plurality of sub-queries as a function of a classification type associated with each of the classified query portions, at least one classified query portion being classified by a type of dataset;

determine a type of triple store loaded based on the type of dataset;

execute the query and the sub-queries as a federated query to one or more of the plurality of datasets using one or more triple stores stored in one or more distributed data repositories configured to be stored and accessed by the collaborative data consolidation system, each of the one or more triple stores being further formatted to query using a format associated with at least one of the one or more of the plurality of data sets, at least one of the one or more triple stores being the type of triple store;

retrieve responsive data representing a query result returned in response to the query or sub-queries from at least one of the one or more distributed data repositories,

identify the query results generated by the query as a data-related activity associated with a subset of datasets in the plurality of datasets;

identify one or more user account identifiers associated with the subsets of datasets associated with the query; and

disseminate an update as the data-related activity to a community of networked users associated with the one or more user account identifiers.

12. The collaborative dataset consolidation system of claim 11 , wherein the processor configured to disseminate the update to the community of networked users is further configured to:

generate data representing a notification to a user associated with the one or more other account identifiers.

13. The collaborative dataset consolidation system of claim 12 , wherein the processor configured to generate data representing the notification is further configured to:

cause presentation of the notification in an activity feed portion of a user interface of a computing device.

14. The collaborative dataset consolidation system of claim 11 , wherein the processor configured to analyze the query is further configured to:

detect discovery of modified data associated with the data-related activity.

15. The collaborative dataset consolidation system of claim 14 , wherein the processor configured to detect discovery of the modified data associated with the data-related activity is further configured to:

determine an activity of querying the dataset.

16. The collaborative dataset consolidation system of claim 14 , wherein the processor configured to detect discovery of the modified data associated with the data-related activity is further configured to:

determine an activity of modifying the dataset.

17. The collaborative dataset consolidation system of claim 14 , wherein the processor is further configured to:

generate data representing a recommendation.

18. The collaborative dataset consolidation system of claim 11 , wherein the processor is further configured to:

identify a number of atomized datasets associated with the query.

19. The collaborative dataset consolidation system of claim 18 , wherein the processor is further configured to:

implement per-dataset permissions to determine whether a user account is authorized to query each of the atomized datasets.

20. The collaborative dataset consolidation system of claim 18 , wherein the processor is further configured to:

identify an identifier associated with the query;

determine authorization data between the identifier and each of atomized datasets in the number of atomized datasets; and

grant authorization to apply the query to the number of atomized datasets based on authorization to each of the atomized datasets.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2025
From: DATA.WORLD, INC.
To: SERVICENOW, INC.
Reel/Frame 073004/0844 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 31, 2019
From: JACOB, BRYON KRISTEN; LOYENS, JON; GRIFFITH, DAVID LEE; HURT, BRETT A.; LE, TRIET MINH; KEEN, ARTHUR ALBERT
To: DATA.WORLD, INC.
Reel/Frame 049335/0588 →
Continuity (2)
Continuation 15186517 · Jun 19, 2016
Related Publication 20190347258A1 · Nov 14, 2019
Cited By (1)
US 12,681,950