IP Library Granted Patent US 10,963,486
Granted Patent B2
US 10,963,486 · App. 16/292,135 · Granted Mar 30, 2021

Management of collaborative datasets via distributed computer networks

Inventors: Bryon Kristen Jacob (Austin, TX); David Lee Griffith (Austin, TX); Triet Minh Le (Austin, TX); Jon Loyens (Austin, TX); Brett A. Hurt (Austin, TX); Arthur Albert Keen (Austin, TX)
Assignee: data.world, Inc.
G06F16/273G06F16/178G06F16/21G06F16/2365G06F16/2455G06F16/258G06F2221/2141
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,963,486
App. No.
16/292,135
Filed
Mar 4, 2019
Granted
Mar 30, 2021
Kind
B2
Art Unit
2156
USPC
707/608
Abstract

Various embodiments relate generally to data science and data analysis, computer software and systems, and wired and wireless network communications to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby a collaborative data layer and associated logic facilitate, for example, efficient access to, and implementation of, collaborative datasets. In some examples, a method may include receiving a dataset and dataset attributes and identifying a first version of the dataset. The method may include identifying data that varies from a first version of the dataset, and generating a second version of the dataset to include a first subset and a second subset of atomized data. The method may include storing subsets of atomized data points as an atomized dataset.

Claims (71)

1. A method comprising:

receiving via a network data representing a dataset having a data format into a collaborative dataset consolidation system, the collaborative dataset consolidation system including one or more processors and one or more repositories configured to convert differently-formatted datasets into atomized datasets as data arrangements to facilitate interoperability among converted datasets;

receiving data representing attributes associated with the dataset, the attributes including data representing an account identifier;

identifying a first version of the dataset associated with a first subset of atomized data points, at least one atomized data point being a triple associated with non-protected data;

identifying a subset of data that varies from the first version of the dataset;

accessing the subset of data as a protected data, access to which is authorized as a function of data representing a level of authorization for the account identifier;

converting the subset of data including a non-atomized data point to a second subset of atomized data points having a specific format similar to the first subset;

generating a second version of the dataset to include the first subset of atomized data points and the second subset of atomized data points;

storing the first subset of atomized data points and the second subset of atomized data points as an atomized dataset in the one or more repositories;

receiving a query to access the atomized dataset;

classifying at least a portion of the query directed to the atomized dataset to determine a classification type; and

applying the portion of the query to at least one of a number of different data stores, a subset of which includes different types of graph-based databases.

2. The method of claim 1 further comprising:

determining resource requirements data to describe a capability to operate a database configured to access graph-based data to identify at least one resource requirement; and

identifying a data store for selection as a function of the classification type.

3. The method of claim 1 further comprising:

inferring an attribute of the attributes to include data representing the account identifier associated with the identifier, one or more other account identifiers, and one or more data-related activities.

4. The method of claim 3 further comprising:

monitoring updates to the attributes to detect an update; and

disseminating the update to a community of networked users.

5. The method of claim 4 wherein disseminating via the network the update to the community of networked users comprises:

generating data representing a notification to a user associated with the one or more other account identifiers.

6. The method of claim 5 wherein generating data representing the notification comprises:

causing presentation of the notification in an activity feed portion of a user interface of a computing device.

7. The method of claim 4 wherein monitoring the updates comprises:

detecting discovery of modified data associated with a data-related activity.

8. The method of claim 7 wherein detecting discovery of the modified data associated with the data-related activity comprises:

determining an activity of querying the dataset.

9. The method of claim 7 wherein detecting discovery of the modified data associated with the data-related activity comprises:

determining an activity of modifying the dataset.

10. The method of claim 7 further comprising:

generating data representing a recommendation.

11. The method of claim 1 wherein each of the atomized data point is data representing an addressable data unit.

12. The method of claim 1 wherein storing the first subset of atomized data points and the second subset of atomized data points as the atomized dataset comprises:

storing atomized data points as triples.

13. The method of claim 12 wherein at least one triple of the triples are formatted to comply with a Resource Description Framework (“RDF”) data model.

14. The method of claim 1 wherein generating the second version of the dataset to include the first subset of atomized data points comprises:

generating a data pointer to a memory location at which the first subset of atomized data points is stored; and

storing the data pointer as the first subset of atomized data points.

15. The method of claim 1 further comprising

managing dataset attributes associated with the atomized dataset.

16. The method of claim 15 wherein managing the dataset attributes comprises:

analyzing atomized datasets associated with the collaborative dataset consolidation system; and

identifying a number of queries associated with the atomized dataset.

17. The method of claim 15 wherein managing the dataset attributes comprises:

analyzing atomized datasets associated with the collaborative dataset consolidation system;

identifying a subset of other account identifiers that include descriptive data that correlate to the atomized dataset;

generating a data signal specifying information for at least one of the account identifiers that accessed the descriptive data; and

causing presentation of the information in an activity feed portion of a user interface;

analyzing atomized datasets associated with the collaborative dataset consolidation system; and

identifying a subset of other atomized datasets including similar classification types.

18. A system comprising:

a memory including executable instructions; and

a processor, responsive to executing the instructions, is configured to:

receive via a network data representing a dataset having a data format into a collaborative dataset consolidation system, the collaborative dataset consolidation system including one or more processors and one or more repositories configured to convert differently-formatted datasets into atomized datasets as data arrangements to facilitate interoperability among converted datasets;

receive data representing attributes associated with the dataset, the attributes including data representing an account identifier;

identify a first version of the dataset associated with a first subset of atomized data points, at least one atomized data point being a triple associated with non-protected data;

identify a subset of data that varies from the first version of the dataset;

access the subset of data as a protected data, access to which is authorized as a function of data representing a level of authorization for the account identifier;

convert the subset of data including a non-atomized data point to a second subset of atomized data points having a specific format similar to the first subset;

generate a second version of the dataset to include the first subset of atomized data points and the second subset of atomized data points;

store the first subset of atomized data points and the second subset of atomized data points as an atomized dataset in the one or more repositories;

classify at least a portion of the query directed to the atomized dataset to determine a classification type; and

apply the portion of the query to at least one of a number of different data stores, a subset of which includes different types of graph-based databases.

19. The system of claim 18 wherein the processor is further configured to:

determine resource requirements data to describe a capability to operate a database configured to access graph-based data to identify at least one resource requirement; and

identify a data store for selection as a function of the classification type.

20. The system of claim 18 wherein the processor is further configured to:

infer an attribute of the attributes to include data representing the account identifier associated with the identifier, one or more other account identifiers, and one or more data-related activities;

monitor updates to the attributes to detect an update;

disseminate via the network the update to a community of networked users.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2025
From: DATA.WORLD, INC.
To: SERVICENOW, INC.
Reel/Frame 073004/0844 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2019
From: JACOB, BRYON KRISTEN; LOYENS, JON; GRIFFITH, DAVID LEE; HURT, BRETT A.; LE, TRIET MINH; KEEN, ARTHUR ALBERT
To: DATA.WORLD, INC.
Reel/Frame 048535/0687 →
Continuity (3)
Continuation 16271687 · Feb 8, 2019
Continuation 15186520 · Jun 19, 2016
Related Publication 20190370230A1 · Dec 5, 2019