IP Library Granted Patent US 10,984,008
Granted Patent B2
US 10,984,008 · App. 16/287,967 · Granted Apr 20, 2021

Collaborative dataset consolidation via distributed computer networks

Inventors: Bryon Kristen Jacob (Austin, TX); Jon Loyens (Austin, TX); David Lee Griffith (Austin, TX); Brett A. Hurt (Austin, TX); Triet Minh Le (Austin, TX); Shad William Reynolds (Austin, TX); Arthur Albert Keen (Austin, TX); Joseph Boutros (Austin, TX); Alexander John Zelenak (Austin, TX)
Assignee: data.world, Inc.
G06F16/2471G06F16/215G06F16/2465G06F16/252G06F16/256G06F16/258G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,984,008
App. No.
16/287,967
Granted
Apr 20, 2021
Kind
B2
Abstract

Various embodiments relate generally to data science and data analysis, computer software and systems, and wired and wireless network communications to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby a collaborative data layer and associated logic facilitate, for example, efficient access to, and implementation of, collaborative datasets. In some examples, a method may include receiving data representing a query into a collaborative dataset consolidation system, identifying datasets relevant to the query, generating one or more queries to access disparate data repositories, and retrieving data representing query results. In some cases, one or more queries are applied (e.g., as a federated query) to atomized datasets stored in one or more atomized data stores, at least two of which may be different.

Claims (58)

1. A method comprising:

receiving a data file including a dataset into a collaborative dataset consolidation system;

formatting the dataset to form a first atomized dataset including atomized data points each including data representing at least two objects and an association between the two objects;

forming a second atomized dataset including the first atomized dataset and one or more other atomized datasets;

receiving data representing a query into the collaborative dataset consolidation system, the query being associated with an identifier, which includes data representing an account identifier;

identifying a subset of the second atomized dataset relevant to the query, wherein portions of the second atomized dataset are disposed in different data repositories;

deriving at least an inferred attribute for the subset of data;

generating a plurality of sub-queries each of which is configured to access at least one of the different data repositories;

retrieving data representing query results from the at least one of the different data repositories;

monitoring updates to the inferred attribute associated with attributes to detect an update; and

disseminating an update to a community of networked users.

2. The method of claim 1 wherein disseminating the update to the community of networked users comprises:

generating data representing a notification to a user computing device associated with the one or more other account identifiers.

3. The method of claim 2 wherein generating data representing the notification comprises:

causing presentation of the notification in an activity feed portion of a user interface of the user computing device.

4. The method of claim 1 further comprising:

detecting discovery of modified data associated with a data-related activity.

5. The method of claim 4 wherein detecting discovery of the modified data associated with the data-related activity comprises:

determining an activity of querying the dataset.

6. The method of claim 4 wherein detecting discovery of the modified data associated with the data-related activity comprises:

determining an activity of modifying the dataset.

7. The method of claim 4 further comprising:

generating data representing a recommendation.

8. The method of claim 1 wherein generating the plurality of sub-queries comprises:

classifying query portions; and

identifying a classification type for a portion of the query.

9. The method of claim 1 wherein the datasets comprise linked data points including triples.

10. The method of claim 9 wherein at least one triple of the triples are formatted to comply with a Resource Description Framework (“RDF”) data model.

11. An apparatus comprising:

a memory including executable instructions; and

a processor configured to execute the executable instructions to:

receive a data file including a dataset into a collaborative dataset consolidation system;

format the dataset to form a first atomized dataset including atomized data points each including data representing at least two objects and an association between the two objects;

form a second atomized dataset including the first atomized dataset and one or more other atomized datasets;

receive data representing a query into the collaborative dataset consolidation system, the query being associated with an identifier, which includes data representing an account identifier;

identify a subset of the second atomized dataset relevant to the query, wherein portions of the second atomized dataset are disposed in different data repositories;

derive at least an inferred attribute for the subset of data;

generate a plurality of sub-queries each of which is configured to access at least one of the different data repositories;

retrieve data representing query results from the at least one of the different data repositories;

monitor updates to the inferred attribute associated with attributes to detect an update; and

disseminate an update to a community of networked users.

12. The apparatus of claim 1 wherein the processor being configured to disseminate the update to the community of networked users is further configured to:

generate data representing a notification to a user computing device associated with the one or more other account identifiers.

13. The apparatus of claim 12 wherein the processor being configured to generate data representing the notification is further configured to:

cause presentation of the notification in an activity feed portion of a user interface of the user computing device.

14. The apparatus of claim 11 wherein the processor is further configured to:

detect discovery of modified data associated with a data-related activity.

15. The apparatus of claim 14 wherein the processor being configured to detect discovery of the modified data associated with the data-related activity is further configured to:

determine an activity of querying the dataset.

16. The apparatus of claim 14 wherein the processor being configured to detect discovery of the modified data associated with the data-related activity is further configured to:

determine an activity of modifying the dataset.

17. The apparatus of claim 14 wherein the processor is further configured to:

generate data representing a recommendation.

18. The apparatus of claim 11 wherein the processor is further configured to generate the plurality of sub-queries by executing instructions configured to:

classify query portions; and

identify a classification type for a portion of the query.

19. The apparatus of claim 11 wherein the datasets comprise linked data points including triples.

20. The apparatus of claim 19 wherein at least one triple of the triples are formatted to comply with a Resource Description Framework (“RDF”) data model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2025
From: DATA.WORLD, INC.
To: SERVICENOW, INC.
Reel/Frame 073004/0844 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2019
From: JACOB, BRYON KRISTEN; LOYENS, JON; GRIFFITH, DAVID LEE; HURT, BRETT A.; REYNOLDS, SHAD; KEEN, ARTHUR ALBERT; BOUTROS, JOSEPH; ZELENAK, ALEXANDER JOHN
To: DATA.WORLD, INC.
Reel/Frame 048528/0775 →
Continuity (4)
Continuation 15186520 · Jun 19, 2016
Continuation 16120057 · Aug 31, 2018
Continuation 15186514 · Jun 19, 2016
Related Publication 20190266155A1 · Aug 29, 2019
Cited By (2)
US 12,292,870 US 12,608,366