IP Library Granted Patent US 11,816,118
Granted Patent B2
US 11,816,118 · App. 17/893,100 · Granted Nov 14, 2023

Collaborative dataset consolidation via distributed computer networks

Inventors: Bryon Kristen Jacob (Austin, TX); Jon Loyens (Austin, TX); David Lee Griffith (Austin, TX); Brett A. Hurt (Austin, TX); Triet Minh Le (Austin, TX); Shad William Reynolds (Austin, TX); Arthur Albert Keen (Austin, TX); Joseph Boutros (Austin, TX); Alexander John Zelenak (Austin, TX)
Assignee: data.world, Inc.
G06F16/2471G06F16/215G06F16/2465G06F16/252G06F16/256G06F16/258G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,816,118
App. No.
17/893,100
Granted
Nov 14, 2023
Kind
B2
Abstract

Various embodiments relate generally to data science and data analysis, computer software and systems, and wired and wireless network communications to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby a collaborative data layer and associated logic facilitate, for example, efficient access to, and implementation of, collaborative datasets. In some examples, a method may include receiving data representing a query into a collaborative dataset consolidation system, identifying datasets relevant to the query, generating one or more queries to access disparate data repositories, and retrieving data representing query results. In some cases, one or more queries are applied (e.g., as a federated query) to atomized datasets stored in one or more atomized data stores, at least two of which may be different.

Claims (50)

1. A method comprising:

formatting a dataset to form a first atomized dataset including graph-based data associated with metadata including attributes of the dataset, the first atomized dataset being a first version;

forming a second atomized dataset including the first atomized dataset, the second atomized dataset including changed data as a second version from the first version;

receiving data representing a query;

re-writing the query to generate one or more sub-queries configured to access to the second version of the dataset including at least a portion of data stored in at least one of a different data repositories to perform a federated query;

classifying query portions of the one or more sub-queries to identify a classification type for a query portion;

applying the one or more sub-queries to the different data repositories; and

retrieving data responsive to the data representing the query.

2. The method of claim 1 wherein classifying the query portions of the one or more sub-queries to identify the classification type comprises:

identifying the classification type for a repository.

3. The method of claim 1 further comprising:

identifying a repository as including a graph database based on the classification type.

4. The method of claim 1 wherein classifying the query portions further comprises:

identifying a repository as including a triplestore database.

5. The method of claim 1 wherein classifying the query portions further comprises:

identifying a repository as a cloud storage server.

6. The method of claim 1 further comprising:

inferring dataset attributes to generate at least one of the first and second atomized datasets.

7. The method of claim 6 further comprising:

generating an inferred data attribute based on a subset of data of a tabular dataset format associated with the at least one of the first and second atomized datasets.

8. The method of claim 6 wherein generating the inferred data attribute comprises:

inferring a dataset attribute representing columnar representation of data associated with the at least one of the first and second atomized datasets.

9. The method of claim 8 wherein linked data points comprise triples.

10. The method of claim 1 wherein the first and second atomized datasets comprise linked data points.

11. The method of claim 10 wherein at least one triple of the triples is formatted to comply with a Resource Description Framework (“RDF”) data model.

12. A system comprising:

a data store configured to store executable instructions and data; and

a processor configured to execute the executable instructions, the processor being configured to:

format a dataset to form a first atomized dataset including graph-based data associated with metadata including attributes of the dataset, the first atomized dataset being a first version;

form a second atomized dataset including the first atomized dataset, the second atomized dataset including changed data as a second version from the first version;

receive data representing a query;

re-write the query to generate one or more sub-queries configured to access to the second version of the dataset including at least a portion of data stored in at least one of a number of different data repositories to perform a federated query;

classify query portions of the one or more sub-queries to identify a classification type for a query portion;

apply the one or more sub-queries to the different data repositories; and

retrieve data responsive to the data representing the query.

13. The system of claim 12 wherein the processor is further configured to:

identify the classification type for a repository.

14. The system of claim 12 wherein the processor is further configured to:

identify a repository as including a graph database based on the classification type.

15. The system of claim 12 wherein the processor is further configured to:

identify a repository as including a triplestore database.

16. The system of claim 12 wherein the processor is further configured to:

identify a repository as a cloud storage server.

17. The system of claim 12 wherein the processor is further configured to:

infer dataset attributes to generate at least one of the first and second atomized datasets.

18. The system of claim 17 wherein the processor is further configured to:

generate an inferred data attribute based on a subset of data of a tabular dataset format associated with the at least one of the first and second atomized datasets.

19. The system of claim 17 wherein the processor is further configured to:

infer a dataset attribute representing columnar representation of data associated with the at least one of the first and second atomized datasets.

20. The system of claim 12 wherein the first and second atomized datasets comprise linked data points as triples, wherein at least one triple of the triples are formatted to comply with a Resource Description Framework (“RDF”) data model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2025
From: DATA.WORLD, INC.
To: SERVICENOW, INC.
Reel/Frame 073004/0844 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2022
From: JACOB, BRYON KRISTEN; LOYENS, JON; GRIFFITH, DAVID LEE; HURT, BRETT A.; LE, TRIET MINH; REYNOLDS, SHAD; KEEN, ARTHUR ALBERT; BOUTROS, JOSEPH; ZELENAK, ALEXANDER JOHN
To: DATA.WORLD, INC.
Reel/Frame 061272/0570 →
Continuity (4)
Continuation 17037005 · Sep 29, 2020
Continuation 16120057 · Aug 31, 2018
Continuation 15186514 · Jun 19, 2016
Related Publication 20230153312A1 · May 18, 2023