IP Library Granted Patent US 11,675,808
Granted Patent B2
US 11,675,808 · App. 17/589,884 · Granted Jun 13, 2023

Dataset analysis and dataset attribute inferencing to form collaborative datasets

Inventors: Bryon Kristen Jacob (Austin, TX); David Lee Griffith (Austin, TX); Triet Minh Le (Austin, TX); Jon Loyens (Austin, TX); Brett A. Hurt (Austin, TX); Arthur Albert Keen (Austin, TX)
Assignee: data.world, Inc.
G06F16/258G06F16/1794G06F16/9024G06N5/022G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,675,808
App. No.
17/589,884
Granted
Jun 13, 2023
Kind
B2
Abstract

Various embodiments relate generally to data science and data analysis, computer software and systems, and wired and wireless network communications to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby a collaborative data layer and associated logic facilitate, for example, efficient access to, and implementation of, collaborative datasets. In some examples, a method may include receiving a dataset having a data format into a dataset ingestion controller configured to form a collaborative dataset, interpreting data of the dataset against data classifications at an inference engine to derive at least an inferred attribute, associating the data with annotative data identifying the inferred attribute, and converting the dataset at a format converter to form an atomized dataset.

Claims (17)

1. A method, comprising: receiving data associated with a query into a collaborative data consolidation system, the query being executed across a plurality of atomized datasets;

converting the data into one or more triples, the one or more triples being stored in one or more triplestores;

analyzing the query to classify a portion of the query to form classified query portions;

rewriting the query into a plurality of sub-queries based on the classified query portions, each of the plurality of sub-queries being formatted in a data type associated with at least one of the one or more triplestores;

transmitting each of the plurality of queries after the rewriting to one or more distributed data repositories formatted to the data type;

retrieving one or more query results in response to at least one of the plurality of sub-queries;

federating the one or more query results retrieved from the one or more distributed data repositories to generate a federated query result in response to the query;

transmitting each of the plurality of sub-queries to at least one of the one or more triplestores, each of the one or more triplestores being hosted on one or more distributed data repositories formatted according to the data type; and

retrieving one or more query results in response to at least one of the plurality of sub-queries, the one or more query results being transmitted in response to the query as the federated query result.

2. The method of claim 1 , wherein each of the plurality of atomized data sets has a data type associated with at least one repository.

3. The method of claim 1 , wherein receiving the data associated with a query is configured to be ingested by a dataset ingestion controller.

4. The method of claim 1 , wherein the data type associated with the query is SPARQL.

5. The method of claim 1 , wherein the query is executed over a graph comprising the one or more distributed data repositories.

6. The method of claim 1 , wherein a graph is generated using the query, the plurality of sub-queries, and the one or more query results, the graph being configured to execute the query as a federated query.

7. The method of claim 1 , wherein at least one of the one or more distributed data repositories is configured to store atomized triples.

8. The method of claim 1 , wherein the plurality of atomized datasets are transformed into the one or more triples, each of the one or more triples having at least one attribute.

9. The method of claim 1 , wherein each of the plurality of atomized datasets has at least one attribute, the at least one attribute being managed by a dataset attribute manager configured to monitor the plurality of atomized datasets and to update the at least one attribute if a change is detected.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2025
From: DATA.WORLD, INC.
To: SERVICENOW, INC.
Reel/Frame 073004/0844 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: JACOB, BRYON KRISTEN; LOYENS, JON; GRIFFITH, DAVID LEE; HURT, BRETT A.; LE, TRIET MINH; KEEN, ARTHUR ALBERT
To: DATA.WORLD, INC.
Reel/Frame 059103/0975 →
Continuity (13)
Continuation 17114377 · Dec 7, 2020
Continuation 16292120 · Mar 4, 2019
Continuation 16271263 · Feb 8, 2019
Continuation 15186516 · Jun 19, 2016
Continuation 15186516 · Jun 19, 2016
Continuation 16271263 · Feb 8, 2019
Continuation 15186516 · Jun 19, 2016
Continuation 16271687 · Feb 8, 2019
Continuation 15186520 · Jun 19, 2016
Continuation 16292135 · Mar 4, 2019
Continuation 16271687 · Feb 8, 2019
Continuation 15186520 · Jun 19, 2016
Related Publication 20220229847A1 · Jul 21, 2022
Cited By (2)
US 12,223,278 US 12,380,243