IP Library Granted Patent US 11,086,896
Granted Patent B2
US 11,086,896 · App. 15/985,705 · Granted Aug 10, 2021

Dynamic composite data dictionary to facilitate data operations via computerized tools configured to access collaborative datasets in a networked computing platform

Inventors: Joseph Boutros (Austin, TX); Sharon Brener (Austin, TX); Alexander John Zelenak (Austin, TX); Robert Thomas Grochowicz (Austin, TX); Mark Joseph DiMarco (Austin, TX); Bryon Kristen Jacob (Austin, TX); David Lee Griffith (Austin, TX); Shad William Reynolds (Austin, TX)
Assignee: data.world, Inc.
G06F16/258G06F16/212G06F16/9024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,086,896
App. No.
15/985,705
Granted
Aug 10, 2021
Kind
B2
Abstract

Various embodiments relate generally to data science and data analysis, computer software and systems, network communications to interface among repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform configured to provide one or more computerized tools that facilitate data projects by providing an interactive, project-centric workspace interface that may include, for example, a unified view in which to identify data sources, generate transformative datasets, and/form queries over a composite data dictionary coupled to collaborative computing devices and user accounts. For example, a method may include forming a first data dictionary, linking a dataset associated with the first data dictionary to another dataset, which may be associated with a second data dictionary, and forming a dynamic composite data dictionary.

Claims (78)

1. A method comprising:

receiving data representing a dataset into dataset ingestion controller, the data being converted into an atomized dataset comprising a triple, the data also comprising an executable command to the atomized dataset to generate one or more results;

identifying a first data arrangement in which the data representing the dataset has a first format;

analyzing the triple associated with the data representing the dataset to determine a first subset of identifiers for subsets of data;

forming a first data dictionary including the first subset of identifiers for the subsets of data in the dataset;

formatting the dataset into a second data arrangement having a second format;

receiving the triple associated with the data originating at a data project user interface to link the dataset to another dataset, which is associated with a second data dictionary; and

forming a composite data dictionary including the first data dictionary and the second data dictionary using the one or more results generated by applying the executable command to the atomized dataset.

2. The method of claim 1 further comprising:

receiving data representing a request to import the another dataset into a data project;

associating the first set of identifiers of the dataset and the second set of identifiers of the another dataset in a data arrangement representing the data project; and

forming a combined dataset that is associated with the composite data dictionary.

3. The method of claim 1 further comprising:

determining whether to store data of the another dataset at a first computing device.

4. The method of claim 3 further comprising:

forming a link identifier to reference data of the another dataset locally;

storing the data of the another dataset in a data repository at the first computing device; and

implementing a dataset identifier in the composite data dictionary,

wherein the first computing device is disposed locally relative to a collaborative dataset consolidation system.

5. The method of claim 3 further comprising:

detecting linkage to a remote dataset at a second computing device;

forming a link identifiers to reference data of the another dataset remotely to identify a remote dataset; and

transforming a dataset identifier associated with the remote dataset into a local namespace;

wherein the second computing device is disposed externally relative to a collaborative dataset consolidation system.

6. The method of claim 5 further comprising:

implementing the dataset identifier in the composite data dictionary.

7. The method of claim 1 further comprising:

detecting activation of user input to form a query operation; and

detecting selection of a dataset identifier associated with the composite data dictionary, the dataset identifier referencing a subset of data in the another dataset.

8. The method of claim 7 further comprising:

determining the another dataset is disposed remotely to form a remote dataset.

9. The method of claim 8 further comprising:

determining a transformed link identifier is available; and

applying the query operation via the transformed link identifier against the remote dataset as an implicit federated query.

10. The method of claim 8 further comprising:

applying the query operation via a path identifier against the remote dataset as an explicit federated query.

11. The method of claim 8 further comprising:

generating a service graph call to access the remote dataset.

12. The method of claim 1 wherein formatting the dataset into the second data arrangement comprises:

converting the dataset into an atomized dataset having a graph data arrangement.

13. The method of claim 12 further comprising:

associating an identifier of the data dictionary to a subset of the atomized dataset.

14. The method of claim 12 wherein analyzing the data representing the dataset to determine identifiers for subsets of data comprises:

determining a subset of dataset attributes for a subset of the atomized dataset; and

deriving an annotation to form a derived annotation for the subset of the atomized dataset; and

forming the identifier as a function of the derived annotation.

15. The method of claim 14 wherein deriving the annotation comprises:

extracting data representing a column header associated with the first data arrangement; and

implementing the data representing the column header as an identifier.

16. The method of claim 13 wherein analyzing the data representing the dataset to determine identifiers for subsets of data comprises:

generating data representing a request for an identifier;

receiving data responsive to the request for the identifier; and

applying the data as the identifier.

17. An apparatus comprising:

a memory including executable instructions; and

a processor, responsive to executing the instructions, is configured to:

receive data representing a dataset into dataset ingestion controller, the data being converted into an atomized dataset comprising a triple, the data also comprising an executable command to the atomized dataset to generate one or more results;

identify a first data arrangement in which the data representing the dataset has a first format;

analyze the triple associated with the data representing the dataset to determine a first subset of identifiers for subsets of data;

form a first data dictionary including the first subset of identifiers for the subsets of data in the dataset;

format the dataset into a second data arrangement having a second format;

receive the triple associated with the data originating at a data project user interface to link the dataset to another dataset, which is associated with a second data dictionary; and

form a composite data dictionary including the first data dictionary and the second data dictionary using the one or more results generated by applying the executable command to the atomized dataset.

18. The apparatus of claim 17 wherein a subset of the instructions further causes the processor to:

receive data representing a request to import the another dataset into a data project;

associate the first set of identifiers of the dataset and the second set of identifiers of the another dataset in a data arrangement representing the data project; and

form a combined dataset that is associated with the composite data dictionary.

19. The apparatus of claim 17 wherein a subset of the instructions further causes the processor to:

determine whether to store data of the another dataset at a first computing device;

form a link identifiers to reference data of the another dataset locally;

store the data of the another dataset in a data repository at the first computing device; and

implement a dataset identifier in the composite data dictionary,

wherein the first computing device is disposed locally relative to a collaborative dataset consolidation system.

20. The apparatus of claim 17 wherein a subset of the instructions further causes the processor to:

form a link identifier to reference data of the another dataset locally;

store the data of the another dataset in a data repository at a first computing device; and

implement a dataset identifier in the composite data dictionary,

wherein the first computing device is disposed locally relative to a collaborative dataset consolidation system.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2025
From: DATA.WORLD, INC.
To: SERVICENOW, INC.
Reel/Frame 073004/0844 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2018
From: BOUTROS, JOSEPH; BRENER, SHARON; ZELENAK, ALEXANDER JOHN; GROCHOWICZ, ROBERT THOMAS; DIMARCO, MARK JOSEPH; JACOB, BRYON KRISTEN; GRIFFITH, DAVID LEE; REYNOLDS, SHAD WILLIAM
To: DATA.WORLD, INC.
Reel/Frame 045981/0994 →
Continuity (6)
Continuation In Part 15186514 · Jun 19, 2016
Continuation In Part 15186516 · Jun 19, 2016
Continuation In Part 15454923 · Mar 9, 2017
Continuation In Part 15926999 · Mar 20, 2018
Continuation In Part 15927004 · Mar 20, 2018
Related Publication 20190065569A1 · Feb 28, 2019