IP Library Granted Patent US 11,366,824
Granted Patent B2
US 11,366,824 · App. 17/114,377 · Granted Jun 21, 2022

Dataset analysis and dataset attribute inferencing to form collaborative datasets

Inventors: Bryon Kristen Jacob (Austin, TX); David Lee Griffith (Austin, TX); Triet Minh Le (Austin, TX); Jon Loyens (Austin, TX); Brett A. Hurt (Austin, TX); Arthur Albert Keen (Austin, TX)
Assignee: data.world, Inc.
G06F16/258G06F16/9024G06N5/022G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,366,824
App. No.
17/114,377
Granted
Jun 21, 2022
Kind
B2
Abstract

Various embodiments relate generally to data science and data analysis, computer software and systems, and wired and wireless network communications to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby a collaborative data layer and associated logic facilitate, for example, efficient access to, and implementation of, collaborative datasets. In some examples, a method may include receiving a dataset having a data format into a dataset ingestion controller configured to form a collaborative dataset, interpreting data of the dataset against data classifications at an inference engine to derive at least an inferred attribute, associating the data with annotative data identifying the inferred attribute, and converting the dataset at a format converter to form an atomized dataset.

Claims (61)

1. A method comprising:

receiving ingested data at a computing system including one or more processors configured to execute instructions to implement one or more executable modules, the ingested data representing at least one dataset of a plurality of datasets, each dataset having a different data format as ingested into a dataset ingestion controller module;

implementing one or more modules in a collaborative dataset consolidation system configured to form a collaborative dataset independent of each of the different data formats prior to ingestion;

determining attributes of the collaborative dataset, a first attribute being associated with an account identifier and a second attribute identifying a first version of the collaborative dataset;

interpreting a subset of ingested data of the dataset against one or more data classifications at an inference engine module to derive at least an inferred attribute for the subset of data;

updating the collaborative dataset to link to data associated with the ingested data;

generating a second version of the collaborative dataset responsive to the link to the data associated with the ingested data including the inferred attribute;

transmitting activity feed data to disseminate an update to a community of networked users associated with the account identifier, the update including data representing generation of the second version of the collaborative dataset

deducing a data classification for the subset of ingested data of the dataset to form a deduced data classification;

analyzing other annotative data associated with equivalently classified data in other datasets;

identifying that the subset of ingested data omits an association to annotative data representing an annotative description; and

associating the subset of the ingested data with the annotative data based on the deduced data classification to identify the inferred attribute that includes data representing an account identifier associated with the identifier, one or more other account identifiers, and one or more data-related activities.

2. The method of claim 1 further comprising:

identifying that the subset of ingested data is associated with a protected dataset;

determining access to the protected dataset is authorized in association with the account identifier; and

forming the second version of the collaborative dataset.

3. The method of claim 1 further comprising:

generating a data signal specifying information for at least one of the account identifiers that accesses data associated with annotative data; and

causing presentation of the information in an activity feed portion of a user interface.

4. The method of claim 1 further comprising:

converting the plurality of datasets from the different data formats at a format converter module associated with the computing system to form an atomized dataset having a specific format.

5. A system comprising:

one or more computer systems including a memory including executable instructions, and a processor that is configured, responsive to executing the instructions, to:

receive ingested data at a computing system including one or more processors configured to execute instructions to implement one or more executable modules, the ingested data representing at least one dataset of a plurality of datasets, each dataset having a different data format as ingested into a dataset ingestion controller module; implement one or more modules in a collaborative dataset consolidation system configured to form a collaborative dataset independent of each of the different data formats prior to ingestion;

determine attributes of the collaborative dataset, a first attribute being associated with an account identifier and a second attribute identifying a first version of the collaborative dataset;

interpret a subset of ingested data of the dataset against one or more data classifications at an inference engine module to derive at least an inferred attribute for the subset of data;

update the collaborative dataset to link to data associated with the ingested data;

generate a second version of the collaborative dataset responsive to the link to the data associated with the ingested data including the inferred attribute;

transmit activity feed data to disseminate an update to a community of networked users associated with the account identifier, the update including data representing generation of the second version of the collaborative dataset;

deduce a data classification for the subset of ingested data of the dataset to form a deduced data classification;

analyze other annotative data associated with equivalently classified data in other datasets;

identify that the subset of ingested data omits an association to annotative data representing an annotative description; and

associate the subset of the ingested data with the annotative data based on the deduced data classification to identify the inferred attribute that includes data representing an account identifier associated with the identifier, one or more other account identifiers, and one or more data-related activities.

6. The system of claim 5 wherein the processor is further configured to:

identify that the subset of ingested data is associated with a protected dataset;

determine access to the protected dataset is authorized in association with the account identifier; and

form the second version of the collaborative dataset.

7. The system of claim 5 wherein the processor is further configured to:

generate a data signal specifying information for at least one of the account identifiers that accesses data associated with annotative data; and

cause presentation of the information in an activity feed portion of a user interface.

8. The system of claim 5 wherein the processor is further configured to:

convert the plurality of datasets from the different data formats at a format converter module associated with the computing system to form an atomized dataset having a specific format.

9. A computer readable storage device comprised of data representing computer instructions configured to cause a processor to initiate a computer-implemented method comprising:

receiving ingested data at a computing system including one or more processors configured to execute instructions to implement one or more executable modules, the ingested data representing at least one dataset of a plurality of datasets, each dataset having a different data format as ingested into a dataset ingestion controller module;

implementing one or more modules in a collaborative dataset consolidation system configured to form a collaborative dataset independent of each of the different data formats prior to ingestion;

determining attributes of the collaborative dataset, a first attribute being associated with an account identifier and a second attribute identifying a first version of the collaborative dataset;

interpreting a subset of ingested data of the dataset against one or more data classifications at an inference engine module to derive at least an inferred attribute for the subset of data;

updating the collaborative dataset to link to data associated with the ingested data;

generating a second version of the collaborative dataset responsive to the link to the data associated with the ingested data including the inferred attribute;

transmitting activity feed data to disseminate an update to a community of networked users associated with the account identifier, the update including data representing generation of the second version of the collaborative dataset

deducing a data classification for the subset of ingested data of the dataset to form a deduced data classification;

analyzing other annotative data associated with equivalently classified data in other datasets;

identifying that the subset of ingested data omits an association to annotative data representing an annotative description; and

associating the subset of the ingested data with the annotative data based on the deduced data classification to identify the inferred attribute that includes data representing an account identifier associated with the identifier, one or more other account identifiers, and one or more data-related activities.

10. The computer-implemented method of claim 9 further comprising:

identifying that the subset of ingested data is associated with a protected dataset;

determining access to the protected dataset is authorized in association with the account identifier; and

forming the second version of the collaborative dataset.

11. The computer-implemented method of claim 9 further comprising:

generating a data signal specifying information for at least one of the account identifiers that accesses data associated with annotative data; and

causing presentation of the information in an activity feed portion of a user interface.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2025
From: DATA.WORLD, INC.
To: SERVICENOW, INC.
Reel/Frame 073004/0844 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2021
From: JACOB, BRYON KRISTEN; LOYENS, JON; GRIFFITH, DAVID LEE; HURT, BRETT A.; LE, TRIET MINH; KEEN, ARTHUR ALBERT
To: DATA.WORLD, INC.
Reel/Frame 055085/0253 →
Continuity (15)
Continuation 16271263 · Feb 8, 2019
Continuation 15186516 · Jun 19, 2016
Continuation 17114377
Continuation 16292120 · Mar 4, 2019
Continuation 16271263 · Feb 8, 2019
Continuation 15186516 · Jun 19, 2016
Continuation 15186516 · Jun 19, 2016
Continuation 17114377
Continuation 16271687 · Feb 8, 2019
Continuation 15186520 · Jun 19, 2016
Continuation 17114377
Continuation 16292135 · Mar 4, 2019
Continuation 16271687 · Feb 8, 2019
Continuation 15186520 · Jun 19, 2016
Related Publication 20210173848A1 · Jun 10, 2021
Cited By (2)
US 12,608,394 US 12,681,950