IP Library Patent Application 15993530
Patent Application
App. No. 15/993,530

PROVIDING FULL DATA PROVENANCE VISUALIZATION FOR VERSIONED DATASETS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
15/993,530
Abstract

Systems and methods for providing full data provenance visualization for versioned datasets. A method includes receiving selection of a versioned dataset that is within a data pipeline system. The method also includes determining the full data provenance of the selected versioned dataset. The full data provenance may comprise a set of versioned datasets. The method further includes providing for display of a visualization of the full data provenance of the selected versioned dataset. The visualization comprises a graph. The graph comprises a compound node for the selected versioned dataset and for each versioned dataset in the set of versioned datasets. The graph further comprises edges connecting the compounds nodes. Each edge represents a derivation dependency between versions of the versioned datasets represented by the compound nodes connected by the edge.

Claims (33)

1 . A method, comprising:

at one or more computing devices having one or more processors and memory storing one or more programs executed by the one or more processors to perform the method, performing the operations of:

receiving selection of a versioned dataset that is within a data pipeline system;

determining full data provenance of the selected versioned dataset, the full data provenance comprising a set of versioned datasets;

providing for display of a visualization of the full data provenance of the selected versioned dataset, the visualization comprising a graph, the graph comprising a compound node for the selected versioned dataset and a compound node for each versioned dataset in the set of versioned datasets, the graph further comprising edges connecting the compounds nodes, each edge representing a derivation dependency between versions of the versioned datasets represented by the compound nodes connected by the edge, the graph further comprising one or more compound nodes with a plurality of sub-entries for different versions of a dataset of the one or more compound nodes.

2 . The method of claim 1 , wherein the compound node of the selected versioned dataset indicates a name or identifier of the selected version dataset; and wherein the compound node for each versioned dataset in the set of versioned datasets indicates a name or identifier of the each versioned dataset.

3 . The method of claim 1 , wherein the compound node of the selected versioned dataset comprises a sub-entry representing a particular version of the selected versioned dataset.

4 . The method of claim 1 , wherein the compound node for each versioned dataset in the set of versioned datasets comprises at least one sub-entry representing a version of the each versioned dataset in the full data provenance of the selected versioned dataset.

5 . The method of claim 1 , wherein a sub-entry of the compound node for a particular versioned dataset in the set of versioned datasets is visually distinguished in the graphical user interface from other compound node sub-entries of the graph to indicate that a version, of the particular versioned dataset, represented by the sub-entry has been flagged in a database as containing invalid data.

6 . The method of claim 1 , wherein an edge in the graph representing a derivation dependency of a first version of a first versioned dataset in the set of versioned datasets on a second version of a second versioned dataset in the set of versioned datasets is visually distinguished from other edges in the graph to indicate that the first version of the first versioned dataset potentially contains invalid data as a result of the derivation dependency.

7 . The method of claim 1 , wherein at least one version of a versioned dataset in the set of versioned datasets contains data generated as a result of one or more Spark systems executing a derivation program taking at least one version of another versioned dataset as input to the derivation program.

8 . The method of claim 1 , wherein at least one version of a versioned dataset in the set of versioned datasets contains data generated as a result of one or more MapReduce systems executing a derivation program taking at least one version of another versioned dataset as input as input to the derivation program.

9 . One or more non-transitory computer-readable media storing one or more programs, the one or more programs comprising instructions for:

receiving selection of a versioned dataset that is within a data pipeline system;

determining full data provenance of the selected versioned dataset, the full data provenance comprising a set of versioned datasets;

providing for display of a visualization of the full data provenance of the selected versioned dataset, the visualization comprising a graph, the graph comprising a compound node for the selected versioned dataset and a compound node for each versioned dataset in the set of versioned datasets, the graph further comprising edges connecting the compounds nodes, each edge representing a derivation dependency between versions of the versioned datasets represented by the compound nodes connected by the edge, the graph further comprising one or more compound nodes with a plurality of sub-entries for different versions of a dataset of the one or more compound nodes.

10 . The one or more non-transitory computer-readable media of claim 9 , wherein the compound node of the selected versioned dataset indicates a name or identifier of the selected version dataset; and wherein the compound node for each versioned dataset in the set of versioned datasets indicates a name or identifier of the each versioned dataset.

11 . The one or more non-transitory computer-readable media of claim 9 , wherein the compound node of the selected versioned dataset comprises a sub-entry representing a particular version of the selected versioned dataset.

12 . The one or more non-transitory computer-readable media of claim 9 , wherein the compound node for each versioned dataset in the set of versioned datasets comprises at least one sub-entry representing a version of the each versioned dataset in the full data provenance of the selected versioned dataset.

13 . The one or more non-transitory computer-readable media of claim 9 , wherein a sub-entry of the compound node for a particular versioned dataset in the set of versioned datasets is visually distinguished in the graphical user interface from other compound node sub-entries of the graph to indicate that a version, of the particular versioned dataset, represented by the sub-entry has been flagged in a database as containing invalid data.

14 . The one or more non-transitory computer-readable media of claim 9 , wherein an edge in the graph representing a derivation dependency of a first version of a first versioned dataset in the set of versioned datasets on a second version of a second versioned dataset in the set of versioned datasets is visually distinguished from other edges in the graph to indicate that the first version of the first versioned dataset potentially contains invalid data as a result of the derivation dependency.

15 . The one or more non-transitory computer-readable media of claim 9 , wherein at least one version of a versioned dataset in the set of versioned datasets contains data generated as a result of one or more Spark systems executing a derivation program taking at least one version of another versioned dataset as input to the derivation program.

16 . The one or more non-transitory computer-readable media of claim 9 , wherein at least one version of a versioned dataset in the set of versioned datasets contains data generated as a result of one or more MapReduce systems executing a derivation program taking at least one version of another versioned dataset as input as input to the derivation program.

17 . A system comprising:

memory;

one or more processors;

one or more programs stored in the memory and configured for execution by the one or more processors, the one or more programs comprising instructions for:

receiving selection of a versioned dataset that is within a data pipeline system;

determining full data provenance of the selected versioned dataset, the full data provenance comprising a set of versioned datasets;

providing for display of a visualization of the full data provenance of the selected versioned dataset, the visualization comprising a graph, the graph comprising a compound node for the selected versioned dataset and a compound node for each versioned dataset in the set of versioned datasets, the graph further comprising edges connecting the compounds nodes, each edge representing a derivation dependency between versions of the versioned datasets represented by the compound nodes connected by the edge, the graph further comprising one or more compound nodes with a plurality of sub-entries for different versions of a dataset of the one or more compound nodes.

18 . The system of claim 17 , wherein the compound node of the selected versioned dataset indicates a name or identifier of the selected version dataset; and wherein the compound node for each versioned dataset in the set of versioned datasets indicates a name or identifier of the each versioned dataset.

19 . The system of claim 17 , wherein the compound node of the selected versioned dataset comprises a sub-entry representing a particular version of the selected versioned dataset.

20 . The system of claim 17 , wherein the compound node for each versioned dataset in the set of versioned datasets comprises at least one sub-entry representing a version of the each versioned dataset in the full data provenance of the selected versioned dataset.

Assignments (7)
SECURITY INTEREST Recorded Jul 3, 2022
From: PALANTIR TECHNOLOGIES INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0506 →
ASSIGNMENT OF INTELLECTUAL PROPERTY SECURITY AGREEMENTS Recorded Jul 3, 2022
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0640 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ERRONEOUSLY LISTED PATENT BY REMOVING APPLICATION NO. 16/832267 FROM THE RELEASE OF SECURITY INTEREST PREVIOUSLY RECORDED ON REEL 052856 FRAME 0382. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Aug 26, 2021
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 057335/0753 →
SECURITY INTEREST Recorded Jun 4, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 052856/0817 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2020
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 052856/0382 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 051709/0471 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
Reel/Frame 051713/0149 →