IP Library Granted Patent US 10,922,291
Granted Patent B2
US 10,922,291 · App. 16/230,769 · Granted Feb 16, 2021

Data pipeline branching

Inventors: Vipul Shekhawat (Brooklyn, NY); Eliot Ball (London, GB); Mikhail Proniushkin (New York, NY); Meghan Nayan (New York, NY); Mihir Rege (London, GB)
Assignee: Palantir Technologies Inc.
G06F16/219G06F11/1451G06F16/2379G06F2201/80G06F2201/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,922,291
App. No.
16/230,769
Granted
Feb 16, 2021
Kind
B2
Abstract

A workbook management system provides a master branch of a data pipeline comprising a pointer(s) to a snapshot(s) of an initial dataset(s), a first logic, and a pointer(s) to a snapshot(s) of a first derived dataset(s) resulting from applying the first logic to the initial dataset(s). Responsive to user input requesting a test branch corresponding to the master branch, the system creates the test branch comprising the pointer(s) to the snapshot(s) of the initial dataset(s) and a copy of the first logic. The system receives a request to modify the test branch comprising at least one change to the copy of the first logic, and modifies the test branch independently of the master branch to include second logic reflecting the at least one change to the copy of the first logic, the pointer(s) to the snapshot(s) of the initial dataset(s), and a pointer(s) to snapshot(s) of a second derived dataset(s) resulting from applying the second logic to the initial dataset(s). Responsive to user input requesting a merge of the modified test branch into the master branch, the system updates the master branch to replace the first logic with the second logic and to replace the pointer(s) to the snapshot(s) of the first derived dataset(s) with the pointer(s) to the snapshot(s) of the second derived dataset(s).

Claims (94)

1. A method comprising:

identifying a master branch of a data pipeline comprising ordered data transformation operations, the master branch having a master branch entry in a branch data structure, the master branch entry comprising a first reference to a snapshot of an initial dataset, a first logic implementing a data transformation operation, and a second reference to a snapshot of a first derived dataset resulting from applying the first logic to the initial dataset;

creating a first test branch having a first test branch entry in the branch data structure, wherein creating the first test branch comprises:

storing, in the first test branch entry of the branch data structure, the first reference to the snapshot of the initial dataset, and

storing, in the first test branch entry of the branch data structure, a first copy of the first logic and the second reference to the snapshot of the first derived dataset;

receiving a request to modify the first test branch, the request comprising at least one change to the first copy of the first logic;

modifying the first test branch independently of the master branch to include a second logic reflecting the at least one change to the first copy of the first logic, wherein modifying the first test branch comprises:

updating the first logic with the second logic in the first test branch entry of the branch data structure, and

generating a second derived dataset by applying the second logic to the snapshot of the initial dataset; and

responsive to receiving user input requesting a merge of the modified first test branch into the master branch:

replacing, in the master branch entry of the branch data structure, the first logic with the second logic, and

replacing, in the master branch entry of the branch data structure, the second reference to the snapshot of the first derived dataset with a third reference to a snapshot of the second derived dataset;

wherein the method is performed using one or more processors.

2. The method of claim 1 , further comprising:

prior to updating the master branch to replace the first logic with the second logic, determining one or more differences between the first logic and the second logic;

generating an indication of the one or more differences between the first logic and the second logic; and

receiving user input confirming that the one or more differences between the first logic and the second logic are approved.

3. The method of claim 1 , further comprising:

prior to updating the master branch to replace the second reference to the snapshot of the first derived dataset with the third reference to the snapshot of the second derived dataset, determining one or more differences between the first derived dataset and the second derived dataset;

generating an indication of the one or more differences between the first derived dataset and the second derived dataset; and

receiving user input confirming that the one or more differences between the first derived dataset and the second derived dataset are approved.

4. The method of claim 1 , further comprising:

responsive to receiving user input requesting a second test branch corresponding to the master branch, creating the second test branch having a second test branch entry in the branch data structure, the second test branch entry comprising the first reference to the snapshot of the initial dataset, and a second copy of the first logic;

receiving a request to modify the second test branch, the request comprising at least one change to the second copy of the first logic;

modifying the second test branch independently of the master branch to include third logic reflecting the at least one change to the second copy of the first logic, the third logic to be applied to the initial dataset to produce a third derived dataset, wherein modifying the second test branch comprises updating the second logic with the third logic in the second test branch entry in the branch data structure; and

responsive to user input requesting a merge of the modified second test branch into the updated master branch, updating the updated master branch entry in the branch data structure to replace the second logic with the third logic and to replace the third reference to the snapshot of the second derived dataset with a fourth reference to a snapshot of the third derived dataset.

5. The method of claim 4 , further comprising:

prior to updating the updated master branch, determining whether a merge conflict exists between the second logic and the third logic; and

responsive to determining that the merge conflict exists, receiving user input comprising a selection of the third logic to resolve the merge conflict.

6. The method of claim 1 , further comprising:

prior to updating the master branch to replace the second reference to the snapshot of the first derived dataset with the third reference to the snapshot of the second derived dataset, executing a data health check operation on the second derived dataset to determine whether the second derived dataset satisfies one or more conditions, the one or more conditions comprising a verification that a creation of the second derived dataset completed successfully and a verification that the second derived dataset is not stale.

7. The method of claim 1 , further comprising: responsive to user input requesting protection of the modified first test branch,

preventing other users from further modifying the modified first test branch; and

responsive to a request from another user to further modify the modified first test branch, creating a child test branch associated with the modified first test branch, the child test branch comprising the first reference to the snapshot of the initial dataset and a copy of the second logic.

8. The method of claim 7 , further comprising:

responsive to updating the master branch, deleting the modified first test branch and associating the child test branch with the master branch.

9. The method of claim 1 , wherein the first logic is part of the data pipeline, the data pipeline further comprising additional logic to apply to the first derived dataset to produce one or more first additional derived datasets, the method further comprising:

replacing the first logic in the data pipeline with the second logic to derive the second derived dataset;

applying the additional logic to the second derived dataset to derive one or more second additional derived datasets;

identifying one or more differences between the one or more second additional derived datasets and the one or more first additional derived datasets; and

generating an indication of the differences between the one or more second additional derived datasets and the one or more first additional derived datasets.

10. The method of claim 1 , further comprising: displaying, via a graphical user interface (GUI), a visual representation of the data pipeline, including a first graph corresponding to the master branch and a second graph corresponding to the first test branch,

wherein the first graph includes a first node representing the initial dataset, a second node representing the first derived dataset, and a first edge connecting the first node and the second node, wherein the first edge references the first logic to be applied to the initial dataset in order to produce the first derived dataset, and

wherein the second graph includes a third node representing the initial dataset, a fourth node representing the second derived dataset, and a second edge connecting the third node and the fourth node, wherein the second edge references the second logic to be applied to the initial dataset in order to produce the second derived dataset.

11. A system comprising: memory; and

one or more processors coupled to the memory, the one or more processors to execute instructions to cause the one or more processors to perform operations comprising:

identifying a master branch of a data pipeline comprising ordered data transformation operations, the master branch having a master branch entry in a branch data structure, the master branch entry comprising a first reference to a snapshot of an initial dataset, a first logic implementing a data transformation operation, and a second reference to a snapshot of a first derived dataset resulting from applying the first logic to the initial dataset;

creating a first test branch having a first test branch entry in the branch data structure, wherein creating the first test branch comprises:

storing, in the first test branch entry of the branch data structure, the first to the snapshot of the initial dataset, and

storing, in the first test branch entry of the branch data structure, a first copy of the first logic and the second reference to the snapshot of the first derived dataset;

receiving a request to modify the first test branch, the request comprising at least one change to the first copy of the first logic;

modifying the first test branch independently of the master branch to include a second logic reflecting the at least one change to the first copy of the first logic, wherein modifying the first test branch comprises:

updating the first logic with the second logic in the first test branch entry in the branch data structure, and

generating a second derived dataset by applying the second logic to the snapshot of the initial dataset; and

responsive to user input requesting a merge of the modified first test branch into the master branch:

replacing, in the master branch entry of the branch data structure, the first logic with the second logic, and

replacing, in the master branch entry of the branch data structure, the second reference to the snapshot of the first derived dataset with a third reference to a snapshot of the second derived dataset.

12. The system of claim 11 , wherein the operations further comprise:

prior to updating the master branch to replace the first logic with the second logic, determining one or more differences between the first logic and the second logic;

generating an indication of the one or more differences between the first logic and the second logic; and

receiving user input confirming that the one or more differences between the first logic and the second logic are approved.

13. The system of claim 11 , wherein the operations further comprise:

prior to updating the master branch to replace the second reference to the snapshot of the first derived dataset with the third reference to the snapshot of the second derived dataset, determining one or more differences between the first derived dataset and the second derived dataset;

generating an indication of the one or more differences between the first derived dataset and the second derived dataset; and

receiving user input confirming that the one or more differences between the first derived dataset and the second derived dataset are approved.

14. The system of claim 11 , wherein the operations further comprise:

responsive to user input requesting a second test branch corresponding to the master branch, creating the second test branch having a second test branch entry in the branch data structure, the second test branch entry comprising the first reference to the snapshot of the initial dataset, and a second copy of the first logic;

receiving a request to modify the second test branch, the request comprising at least one change to the second copy of the first logic;

modifying the second test branch independently of the master branch to include third logic reflecting the at least one change to the second copy of the first logic, the third logic to be applied to the initial dataset to produce a third derived dataset, wherein modifying the second test branch comprises updating the second logic with the third logic in the second test branch entry in the branch data structure; and

responsive to user input requesting a merge of the modified second test branch into the updated master branch, updating the updated master branch entry in the branch data structure to replace the second logic with the third logic and to replace the third reference to the snapshot of the second derived dataset with a fourth reference to a snapshot of the third derived dataset.

15. The system of claim 14 , wherein the operations further comprise:

prior to updating the updated master branch, determining whether a merge conflict exists between the second logic and the third logic; and

responsive to determining that the merge conflict exists, receiving user input comprising a selection of the third logic to resolve the merge conflict.

16. The system of claim 11 , wherein the operations further comprise:

prior to updating the master branch to replace the second reference to the snapshot of the first derived dataset with the third reference to the snapshot of the second derived dataset, executing a data health check operation on the second derived dataset to determine whether the second derived dataset satisfies one or more conditions, the one or more conditions comprising a verification that a creation of the second derived dataset completed successfully and a verification that the second derived dataset is not stale.

17. The system of claim 11 , wherein the operations further comprise:

responsive to user input requesting protection of the modified first test branch, preventing other users from further modifying the modified first test branch; and

responsive to a request from another user to further modify the modified first test branch, creating a child test branch associated with the modified first test branch, the child test branch comprising the first reference to the snapshot of the initial dataset and a copy of the second logic.

18. A non-transitory computer readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

identifying a master branch of a data pipeline comprising ordered data transformation operations, the master branch having a master branch entry in a branch data structure, the master branch entry comprising a first reference to a snapshot of an initial dataset, a first logic implementing a data transformation operation, and a second reference to a snapshot of a first derived dataset resulting from applying the first logic to the initial dataset;

creating a first test branch having a first test branch entry in the branch data structure, wherein creating the first test branch comprises:

storing, in the first test branch entry of the branch data structure, the first reference to the snapshot of the initial dataset, and

storing, in the first test branch entry of the branch data structure, a first copy of the first logic and the second reference to the snapshot of the first derived dataset;

receiving a request to modify the first test branch, the request comprising at least one change to the first copy of the first logic;

modifying the first test branch independently of the master branch to include a second logic reflecting the at least one change to the first copy of the first logic, wherein modifying the first test branch comprises:

updating the first logic with the second logic in the first test branch entry in the branch data structure, and

generating a second derived dataset by applying the second logic to the snapshot of the initial dataset; and

responsive to user input requesting a merge of the modified first test branch into the master branch:

replacing, in the master branch entry of the branch data structure, the first logic with the second logic, and

replacing, in the master branch entry of the branch data structure, the second reference to the snapshot of the first derived dataset with a third reference to a snapshot of the second derived dataset.

19. The non-transitory computer readable storage medium of claim 18 , wherein the operations further comprise:

prior to updating the master branch to replace the first logic with the second logic, determining one or more first differences between the first logic and the second logic and one or more second differences between the first derived dataset and the second derived dataset;

generating an indication of the first one or more differences between the first logic and the second logic and of the second one or more differences between the first derived dataset and the second derived dataset; and

receiving user input confirming that the first one or more differences between the first logic and the second logic and the second one or more differences between the first derived dataset and the second derived dataset are approved.

Assignments (2)
SECURITY INTEREST Recorded Jul 3, 2022
From: PALANTIR TECHNOLOGIES INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0506 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2019
From: SHEKHAWAT, VIPUL; BALL, ELIOT; PRONIUSHKIN, MIKHAIL; NAYAN, MEGHAN; REGE, MIHIR
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 048204/0146 →