IP Library Granted Patent US 10,853,338
Granted Patent B2
US 10,853,338 · App. 16/240,507 · Granted Dec 1, 2020

Universal data pipeline

Inventors: Jacob Meacham (Sunnyvale, CA); Michael Harris (Palo Alto, CA); Gustav Brodman (Palo Alto, CA); Lynn Cuthriell (San Francisco, CA); Hannah Korus (Palo Alto, CA); Brian Toth (Palo Alto, CA); Jonathan Hsiao (Palo Alto, CA); Mark Elliot (Arlington, VA); Brian Schimpf (Vienna, VA); Michael Garland (Palo Alto, CA); Evelyn Nguyen (Menlo Park, CA)
Assignee: PALANTIR TECHNOLOGIES INC.
G06F16/219G06F16/1865G06F16/1873G06F16/211G06F16/2386G06F16/254G06F16/2365
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,853,338
App. No.
16/240,507
Granted
Dec 1, 2020
Kind
B2
Abstract

A history preserving data pipeline computer system and method. In one aspect, the history preserving data pipeline system provides immutable and versioned datasets. Because datasets are immutable and versioned, the system makes it possible to determine the data in a dataset at a point in time in the past, even if that data is no longer in the current version of the dataset.

Claims (44)

1. A method comprising:

at one or more computing devices comprising one or more processors and one or more storage media storing one or more computer programs executed by the one or more processors to perform the method, performing operations comprising:

maintaining a build catalog comprising a plurality of build catalog entries, each build catalog entry comprising:

an identifier of a version of a derived dataset corresponding to a build catalog entry,

one or more dataset build dependencies of the version of the derived dataset corresponding to the build catalog entry, each of the one or more dataset build dependencies comprising an identifier of a version of a child dataset from which the version of the derived dataset corresponding to the build catalog entry is derived, and

a derivation program build dependency that is executable to generate the version of the derived dataset corresponding to the build catalog entry;

creating a new version of a particular derived dataset based on providing one or more particular child dataset versions as input to executing a particular version of a particular derivation program; and

adding a new build catalog entry to the build catalog, the new build catalog entry comprising an identifier of each of the one or more particular child dataset versions, and at least one identifier of one or more particular child dataset versions that were provided as input to the particular derivation program.

2. The method of claim 1 , wherein the derivation program build dependency of a version of the derived dataset corresponding to the build catalog entry comprises an identifier of a version of a derivation program executed to generate the version of the derived dataset corresponding to the build catalog entry.

3. The method of claim 2 , further comprising:

storing a first version of the derived dataset using a data lake;

updating another dataset to produce a second version of the derived dataset,

storing the second version of the derived dataset in the data lake in context of a successful transaction; and

wherein the data lake comprises a distributed file system.

4. The method of claim 3 , wherein an identifier of the first version of the derived dataset is an identifier assigned to a commit of a transaction that stored the first version of the derived dataset.

5. The method of claim 3 , wherein an identifier of the second version of the derived dataset is an identifier assigned to a commit of a transaction that stored the second version of the derived dataset.

6. The method of claim 3 , wherein the first version of the derived dataset is stored in a first set of one or more data containers.

7. The method of claim 3 , wherein the second version of the derived dataset is stored in a second set of one or more data containers.

8. The method of claim 7 , wherein the second set of one or more data containers comprises delta encodings reflecting deltas between the first version of the derived dataset and the second version of the derived dataset.

9. The method of claim 3 , wherein the first version of the derivation program is executed to produce the first version of the derived dataset.

10. The method of claim 3 , wherein the first version of the derivation program is executed to produce the second version of the derived dataset.

11. A computer system comprising:

one or more hardware processors;

one or more computer programs; and

one or more storage media storing the one or more computer programs for execution by the one or more hardware processors, the one or more computer programs comprising instructions for performing operations comprising:

maintaining a build catalog comprising a plurality of build catalog entries, each build catalog entry comprising:

an identifier of a version of a derived dataset corresponding to a build catalog entry,

one or more dataset build dependencies of the version of the derived dataset corresponding to the build catalog entry, each of the one or more dataset build dependencies comprising an identifier of a version of a child dataset from which the version of the derived dataset corresponding to the build catalog entry is derived, and

a derivation program build dependency that is executable to generate the version of the derived dataset corresponding to the build catalog entry;

creating a new version of a particular derived dataset based on providing one or more particular child dataset versions as input to executing a particular version of a particular derivation program; and

adding a new build catalog entry to the build catalog, the new build catalog entry comprising an identifier of each of the one or more particular child dataset versions, and at least one identifier of one or more particular child dataset versions that were provided as input to the particular derivation program.

12. The computer system of claim 11 , wherein the derivation program build dependency of a version of the derived dataset corresponding to the build catalog entry comprises an identifier of a version of a derivation program executed to generate the version of the derived dataset corresponding to the build catalog entry.

13. The computer system of claim 12 , wherein the one or more storage media store additional computer programs for performing operations comprising:

storing a first version of the derived dataset using a data lake;

updating another dataset to produce a second version of the derived dataset;

storing the second version of the derived dataset in the data lake in context of a successful transaction; and

wherein the data lake comprises a distributed file system.

14. The computer system of claim 13 , wherein an identifier of the first version of the derived dataset is an identifier assigned to a commit of a transaction that stored the first version of the derived dataset.

15. The computer system of claim 13 , wherein an identifier of the second version of the derived dataset is an identifier assigned to a commit of a transaction that stored the second version of the derived dataset.

16. The computer system of claim 13 , wherein the first version of the derived dataset is stored in a first set of one or more data containers.

17. The computer system of claim 13 , wherein the second version of the derived dataset is stored in a second set of one or more data containers.

18. The computer system of claim 17 , wherein the second set of one or more data containers comprises delta encodings reflecting deltas between the first version of the derived dataset and the second version of the derived dataset.

19. The computer system of claim 13 , wherein the first version of the derivation program is executed to produce the first version of the derived dataset.

20. The computer system of claim 13 , wherein the first version of the derivation program is executed to produce the second version of the derived dataset.

Assignments (7)
SECURITY INTEREST Recorded Jul 3, 2022
From: PALANTIR TECHNOLOGIES INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0506 →
ASSIGNMENT OF INTELLECTUAL PROPERTY SECURITY AGREEMENTS Recorded Jul 3, 2022
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0640 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ERRONEOUSLY LISTED PATENT BY REMOVING APPLICATION NO. 16/832267 FROM THE RELEASE OF SECURITY INTEREST PREVIOUSLY RECORDED ON REEL 052856 FRAME 0382. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Aug 26, 2021
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 057335/0753 →
SECURITY INTEREST Recorded Jun 4, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 052856/0817 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2020
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 052856/0382 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 051709/0471 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
Reel/Frame 051713/0149 →
Cited By (4)
US 12,306,853 US 12,321,354 US 12,602,392 US 12,613,881