IP Library Granted Patent US 9,229,952
Granted Patent B1
US 9,229,952 · App. 14/533,433 · Granted Jan 5, 2016

History preserving data pipeline system and method

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,229,952
App. No.
14/533,433
Granted
Jan 5, 2016
Kind
B1
Abstract

A history preserving data pipeline computer system and method. In one aspect, the history preserving data pipeline system provides immutable and versioned datasets. Because datasets are immutable and versioned, the system makes it possible to determine the data in a dataset at a point in time in the past, even if that data is no longer in the current version of the dataset.

Claims (51)

1. A method for preserving history of a derived dataset, the method comprising:

at one or more computing devices comprising one or more processors and storage media storing one or more computer programs executed by the one or more processors to perform the method, perform operations of:

storing a first version of a derived dataset;

wherein the first version of the derived dataset is derived from at least a first version of another dataset by executing a first version of derivation program associated with the derived dataset;

storing a first build catalog entry, the first build catalog entry associated with the derived dataset and comprising an identifier of the first version of the other dataset and comprising an identifier of the first version of the derivation program;

wherein the first build catalog entry comprises a name of the derived dataset and an identifier of the first version of the derived dataset;

updating the other dataset to produce a second version of the other dataset;

storing a second version of the derived dataset;

wherein the second version of the derived dataset is derived from at least the second version of the other dataset by executing the first version of the derivation program associated with the derived dataset;

storing a second build catalog entry, the second build catalog entry associated with the derived dataset and comprising an identifier of the second version of the other dataset and comprising an identifier of the first version of the derivation program; and

wherein the second build catalog entry comprises the name of the derived dataset and an identifier of the second version of the derived dataset.

2. The method of claim 1 , further comprising storing the first version of the derived dataset and the second version of the derived dataset in a data lake.

3. The method of claim 2 , wherein the data lake comprises a distributed file system.

4. The method of claim 1 , wherein the identifier of the first version of the derived dataset is an identifier assigned to a commit of a transaction that stored the first version of the derived dataset.

5. The method of claim 1 , wherein the identifier of the second version of the derived dataset is an identifier assigned to a commit of a transaction that stored the second version of the derived dataset.

6. The method of claim 1 , wherein the first version of the derived dataset is stored in a first set of one or more data containers and the second version of the derived dataset is stored in a second set of one or more data containers.

7. The method of claim 5 , wherein the second set of one or more data containers comprises delta encodings reflecting deltas between the first version of the derived dataset and the second version of the derived dataset.

8. The method of claim 1 , wherein the first version of the derivation program, when executed to produce the first version of the derived dataset, transforms data of the first version of the other dataset to produce data of the first version of the derived dataset.

9. The method of claim 1 , wherein the first version of the derivation program, when executed to produce the second version of the derived dataset, transforms data of the second version of the other dataset to produce data of the second version of the derived dataset.

10. The method of claim 1 , wherein the operations of storing the first version of the derived dataset and storing the second version of the derived dataset are performed by a data lake.

11. The method of claim 1 , wherein the operations of storing the first build catalog entry and storing the second build catalog entry are performed by a build service.

12. The method of claim 1 , wherein the operation of updating the other dataset to produce the second version of the other dataset is performed by a transaction service.

13. The method of claim 1 , wherein the first build catalog entry and the second build catalog entry are stored in a database.

14. The method of claim 1 , further comprising:

storing a transaction entry in a database comprising a transaction commit identifier of the first version of the derived dataset;

wherein the first build catalog entry comprises the transaction commit identifier.

15. The method of claim 1 , further comprising:

storing a transaction entry in a database comprising a transaction commit identifier of the second version of the derived dataset;

wherein the second build catalog entry comprises the transaction commit identifier.

16. The method of claim of 1 , further comprising:

storing a transaction entry in a database comprising a transaction commit identifier of the first version of the other dataset;

wherein the identifier of the first version of the other dataset in the first build catalog entry is the transaction commit identifier.

17. The method of claim 1 , further comprising:

storing a transaction entry in a database comprising a transaction commit identifier of the second version of the other dataset;

wherein the identifier of the second version of the other dataset in the second build catalog entry is the transaction commit identifier.

18. A history preserving data pipeline system comprising:

one or more computing devices having one or more processors and memory;

means for storing a first version of a derived dataset;

wherein the first version of the derived dataset is derived from at least a first version of another dataset by executing a first version of derivation program associated with the derived dataset;

means for storing a first build catalog entry, the first build catalog entry associated with the derived dataset and comprising an identifier of the first version of the other dataset and comprising an identifier of the first version of the derivation program;

wherein the first build catalog entry comprises a name of the derived dataset and an identifier of the first version of the derived dataset;

means for updating the other dataset to produce a second version of the other dataset; means for storing a second version of the derived dataset;

wherein the second version of the derived dataset is derived from at least the second version of the other dataset by executing the first version of the derivation program associated with the derived dataset;

means for storing a second build catalog entry, the second build catalog entry associated with the derived dataset and comprising an identifier of the second version of the other dataset and comprising an identifier of the first version of the derivation program; and

wherein the second build catalog entry comprises the name of the derived dataset and an identifier of the second version of the derived dataset.

19. A history preserving data pipeline system comprising:

one or more computing devices having one or more processors and memory;

a data lake for persistently storing a first version of a derived dataset, a second version of the derived dataset, a first version of another dataset, and a second version of the other dataset;

a build service for deriving the first version of the derived dataset from at least the first version of the other dataset by executing a first version of derivation program associated with the derived dataset, and for deriving the second version of the derived dataset from at least the second version of the other dataset by executing the first version of derivation program associated with the derived dataset;

a build database comprising a first build catalog entry and a second build catalog entry, the first build catalog entry and the second build catalog entry associated with the derived dataset, the first build catalog entry comprising a first transaction commit identifier of the first version of the other dataset and comprising an identifier of the first version of the derivation program, the second build catalog entry comprising a second transaction commit identifier of the second version of the other dataset and comprising the identifier of the first version of the derivation program; and

a transaction service for assigning the first transaction commit identifier of the first version of the other dataset to a first transaction that successfully commits the first version of the other database, for assigning the second transaction commit identifier of the second version of the other dataset to a second transaction that successfully commits the second version of the other database, for atomically creating a first entry in a transaction database responsive to successfully committing the first transaction, and for atomically creating a second entry in the transaction database responsive to successfully committing the second transaction, the first entry comprising the first transaction commit identifier, the second entry comprising the second transaction commit identifier.

Assignments (8)
ASSIGNMENT OF INTELLECTUAL PROPERTY SECURITY AGREEMENTS Recorded Jul 3, 2022
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0640 →
SECURITY INTEREST Recorded Jul 3, 2022
From: PALANTIR TECHNOLOGIES INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0506 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ERRONEOUSLY LISTED PATENT BY REMOVING APPLICATION NO. 16/832267 FROM THE RELEASE OF SECURITY INTEREST PREVIOUSLY RECORDED ON REEL 052856 FRAME 0382. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Aug 26, 2021
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 057335/0753 →
SECURITY INTEREST Recorded Jun 4, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 052856/0817 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2020
From: ROYAL BANK OF CANADA
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 052856/0382 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS ADMINISTRATIVE AGENT
Reel/Frame 051713/0149 →
SECURITY INTEREST Recorded Jan 27, 2020
From: PALANTIR TECHNOLOGIES INC.
To: ROYAL BANK OF CANADA, AS ADMINISTRATIVE AGENT
Reel/Frame 051709/0471 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2014
From: MEACHAM, JACOB; HARRIS, MICHAEL; BRODMAN, GUSTAV; CUTHRIELL, LYNN; KORUS, HANNAH; TOTH, BRIAN; HSIAO, JONATHAN; ELLIOT, MARK; SCHIMPF, BRIAN; GARLAND, MICHAEL; NGUYEN, EVELYN
To: PALANTIR TECHNOLOGIES, INC.
Reel/Frame 034168/0229 →