IP Library Granted Patent US 10,970,189
Granted Patent B2
US 10,970,189 · App. 16/362,129 · Granted Apr 6, 2021

Configuring data processing pipelines

Inventors: Saurabh Shukla (London, GB); Subbanarasimhiah Harish (London, GB); Harsh Pandey (New York, NY); Thomas Boam (Hertfordshire, GB); Vinoo Ganesh (New York, NY)
Assignee: Palantir Technologies Inc.
G06F11/3452G06F11/302G06F16/2456G06F16/258
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,970,189
App. No.
16/362,129
Granted
Apr 6, 2021
Kind
B2
Abstract

Systems and methods are provided that are useful for configuring data processing pipelines. During building of a dataset in a data processing pipeline, statistics can be calculated relating to the dataset.

Claims (40)

1. A method performed by one or more processors of a data processing pipeline, the method comprising:

receiving a source dataset from a source file storage system;

compiling a dataset from the source dataset, wherein compiling the dataset from the source dataset comprises:

applying one or more transformations to each row of the source dataset, wherein at least one transformation includes an encryption to a row of the source dataset;

determining, based on the one or more transformations to each row of the source dataset, statistics for each corresponding row of the dataset as the dataset being compiled; and

storing the dataset and the statistics for each row of the dataset to a database; and

during compilation of the dataset in the data processing pipeline, determining statistics relating to the dataset.

2. The method of claim 1 , further comprising updating the statistics relating to the dataset as each row of the dataset is being compiled.

3. The method of claim 1 , further comprising updating the statistics for each row of the dataset to a least one cell of a row of the dataset as the row of the dataset is being compiled.

4. The method of claim 1 , further comprising configuring, based on the statistics relating to the dataset, one or more downstream transformations associated with the data processing pipeline.

5. The method of claim 4 , further comprising optimizing, based on the statistics relating to the dataset, a join operation associated with the data processing pipeline.

6. The method of claim 1 , further comprising presenting the statistics relating to the dataset in a user interface, and providing a control to allow a user to configure the data processing pipeline based on the statistics.

7. The method of claim 1 , further comprising presenting the statistics relating to the dataset in a user interface, and providing a control to allow a user to cease further processing of the dataset by the data processing pipeline.

8. The method of claim 1 , further comprising generating a statistics configuration, and determining, based on the statistics configuration, the statistics relating to the dataset as the dataset is being compiled.

9. The method of claim 8 , further comprising optimizing the one or more transformations based on the statistics configuration.

10. The method of claim 8 , further comprising generating the statistics configuration based on a configuration input received through a user interface.

11. The method of claim 1 , further comprising changing an execution order of data processing stages associated with the data processing pipeline based on the statistics relating to the dataset, wherein changing the executing order of the data processing stages comprises rewriting user code based on the statistics relating to the dataset.

12. A system configured for a data processing pipeline, the system comprising:

one or more processors; and

a memory storing instructions that, when executed by the one or more processors, cause the system to perform:

receiving a source dataset from a source file storage system;

compiling a dataset from the source dataset, wherein compiling the dataset from the source dataset comprises:

applying one or more transformations to each row of the source dataset, wherein at least one transformation includes an encryption to a row of the source dataset:

determining, based on the one or more transformations to each row of the source dataset, statistics for each corresponding row of the dataset as the dataset being compiled; and

storing the dataset and the statistics for each row of the dataset to a database; and

during compilation of the dataset in the data processing pipeline, determining statistics relating to the dataset.

13. The system of claim 12 , wherein the instructions, when executed, further cause the system to perform updating the statistics relating to the dataset as each row of the dataset is being compiled.

14. The system of claim 12 , wherein the instructions, when executed, further cause the system to perform updating the statistics for each row of the dataset to a least one cell of a row of the dataset as the row of the dataset is being compiled.

15. The system of claim 12 , wherein the instructions, when executed, further cause the system to perform configuring based on the statistics relating to the dataset, one or more downstream transformations associated with the data processing pipeline.

16. The system of claim 15 , wherein the instructions, when executed, further cause the system to perform optimizing, based on the statistics relating to the dataset, a join operation associated with the data processing pipeline.

17. The system of claim 12 , wherein the instructions, when executed, further cause the system to perform presenting the statistics relating to the dataset in a user interface, and providing a control to allow a user to configure the data processing pipeline based on the statistics.

18. A non-transitory computer readable medium of a computing system configured for a data processing pipeline storing instructions that, when executed by one or more processors the computing system, cause the computing system to perform:

receiving a source dataset from a source file storage system;

compiling a dataset from the source dataset, wherein compiling the dataset from the source dataset comprises:

applying one or more transformations to each row of the source dataset, wherein at least one transformation includes an encryption to a row of the source dataset;

determining, based on the one or more transformations to each row of the source dataset, statistics for each corresponding row of the dataset as the dataset being compiled; and

storing the dataset and the statistics for each row of the dataset to a database; and

during compilation of the dataset in the data processing pipeline, determining statistics relating to the dataset.

19. The non-transitory computer readable medium of claim 18 , wherein the instructions, when executed, further cause the computing system to perform updating the statistics relating to the dataset as each row of the dataset is being compiled.

20. The non-transitory computer readable medium of claim 18 , wherein the instructions, when executed, further cause the computing system to perform updating the statistics for each row of the dataset to a least one cell of a row of the dataset as the row of the dataset is being compiled.

Assignments (2)
SECURITY INTEREST Recorded Jul 3, 2022
From: PALANTIR TECHNOLOGIES INC.
To: WELLS FARGO BANK, N.A.
Reel/Frame 060572/0506 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2019
From: SHUKLA, SAURABH; HARISH, SUBBANARASIMHIAH; PANDY, HARSH; BOAM, THOMAS; GANESH, VINOO
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 048983/0343 →