IP Library › Granted Patent US 12,417,215
Granted Patent B2
US 12,417,215 · App. 18/435,216 · Granted Sep 16, 2025

Output validation of data processing systems

Inventors: Sharon Fridman (London, GB); Andrei Spatariu (London, GB)
Assignee: Palantir Technologies Inc.
G06F16/214G06F16/213G06F16/2282G06F16/244G06F16/24558G06F16/258
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,215
App. No.
18/435,216
Granted
Sep 16, 2025
Kind
B2
Abstract

A method is provided for output validation of data processing systems, performed by one or more processors. The method comprises performing a data comparison between a first data table and a second data table to determine a data differentiating table, wherein the first data table is based on an output of a first data pipeline, and wherein the second data table is based on an output of a second data pipeline; performing a schema comparison between the first data table and the second data table to determine a schema differentiating table; generating a first output validation score based on the data differentiating table; generating a second output validation score based on the schema differentiating table; and generating a summary comprising both the first and second output validation scores.

Claims (42)

1. A computer-implemented method comprising:

generating a first aggregated data table based on an output of a first data pipeline, wherein generating the first aggregated data table includes performing aggregations on the output of the first data pipeline to determine aggregation values, wherein performing aggregations on the output of the first data pipeline includes calculating first hash values for respective columns or rows of the output of the first data pipeline, and wherein the aggregation values of the first aggregated data table include the calculated first hash values;

generating a second aggregated data table based on an output of a second data pipeline, wherein generating the second aggregated data table performing aggregations on the output of the second data pipeline to determine aggregation values, wherein performing aggregations on the output of the second data pipeline includes calculating second hash values for respective columns or rows of the output of the second data pipeline, and wherein the aggregation values of the second aggregated data table include the calculated second hash values;

performing a data comparison between the first aggregated data table and the second aggregated data table to determine a data differentiating table, wherein the data comparison includes comparing respective pairs of the calculated first hash values and the calculated second hash values;

performing a schema comparison between the first aggregated data table and the second aggregated data table to determine a schema differentiating table;

generating a first output validation score based on the data differentiating table;

generating a second output validation score based on the schema differentiating table; and

generating a summary comprising both the first and second output validation scores.

2. The computer-implemented method of claim 1 , wherein the first and second aggregated data tables comprise aggregations of the outputs of the respective first and second data pipelines.

3. The computer-implemented method of claim 1 , wherein the aggregation values of the first and second aggregated data tables are determined by at least one of: summing, averaging, determining a median, determining a minimum, determining a maximum, determining a variance, determining a kurtosis, determining a standard deviation for numeric values, calculating a hash value for concatenated string values, calculating hash values for numeric values, or using a histogram of characters in a string value.

4. The computer-implemented method of claim 1 , wherein the first and second aggregated data tables include aggregation values in each of the rows of the first and second aggregated data tables, and wherein generating the first and second aggregated data tables includes performing aggregations column-wise or row-wise on the outputs of the respective first and second data pipelines to obtain one aggregation value per column or row, respectively, of the outputs of the respective first and second data pipelines.

5. The computer-implemented method of claim 1 , wherein the data differentiating table includes indications of, for each of the columns of the first and second aggregated data tables, a plurality of data comparison characteristics.

6. The computer-implemented method of claim 1 , wherein the schema differentiating table includes indications of, for each of the columns of the first and second aggregated data tables, a plurality of schema comparison characteristics.

7. The computer-implemented method of claim 1 , wherein the first data pipeline is a legacy data pipeline and the second data pipeline is a target data pipeline that is intended to replace the first data pipeline, and wherein the first aggregated data table is transferred to a second data processing system associated with the second data pipeline, and wherein the data comparison and the schema comparison are performed at the second data processing system.

8. The computer-implemented method of claim 1 , wherein the schema differentiating table includes indications of at least: names of each column in the first aggregated data table, any columns present in the first aggregated data table but not in the second aggregated data table, and any columns with a type mismatch between the first and second aggregated data tables.

9. The computer-implemented method of claim 8 , wherein the schema comparison is based on a comparison of datatypes and/or names of one or more columns of the first aggregated data table and the second aggregated data table.

10. The computer-implemented method of claim 1 , wherein the generating the summary comprises using weights to obtain a use case aware summary, and wherein the summary indicates a similarity between the output of the first data pipeline and the output of the second data pipeline.

11. The computer-implemented method of claim 10 , wherein at least one of: a user can determine how columns and/or rows are to be weighted; or weights of columns and/or rows are automatically determined based on user input in one or more dashboards.

12. The computer-implemented method of claim 1 further comprising:

causing display of the summary as a summary table.

13. A system comprising one or more processors and memory storing instructions that, when executed by the one or more processors, cause the system to:

generate a first aggregated data table based on an output of a first data pipeline, wherein generating the first aggregated data table includes performing aggregations on the output of the first data pipeline to determine aggregation values, wherein performing aggregations on the output of the first data pipeline includes calculating first hash values for respective columns or rows of the output of the first data pipeline, and wherein the aggregation values of the first aggregated data table include the calculated first hash values;

generate a second aggregated data table based on an output of a second data pipeline, wherein generating the second aggregated data table performing aggregations on the output of the second data pipeline to determine aggregation values, wherein performing aggregations on the output of the second data pipeline includes calculating second hash values for respective columns or rows of the output of the second data pipeline, and wherein the aggregation values of the second aggregated data table include the calculated second hash values;

perform a data comparison between the first aggregated data table and the second aggregated data table to determine a data differentiating table, wherein the data comparison includes comparing respective pairs of the calculated first hash values and the calculated second hash values;

perform a schema comparison between the first aggregated data table and the second aggregated data table to determine a schema differentiating table;

generate a first output validation score based on the data differentiating table;

generate a second output validation score based on the schema differentiating table; and

generate a summary comprising both the first and second output validation scores.

14. The system of claim 13 , wherein the first and second aggregated data tables comprise aggregations of the outputs of the respective first and second data pipelines.

15. The system of claim 13 , wherein the aggregation values of the first and second aggregated data tables are determined by at least one of: summing, averaging, determining a median, determining a minimum, determining a maximum, determining a variance, determining a kurtosis, determining a standard deviation for numeric values, calculating a hash value for concatenated string values, calculating hash values for numeric values, or using a histogram of characters in a string value.

16. The system of claim 13 , wherein the first and second aggregated data tables include aggregation values in each of the rows of the first and second aggregated data tables, and wherein generating the first and second aggregated data tables includes performing aggregations column-wise or row-wise on the outputs of the respective first and second data pipelines to obtain one aggregation value per column or row, respectively, of the outputs of the respective first and second data pipelines.

17. The system of claim 13 , wherein the data differentiating table includes indications of, for each of the columns of the first and second aggregated data tables, a plurality of data comparison characteristics.

18. The system of claim 13 , wherein the schema differentiating table includes indications of, for each of the columns of the first and second aggregated data tables, a plurality of schema comparison characteristics.

19. The system of claim 13 , wherein the first data pipeline is a legacy data pipeline and the second data pipeline is a target data pipeline that is intended to replace the first data pipeline, and wherein the first aggregated data table is transferred to a second data processing system associated with the second data pipeline, and wherein the data comparison and the schema comparison are performed at the second data processing system.

20. One or more non-transitory computer-readable mediums comprising computer executable instructions stored thereon which, when executed by one or more processors, cause the one or more processors to:

generate a first aggregated data table based on an output of a first data pipeline, wherein generating the first aggregated data table includes performing aggregations on the output of the first data pipeline to determine aggregation values, wherein performing aggregations on the output of the first data pipeline includes calculating first hash values for respective columns or rows of the output of the first data pipeline, and wherein the aggregation values of the first aggregated data table include the calculated first hash values;

generate a second aggregated data table based on an output of a second data pipeline, wherein generating the second aggregated data table performing aggregations on the output of the second data pipeline to determine aggregation values, wherein performing aggregations on the output of the second data pipeline includes calculating second hash values for respective columns or rows of the output of the second data pipeline, and wherein the aggregation values of the second aggregated data table include the calculated second hash values;

perform a data comparison between the first aggregated data table and the second aggregated data table to determine a data differentiating table, wherein the data comparison includes comparing respective pairs of the calculated first hash values and the calculated second hash values;

perform a schema comparison between the first aggregated data table and the second aggregated data table to determine a schema differentiating table;

generate a first output validation score based on the data differentiating table;

generate a second output validation score based on the schema differentiating table; and

generate a summary comprising both the first and second output validation scores.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2024
From: FRIDMAN, SHARON; SPATARIU, ANDREI
To: PALANTIR TECHNOLOGIES INC.
Reel/Frame 066529/0938 →
Priority Claims (1)
GB 2012816 · Aug 17, 2020 · national
Continuity (3)
Continuation 18058568 · Nov 23, 2022
Continuation 17064947 · Oct 7, 2020
Related Publication 20240211451A1 · Jun 27, 2024
References Cited (31)
US 6233578B1 · Machihara · 2001 [cited by examiner]
US 6278452B1 · Huberman et al. · 2001 [cited by applicant]
US 8037056B2 · Naicken et al. · 2011 [cited by applicant]
US 10484326B2 · Norwood · 2019 [cited by examiner]
US 11550764B2 · Fridman et al. · 2023 [cited by applicant]
US 11934363B2 · Fridman et al. · 2024 [cited by applicant]
US 20080162509A1 · Becker · 2008 [cited by applicant]
US 20090164496A1 · Carnathan · 2009 [cited by examiner]
US 20100121813A1 · Cui et al. · 2010 [cited by applicant]
US 20140310231A1 · Sampathkumaran et al. · 2014 [cited by applicant]
US 20140372374A1 · Bourbonnais · 2014 [cited by examiner]
US 20150293968A1 · Dickie · 2015 [cited by examiner]
US 20160063214A1 · Blue · 2016 [cited by applicant]
US 20160275150A1 · Brurnonnais et al. · 2016 [cited by applicant]
US 20170061027A1 · Chesla et al. · 2017 [cited by applicant]
US 20180307715A1 · Anand · 2018 [cited by examiner]
US 20190205429A1 · Lee · 2019 [cited by applicant]
US 20190332697A1 · Williams et al. · 2019 [cited by applicant]
US 20190377713A1 · Lankford et al. · 2019 [cited by applicant]
US 20200133257A1 · Cella · 2020 [cited by examiner]
US 20210149896A1 · Yu et al. · 2021 [cited by applicant]
US 20210397972A1 · Walters · 2021 [cited by examiner]
US 20230098701A1 · Fridman et al. · 2023 [cited by applicant]
EP 3958140 · 2022 [cited by applicant]
WO WO2019035903 · 2019 [cited by applicant]
Oracle® Database Database Administrator's Guide 12c Release 2 (12.2) E85760-09 (Year: 2020). [cited by examiner]
Oracle® Database Data Warehousing Guide 10g Release 1 (10.1) Part No. B10736-01 Dec. 2003 (Year: 2003). [cited by examiner]
U.S. Appl. No. 18/058,568, Output Validation of Data Processing Systems, filed Nov. 23, 2022. [cited by applicant]
OraRep: “Oracle7™ Server Distributed Systems”, vol. II: Replicated Data, Release 7.3, primary author: Maria Pratt, Feb. 1996, Oracle® Corporation (1996). [cited by applicant]
SQLCompTool: IDERA—“SQL Comparison Toolset Version 10.0” (2019). [cited by applicant]
Official Communication for European Patent Application No. 20200605.2 dated Mar. 25, 2021, 8 pages. [cited by applicant]