IP Library Granted Patent US 12,450,248
Granted Patent B2
US 12,450,248 · App. 18/308,050 · Granted Oct 21, 2025

Data correctness and validation using validation definition language

Inventors: Kalapriya Kannan (Karnataka, IN); Chirag Talreja (Karnataka, IN); Chinmay Chaturvedi (Karnataka, IN); Sagar Venkappa Nyamagouda (Karnataka, IN); Jayasankar Nallasamy (Karnataka, IN); Prasad Pimplaskar (Spring, TX)
Assignee: Hewlett Packard Enterprise Development LP
G06F16/254G06F16/215G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,248
App. No.
18/308,050
Granted
Oct 21, 2025
Kind
B2
Abstract

Systems and methods are provided for generating extract-transform-load (“ETL”) machine learning (“ML”) pipeline validation rules based on user-input, wherein the ETL ML pipeline validation rules may be applicable to validate an ETL ML pipeline against multiple test datasets. The ETL ML pipeline validation rules may comprise compute-type validation rules for computing expected values of data structures within a dataset output by the ETL ML pipeline. The ETL ML pipeline validation rules may comprise check-type validation rules for checking whether data structures within a dataset output by the ETL ML pipeline have intended characteristics.

Claims (42)

1. A method comprising:

receiving a validator rule description, wherein:

the validator rule description denotes a data field within an extract-transform-load (“ETL”) machine learning (“ML”) model as an intake data field;

the validator rule description denotes a data field within a referenced output schema as a result data field; and

based on the validator rule description, generating a validator rule comprising instructions; wherein when executed, the instructions cause a processor to:

compute a result data value based on an intake data value, wherein the intake data value is a value stored at a reappearance of the intake data field within a test dataset, and wherein a flag associated with the result data value is set, such that the result data value will not affect a later validation determination; and

write the result data value to a reappearance of the result data field within a received output schema.

2. The method of claim 1 , wherein when executing the instructions to compute the result data value, the instructions cause the processor to set a result data value to a specified value.

3. The method of claim 1 , wherein when executing the instructions to compute the result data value, the instructions cause the processor to set the result data value to the intake data value.

4. The method of claim 1 , wherein when executing the instructions to compute the result data value, the instructions cause the processor to set the result data value to a specified value if the intake data value is null.

5. The method of claim 1 , further wherein:

the validator rule description denotes a second data field within the ETL ML model as a second intake data field; and

when executing the instructions to compute the result data value, the instructions cause the processor to set the result data value to the intake data value added to a value stored at a reappearance of the second intake data field within the test dataset.

6. The method of claim 1 , further wherein:

the validator rule description denotes a second data field within the ETL ML model as a second intake data field; and

when executing the instructions to compute the result data value, the instructions cause the processor to set the result data value to the intake data value subtracted from a value stored at a reappearance of the second intake data field within the test dataset.

7. The method of claim 1 , further wherein:

the validator rule description denotes a second data field within the ETL ML model as a second intake data field; and

when executing the instructions to compute the result data value, the instructions cause the processor to set the result data value to the intake data value divided by a value stored at a reappearance of the second intake data field within the test dataset.

8. The method of claim 1 , further wherein:

the validator rule description denotes a second data field within the ETL ML model as a second intake data field; and

when executing the instructions to compute the result data value, the instructions cause the processor to set the result data value to the intake data value multiplied by a value stored at a reappearance of the second intake data field within the test dataset.

9. The method of claim 1 , wherein when executing the instructions to compute the result data value, the instructions cause the processor to set the result data value to a specified value if the intake data value satisfies a specified test.

10. The method of claim 1 , wherein when executing the instructions to compute the result data value, the instructions cause the processor to:

compute a converted intake data value, wherein the converted intake data value is the intake data value formatted according to a specified date format; and

set the result data value to the converted intake data value.

11. A method comprising:

receiving a validator rule description, wherein:

the validator rule description denotes a data field within an extract-transform-load (“ETL”) machine learning (“ML”) model as an intake data field;

the validator rule description denotes a data field within a referenced output schema as a result data field; and

based on the validator rule description, generating a validator rule comprising instructions; wherein when executed, the instructions cause a processor to:

compute a result data value based on an intake data value, wherein the intake data value is a value stored at a reappearance of the intake data field within a test dataset, and wherein the result data value is set to a specified value if the intake data value satisfies a specified test; and

write the result data value to a reappearance of the result data field within a received output schema.

12. The method of claim 11 , wherein the specified test is a logical comparison.

13. A method comprising:

receiving a validator rule description, wherein:

the validator rule description denotes a data field within an extract-transform-load (“ETL”) machine learning (“ML”) model as an intake data field;

the validator rule description denotes a data field within a referenced output schema as a result data field; and

based on the validator rule description, generating a validator rule comprising instructions; wherein when executed, the instructions cause a processor to:

compute a result data value based on an intake data value and a converted intake data value, wherein the intake data value is a value stored at a reappearance of the intake data field within a test dataset, and wherein the converted intake data value is the intake data value formatted according to a specified date format; and

set the result data value to the converted intake data value; and

write the result data value to a reappearance of the result data field within a received output schema.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2023
From: KANNAN, KALAPRIYA; TALREJA, CHIRAG; CHATURVEDI, CHINMAY; NYAMAGOUDA, SAGAR VENKAPPA; NALLASAMY, JAYASANKAR; PIMPLASKAR, PRASAD
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 063646/0243 →
Continuity (1)
Related Publication 20240362246A1 · Oct 31, 2024
References Cited (24)
US 9582541B2 · Sareen et al. · 2017 [cited by applicant]
US 9824108B2 · Taylor et al. · 2017 [cited by applicant]
US 10467039B2 · Bailey et al. · 2019 [cited by applicant]
US 20130166515A1 · Kung · 2013 [cited by examiner]
US 20190073388A1 · Desmarets · 2019 [cited by applicant]
US 20200380417A1 · Briancon et al. · 2020 [cited by applicant]
US 20210200747A1 · Aleksandrovich et al. · 2021 [cited by applicant]
US 20210232603A1 · Sundaram et al. · 2021 [cited by applicant]
US 20210232604A1 · Sundaram et al. · 2021 [cited by applicant]
US 20210303585A1 · Fan et al. · 2021 [cited by applicant]
US 20220114483A1 · Sabharwal et al. · 2022 [cited by applicant]
US 20220237101A1 · Singh et al. · 2022 [cited by applicant]
US 20220237102A1 · Bugdayci et al. · 2022 [cited by applicant]
US 20220374442A1 · Kaspa · 2022 [cited by examiner]
US 20240037077A1 · Verma et al. · 2024 [cited by applicant]
“Apache Avro™—a data serialization system”, available online at <https://avro.apache.org/>, 2022, 59 pages. [cited by applicant]
Accenture Consulting, “Model Behavior Nothing Artificial”, 2017, 20 pages. [cited by applicant]
C3.ai., “Model Validation”, available online at <https://c3.ai/glossary/data-science/model-validation/>, 2023, 4 pages. [cited by applicant]
Deepchecks, “AI Model Validation”, available online at <https://deepchecks.com/glossary/ai-model-validation/>, 2023, 7 pages. [cited by applicant]
Google, “Overview of ML Pipelines”, available online at <https://developers.google.com/machine-learning/testing-debugging/pipeline/overview>, 2022, 2 pages. [cited by applicant]
IBM, “What is data modeling?”, available online at <https://web.archive.org/web/20230224113943/https://www.ibm.com/topics/data-modeling>, Feb. 24, 2023, 9 pages. [cited by applicant]
Oreilly, “Chapter 4. Data Validation”, available online at <https://www.oreilly.com/library/view/building-machine-learning/9781492053187/ch04.html>, 2023, 27 pages. [cited by applicant]
Kristina Young, “NWU Data Science Earn your degree entirely online. Testing Your Machine Learning Pipelines”, available online at <https://web.archive.org/web/20231004235634/https://www.kdnuggets.com/2019/11/testing-mac… [cited by applicant]
Misbah Uddin, “Testing Machine Learning Pipelines”, available online at <https://towardsdatascience.com/testing-machine-learning-pipelines-22e59d7b5b56/>, Mar. 30, 2021, 11 pages. [cited by applicant]