IP Library › Granted Patent US 11,429,614
Granted Patent B2
US 11,429,614 · App. 16/824,207 · Granted Aug 30, 2022

Systems and methods for data quality monitoring

Inventor: J. Mitchell Haile (Somerville, MA)
Assignee: Data Culpa Inc.
G06F16/24558G06F16/254
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,429,614
App. No.
16/824,207
Filed
Mar 19, 2020
Granted
Aug 30, 2022
Kind
B2
Examiner
VY, HUNG T
Art Unit
2163
USPC
707/602
Abstract

Systems and methods for data quality monitoring are provided. Various embodiments include a data monitoring system that integrates into a data pipeline. The data monitoring system may receive a call from the data pipeline to analyze data inputs entering the data pipeline. The monitoring system can generate metadata describing the data inputs and compare the generated metadata with previously generated metadata to determine if the data inputs are historically consistent. The data monitoring system may return a consistency measure to the data pipeline. In further embodiments, the data monitoring system can generate metadata describing data outputs from the data pipeline and compare the output metadata to previously generated output metadata. In further embodiments, the data monitoring system may operate as a read only entity in a database. The monitoring system may monitor for changes in the database and determine when adverse changes occur in the database.

Claims (20)

1. A method to facilitate data monitoring in a data pipeline computing system, the method comprising:

receiving a call from a data pipeline to ingest unprocessed data intended for the data pipeline from one or more data input streams;

ingesting the unprocessed data from the one or more data input streams;

generating metadata using the unprocessed data, determining a value distribution of the unprocessed data, checking data types of the unprocessed data, and identifying a schema for the unprocessed data;

computing one or more expected data outputs of the data pipeline based on the metadata, the value distribution, the data types, and the data schema of the unprocessed data;

receiving a call from the data pipeline to ingest processed data generated by the data pipeline from one or more data output streams wherein the processed data comprises one or more actual data outputs;

ingesting the processed data from the one or more data output streams;

generating output metadata, determining an output value distribution, checking output data types, and identifying an output data schema for the one of more actual data outputs of the processed data;

comparing the metadata, the value distribution, the data types, and the data schema for the one or more expected data outputs with the output metadata, the output value distribution, the output data types, and the output data schema for the one or more actual data outputs;

determining that at least one of the output metadata, the output value distribution, the output data types, and the output data schema for the one or more actual data outputs does not align with the metadata, the value distribution, the data types, and the data schema for the one or more expected data outputs;

generating an alert signifying that at least one of the one or more expected data outputs does not align with the one or more actual data outputs; and

sending the alert to a client.

2. The method of claim 1 , wherein generating the alert further comprises generating a visual error report, wherein the visual error report highlights which of the one or more actual data outputs does not align with the one or more expected data outputs.

3. The method of claim 2 , wherein:

the visual error report comprises at least one or a chart, a graph, a table, a gif, or an animation; and wherein:

the visual error report further highlights which ones of the output metadata, the output value distribution, the output data types, and the output data schema for the one or more actual data outputs do not align with the metadata, the value distribution, the data types, and the data schema for the one or more expected data outputs.

4. The method of claim 1 , further comprising tracking format changes to the unprocessed data and notifying the client of the format changes to the unprocessed data.

5. The method of claim 4 , further comprising detecting, in real time, changes to object records of the unprocessed data and notifying the client of changes to the object records of the unprocessed data.

6. The method of claim 1 , further comprising generating a metadata confidence level wherein the metadata confidence level indicates at least an accuracy of the metadata.

7. The method of claim 1 , wherein the data output stream originates from an extract/transform/load (ETL) orchestrated environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2020
From: HAILE, J. MITCHELL
To: DATA CULPA INC.
Reel/Frame 052174/0686 →
Continuity (2)
Provisional Application 62978291 · Feb 10, 2020
Related Publication 20210248144A1 · Aug 12, 2021