IP Library › Granted Patent US 12,748,736
Granted Patent B2
US 12,748,736 · App. 18/754,854 · Granted Sep 29, 2026

Automated data observability system

Inventors: Sanjay Agrawal (Sammamish, WA); Shashank Gupta (Sammamish, WA); Pramod Kalipatnapu (Sammamich, WA)
Assignee: Revefi, Inc.
G06F16/215G06F11/079G06F16/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,736
App. No.
18/754,854
Granted
Sep 29, 2026
Kind
B2
Abstract

An automated data observation system, platform, and corresponding methods and non-transitory computer readable medium include periodically obtaining metadata, automatically generating data quality metrics, obtaining predicted values for the data quality metrics, determining a first anomaly, generating a data pipeline graph, and estimating a root cause relating to the first anomaly by traversing the data pipeline graph. Obtaining predicted values for the data quality metrics includes executing a plurality of candidate machine learning models over time series data relating to the data quality metrics to generate candidate predicted values and selecting a machine learning model for a given data quality metric based on the candidate predicted values and measured values.

Claims (57)

1 . An automated data observation system, comprising a processor and a memory storing instructions that when executed by the processor cause the processor to:

periodically obtain metadata relating to respective nodes of a plurality of nodes;

automatically generate data quality metrics for respective nodes of the plurality of nodes based on the metadata;

obtain predicted values for the data quality metrics including by:

executing a plurality of candidate machine learning models over time series data including measured values corresponding to respective ones of the data quality metrics to generate candidate predicted values for respective ones of the data quality metrics,

selecting a selected machine learning model from the plurality of candidate machine learning models for respective ones of the data quality metrics based on a comparison between at least some of the measured values and at least some of the predicted values, and

obtaining the predicted values based on the selected machine learning model for respective ones of the data quality metrics;

determine a first anomaly relating to a first one of the data quality metrics for a first node of the plurality of nodes including by comparing a first measured value at a first time relating to the first one of the data quality metrics and a first predicted value at the first time relating to the first one of the data quality metrics;

generate a data pipeline graph corresponding to the at least some of the plurality of nodes including the first node by parsing logs including information relating to usage of at least some of the plurality of nodes to identify relationships between at least some of the plurality of nodes, including creating a directional edge from a first table to a second table based on an identification of a pattern of inserts into the second table based on data selected from the first table, wherein the data pipeline graph includes information indicative of an importance of the directional edge based on the pattern of inserts;

estimate a root cause relating to the first anomaly by traversing the data pipeline graph from the first node;

determine a priority of the estimated root cause based on information indicative of the importance of directional edges between a second node associated with the root cause and the first node, wherein the priority is further determined based on an identity of users accessing the first node, second node, or nodes between the first node and second node and a frequency of users accessing the first node, second node, or nodes between the first node and second node; and

transmit, to a client device, information relating to the estimated root cause for display by the client device based on the priority,

wherein the nodes of the plurality of nodes are tables in a data warehouse system, the logs are query logs including queries executed by the data warehouse system.

2 . The system of claim 1 , wherein the plurality of nodes includes all tables in a data warehouse system, data quality metrics are generated for all of the plurality of nodes, and predicted values are obtained for all of the data quality metrics.

3 . The system of claim 1 , wherein the instructions to estimate the root cause includes comparing the information indicative of the importance to a threshold.

4 . The system of claim 1 , wherein the memory further comprises instructions that when executed by the processor cause the processor to:

detect that the first node has been deleted and that a second node has been created to replace the first node; and

in response to the detection, associate data quality metrics and time series data of the first node with the second node.

5 . The system of claim 1 , wherein the memory further comprises instructions that when executed by the processor cause the processor to:

detect that the first node has not been updated for a time period exceeding a time period threshold; and

based on the detection, suspend obtaining predicted values for data quality metrics associated with the first node.

6 . The system of claim 1 , wherein the data quality metrics include a last modification metric, a quantity of updates per time period metric, a volume of data per time period metric, and a node size metric.

7 . The system of claim 1 , wherein a subset of the candidate machine learning models of the plurality of candidate machine learning models are removed over time by eliminating candidate machine learning models that are measured to have a higher error over time.

8 . The system of claim 7 , wherein the plurality of candidate machine learning models includes a minimum of three candidate machine learning models.

9 . The system of claim 7 , wherein, after a reset time period has elapsed, candidate machine learning models are re-added to the plurality of candidate machine learning models.

10 . A method comprising:

periodically obtaining metadata relating to respective nodes of a plurality of nodes;

automatically generating data quality metrics for respective nodes of the plurality of nodes based on the metadata;

obtaining predicted values for the data quality metrics including by:

executing a plurality of candidate machine learning models over time series data including measured values corresponding to respective ones of the data quality metrics to generate candidate predicted values for respective ones of the data quality metrics,

selecting a selected machine learning model from the plurality of candidate machine learning models for respective ones of the data quality metrics based on a comparison between at least some of the measured values and at least some of the predicted values, and

obtaining the predicted values based on the selected machine learning model for respective ones of the data quality metrics;

determining a first anomaly relating to a first one of the data quality metrics for a first node of the plurality of nodes including by comparing a first measured value at a first time relating to the first one of the data quality metrics and a first predicted value at the first time relating to the first one of the data quality metrics;

generating a data pipeline graph corresponding to the at least some of the plurality of nodes including the first node by parsing logs including information relating to usage of at least some of the plurality of nodes to identify relationships between at least some of the plurality of nodes, including creating a directional edge from a first table to a second table based on an identification of a pattern of inserts into the second table based on data selected from the first table, wherein the data pipeline graph includes information indicative of an importance of the directional edge based on the pattern of inserts;

estimating a root cause relating to the first anomaly by traversing the data pipeline graph from the first node;

determining a priority of the estimated root cause based on information indicative of the importance of directional edges between a second node associated with the root cause and the first node, wherein the priority is further determined based on an identity of users accessing the first node, second node, or nodes between the first node and second node and a frequency of users accessing the first node, second node, or nodes between the first node and second node; and

transmitting information relating to the relating to the estimated root cause to a client device for display on the client device based on the priority,

wherein the nodes of the plurality of nodes are tables in a data warehouse system, the logs are query logs including queries executed by the data warehouse system.

11 . The method of claim 10 , further comprising:

generating a data pipeline graph corresponding to at least some of the plurality of nodes by parsing logs including information relating to usage of at least some of the plurality of nodes to identify relationships between at least some of the plurality of nodes; and

estimating a root cause relating to the first anomaly by traversing the data pipeline graph from the first node.

12 . The method of claim 11 , wherein candidate machine learning models are removed from the plurality of candidate machine learning models over time provided that the plurality of candidate machine learning models includes at least three candidate machine learning models, and wherein removed candidate machine learning models are periodically re-added to the plurality of candidate machine learning models.

13 . The method of claim 12 , wherein the data quality metrics include a last modification metric, a quantity of updates per time period metric, a volume of data per time period metric, and a node size metric, and obtaining predicted values for the data quality metrics relating to the first node is suspended based on a determination that measured values relating to the last modification metric indicates that the first node has not been modified for a time period exceeding a time period threshold.

14 . A non-transitory computer readable medium storing instructions that when executed by a processor cause the processor to:

periodically obtain metadata relating to respective nodes of a plurality of nodes;

automatically generate data quality metrics for respective nodes of the plurality of nodes based on the metadata;

obtain predicted values for the data quality metrics including by:

executing a plurality of candidate machine learning models over time series data including measured values corresponding to respective ones of the data quality metrics to generate candidate predicted values for respective ones of the data quality metrics,

selecting a selected machine learning model from the plurality of candidate machine learning models for respective ones of the data quality metrics based on a comparison between at least some of the measured values and at least some of the predicted values, and

obtaining the predicted values based on the selected machine learning model for respective ones of the data quality metrics;

determine a first anomaly relating to a first one of the data quality metrics for a first node of the plurality of nodes including by comparing a first measured value at a first time relating to the first one of the data quality metrics and a first predicted value at the first time relating to the first one of the data quality metrics;

generate a data pipeline graph corresponding to at least some of the plurality of nodes by parsing logs including information relating to usage of at least some of the plurality of nodes to identify relationships between at least some of the plurality of nodes, including creating a directional edge from a first table to a second table based on an identification of a pattern of inserts into the second table based on data selected from the first table, wherein the data pipeline graph includes information indicative of an importance of the directional edge based on the pattern of inserts;

estimate a root cause relating to the first anomaly by traversing the data pipeline graph from the first node;

determine a priority of the estimated root cause based on information indicative of the importance of directional edges between a second node associated with the root cause and the first node, wherein the priority is further determined based on an identity of users accessing the first node, second node, or nodes between the first node and second node and a frequency of users accessing the first node, second node, or nodes between the first node and second node; and

transmit, to a client device, information relating to the estimated root cause for display by the client device based on the priority,

wherein the nodes of the plurality of nodes are tables in a data warehouse system, the logs are query logs including queries executed by the data warehouse system.

15 . The non-transitory computer readable medium of claim 14 , wherein an importance of the estimated root cause is determined based on the information indicative of importance of the nodes and edges traversed when estimating the root cause.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2024
From: AGRAWAL, SANJAY; GUPTA, SHASHANK; KALIPATNAPU, PRAMOD
To: REVEFI, INC.
Reel/Frame 067858/0767 →
Continuity (2)
Provisional Application 63510340 · Jun 26, 2023
Related Publication 20240427747A1 · Dec 26, 2024
References Cited (12)
US 11451842B2 · Gershey · 2022 [cited by examiner]
US 11537627B1 · Baskaran · 2022 [cited by examiner]
US 20200104401A1 · Burnett et al. · 2020 [cited by applicant]
Marwa Elsayed; PredictDeep: Security Analytics as a Service for Anomaly Detection and Prediction; IEEE; 2020 pp. 45184-45197. [cited by examiner]
Md Salik Parwez; Big Data Analytics for User-Activity Analysis and User-Anomaly Detection in Mobile Wireless Network; IEEE; (Year: 2017). [cited by examiner]
Using Machine Learning to Detect Traffic Anomalies by Jim Goodrich; Dated 2019; Accessed Apr. 24, 2026; 34 pages; https://conf.splunk.com/files/2019/slides/FN1390.pdf. [cited by applicant]
Deep dive: Using ML to identify network traffic anomalies; Dated 2024; Accessed Apr. 24, 2026; 5 pages; https://docs.splunk.com/Documentation/MLApp/5.4.1/User/IDnetworktrafficanoms. [cited by applicant]
Anomaly Detection Using Program Control Flow Graph Mining from Execution Logs by Animesh Nandi et al.; Dated 2016; Accessed Apr. 24, 2026; 10 pages: https://netman.aiops.org/~peidan/ANM2019/7.TraceAnomalyDetection/Lectu… [cited by applicant]
A Survey of Graph-based Deep Learning for Anomaly Detection in Distributed Systems by A. Danesh Pazho et al: Dated Jun. 1, 2023; Accessed Apr. 24, 2026; 20 pages; https://arxiv.org/pdf/2206.04149. [cited by applicant]
Data-Driven Construction of Data Center Graph of Thingsfor Anomaly Detection by Hao Zhang et al.; Dated Apr. 27, 2020; Accessed Apr. 24, 2026; 17 pages; https://arxiv.org/pdf/2004.12540. [cited by applicant]
ANEMONE: Graph Anomaly Detection with Multi-Scale Contrastive Learning by Ming Jin et al.; Dated Nov. 1-5, 2021; Accessed Apr. 24, 2026; 5 pages; https://shiruipan.github.io/publication/cikm-21-jin/cikm-21-jin.pdf. [cited by applicant]
Graph Contrastive Learning for Anomaly Detection by Bo Chen et al; Dated Aug. 29, 2022; Accessed Apr. 24, 2026; 14 pages; https://arxiv.org/pdf/2108.07516. [cited by applicant]