IP Library › Granted Patent US 12,189,608
Granted Patent B2
US 12,189,608 · App. 17/414,921 · Granted Jan 7, 2025

System and method of identifying event as root cause of data quality anomaly

Inventors: Shashwat Mishra (NawabGanj, IN); Shreyash Rammohan Hisariya (Mahadevpura, IN); Mrigank Mrigank (Bengaluru, IN); Nanada Kishore Thatikonda (Bengaluru, IN)
Assignee: Visa International Service Association
G06F16/2365G06F16/215
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,189,608
App. No.
17/414,921
Granted
Jan 7, 2025
Kind
B2
Abstract

Embodiments detect and predict data disparity issues in data warehouses. Embodiments derive meaningful insights about the events occurred prior to the data disparity and correlate the events to understand the root cause of the data disparity (or the root cause of an alert generated as a result of detecting the data disparity). Embodiments either take or recommend actionable measures to prevent further occurrences of the event identified as the root cause. According to various embodiments, when the monitored data is transaction data (e.g. transaction volume, transaction amount, transaction processing speed, etc.), internal events (e.g. data job failures, job delays, job server maintenances) or external events (e.g. seasonal holiday events, natural calamities) may cause a dip or spike in the transaction data resulting in a data quality anomaly (i.e. a data disparity).

Claims (63)

1. A method for identifying a cause of data disparity among monitored data, the method comprising:

monitoring, using a server computer, parameters associated with data collected in connection with a processing computer;

detecting, using the server computer, a data disparity among the data;

identifying, using the server computer, a first event associated with the data disparity;

determining, using the server computer, a set of events associated with the first event, wherein an event among the set of events represents an occurrence that impacts at least one of an amount of data processed by, or a processing speed of, the processing computer, wherein the set of events include one or more internal events internal to an organization comprising the processing computer and one or more external events external to the organization;

calculating, using the server computer, a score for each event among the set of events as a function of a weight assigned to each event among the set of events and occurrence score determined for each event among the set of events, wherein the weight assigned to each event among the set of events is a measure of probability of that event being the cause of the data disparity, wherein the occurrence score for the event includes a sum of occurrence scores of child events of the event;

identifying, using the server computer, a second event among the set of events as the cause of the data disparity, wherein the second event has the highest score among the set of events; and

taking remedial or preventive actions to address the data disparity in view of the second event being the identified cause of the data disparity.

2. The method of claim 1 , wherein the occurrence score of a selected event is determined based on runtime characteristics of all child events and parent events of the selected event, wherein the selected event occurred prior to the all child events of the selected event, and all parent events of the selected event occurred prior to the selected event.

3. The method of claim 1 , further comprising, prior to taking the remedial or preventive actions:

identifying a third event, different than the second event, as an actual cause of the data disparity;

determining that the third event is included in the set of events;

adjusting the weight of each event among the set of events by a predetermined amount, wherein adjusting includes increasing the weight of the third event; and

recalculating the score for each event among the set of events.

4. The method of claim 3 , further comprising two or more iterations of adjusting and recalculating, wherein the score of the third event increases at each iteration such that the third event has the highest score among the set of events at conclusion of all iterations.

5. The method of claim 1 , wherein the weight of a given event is stored along with a history of the given event being an actual cause of the data disparity.

6. The method of claim 1 , further comprising, prior to taking the remedial or preventive actions:

identifying a third event, different than the second event, as an actual cause of the data disparity;

determining that the third event is not included in the set of events;

adding the third event to the set of events;

adjusting the weight of each event among the set of events by a predetermined amount, wherein adjusting includes increasing the weight of the third event; and

recalculating the score for each event among the set of events.

7. The method of claim 6 , further comprising two or more iterations of adjusting and recalculating, wherein the score of the third event increases at each iteration such that the third event has the highest score among the set of events at conclusion of all iterations.

8. The method of claim 1 , further comprising:

receiving an alert associated with the data disparity;

in response to the alert, identifying the first event associated with the data disparity.

9. The method of claim 1 , wherein the set of events associated with the first event includes one or more parent events of the first event, wherein the one or more parent events occurred prior to the first event.

10. The method of claim 1 , wherein the first event and the set of events form a dependency graph, the method further comprising:

adding a new event to the dependency graph, wherein the new event is associated with a third event and a fourth event in the set of events, wherein the fourth event is a descendent of the third event;

associating the new event with the third event without associating with the fourth event.

11. The method of claim 1 , wherein the weight assigned to the first event among the set of events has a first value for the data disparity and a second value for another data disparity.

12. A computer comprising:

a processor; and

a computer readable medium, the computer readable medium comprising code that, when executed by the processor, cause the processor to:

monitor parameters associated with data collected in connection with a processing computer;

detect a data disparity among the data;

identify a first event associated with the data disparity;

determine a set of events associated with the first event, wherein an event among the set of events represents an occurrence that impacts at least one of an amount of data processed by, or a processing speed of, the processing computer, wherein the set of events include one or more internal events internal to an organization comprising the processing computer and one or more external events external to the organization;

calculate a score for each event among the set of events as a function of a weight assigned to each event among the set of events and occurrence score determined for each event among the set of events, wherein the weight assigned to each event among the set of events is a measure of probability of that event being the cause of the data disparity, wherein the occurrence score for the event includes a sum of occurrence scores of child events of the event;

identify a second event among the set of events as the cause of the data disparity, wherein the second event has the highest score among the set of events; and

take remedial or preventive actions to address the data disparity in view of the second event being the identified cause of the data disparity.

13. The computer of claim 12 , wherein the occurrence score of a selected event is determined based on runtime characteristics of all child events and parent events of the selected event, wherein the selected event occurred prior to the all child events of the selected event, and all parent events of the selected event occurred prior to the selected event.

14. The computer of claim 12 , wherein the code, when executed by the processor, further causes the processor to:

prior to taking the remedial or preventive actions:

identify a third event, different than the second event, as an actual cause of the data disparity;

determine that the third event is included in the set of events;

adjust the weight of each event among the set of events by a predetermined amount, wherein adjusting includes increasing the weight of the third event; and

recalculate the score for each event among the set of events.

15. The computer of claim 12 , wherein the weight of a given event is stored along with a history of the given event being an actual cause of the data disparity.

16. The computer of claim 12 , wherein the code, when executed by the processor, further causes the processor to:

prior to taking the remedial or preventive actions:

identify a third event, different than the second event, as an actual cause of the data disparity;

determine that the third event is not included in the set of events;

add the third event to the set of events;

adjust the weight of each event among the set of events by a predetermined amount, wherein adjusting includes increasing the weight of the third event; and

recalculate the score for each event among the set of events.

17. The computer of claim 12 , wherein the code, when executed by the processor, further causes the processor to:

receive an alert associated with the data disparity;

in response to the alert, identify the first event associated with the data disparity.

18. The computer of claim 12 , wherein the set of events associated with the first event includes one or more parent events of the first event, wherein the one or more parent events occurred prior to the first event.

19. The computer of claim 12 , wherein the first event and the set of events form a dependency graph, and wherein the code, when executed by the processor, further causes the processor to:

add a new event to the dependency graph, wherein the new event is associated with a third event and a fourth event in the set of events, wherein the fourth event is a descendent of the third event;

associate the new event with the third event without associating with the fourth event.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2021
From: MISHRA, SHASHWAT; HISARIYA, SHREYASH RAMMOHAN; MRIGANK, MRIGANK; THATIKONDA, NANADA KISHORE
To: VISA INTERNATIONAL SERVICE ASSOCIATION
Reel/Frame 056568/0982 →
Continuity (1)
Related Publication 20220067022A1 · Mar 3, 2022
References Cited (32)
US 7111205B1 · Jahn et al. · 2006 [cited by applicant]
US 7310590B1 · Bansal · 2007 [cited by applicant]
US 20100100775A1 · Slutsman · 2010 [cited by applicant]
US 20130061095A1 · Caffrey · 2013 [cited by applicant]
US 20130097463A1 · Marvasti · 2013 [cited by applicant]
US 20130305356A1 · Cohen-Ganor · 2013 [cited by examiner]
US 20170070414A1 · Bell · 2017 [cited by examiner]
US 20170075749A1 · Ambichl et al. · 2017 [cited by applicant]
US 20170235626A1 · Zhang et al. · 2017 [cited by applicant]
US 20180039895A1 · Wang et al. · 2018 [cited by applicant]
US 20180113773A1 · Krishnan et al. · 2018 [cited by applicant]
US 20180276063A1 · Mendes · 2018 [cited by examiner]
CN 102129372A · 2011 [cited by applicant]
CN 103955505A · 2014 [cited by applicant]
CN 106294865A · 2017 [cited by applicant]
CN 108989132A · 2018 [cited by applicant]
WO 2015168071A1 · 2015 [cited by applicant]
WO 2018160177 · 2018 [cited by applicant]
Vassiliadis, Panos. “Data warehouse modeling and quality issues.” National Technical University of Athens Zographou, Athens, Greece (2000). [cited by examiner]
Singh, Ranjit, and Kawaljeet Singh. “A descriptive classification of causes of data quality problems in data warehousing.” International Journal of Computer Science Issues (IJCSI) 7.3 (2010): 41. [cited by examiner]
Application No. EP18943566.2 , Extended European Search Report, Mailed on Nov. 25, 2021, 10 pages. [cited by applicant]
Application No. PCT/US2018/066540 , International Search Report and Written Opinion, Mailed on Sep. 18, 2019, 11 pages. [cited by applicant]
Harper, et al., “The Application of Neural Networks to Predicting the Root Cause of Service Failures”, 2017, pp. 953-958. [cited by applicant]
“Alert Correlation Rules”, Service Now Docs, http:/docs.servicenow.com/bundle/kingston-it-operationns-management/page/product/eve, Aug. 31, 2018, 3 pages. [cited by applicant]
Costa, et al. “Forecasting Time Series Combining Holt-Winters and Bootstrap Approaches”, AIP Conference Proceedings, 1648, 110005, (2015), 5 pages. [cited by applicant]
Application No. EP18943566.2 , Office Action, Mailed on Feb. 23, 2023, 9 pages. [cited by applicant]
Application No. SG11202106336V, Written Opinion, Mailed on Nov. 15, 2022, 12 pages. [cited by applicant]
Application No. CN201880100130.9 , Office Action, Mailed on Jan. 4, 2024, 14 pages. [cited by applicant]
Mao et al., “M-TAEDA: Temporal Abnormal Event Detection Algorithm for Multivariate Time-series Data of Water Quality”, Journal of Computer Applications, vol. 37, No. 1, Jan. 10, 2017, pp. 138-144. [cited by applicant]
EP18943566.2, “Summons to Attend Oral Proceedings”, May 8, 2024, 10 pages. [cited by applicant]
Application No. SG11202106336V, Further Written Opinion, Mailed on Apr. 22, 2024, 8 pages. [cited by applicant]
Application No. CN201880100130.9 , Notice of Decision to Grant, Mailed on Jul. 8, 2024, 6 pages. [cited by applicant]