IP Library › Granted Patent US 12,367,088
Granted Patent B2
US 12,367,088 · App. 18/533,516 · Granted Jul 22, 2025

System and method for automated data pipeline degradation detection

Inventors: Stephanie Margaret Pirman (Chicago, IL); Jeffrey Wayne Texada (Carrollton, TX); Eric Joseph DePree (Evanston, IL)
Assignee: Bank of America Corporation
G06F11/0757G06F11/0709G06F11/0727
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,088
App. No.
18/533,516
Filed
Dec 8, 2023
Granted
Jul 22, 2025
Kind
B2
Art Unit
2114
USPC
714/47.2
Abstract

A method analyzes dependency information for a first data store. Upon determining that a data pipeline associates a first log type generated by the first data store with a second log type generated by a second data store, a first number of logs of the first log type that are generated at a first time, a first baseline number, and a first threshold value are determined. Upon determining that the first number of logs differs from the first baseline number by less than the first threshold value, a second number of logs of the second log type that are generated at a second time, a second baseline number, and a second threshold value are determined. Upon determining that the second number of logs differs from the second baseline number by more than the second threshold value, the data pipeline is identified as degraded.

Claims (169)

1. A system comprising:

a memory configured to store:

a first database comprising:

a first profile for a first data store of a plurality of data stores, wherein the first profile comprises:

a first plurality of log types that are generated by the first data store;

a first plurality of impact scores associated with the first plurality of log types;

first dependency information for the first data store, wherein the first dependency information identifies a first data pipeline that associates logs of a first log type generated by the first data store with logs of a second log type generated by a second data store of the plurality of data stores; and

a first time delay associated with the first data pipeline; and

a second profile for a second data store of the plurality of data stores, wherein the second profile comprises:

a second plurality of log types that are generated by the second data store;

a second plurality of impact scores associated with the second plurality of log types; and

second dependency information for the second data store, wherein the second dependency information identifies that no data store receives data items from the second data store; and

a second database comprising:

a first plurality of numbers;

a first plurality of timestamps associated with the first plurality of numbers, wherein each of the first plurality of numbers is a number of logs of the first log type that were generated by the first data store at a time defined by a respective timestamp for a first time interval;

a second plurality of numbers; and

a second plurality of timestamps associated with the second plurality of numbers, wherein each of the second plurality of numbers is a number of logs of the second log type that were generated by the second data store at a time defined by a respective timestamp for the first time interval; and

a processor operably coupled to the memory and configured to:

analyze the first dependency information for the first data store; and

in response to determining that the first data pipeline associates the first log type generated by the first data store with the second log type generated by the second data store:

determine a first number of logs of the first log type that are generated by the first data store at a first time for the first time interval;

determine a first baseline number and a first threshold value based on the first plurality of numbers, wherein the first baseline number is an expected number of logs of the first log type generated by the first data store at the first time for the first time interval;

compare the first number of logs of the first log type to the first baseline number; and

in response to determining that the first number of logs differs from the first baseline number by less than the first threshold value:

determine that a second number of logs of the second log type are generated by the second data store at a second time for the first time interval, wherein the second time is later than the first time by the first time delay;

determine a second baseline number and a second threshold value based on the second plurality of numbers, wherein the second baseline number is an expected number of logs of the second log type generated by the second data store at the second time for the first time interval;

compare the second number of logs of the second log type to the second baseline number; and

in response to determining that the second number of logs of the second log type differs from the second baseline number by more than the second threshold value:

 identify the first data pipeline as degraded;

 identify the second data store as degraded;

 identify the second log type as degraded;

 determine a second impact score of the second log type;

analyze the second dependency information for the second data store; and

 in response to determining that no data store receives data items from the second data store, generate a report, wherein the report comprises:

 an identification that the first data pipeline is degraded;

 an identification that the second data store is degraded;

 an identification that the second log type is degraded; and

 the second impact score of the second log type.

2. The system of claim 1 , wherein the processor is further configured to, in response to determining that the second number of logs of the second log type does not differ from the second baseline number by more than the second threshold value, identify the first data pipeline as not degraded.

3. The system of claim 1 , wherein:

the first dependency information further identifies a second data pipeline that associates logs of the first log type generated by the first data store with logs of a third log type generated by a third data store of the plurality of data stores;

the first profile of the first data store further comprises a second time delay corresponding to the second data pipeline;

the first database further comprises a third profile for the third data store, wherein the third profile comprises:

a third plurality of log types that are generated by the third data store;

a third plurality of impact scores associated with the third plurality of log types; and

third dependency information for the third data store, wherein the third dependency information identifies that no data store receives data items from the third data store;

the second database further comprises:

a third plurality of numbers; and

a third plurality of timestamps associated with the third plurality of numbers, wherein each of the third plurality of numbers is a number of logs of the third log type that were generated by the third data store at a time defined by a respective timestamp for the first time interval; and

the processor is further configured to, in response to determining that the second data pipeline associates the logs of the first log type generated by the first data store with the logs of the third log type generated by the third data store:

determine that a third number of logs of the third log type are generated by the third data store at a third time for the first time interval, wherein the third time is later than the first time by the second time delay;

determine a third baseline number and a third threshold value based on the third plurality of numbers, wherein the third baseline number is an expected number of logs of the third log type generated by the third data store at the third time for the first time interval; compare the third number of logs of the third log type to the third baseline number; and

in response to determining that the third number of logs of the third log type differs from the third baseline number by more than the third threshold value:

identify the second data pipeline as degraded;

identify the third data store as degraded;

identify the third log type as degraded;

determine a third impact score of the third log type;

analyze the third dependency information for the third data store; and

in response to determining that no data store receives data items from the third data store, update the report, wherein the report further comprises:

an identification that the second data pipeline is degraded;

an identification that the third data store is degraded;

an identification that the third log type is degraded; and

the third impact score of the third log type.

4. The system of claim 3 , wherein the processor is further configured to, in response to determining that the third number of logs of the third log type does not differ from the third baseline number by more than the third threshold value, identify the second data pipeline as not degraded.

5. The system of claim 3 , wherein the processor is further configured to:

determine a priority of the report based at least in part on the second impact score and the third impact score; and

in response to determining that the priority is “high”:

generate an alert that an immediate response to the report is needed;

send the report to a data store maintenance team; and

send the alert to the data store maintenance team.

6. The system of claim 3 , wherein the processor is further configured to:

determine a priority of the report based at least in part on the second impact score and the third impact score; and

in response to determining that the priority is “low,” send the report to a data store maintenance team.

7. The system of claim 3 , wherein the first time delay is different from the second time delay.

8. A method comprising:

analyzing first dependency information for a first data store; and

in response to determining that a first data pipeline associates a first log type generated by the first data store with a second log type generated by a second data store:

determining a first number of logs of the first log type that are generated by the first data store at a first time for a first time interval;

determining a first baseline number and a first threshold value based on a first plurality of numbers, wherein each of the first plurality of numbers is a number of logs of the first log type that were generated by the first data store at a time defined by a respective timestamp for the first time interval, and wherein the first baseline number is an expected number of logs of the first log type generated by the first data store at the first time for the first time interval;

comparing the first number of logs of the first log type to the first baseline number; and

in response to determining that the first number of logs differs from the first baseline number by less than the first threshold value:

determining that a second number of logs of the second log type are generated by the second data store at a second time for the first time interval, wherein the second time is later than the first time by a first time delay;

determining a second baseline number and a second threshold value based on a second plurality of numbers, wherein each of the second plurality of numbers is a number of logs of the second log type that were generated by the second data store at a time defined by a respective timestamp for the first time interval, wherein the second baseline number is an expected number of logs of the second log type generated by the second data store at the second time for the first time interval;

comparing the second number of logs of the second log type to the second baseline number; and

in response to determining that the second number of logs of the second log type differs from the second baseline number by more than the second threshold value:

identifying the first data pipeline as degraded;

identifying the second data store as degraded;

identifying the second log type as degraded;

determining a second impact score of the second log type; and

analyzing second dependency information for the second data store; and

in response to determining that no data store receives data items from the second data store, generating a report, wherein the report comprises:

 an identification that the first data pipeline is degraded;

 an identification that the second data store is degraded;

 an identification that the second log type is degraded; and

 the second impact score of the second log type.

9. The method of claim 8 , further comprising, in response to determining that the second number of logs of the second log type does not differ from the second baseline number by more than the second threshold value, identifying the first data pipeline as not degraded.

10. The method of claim 8 , further comprising, in response to determining that a second data pipeline associates the logs of the first log type generated by the first data store with logs of a third log type generated by a third data store:

determining that a third number of logs of the third log type are generated by the third data store at a third time for the first time interval, wherein the third time is later than the first time by a second time delay;

determining a third baseline number and a third threshold value based on a third plurality of numbers, wherein each of the third plurality of numbers is a number of logs of the third log type that were generated by the third data store at a time defined by a respective timestamp for the first time interval, and wherein the third baseline number is an expected number of logs of the third log type generated by the third data store at the third time for the first time interval;

comparing the third number of logs of the third log type to the third baseline number; and

in response to determining that the third number of logs of the third log type differs from the third baseline number by more than the third threshold value:

identifying the second data pipeline as degraded;

identifying the third data store as degraded;

identifying the third log type as degraded;

determining a third impact score of the third log type;

analyzing third dependency information for the third data store; and

in response to determining that no data store receives data items from the third data store, updating the report, wherein the report further comprises:

an identification that the second data pipeline is degraded;

an identification that the third data store is degraded;

an identification that the third log type is degraded; and

the third impact score of the third log type.

11. The method of claim 10 , further comprising, in response to determining that the third number of logs of the third log type does not differ from the third baseline number by more than the third threshold value, identifying the second data pipeline as not degraded.

12. The method of claim 10 , further comprising:

determining a priority of the report based at least in part on the second impact score and the third impact score; and

in response to determining that the priority is “high”:

generating an alert that an immediate response to the report is needed;

sending the report to a data store maintenance team; and

sending the alert to the data store maintenance team.

13. The method of claim 10 , further comprising:

determining a priority of the report based at least in part on the second impact score and the third impact score; and

in response to determining that the priority is “low,” sending the report to a data store maintenance team.

14. The method of claim 10 , wherein the first time delay is different from the second time delay.

15. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to:

analyze first dependency information for a first data store; and

in response to determining that a first data pipeline associates a first log type generated by the first data store with a second log type generated by a second data store:

determine a first number of logs of the first log type that are generated by the first data store at a first time for a first time interval;

determine a first baseline number and a first threshold value based on a first plurality of numbers, wherein each of the first plurality of numbers is a number of logs of the first log type that were generated by the first data store at a time defined by a respective timestamp for the first time interval, and wherein the first baseline number is an expected number of logs of the first log type generated by the first data store at the first time for the first time interval;

compare the first number of logs of the first log type to the first baseline number; and

in response to determining that the first number of logs differs from the first baseline number by less than the first threshold value:

determine that a second number of logs of the second log type are generated by the second data store at a second time for the first time interval, wherein the second time is later than the first time by a first time delay;

determine a second baseline number and a second threshold value based on a second plurality of numbers, wherein each of the second plurality of numbers is a number of logs of the second log type that were generated by the second data store at a time defined by a respective timestamp for the first time interval, wherein the second baseline number is an expected number of logs of the second log type generated by the second data store at the second time for the first time interval;

compare the second number of logs of the second log type to the second baseline number; and

in response to determining that the second number of logs of the second log type differs from the second baseline number by more than the second threshold value:

identify the first data pipeline as degraded;

identify the second data store as degraded;

identify the second log type as degraded;

determine a second impact score of the second log type; and

analyze second dependency information for the second data store; and

in response to determining that no data store receives data items from the second data store, generate a report, wherein the report comprises:

 an identification that the first data pipeline is degraded;

 an identification that the second data store is degraded;

 an identification that the second log type is degraded; and

 the second impact score of the second log type.

16. The non-transitory computer-readable medium of claim 15 , wherein the instructions, when executed by the one or more processors, further cause the one or more processors to, in response to determining that the second number of logs of the second log type does not differ from the second baseline number by more than the second threshold value, identify the first data pipeline as not degraded.

17. The non-transitory computer-readable medium of claim 15 , wherein the instructions, when executed by the one or more processors, further cause the one or more processors to, in response to determining that a second data pipeline associates the logs of the first log type generated by the first data store with logs of a third log type generated by a third data store:

determine that a third number of logs of the third log type are generated by the third data store at a third time for the first time interval, wherein the third time is later than the first time by a second time delay;

determine a third baseline number and a third threshold value based on a third plurality of numbers, wherein each of the third plurality of numbers is a number of logs of the third log type that were generated by the third data store at a time defined by a respective timestamp for the first time interval, and wherein the third baseline number is an expected number of logs of the third log type generated by the third data store at the third time for the first time interval;

compare the third number of logs of the third log type to the third baseline number; and

in response to determining that the third number of logs of the third log type differs from the third baseline number by more than the third threshold value:

identify the second data pipeline as degraded;

identify the third data store as degraded;

identify the third log type as degraded;

determine a third impact score of the third log type;

analyze third dependency information for the third data store; and

in response to determining that no data store receives data items from the third data store, update the report, wherein the report further comprises:

an identification that the second data pipeline is degraded;

an identification that the third data store is degraded;

an identification that the third log type is degraded; and

the third impact score of the third log type.

18. The non-transitory computer-readable medium of claim 17 , wherein the instructions, when executed by the one or more processors, further cause the one or more processors to, in response to determining that the third number of logs of the third log type does not differ from the third baseline number by more than the third threshold value, identify the second data pipeline as not degraded.

19. The non-transitory computer-readable medium of claim 17 , wherein the instructions, when executed by the one or more processors, further cause the one or more processors to:

determine a priority of the report based at least in part on the second impact score and the third impact score; and

in response to determining that the priority is “high”:

generate an alert that an immediate response to the report is needed;

send the report to a data store maintenance team; and

send the alert to the data store maintenance team.

20. The non-transitory computer-readable medium of claim 17 , wherein the instructions, when executed by the one or more processors, further cause the one or more processors to:

determine a priority of the report based at least in part on the second impact score and the third impact score; and

in response to determining that the priority is “low,” send the report to a data store maintenance team.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2023
From: PIRMAN, STEPHANIE MARGARET; TEXADA, JEFFREY WAYNE; DEPREE, ERIC JOSEPH
To: BANK OF AMERICA CORPORATION
Reel/Frame 065809/0802 →
Continuity (1)
Related Publication 20250190293A1 · Jun 12, 2025
References Cited (38)
US 6694362B1 · Secor et al. · 2004 [cited by applicant]
US 7685083B2 · Fairweather · 2010 [cited by applicant]
US 8028199B1 · Guruprasad et al. · 2011 [cited by applicant]
US 8194646B2 · Elliott et al. · 2012 [cited by applicant]
US 8468244B2 · Redlich et al. · 2013 [cited by applicant]
US 8676753B2 · Sivasubramanian et al. · 2014 [cited by applicant]
US 8725667B2 · Kaushal et al. · 2014 [cited by applicant]
US 8887286B2 · Dupont et al. · 2014 [cited by applicant]
US 8909604B1 · Holenstein et al. · 2014 [cited by applicant]
US 9110898B1 · Chamness et al. · 2015 [cited by applicant]
US 9438648B2 · Asenjo et al. · 2016 [cited by applicant]
US 9665437B2 · Bhargava et al. · 2017 [cited by applicant]
US 9799017B1 · Vermeulen et al. · 2017 [cited by applicant]
US 9860152B2 · Xia et al. · 2018 [cited by applicant]
US 10417108B2 · Tankersley et al. · 2019 [cited by applicant]
US 10608911B2 · Nickolov et al. · 2020 [cited by applicant]
US 10740358B2 · Chan et al. · 2020 [cited by applicant]
US 10904276B2 · Phadke et al. · 2021 [cited by applicant]
US 11038784B2 · Nickolov et al. · 2021 [cited by applicant]
US 11356320B2 · Côté et al. · 2022 [cited by applicant]
US 11386058B2 · Hung et al. · 2022 [cited by applicant]
US 11397709B2 · Vermeulen et al. · 2022 [cited by applicant]
US 11494295B1 · Sirianni et al. · 2022 [cited by applicant]
US 11663097B2 · McAuliffe et al. · 2023 [cited by applicant]
US 11687418B2 · Baker et al. · 2023 [cited by applicant]
US 11741114B2 · Hayes et al. · 2023 [cited by applicant]
US 11792217B2 · Côté et al. · 2023 [cited by applicant]
US 20080077825A1 · Bello et al. · 2008 [cited by applicant]
US 20080155336A1 · Joshi et al. · 2008 [cited by applicant]
US 20190265082A1 · Zafar et al. · 2019 [cited by applicant]
US 20200136943A1 · Banyai et al. · 2020 [cited by applicant]
US 20210081432A1 · Grunwald et al. · 2021 [cited by applicant]
US 20220156173A1 · Chandrasekaran et al. · 2022 [cited by applicant]
US 20230118563A1 · Yadav et al. · 2023 [cited by applicant]
US 20240202059A1 · Haile · 2024 [cited by examiner]
EP 3916556A1 · 2021 [cited by examiner]
U.S. Appl. No. 18/533,347, “System and method for automated data item degradation detection”, filed Dec. 8, 2023. [cited by applicant]
U.S. Appl. No. 18/533,294, “System and method for automated data source degradation detection”, filed Dec. 8, 2023. [cited by applicant]