IP Library › Granted Patent US 12,169,434
Granted Patent B2
US 12,169,434 · App. 18/105,052 · Granted Dec 17, 2024

System, method, and computer program to improve site reliability engineering observability

Inventors: Mahesh Napa (Secaucus, NJ); Gordon Robert MacDonald (Glasgow, GB); Mark Leslie Gibbons (Kent, GB)
Assignee: JPMORGAN CHASE BANK, N.A.
G06F11/079G06F9/542G06F11/0778
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,169,434
App. No.
18/105,052
Granted
Dec 17, 2024
Kind
B2
Abstract

Various methods, apparatuses/systems, and media for improving SRE observability are disclosed. A processor defines a schema in a common manner; causes any application included across a distributed set of applications to utilize the schema to describe an error associated with a downstream application such that root failing component associated with the error is always at a bottom error frame in a response; implements a common structure for distributed error propagation in a chain of applications across the distributed set of applications in connection with the error message; generates error logs received from the chain of applications; stores the error logs in a centralized location accessible by all SRE users and application owners; calls a corresponding application programing interface (API) to access the error logs from the centralized location for utilizing in remediation.

Claims (60)

1. A method for improving site reliability engineering (SRE) observability by utilizing one or more processors along with allocated memory, the method comprising:

defining a schema in a common manner;

causing any application included across a distributed set of applications to utilize the schema to describe an error associated with a downstream application such that a root failing component associated with the error is always at a bottom error frame in a response;

implementing a common structure for distributed error propagation in a chain of applications across the distributed set of applications in connection with an error message;

generating error logs received from the chain of applications;

storing the error logs in a centralized location accessible by all SRE users and application owners;

calling a corresponding application programing interface (API) to access the error logs from the centralized location; and

automatically implementing a remedial algorithm to correct the root failing component of the error message identified in the error logs.

2. The method according to claim 1 , wherein in defining the schema in a common manner, the method further comprising:

defining a standardized error schema independently by each application across the distributed set of applications.

3. The method according to claim 2 , wherein the standardized error schema provides a set of guidelines and guardrails for standardizing error codes while still providing a flexibility to an application owner to define and manage the application owner's error codes.

4. The method according to claim 1 , further comprising:

calling corresponding API by each application to communicate an error response to each other among the distributed set of applications.

5. The method according to claim 1 , further comprising:

standardizing the common structure for the distributed error propagation in a manner such that each application among the distributed set of applications can participate in a distributed error reporting while providing its own contextualization of the error message.

6. The method according to claim 1 , wherein the centralized location is a centralized server, or a centralized database, or a centralized memory.

7. The method according to claim 1 , the method further comprising:

receiving consistent error frames from all downstream applications across the distributed set of applications that provide the error logs that describe what, where and when the error has occurred; and

propagating the error to upstream applications across the distributed set of applications.

8. The method according to claim 1 , wherein the method further comprising:

implementing an artificial intelligence (AI)/machine learning (ML) algorithm to capture all error codes from downstream applications in the chain of applications;

generating the error logs based on the captured error codes; and

implementing a self-healing algorithm to correct the root failing component of the error message identified in the error logs.

9. The method according to claim 1 , wherein the method further comprising: implementing an artificial intelligence (AI)/machine learning (ML) algorithm to capture all error logs, error codes, from downstream applications in the chain of applications; and implementing an inventory of previous remediations using a self-healing algorithm; and generating patterns of known error codes and error logs; and implementing a AL/ML logic which will use those known patterns to predict future error conditions and take preventative remedial steps to prevent the errors from occurrence.

10. A system for improving site reliability engineering (SRE) observability, the system comprising:

a processor; and

a memory operatively connected to the processor via a communication interface, the memory storing computer readable instructions, when executed, causes the processor to:

define a schema in a common manner;

cause any application included across a distributed set of applications to utilize the schema to describe an error associated with a downstream application such that a root failing component associated with the error is always at a bottom error frame in a response;

implement a common structure for distributed error propagation in a chain of applications across the distributed set of applications in connection with an error message;

generate error logs received from the chain of applications;

store the error logs in a centralized location accessible by all SRE users and application owners;

call a corresponding application programing interface (API) to access the error logs from the centralized location; and

automatically implement a remedial algorithm to correct the root failing component of the error message identified in the error logs.

11. The system according to claim 10 , wherein in defining the schema in a common manner, the processor is further configured to:

define a standardized error schema independently by each application across the distributed set of applications.

12. The system according to claim 11 , wherein the standardized error schema provides a set of guidelines and guardrails for standardizing error codes while still providing a flexibility to an application owner to define and manage the application owner's error codes.

13. The system according to claim 10 , wherein the processor is further configured to:

call corresponding API by each application to communicate an error response to each other among the distributed set of applications.

14. The system according to claim 10 , wherein the processor is further configured to:

standardize the common structure for the distributed error propagation in a manner such that each application among the distributed set of applications can participate in a distributed error reporting while providing its own contextualization of the error message.

15. The system according to claim 10 , wherein the centralized location is a centralized server, or a centralized database, or a centralized memory.

16. The system according to claim 10 , wherein the processor is further configured to:

receive consistent error frames from all downstream applications across the distributed set of applications that provide the error logs that describe what, where and when the error has occurred; and

propagate the error to upstream applications across the distributed set of applications.

17. The system according to claim 10 , wherein the processor is further configured to:

implement an artificial intelligence (AI)/machine learning (ML) algorithm to capture all error codes from downstream applications in the chain of applications;

generate the error logs based on the captured error codes; and

implement a self-healing algorithm to correct the root failing component of the error message identified in the error logs.

18. A non-transitory computer readable medium configured to store instructions for improving site reliability engineering (SRE) observability, the instructions, when executed, cause a processor to perform the following:

defining a schema in a common manner;

causing any application included across a distributed set of applications to utilize the schema to describe an error associated with a downstream application such that a root failing component associated with the error is always at a bottom error frame in a response;

implementing a common structure for distributed error propagation in a chain of applications across the distributed set of applications in connection with an error message;

generating error logs received from the chain of applications;

storing the error logs in a centralized location accessible by all SRE users and application owners;

calling a corresponding application programing interface (API) to access the error logs from the centralized location; and

automatically implementing a remedial algorithm to correct the root failing component of the error message identified in the error logs.

19. The non-transitory computer readable medium according to claim 18 , wherein in defining the schema in a common manner, the instructions, when executed, cause the processor to further perform the following:

defining a standardized error schema independently by each application across the distributed set of applications.

20. The non-transitory computer readable medium according to claim 19 , wherein the standardized error schema provides a set of guidelines and guardrails for standardizing error codes while still providing a flexibility to an application owner to define and manage the application owner's error codes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2023
From: NAPA, MAHESH; MACDONALD, GORDON ROBERT; GIBBONS, MARK LESLIE
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 063427/0635 →
Continuity (1)
Related Publication 20240264894A1 · Aug 8, 2024