IP Library › Granted Patent US 12,411,726
Granted Patent B2
US 12,411,726 · App. 18/142,254 · Granted Sep 9, 2025

Method and system for solving for seamless resiliency in a multi-tier network for grey failures

Inventors: Rajeshwari Edamadaka (Allentown, NJ); Diarmuid Leonard (Galway, IE); Nigel T Cook (Boulder, CO)
Assignee: JPMORGAN CHASE BANK, N.A.
G06F11/0793G06F11/0709G06F11/079
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,411,726
App. No.
18/142,254
Granted
Sep 9, 2025
Kind
B2
Abstract

An automatically self-healing multi-tier system for providing seamless resiliency for end users is provided. The system includes a plurality of tiers of elements; a processor; a memory; and a communication interface. The processor is configured to determine whether each respective tier of elements satisfies each of a plurality of intrinsic observer capabilities, a plurality of intrinsic reactor capabilities, a plurality of first health checks received from an internal tier, and a plurality of second health checks received from an external tier. When any of the intrinsic observer capabilities and the intrinsic reactor capabilities are not satisfied, an extrinsic observer capability and/or an extrinsic reactor capability is used to compensate for the unsatisfied capability. When any of health checks discover a degradation of service, communication flows are routed so as to fully or partially avoid the affected tier of elements.

Claims (58)

1. A method for mitigating a failure in a distributed cloud-based data center that includes at least two availability zones, each respective availability zone including a corresponding plurality of components, the method being implemented by at least one processor, the method comprising:

checking, by the at least one processor, each respective component from among a first plurality of components included in a first availability zone;

detecting, by the at least one processor, at least one partial failure that is associated with at least one component from among the first plurality of components;

routing, by the at least one processor, at least one communication flow so as to avoid the at least one component for which the at least one partial failure has been detected;

generating, by the at least one processor, a notification message that includes information that relates to the detected at least one partial failure; and

transmitting, by the at least one processor, the notification message to a predetermined destination,

wherein the routing comprises routing the at least one communication flow so as to avoid the first availability zone, and

wherein the method further comprises:

determining, based on a result of checking each respective component from among a second plurality of components included in a second availability zone, that all components in the second plurality of components are functioning normally,

wherein the routing further comprises routing the at least one communication flow so as to propagate via the second availability zone.

2. The method of claim 1 , wherein the checking comprises measuring, for each respective component, at least one health metric that indicates whether the respective component is functioning normally.

3. The method of claim 2 , wherein the checking further comprises:

generating a synthetic service request;

routing the synthetic service request so as to propagate via each of the first plurality of components;

receiving, from each respective component, a corresponding response to the synthetic service request; and

using each received response to measure each corresponding at least one health metric.

4. The method of claim 1 , further comprising:

receiving information indicating that an operational functionality has been restored for the at least one component for which the at least one partial failure has been detected; and

rerouting the at least one communication flow so as to propagate through the at least one component for which the operational functionality has been restored.

5. The method of claim 1 , further comprising updating a database with information that relates to a number of partial failures that are detected within a predetermined time interval.

6. The method of claim 5 , further comprising updating the database with information that relates to a number of components that are determined as functioning normally within the predetermined time interval.

7. The method of claim 1 , wherein the at least one partial failure includes at least one from among a hardware failure, a network disruption, a communication overload, a performance degradation, a random packet loss, an input/output glitch, a memory thrashing, a capacity pressure, and a non-fatal exception.

8. A computing apparatus for mitigating a failure in a distributed cloud-based data center that includes at least two availability zones, each respective availability zone including a corresponding plurality of components, the computing apparatus comprising:

a processor;

a memory; and

a communication interface coupled to each of the processor and the memory,

wherein the processor is configured to:

check each respective component from among a first plurality of components included in a first availability zone;

detect at least one partial failure that is associated with at least one component from among the first plurality of components;

route at least one communication flow so as to avoid the at least one component for which the at least one partial failure has been detected;

generate a notification message that includes information that relates to the detected at least one partial failure; and

transmit, via the communication interface, the notification message to a predetermined destination,

wherein the processor is further configured to route the at least one communication flow so as to avoid the first availability zone, and

wherein the processor is further configured to:

determine, based on a result of checking each respective component from among a second plurality of components included in a second availability zone, that all components in the second plurality of components are functioning normally; and

route the at least one communication flow so as to propagate via the second availability zone.

9. The computing apparatus of claim 8 , wherein the processor is further configured to measure, for each respective component, at least one health metric that indicates whether the respective component is functioning normally.

10. The computing apparatus of claim 9 , wherein the processor is further configured to:

generate a synthetic service request;

route the synthetic service request so as to propagate via each of the first plurality of components;

receive, from each respective component, a corresponding response to the synthetic service request; and

use each received response to measure each corresponding at least one health metric.

11. The computing apparatus of claim 8 , wherein the processor is further configured to:

receive information indicating that an operational functionality has been restored for the at least one component for which the at least one partial failure has been detected; and

reroute the at least one communication flow so as to propagate through the at least one component for which the operational functionality has been restored.

12. The computing apparatus of claim 8 , wherein the processor is further configured to update a database with information that relates to a number of partial failures that are detected within a predetermined time interval.

13. The computing apparatus of claim 12 , wherein the processor is further configured to update the database with information that relates to a number of components that are determined as functioning normally within the predetermined time interval.

14. The computing apparatus of claim 8 , wherein the at least one partial failure includes at least one from among a hardware failure, a network disruption, a communication overload, a performance degradation, a random packet loss, an input/output glitch, a memory thrashing, a capacity pressure, and a non-fatal exception.

15. An automatically self-healing multi-tier system for providing seamless resiliency, comprising:

a plurality of tiers of elements;

a processor;

a memory; and

a communication interface coupled to the processor, the memory, and each respective tier of elements from among the plurality of tiers of elements,

wherein the processor is configured to:

determine whether each respective tier of elements satisfies each of a plurality of intrinsic observer capabilities, a plurality of intrinsic reactor capabilities, a plurality of first health checks received from an internal tier, and a plurality of second health checks received from an external tier;

when at least one from among the plurality of intrinsic observer capabilities and the plurality of intrinsic reactor capabilities is not satisfied, use at least one from among a plurality of extrinsic observer capabilities and a plurality of extrinsic reactor capabilities to compensate for the unsatisfied capability; and

when at least one from among the plurality of first health checks and the plurality of second health checks is not satisfied, route at least one communication flow so as to avoid the tier of elements for which the corresponding health check is not satisfied.

16. The system of claim 15 , wherein the processor is further configured to measure, for each respective tier of elements, at least one health metric that indicates whether the respective tier of elements is well-behaved.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2023
From: EDAMADAKA, RAJESHWARI; LEONARD, DIARMUID; COOK, NIGEL T
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 063518/0601 →
Continuity (2)
Provisional Application 63364092 · May 3, 2022
Related Publication 20230359520A1 · Nov 9, 2023
References Cited (11)
US 20210157694A1 · Dye et al. · 2021 [cited by applicant]
Haken, Michael, “Advanced Multi-AZ Resilience Patterns: Detecting and Mitigating Gray Failures”, Mar. 2, 2022, Amazon Web Services (Year: 2022). [cited by examiner]
“Amazon CloudWatch User Guide”, Nov. 17, 2021, Amazon Web Services, pp. i-xii, 1-9, 154-231, 937-959 (Year: 2021). [cited by examiner]
Haken, Michael, “Advanced Multi-AZ Resilience Patterns: Detecting and Mitigating Gray Failures”, Jul. 11, 2023, Amazon Web Services (Year: 2023). [cited by examiner]
Jia et al., “Rapid Detection and Localization of Gray Failures in Data Centers via In-band Network Telemetry”, NOMS 2020-2020 IEEE/IFIP Network Operations and Management Symposium (Year: 2020). [cited by examiner]
Huang et al., “Gray Failure: The Achilles' Heel of Cloud-Scale Systems”, May 8-10, 2017, HotOS '17, pp. 150-155 (Year: 2017). [cited by examiner]
Huang et al., “Capturing and Enhancing In Situ System Observability for Failure Detection”, Oct. 8-10, 2018, 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI '18), pp. 1-15 (Year: 2018). [cited by examiner]
International Search Report and Written Opinion in corresponding International Application No. PCT/US2023/020673, dated Jul. 27, 2023. [cited by applicant]
AWS, Reliability Pillar AWS Well-Architected Framework, Amazon Web Services, Feb. 19, 2021. [cited by applicant]
Berenberg A. et al., Deployment Archetypes for Cloud Applications, ACM Computing Surveys, vol. 55, No. 3, Article 61 (pp. 61:1-61:48), Feb. 3, 2022. [cited by applicant]
Benson T. et al., CloudNaaS: A Cloud Networking Platform for Enterprise Applications, Proceedings of the 2 [cited by applicant]