IP Library › Granted Patent US 12,730,948
Granted Patent B2
US 12,730,948 · App. 18/130,767 · Granted Sep 8, 2026

Method and device for dynamic failure mode effect analysis and recovery process recommendation for cloud computing applications

Inventors: Sankar Narayan Das (Barrackpore, IN); Kuntal Dey (Birbhum, IN); Kapil Singi (Bangalore, IN); Vikrant Kaulgud (Pune, IN); Manish Ahuja (Bengaluru, IN); Reuben Rajan George (Enathu, IN); Mallika Fernandes (Bangalore, IN); Mahesh Venkata Raman (Bangalore, IN)
Assignee: Accenture Global Solutions Limited
G06F30/27G06F2119/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,730,948
App. No.
18/130,767
Granted
Sep 8, 2026
Kind
B2
Abstract

Aspects of the present disclosure provide methods, devices, and computer-readable storage media that support detection, effect monitoring, and recovery from failure modes in cloud computing application using a failure mode effect analysis (FMEA) engine. Historical metadata related to operation of a hierarchy of devices may be used as training data to train the FMEA engine to identify failure modes experienced by the hierarchy of devices. After training the FMEA engine, metadata from the hierarchy of devices may be input to the FMEA engine to identify a failure mode that may have occurred, and the FMEA engine may select a recovery process to recommend for addressing or mitigating the identified failure mode. In some implementations, the FMEA engine may output an indication of the recommended recovery process and/or initiate performance of one or more operations at the hierarchy of devices to recover from the failure event.

Claims (56)

1 . A method for determining a recovery process associated with a cloud computing application failure mode, the method comprising:

receiving, by one or more processors, historical metadata associated with a hierarchy of devices associated with a cloud computing application;

providing, by the one or more processors, the historical metadata as training data to one or more machine learning (ML) applications to train the one or more ML applications to determine one or more failure modes associated with the hierarchy of devices based on input metadata;

providing, by the one or more processors, second metadata associated with the hierarchy of devices as input to the one or more ML applications to determine a failure mode occurring at one or more of the hierarchy of devices;

determining, by the one or more processors based on the failure mode and the second metadata, a recommended recovery process that corresponds to the failure mode;

outputting, by the one or more processors, a message indicating the recommended recovery process;

predicting, by the one or more processors, efficiency scores corresponding to multiple candidate recovery processes associated with the failure mode; and

selecting, by the one or more processors, the recommended recovery process from among the multiple candidate recovery processes based on the efficiency scores,

wherein the recommended recovery process is selected from among the multiple candidate recovery processes based on a first metadata profile associated with the recommended recovery process and further based on a second metadata profile associated with a second recovery process of the multiple candidate recovery processes.

2 . The method of claim 1 , further comprising automatically initiating, by the one or more processors, one or more operations of the recommended recovery process at one or more devices of the hierarchy of devices to mitigate an effect of the failure mode.

3 . The method of claim 2 , wherein the one or more operations include one or more of:

selecting among a first communication link and a second communication link for data transmission;

selecting among the first communication link and the second communication link for data reception;

selecting among a first device and a second device for performance of a first service; or

selecting among the first service and a second service.

4 . The method of claim 1 , wherein the one or more failure modes include at least one of a delay associated with a response to a request, a failure to respond to the request, an erroneous response to the request, a crash event, a hardware failure event, or a disruption of communication through a communication link.

5 . The method of claim 1 , wherein the recommended recovery process is selected based on the first metadata profile indicating that a first load increase associated with the recommended recovery process is less than a second load increase associated with the second recovery process.

6 . The method of claim 1 , wherein the recommended recovery process is selected based on the first metadata profile indicating that a first delay increase associated with the recommended recovery process is less than a second delay increase associated with the second recovery process.

7 . A device for determining a recovery process associated with a cloud computing application failure mode, the device comprising:

a memory; and

one or more processors communicatively coupled to the memory, the one or more processors configured to:

receive historical metadata associated with a hierarchy of devices associated with a cloud computing application;

provide the historical metadata as training data to one or more machine learning (ML) applications to train the one or more ML applications to determine one or more failure modes associated with the hierarchy of devices based on input metadata;

provide second metadata associated with the hierarchy of devices as input to the one or more ML applications to determine a failure mode occurring at one or more of the hierarchy of devices;

determine, based on the failure mode and the second metadata, a recommended recovery process that corresponds to the failure mode; and

output a message indicating the recommended recovery process;

predict efficiency scores corresponding to multiple candidate recovery processes associated with the failure mode; and

select the recommended recovery process from among the multiple candidate recovery processes based on the efficiency scores,

wherein the recommended recovery process is selected from among the multiple candidate recovery processes based on a first metadata profile associated with the recommended recovery process and further based on a second metadata profile associated with a second recovery process of the multiple candidate recovery processes.

8 . The device of claim 7 , wherein the historical metadata indicates one or more of a component identifier associated with at least one device of the hierarchy of devices, a geographic location associated with the at least one device, a timestamp of an event associated with the at least one device, a request servicing rate associated with the at least one device, a request servicing error rate associated with the at least one device, a request servicing duration associated with the at least one device, or a resource utilization associated with the at least one device.

9 . The device of claim 7 , wherein the one or more processors are further configured to determine a priority scheme that indicates a plurality of risk priority numbers (RPNs) associated with the one or more failure modes.

10 . The device of claim 9 , wherein the one or more processors are further configured to generate a knowledgebase associated with the failure mode based on an RPN of the plurality of RPNs associated with the failure mode exceeding a threshold RPN.

11 . The device of claim 9 , wherein:

the plurality of RPNs includes at least a first RPN associated with the failure mode, and

the first RPN is based on one or more of a severity value associated with the failure mode, a probability of occurrence associated with the failure mode, or a detectability metric associated with the failure mode.

12 . The device of claim 11 , wherein the severity value associated with the failure mode is based on one or more of a number of occurrences of the failure mode, a recovery time associated with recovering from the failure mode, a data loss event associated with the failure mode, a loss of functionality associated with the failure mode, or a rate of occurrence associated with the failure mode.

13 . The device of claim 11 , wherein the probability of occurrence associated with the failure mode is based on one or more of a number of occurrences of the failure mode or a total number of occurrences among the one or more failure modes.

14 . The device of claim 11 , wherein the detectability metric is based on one or more of an accuracy associated with a detection model associated with the failure mode or an error tolerance value associated with the detection model.

15 . A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations for determining a recovery process associated with a cloud computing application failure mode, the operations comprising:

receiving, by one or more processors, historical metadata associated with a hierarchy of devices associated with a cloud computing application;

providing, by the one or more processors, the historical metadata as training data to one or more machine learning (ML) applications to train the one or more ML applications to determine one or more failure modes associated with the hierarchy of devices based on input metadata;

providing, by the one or more processors, second metadata associated with the hierarchy of devices as input to the one or more ML applications to determine a failure mode occurring at one or more of the hierarchy of devices;

determining, by the one or more processors based on the failure mode and the second metadata, a recommended recovery process that corresponds to the failure mode; and

outputting, by the one or more processors, a message indicating the recommended recovery process;

predicting, by the one or more processors, efficiency scores corresponding to multiple candidate recovery processes associated with the failure mode; and

selecting, by the one or more processors, the recommended recovery process from among the multiple candidate recovery processes based on the efficiency scores,

wherein the recommended recovery process is selected from among the multiple candidate recovery processes based on a first metadata profile associated with the recommended recovery process and further based on a second metadata profile associated with a second recovery process of the multiple candidate recovery processes.

16 . The non-transitory computer-readable storage medium of claim 15 , wherein the operations further comprise:

determining a first trend associated with the failure mode based on evaluation of subsets of the historical metadata using a first time interval; and

determining a second trend associated with the failure mode based on the first trend, the second trend associated with a second time interval greater than the first time interval.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein:

the first trend is associated with a first frequency of occurrence of the failure mode and with a first mean time between failures (MTBF), and

the second trend is associated with a second frequency of occurrence of the failure mode and with a second MTBF.

18 . The non-transitory computer-readable storage medium of claim 15 , wherein:

the cloud computing application includes a cloud continuum application, and

the one or more ML applications are integrated in a failure mode effect analysis (FMEA) engine.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2023
From: DAS, SANKAR NARAYAN; DEY, KUNTAL; SINGI, KAPIL; KAULGUD, VIKRANT; AHUJA, MANISH; GEORGE, REUBEN RAJAN; FERNANDES, MALLIKA; VENKATARAMAN, MAHESH
To: ACCENTURE GLOBAL SOLUTIONS LIMITED
Reel/Frame 064305/0052 →
Continuity (1)
Related Publication 20230315954A1 · Oct 5, 2023
References Cited (4)
Jishan Ahmed, et. al. Predicting severely imbalanced data disk drive failures with machine learning models, Machine Learning with Applications 9 (2022) 100361, pp. 1-12 (Year: 2022). [cited by examiner]
Oliveira, J. et al., “Failure Mode and Effect Analysis for Cyber-Physical Systems,” Future Internet, vol. 12, No. 11, 205, 18 pages, 2020. [cited by applicant]
Ali, N. et al., “Failure detection and prevention for cyber-physical systems using ontology-based knowledge base,” Computers, vol. 7, No. 4, 68, 16 pages, 2018. [cited by applicant]
Qin, J. et al., “Failure mode and effects analysis (FMEA) for risk assessment based on interval type-2 fuzzy evidential reasoning method,” Applied Soft Computing, vol. 89, 106134, 14 pages, 2020. [cited by applicant]