IP Library › Granted Patent US 12,743,578
Granted Patent B2
US 12,743,578 · App. 18/952,547 · Granted Sep 22, 2026

Large language models for microservice fault recovery

Inventor: Hui Li (Shanghai, CN)
Assignee: SAP SE
G06F40/216G06F9/547
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,578
App. No.
18/952,547
Granted
Sep 22, 2026
Kind
B2
Abstract

In an example embodiment, an LLM is used to identify similar microservices to allow a microservice that is similar to a dependent microservice that is down to be used instead of the downed dependent microservice until the downed dependent microservice can be brought back online. Specifically, the LLM is utilized in two different manners. First, it is used to generate an embedding for an API of each of multiple microservices in a system. These embeddings may then be used to retrieve similar APIs to the API of a downed microservice. Then the LLM can be further used to select the most qualified of the similar APIs, based on similar functionality and input/output parameters. The most qualified of the APIs can then be tested for final selection of a (temporary) replacement API that can be used in lieu of the API of the downed microservice.

Claims (58)

1 . A system comprising:

at least one hardware processor;

a non-tangible computer-readable medium storing instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform operations comprising:

generating a first prompt comprising a request to generate an embedding for each of a plurality of document fragments of application program interfaces (APIs) for microservices, the embedding comprising a coordinate in a latent n-dimensional space;

sending the first prompt to a large language model (LLM);

receiving a plurality of embeddings from the LLM;

for a first document fragment, calculating cosine correlation coefficients between the first document fragment and each of a plurality of other document fragments, using the plurality of embeddings;

based on the cosine correlation coefficients, selecting a set of candidate document fragments;

generating a second prompt comprising an identification of the set of candidate document fragments and a request to identify qualified APIs corresponding to one or more candidate fragments in the set of candidate document fragments;

sending the second prompt to the large language model (LLM);

receiving an indication of a set of qualified APIs;

determining a first downstream microservice with a first API is down; and

in response to the determining, causing a first upstream microservice to use one API in the set of qualified APIs in conjunction with using a second downstream microservice.

2 . The system of claim 1 , wherein the embedding is a high-dimensional floating point vector.

3 . The system of claim 1 , wherein the calculating cosine correlation coefficients is performed offline prior to a determination that the first downstream microservice is down.

4 . The system of claim 1 , wherein the selecting comprises:

selecting a particular document fragment as part of the set of candidate document fragments if a cosine correlation coefficient corresponding to the particular document fragment exceeds a threshold.

5 . The system of claim 1 , wherein the selecting comprises:

selecting a particular document fragment as part of the set of candidate document fragments if a cosine correlation coefficient corresponding to the particular document fragment has a value in a top K cosine correlation coefficients of the plurality of other document fragments.

6 . The system of claim 4 , wherein the threshold is dynamically determined based on output of a machine learning algorithm trained by a machine learning algorithm to generate the threshold based on contextual information about the microservices.

7 . The system of claim 5 , wherein K is dynamically determined based on output of a machine learning algorithm trained by a machine learning algorithm to generate a value for K based on contextual information about the microservices.

8 . A method comprising:

generating a first prompt comprising a request to generate an embedding for each of a plurality of document fragments of application program interfaces (APIs) for microservices, the embedding comprising a coordinate in a latent n-dimensional space;

sending the first prompt to a large language model (LLM);

receiving a plurality of embeddings from the LLM;

for a first document fragment, calculating cosine correlation coefficients between the first document fragment and each of a plurality of other document fragments, using the plurality of embeddings;

based on the cosine correlation coefficients, selecting a set of candidate document fragments;

generating a second prompt comprising an identification of the set of candidate document fragments and a request to identify qualified APIs corresponding to one or more candidate fragments in the set of candidate document fragments;

sending the second prompt to the large language model (LLM);

receiving an indication of a set of qualified APIs;

determining a first downstream microservice with a first API is down; and

in response to the determining, causing a first upstream microservice to use one API in the set of qualified APIs in conjunction with using a second downstream microservice.

9 . The method of claim 8 , wherein the embedding is a high-dimensional floating point vector.

10 . The method of claim 8 , wherein the calculating cosine correlation coefficients is performed offline prior to a determination that the First downstream microservice is down.

11 . The method of claim 8 , wherein the selecting comprises:

selecting a particular document fragment as part of the set of candidate document fragments if a cosine correlation coefficient corresponding to the particular document fragment exceeds a threshold.

12 . The method of claim 8 , wherein the selecting comprises:

selecting a particular document fragment as part of the set of candidate document fragments if a cosine correlation coefficient corresponding to the particular document fragment has a value in a top K cosine correlation coefficients of the plurality of other document fragments.

13 . The method of claim 11 , wherein the threshold is dynamically determined based on output of a machine learning algorithm trained by a machine learning algorithm to generate the threshold based on contextual information about the microservices.

14 . The method of claim 12 , wherein K is dynamically determined based on output of a machine learning algorithm trained by a machine learning algorithm to generate a value for K based on contextual information about the microservices.

15 . A non-transitory machine-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations comprising:

generating a first prompt comprising a request to generate an embedding for each of a plurality of document fragments of application program interfaces (APIs) for microservices, the embedding comprising a coordinate in a latent n-dimensional space;

sending the first prompt to a large language model (LLM);

receiving a plurality of embeddings from the LLM;

for a first document fragment, calculating cosine correlation coefficients between the first document fragment and each of a plurality of other document fragments, using the plurality of embeddings;

based on the cosine correlation coefficients, selecting a set of candidate document fragments;

generating a second prompt comprising an identification of the set of candidate document fragments and a request to identify qualified APIs corresponding to one or more candidate fragments in the set of candidate document fragments;

sending the second prompt to the large language model (LLM);

receiving an indication of a set of qualified APIs;

determining a first downstream microservice with a first API is down; and

in response to the determining, causing a first upstream microservice to use one API in the set of qualified APIs in conjunction with using a second downstream microservice.

16 . The non-transitory machine-readable medium of claim 15 , wherein the embedding is a high-dimensional floating point vector.

17 . The non-transitory machine-readable medium of claim 15 , wherein the calculating cosine correlation coefficients is performed offline prior to a determination that the First downstream microservice is down.

18 . The non-transitory machine-readable medium of claim 15 , wherein the selecting comprises:

selecting a particular document fragment as part of the set of candidate document fragments if a cosine correlation coefficient corresponding to the particular document fragment exceeds a threshold.

19 . The non-transitory machine-readable medium of claim 15 , wherein the selecting comprises:

selecting a particular document fragment as part of the set of candidate document fragments if a cosine correlation coefficient corresponding to the particular document fragment has a value in a top K cosine correlation coefficients of the plurality of other document fragments.

20 . The non-transitory machine-readable medium of claim 18 , wherein the threshold is dynamically determined based on output of a machine learning algorithm trained by a machine learning algorithm to generate the threshold based on contextual information about the microservices.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2024
From: LI, HUI
To: SAP SE
Reel/Frame 069329/0798 →
Continuity (1)
Related Publication 20260141172A1 · May 21, 2026
References Cited (11)
US 11397575B2 · Wan · 2022 [cited by examiner]
US 11966725B2 · Kanso · 2024 [cited by examiner]
US 20220188104A1 · Wan · 2022 [cited by examiner]
US 20240362093A1 · Zhou · 2024 [cited by examiner]
US 20240403634A1 · Hawes · 2024 [cited by examiner]
US 20250200073A1 · Ragukumar · 2025 [cited by examiner]
US 20250371321A1 · Khafizov · 2025 [cited by examiner]
US 20260017386A1 · Ohayon · 2026 [cited by examiner]
US 20260140773A1 · Vítecek · 2026 [cited by examiner]
Peng, Baolin, et al. “Check your facts and try again: Improving large language models with external knowledge and automated feedback.” arXiv preprint arXiv:2302.12813 (2023). (Year: 2023). [cited by examiner]
Wang, Tingting, and Guilin Qi. “A comprehensive survey on root cause analysis in (micro) services: Methodologies, challenges, and trends.” arXiv preprint arXiv:2408.00803 (2024). (Year: 2024). [cited by examiner]