Large language models for microservice fault recovery
In an example embodiment, an LLM is used to identify similar microservices to allow a microservice that is similar to a dependent microservice that is down to be used instead of the downed dependent microservice until the downed dependent microservice can be brought back online. Specifically, the LLM is utilized in two different manners. First, it is used to generate an embedding for an API of each of multiple microservices in a system. These embeddings may then be used to retrieve similar APIs to the API of a downed microservice. Then the LLM can be further used to select the most qualified of the similar APIs, based on similar functionality and input/output parameters. The most qualified of the APIs can then be tested for final selection of a (temporary) replacement API that can be used in lieu of the API of the downed microservice.
1 . A system comprising:
at least one hardware processor;
a non-tangible computer-readable medium storing instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform operations comprising:
generating a first prompt comprising a request to generate an embedding for each of a plurality of document fragments of application program interfaces (APIs) for microservices, the embedding comprising a coordinate in a latent n-dimensional space;
sending the first prompt to a large language model (LLM);
receiving a plurality of embeddings from the LLM;
for a first document fragment, calculating cosine correlation coefficients between the first document fragment and each of a plurality of other document fragments, using the plurality of embeddings;
based on the cosine correlation coefficients, selecting a set of candidate document fragments;
generating a second prompt comprising an identification of the set of candidate document fragments and a request to identify qualified APIs corresponding to one or more candidate fragments in the set of candidate document fragments;
sending the second prompt to the large language model (LLM);
receiving an indication of a set of qualified APIs;
determining a first downstream microservice with a first API is down; and
in response to the determining, causing a first upstream microservice to use one API in the set of qualified APIs in conjunction with using a second downstream microservice.
2 . The system of claim 1 , wherein the embedding is a high-dimensional floating point vector.
3 . The system of claim 1 , wherein the calculating cosine correlation coefficients is performed offline prior to a determination that the first downstream microservice is down.
4 . The system of claim 1 , wherein the selecting comprises:
selecting a particular document fragment as part of the set of candidate document fragments if a cosine correlation coefficient corresponding to the particular document fragment exceeds a threshold.
5 . The system of claim 1 , wherein the selecting comprises:
selecting a particular document fragment as part of the set of candidate document fragments if a cosine correlation coefficient corresponding to the particular document fragment has a value in a top K cosine correlation coefficients of the plurality of other document fragments.
6 . The system of claim 4 , wherein the threshold is dynamically determined based on output of a machine learning algorithm trained by a machine learning algorithm to generate the threshold based on contextual information about the microservices.
7 . The system of claim 5 , wherein K is dynamically determined based on output of a machine learning algorithm trained by a machine learning algorithm to generate a value for K based on contextual information about the microservices.
8 . A method comprising:
generating a first prompt comprising a request to generate an embedding for each of a plurality of document fragments of application program interfaces (APIs) for microservices, the embedding comprising a coordinate in a latent n-dimensional space;
sending the first prompt to a large language model (LLM);
receiving a plurality of embeddings from the LLM;
for a first document fragment, calculating cosine correlation coefficients between the first document fragment and each of a plurality of other document fragments, using the plurality of embeddings;
based on the cosine correlation coefficients, selecting a set of candidate document fragments;
generating a second prompt comprising an identification of the set of candidate document fragments and a request to identify qualified APIs corresponding to one or more candidate fragments in the set of candidate document fragments;
sending the second prompt to the large language model (LLM);
receiving an indication of a set of qualified APIs;
determining a first downstream microservice with a first API is down; and
in response to the determining, causing a first upstream microservice to use one API in the set of qualified APIs in conjunction with using a second downstream microservice.
9 . The method of claim 8 , wherein the embedding is a high-dimensional floating point vector.
10 . The method of claim 8 , wherein the calculating cosine correlation coefficients is performed offline prior to a determination that the First downstream microservice is down.
11 . The method of claim 8 , wherein the selecting comprises:
selecting a particular document fragment as part of the set of candidate document fragments if a cosine correlation coefficient corresponding to the particular document fragment exceeds a threshold.
12 . The method of claim 8 , wherein the selecting comprises:
selecting a particular document fragment as part of the set of candidate document fragments if a cosine correlation coefficient corresponding to the particular document fragment has a value in a top K cosine correlation coefficients of the plurality of other document fragments.
13 . The method of claim 11 , wherein the threshold is dynamically determined based on output of a machine learning algorithm trained by a machine learning algorithm to generate the threshold based on contextual information about the microservices.
14 . The method of claim 12 , wherein K is dynamically determined based on output of a machine learning algorithm trained by a machine learning algorithm to generate a value for K based on contextual information about the microservices.
15 . A non-transitory machine-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations comprising:
generating a first prompt comprising a request to generate an embedding for each of a plurality of document fragments of application program interfaces (APIs) for microservices, the embedding comprising a coordinate in a latent n-dimensional space;
sending the first prompt to a large language model (LLM);
receiving a plurality of embeddings from the LLM;
for a first document fragment, calculating cosine correlation coefficients between the first document fragment and each of a plurality of other document fragments, using the plurality of embeddings;
based on the cosine correlation coefficients, selecting a set of candidate document fragments;
generating a second prompt comprising an identification of the set of candidate document fragments and a request to identify qualified APIs corresponding to one or more candidate fragments in the set of candidate document fragments;
sending the second prompt to the large language model (LLM);
receiving an indication of a set of qualified APIs;
determining a first downstream microservice with a first API is down; and
in response to the determining, causing a first upstream microservice to use one API in the set of qualified APIs in conjunction with using a second downstream microservice.
16 . The non-transitory machine-readable medium of claim 15 , wherein the embedding is a high-dimensional floating point vector.
17 . The non-transitory machine-readable medium of claim 15 , wherein the calculating cosine correlation coefficients is performed offline prior to a determination that the First downstream microservice is down.
18 . The non-transitory machine-readable medium of claim 15 , wherein the selecting comprises:
selecting a particular document fragment as part of the set of candidate document fragments if a cosine correlation coefficient corresponding to the particular document fragment exceeds a threshold.
19 . The non-transitory machine-readable medium of claim 15 , wherein the selecting comprises:
selecting a particular document fragment as part of the set of candidate document fragments if a cosine correlation coefficient corresponding to the particular document fragment has a value in a top K cosine correlation coefficients of the plurality of other document fragments.
20 . The non-transitory machine-readable medium of claim 18 , wherein the threshold is dynamically determined based on output of a machine learning algorithm trained by a machine learning algorithm to generate the threshold based on contextual information about the microservices.