Knowledge bot as a service
Methods and systems are presented for providing a knowledge bot configurable to interact with users across multiple domains. The knowledge bot includes at least a text-based search engine and a semantic-based search engine. Each of the search engine is configured to retrieve documents from a corpus of documents based on the user query. The user query is in a natural language format. The retrieved documents may be ranked according to how relevant the documents are to the user query. A subset of the documents is used as the search results based on the ranking. The search results from the search engine are combined with the user query to generate a prompt for an artificial intelligence model. Based on the prompt, a response in the natural language format is generated by the artificial intelligence model.
1 . A system comprising:
a non-transitory memory; and
one or more hardware processors coupled to the non-transitory memory and configured to read instructions from the non-transitory memory to cause the system to perform operations comprising:
obtaining a corpus of documents corresponding to a first domain usable to generate a knowledge bot;
providing, for the knowledge bot, a chat interface configured to receive a user query from a user device;
generating, based on the corpus of documents, an inverted index usable by a text-based search model, wherein the text-based search model is configured to produce a first search result comprising a first set of documents from the corpus of documents using the inverted index;
generating, based on the corpus of documents, a vector index usable by a semantic search model, wherein the semantic search model is configured to produce a second search result comprising a second set of documents from the corpus of documents using the vector index;
generating a semantic cache configured to cache embeddings of previously submitted queries and responses generated by the knowledge bot for the previously submitted queries; and
integrating the text-based search model, the semantic search model, the semantic cache, and a machine learning model within the knowledge bot, wherein the knowledge bot is configured to (i) generate a set of embeddings based on the user query received via the chat interface, (ii) determine whether the user query matches any of the previously submitted queries in the semantic cache based on the set of embeddings; and (iii) in response to determining that the user query does not match any of the previously submitted queries in the semantic cache, generate a response to the user query using the machine learning model and based on the first search result and the second search result, wherein the response comprises a plurality of words in a natural language format, and wherein the chat interface is further configured to present the response on the user device.
2 . The system of claim 1 , wherein the operations further comprise:
obtaining a plurality of user queries associated with the first domain;
determining a plurality of target responses corresponding to the plurality of user queries;
generating, using the knowledge bot, a plurality of responses based on the plurality of user queries; and
performing a semantic comparison between the plurality of target responses and the plurality of responses.
3 . The system of claim 2 , wherein the operations further comprise:
adjusting one or more parameters associated with the machine learning model based on the semantic comparison between the plurality of target responses and the plurality of responses.
4 . The system of claim 2 , wherein the operations further comprise:
adjusting one or more parameters associated with at least one of the text-based search model or the semantic search model based on the semantic comparison between the plurality of target responses and the plurality of responses.
5 . The system of claim 2 , wherein the operations further comprise:
determining that the corpus of documents lacks information associated with a particular topic based on the semantic comparison between the plurality of target responses and the plurality of responses;
obtaining a set of documents associated with the particular topic; and
adding the set of documents to the corpus of documents.
6 . The system of claim 2 , wherein the plurality of user queries comprises a first set of user queries below a threshold query length and a second set of user queries above the threshold query length.
7 . The system of claim 1 , wherein the user query is a first user query, wherein the response is a first response, and wherein the operations further comprise:
storing the set of embeddings of the first user query and the first response in the semantic cache;
in response to receiving a second user query from a second user device, performing a second semantic comparison between the second user query and a plurality of keys stored in the semantic cache memory;
determining a match between the second user query and the set of embeddings of the first user query based on the second semantic comparison;
without using the machine learning model to process the second user query, retrieving the first response from the semantic cache memory; and
generating a second response to the second user query based on the first response.
8 . A method comprising:
providing, for a knowledge bot, a chat interface configured to receive a user query associated with a first domain;
accessing, by a computer system, a corpus of documents associated with the first domain;
generating, by the computer system and based on the corpus of documents, an inverted index enabled for use by a text-based search model that is configured to identify, from the corpus of documents, a first set of documents associated with the user query using the inverted index;
generating, by the computer system and based on the corpus of documents, a vector index enabled for use by a semantic search model that is configured to identify, from the corpus of documents, a second set of documents associated with the user query using the vector index;
accessing, by the computer system, a semantic cache configured to cache embeddings of previously submitted queries and responses generated by the knowledge bot for the previously submitted queries; and
integrating, by the computer system, the text-based search model, the semantic search model, and the semantic cache with an artificial intelligence (AI) model within the knowledge bot, wherein the knowledge bot is configured to (i) generate a set of embeddings based on the user query received via the chat interface, (ii) determine whether the user query matches any of the previously submitted queries in the semantic cache based on the set of embeddings; and (iii) in response to determining that the user query does not match any of the previously submitted queries in the semantic cache, generate a response to the user query using the AI model and based on the first and second sets of documents, and wherein the response comprises content that is derived from the first and second sets of documents; and
presenting the response on the chat interface.
9 . The method of claim 8 , further comprising:
modifying the response to the user query based on a set of policies.
10 . The method of claim 9 , wherein the modifying comprises at least one of replacing a first word in the response with a second word, removing one or more words from the response, or modifying at least one word from the response.
11 . The method of claim 8 , further comprising:
receiving a second query from a device;
determining that the second query is not associated with the first domain; and
in response to determining that the second query is not associated with the first domain and without using the AI model to process the second query, providing a default response to the device.
12 . The method of claim 8 , further comprising:
obtaining a plurality of test queries associated with the first domain;
determining a plurality of benchmark responses corresponding to the plurality of test queries;
generating, using the knowledge bot, a plurality of test responses based on the plurality of test queries; and
determining a deviation between the plurality of benchmark responses and the plurality of test responses based on a semantic comparison between the plurality of benchmark responses and the plurality of test responses.
13 . The method of claim 12 , further comprising:
in response to determining that the deviation exceeds a threshold, adjusting one or more parameters associated with the AI model.
14 . The method of claim 12 , further comprising:
in response to determining that the deviation exceeds a threshold, adjusting one or more parameters associated with at least one of the text-based search model or the semantic search model.
15 . A non-transitory machine-readable medium having stored thereon machine-readable instructions executable to cause a machine to perform operations comprising:
accessing a corpus of documents associated with a first domain;
generating, based on the corpus of documents, an inverted index for use by a text-based search model that is configured to produce a first search result comprising a first set of documents from the corpus of documents based on a user query using the inverted index;
generating, based on the corpus of documents, a vector index for use by a semantic search model that is configured to produce a second search result comprising a second set of documents from the corpus of documents based on the user query using the vector index;
generating a semantic cache for caching embeddings of previously submitted queries and responses generated for the previously submitted queries;
integrating the semantic cache, the text-based search model, and the semantic search model with a machine learning model; and
in response to receiving the user query, (i) generating a set of embeddings based on the user query, (ii) determining whether the user query matches any of the previously submitted queries in the semantic cache based on the set of embeddings; and (iii) in response to determining that the user query does not match any of the previously submitted queries in the semantic cache, generating a response to the user query using the machine learning model and based on the first and second search results, and wherein the response comprises a plurality of words in a natural language format.
16 . The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise:
obtaining a plurality of queries associated with the first domain;
determining a plurality of target answers corresponding to the plurality of queries;
generating, using the machine learning model, a plurality of candidate answers based on the plurality of user queries; and
performing a semantic comparison between the plurality of target answers and the plurality of candidate answers.
17 . The non-transitory machine-readable medium of claim 16 , wherein the operations further comprise:
adjusting at least one of a first parameter associated with the machine learning model or a second parameter associated with at least one of the text-based search model or the semantic search model based on the semantic comparison between the plurality of target answers and the plurality of candidate answers.
18 . The non-transitory machine-readable medium of claim 16 , wherein the operations further comprise:
determining that the corpus of documents lacks information associated with a particular topic corresponding to the first domain based on the semantic comparison between the plurality of target answers and the plurality of candidate answers;
obtaining a set of documents associated with the particular topic; and
adding the set of documents to the corpus of documents.
19 . The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise:
modifying the response to the user query based on a set of policies; and
presenting the modified response on an interface.
20 . The non-transitory machine-readable medium of claim 19 , wherein the modifying comprises at least one of replacing a first word in the response with a second word, modifying at least one word from the response, or removing one or more words from the response.