Multi-sided intelligent large language model assistant
In an example embodiment, a system is provided having multiple software assistants act as an interface to one or more LLMs. These assistants share contextual information about an ongoing shared conversation, but otherwise direct their respective LLM(s) to generate content based on the assistants' individual personas. The result is that a single conversation can include generated content from one or more LLMs based on multiple different personas.
1 . A system comprising:
at least one hardware processor; and
a computer-readable medium storing instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform operations comprising:
creating a first persona by a first software assistant feeding a first initial instruction set to a first Large Language Model (LLM);
creating a second persona by a second software assistant, separate from the first software assistant, feeding a second initial instruction set to a second Large Language Model (LLM);
receiving a first request for first generated content from a user;
passing the first request to the first software assistant;
causing the first software assistant to prompt the first LLM to generate first content based on the first request and using the first persona;
causing presentation of the first content generated based on the first request to the user;
receiving a second request for second generated content from the user;
passing the first request, the first content generated based on the first request, and the second request to the second software assistant;
causing the second software assistant to prompt the second LLM to generate second content based on the second request and using the second persona, using the first request and the first content generated based on the first request as context, and
causing presentation of the second content generated based on the second request to the user.
2 . The system of claim 1 , wherein the first LLM and the second LLM are a shared LLM.
3 . The system of claim 1 , wherein the first LLM utilizes additional context information stored as embeddings in a vector database.
4 . The system of claim 3 , wherein the embeddings are generated by passing content through an embedding machine learning model.
5 . The system of claim 1 , wherein the operations further comprise:
passing the first request, the content generated based on the first request, the second request, and the content generated based on the second request to the first software assistant;
causing the first software assistant to prompt the first LLM to generate content based on the content generated based on the second request, using the first request, the content generated based on the first request, and the second request as context, prior to receiving any user input from the user in response to the presentation of the content generated based on the second request;
receiving content generated based on the content generated based on the second request from the first LLM; and
causing presentation of the content generated based on the content generated based on the second request to the user.
6 . The system of claim 1 , wherein the causing presentation of the content generated based on the first request to the user includes displaying text of the content generated based on the first request in a graphical user interface.
7 . The system of claim 1 , wherein the causing presentation of the content generated based on the first request to the user includes converting text of the content generated based on the first request to an audio file and playing the audio file to the user.
8 . A method comprising:
creating a first persona by a first software assistant feeding a first initial instruction set to a first Large Language Model (LLM);
creating a second persona by a second software assistant, separate from the first software assistant, feeding a second initial instruction set to a second Large Language Model (LLM);
receiving a first request for first generated content from a user;
passing the first request to the first software assistant;
causing the first software assistant to prompt the first LLM to generate first content based on the first request and using the first persona;
causing presentation of the first content generated based on the first request to the user;
receiving a second request for second generated content from the user;
passing the first request, the first content generated based on the first request, and the second request to the second software assistant;
causing the second software assistant to prompt the second LLM to generate second content based on the second request and using the second persona, using the first request and the first content generated based on the first request as context, and
causing presentation of the second content generated based on the second request to the user.
9 . The method of claim 8 , wherein the first LLM and the second LLM are a shared LLM.
10 . The method of claim 8 , wherein the first LLM utilizes additional context information stored as embeddings in a vector database.
11 . The method of claim 10 , wherein the embeddings are generated by passing content through an embedding machine learning model.
12 . The method of claim 8 , further comprising:
passing the first request, the content generated based on the first request, the second request, and the content generated based on the second request to the first software assistant;
causing the first software assistant to prompt the first LLM to generate content based on the content generated based on the second request, using the first request, the content generated based on the first request, and the second request as context, prior to receiving any user input from the user in response to the presentation of the content generated based on the second request;
receiving content generated based on the content generated based on the second request from the first LLM; and
causing presentation of the content generated based on the content generated based on the second request to the user.
13 . The method of claim 8 , wherein the causing presentation of the content generated based on the first request to the user includes displaying text of the content generated based on the first request in a graphical user interface.
14 . The method of claim 8 , wherein the causing presentation of the content generated based on the first request to the user includes converting text of the content generated based on the first request to an audio file and playing the audio file to the user.
15 . A non-transitory machine-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operation comprising:
creating a first persona by a first software assistant feeding a first initial instruction set to a first Large Language Model (LLM);
creating a second persona by a second software assistant, separate from the first software assistant, feeding a second initial instruction set to a second Large Language Model (LLM);
receiving a first request for first generated content from a user;
passing the first request to the first software assistant;
causing the first software assistant to prompt the first LLM to generate first content based on the first request and using the first persona;
causing presentation of the first content generated based on the first request to the user;
receiving a second request for second generated content from the user;
passing the first request, the first content generated based on the first request, and the second request to the second software assistant;
causing the second software assistant to prompt the second LLM to generate second content based on the second request and using the second persona, using the first request and the first content generated based on the first request as context, and
causing presentation of the second content generated based on the second request to the user.
16 . The non-transitory machine-readable medium of claim 15 , wherein the first LLM and the second LLM are a shared LLM.
17 . The non-transitory machine-readable medium of claim 15 , wherein the first LLM utilizes additional context information stored as embeddings in a vector database.
18 . The non-transitory machine-readable medium of claim 17 , wherein the embeddings are generated by passing content through an embedding machine learning model.
19 . The non-transitory machine-readable medium of claim 15 , further comprising:
passing the first request, the content generated based on the first request, the second request, and the content generated based on the second request to the first software assistant;
causing the first software assistant to prompt the first LLM to generate content based on the content generated based on the second request, using the first request, the content generated based on the first request, and the second request as context, prior to receiving any user input from the user in response to the presentation of the content generated based on the second request;
receiving content generated based on the content generated based on the second request from the first LLM; and
causing presentation of the content generated based on the content generated based on the second request to the user.
20 . The non-transitory machine-readable medium of claim 15 , wherein the causing presentation of the content generated based on the first request to the user includes displaying text of the content generated based on the first request in a graphical user interface.