Iterative response generation using generation schemas
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for iteratively generating different sections of a textual electronic document (“TED”) using one or more large language models (LLMs). In one aspect, a method comprises receiving a request to generate a TED, generating a generation schema of the TED specifying two or more sections of the TED for separate generation, inputting a first set of commands instructing one or more LLMs to generate a first set of text for a first section of the generation schema, inputting a second set of commands including at least a portion of the first set of text and instructing the LLMs to generate a second set of text for a second section of the generation schema, and creating the TED based on an aggregation of the first set of text and the second text according to the generation schema.
1 . A computer-implemented method comprising:
receiving, by a system that is connected between a client device and one or more large language models (LLMs), a request to generate a textual electronic document (“TED”) from a client device;
generating, based on the request, a generation schema of the TED, wherein the generation schema specifies two or more sections of the TED that will be separately generated;
inputting, by the system and to the one or more LLMs, a first set of data including (i) the generation schema, (ii) data from the request, and (iii) a first set of commands instructing the one or more LLMs to generate a first set of text for a first section of the generation schema of the TED based on the first set of data;
obtaining, by the system and from the one or more LLMs, a first response to the first set of data, wherein the first response includes the first set of text generated for the first section of the generation schema of the TED;
creating, by the system, a second set of data including (i) the generation schema, (ii) the data from the request, (iii) at least a portion of the first set of text generated for the first section of the generation schema, and (iv) a second set of commands instructing the one or more LLMs to generate a second set of text for a second section of the generation schema of the TED based on the second set of data;
submitting, by the system, the second set of data to the one or more LLMs;
obtaining, by the system, a second response to the second set of data, wherein the second response includes the second set of text generated for the second section of the generation schema of the TED;
creating, by the system, the TED based on an aggregation of the first set of text and the second text according to the generation schema; and
providing, by the system, a graphical representation of the TED to the client device in response to the request.
2 . The computer-implemented method of claim 1 , further comprising:
caching a first state of the TED generation after obtaining the first response, wherein the cached state of the TED includes at least the first set of text; and
maintaining the cached first state of the TED generation in data storage.
3 . The computer-implemented method of claim 2 , further comprising:
receiving an indication of a failed execution of the one or more LLMs during processing of the second set of data;
in response to the indication of failed execution, retrieving the cached first state of the TED generation from the data storage; and
submitting, by the system, the cached first state of the TED generation to the one or more LLMs.
4 . The computer-implemented method of claim 1 , further comprising:
generating, for one of the sections among the two or more sections, two or more subsections;
creating a set of commands for generating a particular subsection among the two or more subsections based on a context for the particular subsection, wherein the context for the particular subsection comprises one or more previously generated subsections in the particular subsection;
submitting the set of commands to the one or more LLMs; and
obtaining, from the one or more LLMs, text of the particular subsection generated based on the set of commands.
5 . The computer-implemented method of claim 1 , further comprising:
providing the first set of text to the client device;
obtaining feedback for the first set of text from the client device; and
determining whether to regenerate the first set of text in accordance with the feedback for the first set of text from the client device.
6 . The computer-implemented method of claim 1 , wherein obtaining the first response to the first set of data comprises:
receiving different responses from each LLM among two or more LLMs; and
selecting, as the first response, a particular response from the different responses in accordance with criteria for the first section.
7 . The computer-implemented method of claim 6 , wherein the criteria for the first section are specified by input from the client device.
8 . The computer-implemented method of claim 1 , further comprising selecting, for the first section, a particular LLM as a target LLM for generating the first set of text of the first section based on the first set of commands, wherein the particular LLM has been finetuned in accordance with a specific task.
9 . The computer-implemented method of claim 1 , wherein generating the generation schema of the TED comprises:
submitting an input comprising information from the request to an LLM with an instruction to generate an outline of the sections of the TED based on the request.
10 . The computer-implemented method of claim 9 , wherein the input further comprises a set of one or more example requests and corresponding TED generation schemas.
11 . A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
receiving, by a system that is connected between a client device and one or more large language models (LLMs), a request to generate a textual electronic document (“TED”) from a client device;
generating, based on the request, a generation schema of the TED, wherein the generation schema specifies two or more sections of the TED that will be separately generated;
inputting, by the system and to the one or more LLMs, a first set of data including (i) the generation schema, (ii) data from the request, and (iii) a first set of commands instructing the one or more LLMs to generate a first set of text for a first section of the generation schema of the TED based on the first set of data;
obtaining, by the system and from the one or more LLMs, a first response to the first set of data, wherein the first response includes the first set of text generated for the first section of the generation schema of the TED;
creating, by the system, a second set of data including (i) the generation schema, (ii) the data from the request, (iii) at least a portion of the first set of text generated for the first section of the generation schema, and (iv) a second set of commands instructing the one or more LLMs to generate a second set of text for a second section of the generation schema of the TED based on the second set of data;
submitting, by the system, the second set of data to the one or more LLMs;
obtaining, by the system, a second response to the second set of data, wherein the second response includes the second set of text generated for the second section of the generation schema of the TED;
creating, by the system, the TED based on an aggregation of the first set of text and the second text according to the generation schema; and
providing, by the system, a graphical representation of the TED to the client device in response to the request.
12 . The system of claim 11 , wherein the operations further comprise:
caching a first state of the TED generation after obtaining the first response, wherein the cached state of the TED includes at least the first set of text; and
maintaining the cached first state of the TED generation in data storage.
13 . The system of claim 12 , wherein the operations further comprise:
receiving an indication of a failed execution of the one or more LLMs during processing of the second set of data;
in response to the indication of failed execution, retrieving the cached first state of the TED generation from the data storage; and
submitting, by the system, the cached first state of the TED generation to the one or more LLMs.
14 . The system of claim 11 , wherein the operations further comprise:
generating, for one of the sections among the two or more sections, two or more subsections;
creating a set of commands for generating a particular subsection among the two or more subsections based on a context for the particular subsection, wherein the context for the particular subsection comprises one or more previously generated subsections in the particular subsection;
submitting the set of commands to the one or more LLMs; and
obtaining, from the one or more LLMs, text of the particular subsection generated based on the set of commands.
15 . The system of claim 11 , wherein the operations further comprise:
selecting, for the first section, a particular LLM as a target LLM for generating the first set of text of the first section based on the first set of commands, wherein the particular LLM has been finetuned in accordance with a specific task.
16 . A non-transitory computer readable storage medium encoded with a computer program, the program comprising instructions that are operable, when executed by data processing apparatus, to cause the data processing apparatus to perform operations comprising:
receiving, by a system that is connected between a client device and one or more large language models (LLMs), a request to generate a textual electronic document (“TED”) from a client device;
generating, based on the request, a generation schema of the TED, wherein the generation schema specifies two or more sections of the TED that will be separately generated;
inputting, by the system and to the one or more LLMs, a first set of data including (i) the generation schema, (ii) data from the request, and (iii) a first set of commands instructing the one or more LLMs to generate a first set of text for a first section of the generation schema of the TED based on the first set of data;
obtaining, by the system and from the one or more LLMs, a first response to the first set of data, wherein the first response includes the first set of text generated for the first section of the generation schema of the TED;
creating, by the system, a second set of data including (i) the generation schema, (ii) the data from the request, (iii) at least a portion of the first set of text generated for the first section of the generation schema, and (iv) a second set of commands instructing the one or more LLMs to generate a second set of text for a second section of the generation schema of the TED based on the second set of data;
submitting, by the system, the second set of data to the one or more LLMs;
obtaining, by the system, a second response to the second set of data, wherein the second response includes the second set of text generated for the second section of the generation schema of the TED;
creating, by the system, the TED based on an aggregation of the first set of text and the second text according to the generation schema; and
providing, by the system, a graphical representation of the TED to the client device in response to the request.
17 . The non-transitory computer readable storage medium of claim 16 , wherein the operations further comprise:
caching a first state of the TED generation after obtaining the first response, wherein the cached state of the TED includes at least the first set of text; and
maintaining the cached first state of the TED generation in data storage.
18 . The non-transitory computer readable storage medium of claim 17 , wherein the operations further comprise:
receiving an indication of a failed execution of the one or more LLMs during processing of the second set of data;
in response to the indication of failed execution, retrieving the cached first state of the TED generation from the data storage; and
submitting, by the system, the cached first state of the TED generation to the one or more LLMs.
19 . The non-transitory computer readable storage medium of claim 16 , wherein the operations further comprise:
generating, for one of the sections among the two or more sections, two or more subsections;
creating a set of commands for generating a particular subsection among the two or more subsections based on a context for the particular subsection, wherein the context for the particular subsection comprises one or more previously generated subsections in the particular subsection;
submitting the set of commands to the one or more LLMs; and
obtaining, from the one or more LLMs, text of the particular subsection generated based on the set of commands.
20 . The non-transitory computer readable storage medium of claim 16 , wherein the operations further comprise:
selecting, for the first section, a particular LLM as a target LLM for generating the first set of text of the first section based on the first set of commands, wherein the particular LLM has been finetuned in accordance with a specific task.