Pipeline for generating editable graphic designs from natural language prompts
A device includes a processor, and a memory storing executable instructions which, when executed by the processor, cause the processor alone or in combination with other processors to perform the following functions: receive textual user input from a user describing a design to be generated; implement a first prompt generator to generate a first prompt for a Large Language Model (LLM) to restructure the user input; and implement a second prompt generator to generate a second prompt for a text-to-image model using output of the LLM to produce, the second prompt to prompt the text-to-image model to produce a proposed design based on the user input. The proposed design is provided to the user via an application comprising controls for further editing the proposed design.
1 . A data processing system comprising:
a processor, and
a memory storing executable instructions which, when executed by the processor, cause the processor alone or in combination with other processors to perform operations of:
receive textual user input from a user, the textual user input including a natural language user description of a design to be generated;
generate, by a first prompt generator, a first prompt for a Large Language Model (LLM) by structuring the first prompt that specifies the LLM to extract values from the natural language user description for a set of fields in a semantic frame;
generate, by a second prompt generator, a second prompt for a text-to-image model by inserting the values from the semantic frame into a template selected, at random or based on training, from a template database, wherein the second prompt specifies the text-to-image model to produce a proposed design based on the selected template; and
provide the proposed design to the user via an application comprising controls for further editing the proposed design.
2 . The system of claim 1 , wherein the LLM comprises a generative pre-trained transformer (GPT) model.
3 . The system of claim 1 , wherein the text-to-image model comprises a DALL-E model or a Stable Diffusion model.
4 . The system of claim 1 , the processor further to submit the proposed design and instructions derived from the textual user input to a text placement model so as to prompt the text placement model to provide a position for text in the proposed design.
5 . The system of claim 4 , the processor further to call an Optical Character Recognition (OCR) service to check the proposed design for the text before submitting the proposed design to the text placement model.
6 . The system of claim 5 , the processor further to discard the proposed design when the text is identified in the proposed design by the OCR service.
7 . The system of claim 1 , further comprising a server that comprises the processor, the memory, the first prompt generator, and the second prompt generator to provide a design suggestion service via a network to a user terminal.
8 . A method of providing a design suggestion service based on user input, the method comprising:
receiving textual user input from a user, the textual user input including a natural language user description of a design to be generated;
generating, by a first prompt generator, a first prompt for a Large Language Model (LLM) by structuring the first prompt that specifies the LLM to extract values from the natural language user description for a set of fields in a semantic frame;
generating, by a second prompt generator, a second prompt for a text-to-image model by inserting the values from the semantic frame into a template selected, at random or based on training, from a template database, wherein the second prompt specifies the text-to-image model to produce a proposed design based on the user input; and
providing the proposed design to a productivity application of the user, the productivity application comprising controls for further editing the proposed design.
9 . The method of claim 8 , wherein the LLM comprises a generative pre-trained transformer (GPT) model.
10 . The method of claim 8 , wherein the text-to-image model comprises a DALL-E model or a Stable Diffusion model.
11 . The method of claim 8 , further comprising submitting the proposed design and instructions derived from the user input to a text placement model so as to prompt the text placement model to provide a position for text in the proposed design.
12 . The method of claim 11 , further comprising calling an Optical Character Recognition (OCR) service to check the proposed design for the text before submitting the proposed design to the text placement model.
13 . The method of claim 12 , further comprising discarding the proposed design when the text is identified in the proposed design by the OCR service.
14 . The method of claim 13 , further comprising generating a new proposed design after discarding the proposed design in which the text was identified.
15 . The method of claim 8 , further comprising providing the design suggestion service via a network to a user terminal from where the user input is received.
16 . The method of claim 8 , further comprising providing the design suggestion service with components of the productivity application.
17 . A device comprising:
a processor, and
a memory storing executable instructions which, when executed by the processor, cause the processor alone or in combination with other processors to perform functions of:
receive textual user input from a user describing a design to be generated;
generate, by a first prompt generator, a first prompt for a Large Language Model (LLM) by structuring the first prompt that specifies the LLM to extract values from a natural language user description for a set of fields in a semantic frame;
generate, by a second prompt generator, a second prompt for a text-to-image model by inserting the values from the semantic frame into a template selected, at random or based on training, from a template database, wherein the second prompt specifies the text-to-image model to produce a proposed design based on the selected template;
submit the proposed design along with instructions derived from the textual user input to a text placement model to prompt the text placement model to provide a position for text in the proposed design; and
provide the proposed design with the text to the user via an application comprising controls for further editing the proposed design.
18 . The device of claim 17 , wherein the text placement model further provides layering and bounding boxes of elements of the proposed design with the text to support editing of the proposed design and the text.