IP Library › Granted Patent US 12,639,866
Granted Patent B2
US 12,639,866 · App. 18/484,512 · Granted May 26, 2026

Pipeline for generating editable graphic designs from natural language prompts

Inventors: Sumithra Bhakthavatsalam (Kirkland, WA); Gaurav Vinayak Tendolkar (Reston, VA)
Assignee: Microsoft Technology Licensing, LLC
G06T11/60G06F40/30G06F40/40G06V30/19
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,866
App. No.
18/484,512
Granted
May 26, 2026
Kind
B2
Abstract

A device includes a processor, and a memory storing executable instructions which, when executed by the processor, cause the processor alone or in combination with other processors to perform the following functions: receive textual user input from a user describing a design to be generated; implement a first prompt generator to generate a first prompt for a Large Language Model (LLM) to restructure the user input; and implement a second prompt generator to generate a second prompt for a text-to-image model using output of the LLM to produce, the second prompt to prompt the text-to-image model to produce a proposed design based on the user input. The proposed design is provided to the user via an application comprising controls for further editing the proposed design.

Claims (35)

1 . A data processing system comprising:

a processor, and

a memory storing executable instructions which, when executed by the processor, cause the processor alone or in combination with other processors to perform operations of:

receive textual user input from a user, the textual user input including a natural language user description of a design to be generated;

generate, by a first prompt generator, a first prompt for a Large Language Model (LLM) by structuring the first prompt that specifies the LLM to extract values from the natural language user description for a set of fields in a semantic frame;

generate, by a second prompt generator, a second prompt for a text-to-image model by inserting the values from the semantic frame into a template selected, at random or based on training, from a template database, wherein the second prompt specifies the text-to-image model to produce a proposed design based on the selected template; and

provide the proposed design to the user via an application comprising controls for further editing the proposed design.

2 . The system of claim 1 , wherein the LLM comprises a generative pre-trained transformer (GPT) model.

3 . The system of claim 1 , wherein the text-to-image model comprises a DALL-E model or a Stable Diffusion model.

4 . The system of claim 1 , the processor further to submit the proposed design and instructions derived from the textual user input to a text placement model so as to prompt the text placement model to provide a position for text in the proposed design.

5 . The system of claim 4 , the processor further to call an Optical Character Recognition (OCR) service to check the proposed design for the text before submitting the proposed design to the text placement model.

6 . The system of claim 5 , the processor further to discard the proposed design when the text is identified in the proposed design by the OCR service.

7 . The system of claim 1 , further comprising a server that comprises the processor, the memory, the first prompt generator, and the second prompt generator to provide a design suggestion service via a network to a user terminal.

8 . A method of providing a design suggestion service based on user input, the method comprising:

receiving textual user input from a user, the textual user input including a natural language user description of a design to be generated;

generating, by a first prompt generator, a first prompt for a Large Language Model (LLM) by structuring the first prompt that specifies the LLM to extract values from the natural language user description for a set of fields in a semantic frame;

generating, by a second prompt generator, a second prompt for a text-to-image model by inserting the values from the semantic frame into a template selected, at random or based on training, from a template database, wherein the second prompt specifies the text-to-image model to produce a proposed design based on the user input; and

providing the proposed design to a productivity application of the user, the productivity application comprising controls for further editing the proposed design.

9 . The method of claim 8 , wherein the LLM comprises a generative pre-trained transformer (GPT) model.

10 . The method of claim 8 , wherein the text-to-image model comprises a DALL-E model or a Stable Diffusion model.

11 . The method of claim 8 , further comprising submitting the proposed design and instructions derived from the user input to a text placement model so as to prompt the text placement model to provide a position for text in the proposed design.

12 . The method of claim 11 , further comprising calling an Optical Character Recognition (OCR) service to check the proposed design for the text before submitting the proposed design to the text placement model.

13 . The method of claim 12 , further comprising discarding the proposed design when the text is identified in the proposed design by the OCR service.

14 . The method of claim 13 , further comprising generating a new proposed design after discarding the proposed design in which the text was identified.

15 . The method of claim 8 , further comprising providing the design suggestion service via a network to a user terminal from where the user input is received.

16 . The method of claim 8 , further comprising providing the design suggestion service with components of the productivity application.

17 . A device comprising:

a processor, and

a memory storing executable instructions which, when executed by the processor, cause the processor alone or in combination with other processors to perform functions of:

receive textual user input from a user describing a design to be generated;

generate, by a first prompt generator, a first prompt for a Large Language Model (LLM) by structuring the first prompt that specifies the LLM to extract values from a natural language user description for a set of fields in a semantic frame;

generate, by a second prompt generator, a second prompt for a text-to-image model by inserting the values from the semantic frame into a template selected, at random or based on training, from a template database, wherein the second prompt specifies the text-to-image model to produce a proposed design based on the selected template;

submit the proposed design along with instructions derived from the textual user input to a text placement model to prompt the text placement model to provide a position for text in the proposed design; and

provide the proposed design with the text to the user via an application comprising controls for further editing the proposed design.

18 . The device of claim 17 , wherein the text placement model further provides layering and bounding boxes of elements of the proposed design with the text to support editing of the proposed design and the text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2023
From: BHAKTHAVATSALAM, SUMITHRA; TENDOLKAR, GAURAV VINAYAK
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065177/0752 →
Continuity (1)
Related Publication 20250124622A1 · Apr 17, 2025
References Cited (3)
US 20250111137A1 · Dumitrescu · 2025 [cited by examiner]
US 20250111139A1 · Mironica · 2025 [cited by examiner]
“Microsoft Designer”, Retrieved from: https://designer.microsoft.com/, Retrieved Date: May 25, 2023, 03 Pages. [cited by applicant]