IP Library › Granted Patent US 12,493,739
Granted Patent B2
US 12,493,739 · App. 18/477,978 · Granted Dec 9, 2025

Generating full designs from text

Inventors: Stefan Daniel Dumitrescu (Valencia, ES); Vlad-Constantin Lungu-Stan (Bucharest, RO); Ionut Mironica (Bucharest, RO); Oliver Brdiczka (San Jose, CA)
Assignee: ADOBE INC.
G06F40/186G06F40/126G06F40/30G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,493,739
App. No.
18/477,978
Granted
Dec 9, 2025
Kind
B2
Abstract

Systems and methods for generating full designs from text include retrieving a plurality of document templates based on a design prompt, and filtering the document templates based on an image prompt. A document template is then selected based on the filtering, and a document is generated based on the document template and the image prompt. Embodiments are further configured to generate content from one or more of the prompts, where the content is included in the final design document.

Claims (54)

1 . A method comprising:

retrieving a plurality of document templates based on a design prompt;

encoding an image prompt to obtain an image prompt embedding;

filtering the plurality of document templates by comparing the image prompt embedding to a template embedding for each of the plurality of document templates, respectively;

selecting a document template based on the filtering; and

generating a document based on the document template and the image prompt.

2 . The method of claim 1 , wherein the retrieving of the plurality of document templates further comprises:

encoding the design prompt to obtain a design prompt embedding; and

comparing the design prompt embedding to a template embedding for each of the plurality of document templates, respectively.

3 . The method of claim 1 , further comprising:

encoding a description of a template to obtain a description embedding;

encoding a caption of a visual element of the template to obtain a caption embedding; and

combining the description embedding and the caption embedding to obtain the template embedding.

4 . The method of claim 1 , further comprising:

receiving a text input for the document, wherein the document is based on the text input.

5 . The method of claim 4 , further comprising:

generating a recommended text field based on the design prompt, wherein the text input corresponds to the recommended text field.

6 . The method of claim 4 , further comprising:

sorting the plurality of document templates based on the text input.

7 . The method of claim 4 , further comprising:

computing a location for the text input based on the text input and the document template, wherein the document includes the text input at the location.

8 . The method of claim 1 , further comprising:

displaying the plurality of document templates; and

receiving a selection input, wherein the document template is selected based on the selection input.

9 . The method of claim 1 , wherein:

the image prompt comprises an image, and wherein the document comprises the image.

10 . The method of claim 1 , wherein:

the image prompt comprises a description of an image.

11 . The method of claim 10 , further comprising:

generating an image based on the image prompt, wherein the document includes the generated image.

12 . A non-transitory computer readable medium storing code, the code comprising instructions executable by a processor to:

retrieve a plurality of document templates based on a design prompt;

encode an image prompt to obtain an image prompt embedding;

filter the plurality of document templates by comparing the image prompt embedding to a template embedding for each of the plurality of document templates, respectively;

select a document template based on the filtering; and

generate a document based on the document template and the image prompt.

13 . The non-transitory computer readable medium of claim 12 , the code further comprising instructions executable by a processor to:

encode the design prompt to obtain a design prompt embedding; and

compare the design prompt embedding to a template embedding for each of the plurality of document templates, respectively, wherein the retrieval of the plurality of document templates is based on the comparison.

14 . The non-transitory computer readable medium of claim 12 , the code further comprising instructions executable by a processor to:

receive a text input for the document, wherein the document is based on the text input.

15 . The non-transitory computer readable medium of claim 14 , the code further comprising instructions executable by a processor to:

generate a recommended text field based on the design prompt, wherein the text input corresponds to the recommended text field.

16 . An apparatus comprising:

at least one processor;

at least one memory including instructions executable by the at least one processor;

a template database storing a plurality of document templates;

an intent component configured to encode a design prompt to obtain a design prompt embedding, and further configured to encode an image prompt to obtain an image prompt embedding;

a filtering component configured to filter the plurality of document templates by comparing the image prompt embedding to a template embedding for each of the plurality of document templates, respectively; and

a document generation component configured to generate a document based on an image prompt and a selected document template of the plurality of document templates.

17 . The apparatus of claim 16 , further comprising:

text field recommender configured to generate a recommended text field based on the design prompt.

18 . The apparatus of claim 16 , further comprising:

an image generation model configured to generate an image based on the image prompt.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2023
From: DUMITRESCU, STEFAN DANIEL; LUNGU-STAN, VLAD-CONSTANTIN; MIRONICA, IONUT; BRDICZKA, OLIVER
To: ADOBE INC.
Reel/Frame 065074/0001 →
Continuity (1)
Related Publication 20250111137A1 · Apr 3, 2025
References Cited (27)
US 8037408B2 · Hartmann · 2011 [cited by examiner]
US 10262062B2 · Chang · 2019 [cited by examiner]
US 10339373B1 · Yellapragada · 2019 [cited by examiner]
US 11868714B2 · Modani · 2024 [cited by examiner]
US 12045735B1 · Parasnis · 2024 [cited by examiner]
US 20160357860A1 · Shmiel · 2016 [cited by examiner]
US 20190251159A1 · Guggilla · 2019 [cited by examiner]
US 20220309277A1 · Shu · 2022 [cited by examiner]
US 20230095089A1 · Kaliyaperumal · 2023 [cited by examiner]
US 20240152695A1 · Shukla · 2024 [cited by examiner]
US 20240242037A1 · Heller · 2024 [cited by examiner]
US 20240265274A1 · Parasnis · 2024 [cited by examiner]
US 20240320444A1 · Maschmeyer · 2024 [cited by examiner]
US 20240427984A1 · Ragan · 2024 [cited by examiner]
US 20250061284A1 · Downs · 2025 [cited by examiner]
US 20250103771A1 · Radford · 2025 [cited by examiner]
CN 112749536A · 2021 [cited by examiner]
CN 114942990A · 2022 [cited by examiner]
CN 117313670A · 2023 [cited by examiner]
CN 117493558A · 2024 [cited by examiner]
WO WO2025136437A2 · 2025 [cited by examiner]
1Chung, et al., “Scaling Instruction-Finetuned Language Models”, arXiv preprint arXiv:2210.11416v5 [cs.LG] Dec. 6, 2022, pp. 1-54. [cited by applicant]
2Li, et al., “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models”, arXiv preprint arXiv:2301.12597v3 [cs.CV] Jun. 15, 2023, found on the internet at https://github.com… [cited by applicant]
3Radford, et al., “Learning Transferable Visual Models From Natural Language Supervision”, arXiv preprint arXiv:2103.00020v1 [cs.CV] Feb. 26, 2021, 48 pages. [cited by applicant]
4Reimers, et al., “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks”, arXiv preprint arXiv:1908.10084v1 [cs.CL] Aug. 27, 2019, 11 pages. [cited by applicant]
5Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, arXiv preprint arXiv:2112.10752v2 [cs.CV] Apr. 13, 2022, pp. 1-45. [cited by applicant]
6Meng, et al., “SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations”, arXiv preprint arXiv:2108.01073v2 [cs.CV] Jan. 5, 2022, pp. 1-33. [cited by applicant]