IP Library › Granted Patent US 12,743,606
Granted Patent B2
US 12,743,606 · App. 18/619,587 · Granted Sep 22, 2026

Generating synthetic digital images utilizing a text-to-image generation neural network with localized constraints

Inventors: Weixi Feng (Goleta, CA); Yijun Li (Seattle, WA); Trung Bui (San Jose, CA); Tobias Hinz (Ulm, DE); Scott Cohen (Sunnyvale, CA); Quan Tran (San Jose, CA); Jianming Zhang (Campbell, CA); Handong Zhao (San Jose, CA); Franck Dernoncourt (Seattle, WA)
Assignee: Adobe Inc.
G06N3/0455G06T11/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,606
App. No.
18/619,587
Granted
Sep 22, 2026
Kind
B2
Abstract

Methods, systems, and non-transitory computer readable storage media are disclosed for generating digital images via a generative neural network with localized constraints. The disclosed system generates, utilizing one or more encoder neural networks, a sequence of embeddings comprising a prompt embedding representing a text prompt and an object text embedding representing a phrase indicating an object in the text prompt. The disclosed system generates, utilizing the one or more encoder neural networks, a visual embedding representing an object image corresponding to the object. The disclosed system determines a modified sequence of embeddings by replacing the object text embedding with the visual embedding in the sequence of embeddings. The disclosed system also generates, utilizing a generative neural network, a synthetic digital image from the modified sequence of embeddings comprising the visual embedding.

Claims (74)

1 . A computer-implemented method comprising:

generating, utilizing one or more encoder neural networks, a sequence of embeddings comprising a prompt embedding representing a text prompt and an object text embedding representing a phrase in the text prompt, wherein the phrase indicates an object;

generating, utilizing the one or more encoder neural networks, a visual embedding representing an object image corresponding to the object;

determining a modified sequence of embeddings by replacing the object text embedding with the visual embedding in the sequence of embeddings; and

generating, utilizing a generative neural network, a synthetic digital image from the modified sequence of embeddings comprising the visual embedding.

2 . The computer-implemented method of claim 1 , wherein generating the sequence of embeddings comprises:

determining, from the text prompt, a plurality of phrases indicating a plurality of objects in the text prompt;

generating, based on the plurality of phrases, the object text embedding representing the phrase indicating the object; and

generating, based on the plurality of phrases, an additional object text embedding representing an additional phrase indicating an additional object.

3 . The computer-implemented method of claim 1 , wherein generating the visual embedding comprises generating the object image comprising an example object based on the phrase indicating the object.

4 . The computer-implemented method of claim 3 , further comprising:

generating an additional object image comprising an additional example object based on an additional phrase indicating an additional object in the text prompt; and

generating an additional visual embedding representing the additional object image corresponding to the additional object.

5 . The computer-implemented method of claim 1 , wherein determining the modified sequence of embeddings comprises:

determining a position of the object text embedding in the sequence of embeddings;

removing the object text embedding from the sequence of embeddings; and

inserting the visual embedding into the sequence of embeddings at the position.

6 . The computer-implemented method of claim 1 , wherein:

generating the sequence of embeddings comprises generating the object text embedding in a feature space; and

generating the visual embedding comprises generating the visual embedding in the feature space of the object text embedding.

7 . The computer-implemented method of claim 1 , further comprising:

generating, for a caption corresponding to an additional digital image, an additional sequence of embeddings comprising a plurality of object text embeddings representing a plurality of phrases in the caption, the plurality of phrases indicating a plurality of objects in the additional digital image;

generating a plurality of visual embeddings representing the plurality of objects in the additional digital image;

determining an additional modified sequence of embeddings by replacing the plurality of object text embeddings with the plurality of visual embeddings; and

adjusting parameters of the generative neural network by reducing an output of a loss function based on the additional modified sequence of embeddings.

8 . The computer-implemented method of claim 7 , wherein adjusting the parameters of the generative neural network comprises:

determining a plurality of ground-truth masks corresponding to the plurality of objects of the additional digital image;

determining, utilizing the loss function, a localization loss based on comparisons between the plurality of ground-truth masks and a plurality of cross-attention maps corresponding to the plurality of visual embeddings; and

adjusting the parameters of the generative neural network according to the localization loss.

9 . A system comprising:

one or more memory devices; and

one or more processors configured to cause the system to:

generate, utilizing one or more encoder neural networks, a sequence of embeddings comprising a first object text embedding representing a first phrase in a text prompt, wherein the first phrase indicates a first object, and a second object text embedding representing a second phrase in the text prompt, wherein the second phrase indicates a second object;

generate, utilizing the one or more encoder neural networks, a first visual embedding representing a first object image corresponding to the first object and a second visual embedding representing a second object image corresponding to the second object;

determine a modified sequence of embeddings by replacing, in the sequence of embeddings, the first object text embedding with the first visual embedding and the second object text embedding with the second visual embedding; and

generate, utilizing a generative neural network, a synthetic digital image from the modified sequence of embeddings comprising the first visual embedding and the second visual embedding.

10 . The system of claim 9 , wherein the one or more processors are further configured to generate the sequence of embeddings by generating, utilizing the one or more encoder neural networks, a prompt embedding representing the text prompt in a feature space corresponding to the first object text embedding and the second object text embedding.

11 . The system of claim 9 , wherein the one or more processors are further configured to generate the sequence of embeddings by parsing the text prompt to:

determine the first object and one or more visual attributes of the first object; and

determine the second object and one or more visual attributes of the second object.

12 . The system of claim 11 , wherein the one or more processors are further configured to:

determine, based on the first phrase, the first object image comprising a first example object including the one or more visual attributes of the first object; and

determine, based on the second phrase, the second object image comprising a second example object including the one or more visual attributes of the second object.

13 . The system of claim 9 , wherein the one or more processors are further configured to determine the modified sequence of embeddings by:

determining a first location in the sequence of embeddings corresponding to the first object text embedding and a second location in the sequence of embeddings corresponding to the second object text embedding;

removing the first object text embedding and the second object text embedding from the sequence of embeddings; and

inserting the first visual embedding at the first location and the second visual embedding at the second location.

14 . The system of claim 9 , wherein the one or more processors are further configured to generate the synthetic digital image by providing the modified sequence of embeddings with a noised image embedding to the generative neural network.

15 . The system of claim 9 , wherein the one or more processors are further configured to:

generate a plurality of visual embeddings from a plurality of object text embeddings representing phrases in captions of objects in a training digital image;

determine an additional modified sequence of embeddings by replacing the plurality of object text embeddings with the plurality of visual embeddings in a corresponding sequence of embeddings; and

adjust parameters of the generative neural network by reducing an output of a loss function based on the additional modified sequence of embeddings.

16 . The system of claim 15 , wherein the one or more processors are further configured to adjust the parameters of the generative neural network:

determine a plurality of ground-truth masks corresponding to the objects of the training digital image;

determine cross-attention maps generated by the generative neural network for the plurality of visual embeddings of the additional modified sequence of embeddings; and

adjust the parameters of the generative neural network to reduce the output of the loss function based on comparisons between the plurality of ground-truth masks and the cross-attention maps.

17 . A non-transitory computer readable medium storing executable instructions which, when executed by at least one processing device, cause the at least one processing device to perform operations comprising:

determining, by parsing a text prompt for generating or modifying a digital image, a plurality of phrases corresponding to a plurality of objects;

generating, utilizing one or more encoder neural networks, a sequence of embeddings comprising a plurality of object text embeddings representing the plurality of phrases in the text prompt, wherein the plurality of phrases corresponds to the plurality of objects;

generating, utilizing the one or more encoder neural networks, a plurality of visual embeddings representing a plurality of object images corresponding to the plurality of objects;

determining a modified sequence of embeddings by replacing the plurality of object text embeddings with the plurality of visual embeddings in the sequence of embeddings; and

generating, utilizing a generative neural network, a synthetic digital image from the modified sequence of embeddings comprising the plurality of visual embeddings.

18 . The non-transitory computer readable medium of claim 17 , wherein:

parsing the text prompt comprises:

determining a first phrase corresponding to a first object of the digital image and one or more visual attributes of the first object; and

determining a second phrase corresponding to a second object of the digital image and one or more visual attributes of the second object; and

generating the sequence of embeddings comprises generating a first object text embedding representing the first phrase and a second object text embedding representing the second phrase.

19 . The non-transitory computer readable medium of claim 18 , wherein generating the plurality of visual embeddings comprises:

determining a first object image comprising a first example object including the one or more visual attributes of the first object;

determining a second object image comprising a second example object including the one or more visual attributes of the second object; and

generating, utilizing the one or more encoder neural networks, a first visual embedding representing the first object image and a second visual embedding representing the second object image.

20 . The non-transitory computer readable medium of claim 18 , wherein determining the modified sequence of embeddings comprises:

determining locations corresponding to the plurality of object text embeddings in the sequence of embeddings; and

replacing the plurality of object text embeddings with the plurality of visual embeddings at the locations in the sequence of embeddings.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2024
From: LI, YIJUN; FENG, WEIXI; BUI, TRUNG; HINZ, TOBIAS; COHEN, SCOTT; TRAN, QUAN; ZHANG, JIANMING; ZHAO, HANDONG; DERNONCOURT, FRANCK
To: ADOBE INC.
Reel/Frame 066976/0447 →
Continuity (1)
Related Publication 20250307606A1 · Oct 2, 2025
References Cited (13)
US 20230377226A1 · Saharia · 2023 [cited by examiner]
US 20240169662A1 · Sajjadi · 2024 [cited by examiner]
US 20240282131A1 · Ren · 2024 [cited by examiner]
US 20240378256A1 · Sadr · 2024 [cited by examiner]
US 20240394932A1 · Chemerys · 2024 [cited by examiner]
CN 110866958A · 2020 [cited by applicant]
CN 114140666A · 2022 [cited by applicant]
WO 2023225344A1 · 2023 [cited by applicant]
Aditya Ramesh et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents,” arXiv:2204.06125v1, Apr. 13, 2022. [cited by applicant]
Amir Hertz et al. “Prompt-to-prompt image editing with cross attention control.”arXiv preprint arXiv:2208.01626v1, Aug. 2, 2022. [cited by applicant]
Hila Chefer et al., “Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models,” arXiv:2301.13826v2, May 31, 2023. [cited by applicant]
Royi Rassin et al., “Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment,” arXiv:2306.08877v2, Oct. 20, 2023. [cited by applicant]
Search Report as received in United Kingdom Application No. 2501054.7 dated Jul. 25, 2025. [cited by applicant]