IP Library › Granted Patent US 12,462,449
Granted Patent B2
US 12,462,449 · App. 18/319,808 · Granted Nov 4, 2025

Mask conditioned image transformation based on a text prompt

Inventors: Ambareesh Revanur (San Jose, CA); Debraj Debashish Basu (Sunnvyale, CA); Shradha Agrawal (Milpitas, CA); Dhwanit Agarwal (San Jose, CA); Deepak Pai (Sunnyvale, CA)
Assignee: Adobe Inc.
G06T11/001G06F40/40G06T7/11G06V10/774G06V10/82G06V20/70G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,449
App. No.
18/319,808
Granted
Nov 4, 2025
Kind
B2
Abstract

In accordance with the described techniques, an image transformation system receives an input image and a text prompt, and leverages a generator network to edit the input image based on the text prompt. The generator network includes a plurality of layers configured to perform respective edits. A plurality of masks are generated based on the text prompt that define local edit regions, respectively, of the input image for respective layers of the generator network. Further, the generator network generates an edited image by editing the input image based on the plurality of masks, the respective edits of the respective layers, and the text prompt.

Claims (50)

1 . A method, comprising:

receiving, by a processing device, a text prompt and an input image by a generator network, the generator network including a plurality of layers configured to perform respective edits for the text prompt at different resolutions;

generating, by the processing device, a plurality of masks defining local edit regions, respectively, of the input image for respective layers of the plurality of layers, the plurality of masks based on the text prompt;

generating, by the processing device using the generator network, an edited image by editing the input image based on the plurality of masks, the respective edits of the respective layers, and the text prompt; and

outputting, by the processing device, the edited image.

2 . The method of claim 1 , wherein the generating the plurality of masks includes segmenting, using a segmentation network, the input image into multiple semantic segments that each identify a different portion of a subject depicted in the input image.

3 . The method of claim 2 , wherein the generating the plurality of masks includes generating a matrix having columns that represent different layers of the generator network, rows that represent different semantic segments of the multiple semantic segments, and entries populated with confidence values indicating degrees of likelihood that the respective layers affect corresponding semantic segments based on the text prompt.

4 . The method of claim 3 , wherein the generating the plurality of masks includes selecting, as the local edit regions for the respective layers, one or more semantic segments having confidence values in respective columns of the matrix that exceed a threshold.

5 . The method of claim 1 , wherein the generating the plurality of masks is performed using convolutional neural networks associated with the respective layers, the generating the plurality of masks further including conditioning the convolutional neural networks on the text prompt and unedited features output by the respective layers.

6 . The method of claim 1 , wherein the generating the edited image includes:

determining latent edit vectors for the respective layers based on the text prompt;

generating combined latent vectors for the respective layers by combining the latent edit vectors with a latent vector that defines the input image; and

editing, by the respective layers, the input image based on the combined latent vectors.

7 . The method of claim 6 , wherein the generating the edited image includes:

outputting, by the plurality of layers, unedited features based on the latent vector;

outputting, by the plurality of layers, edited features based on respective combined latent vectors of the combined latent vectors; and

generating blended features for the plurality of layers by blending the edited features and the unedited features based on the plurality of masks, the blended features including respective edited features in the local edit regions and respective unedited features outside the local edit regions, the edited image incorporating the blended features.

8 . The method of claim 7 , wherein the outputting the unedited features and the outputting the edited features includes conditioning the plurality of layers on the blended features output by previous layers of the generator network.

9 . The method of claim 7 , wherein one or more masks generated for one or more layers are zero masks indicating that the one or more layers do not affect the input image based on the text prompt, and the blended features generated for the one or more layers are the unedited features output by the one or more layers.

10 . The method of claim 6 , wherein the determining the latent edit vectors includes determining, using one or more machine learning mapper models, the latent edit vectors based on the text prompt and the latent vector, the latent edit vectors being dependent on the input image.

11 . The method of claim 6 , wherein the determining the latent edit vectors includes determining a global direction for the latent edit vectors, the latent edit vectors being independent of the input image.

12 . The method of claim 6 , wherein the generating the plurality of masks and the determining the latent edit vectors is performed using one or more machine learning models.

13 . The method of claim 12 , further comprising:

generating an additional edited image by editing the input image based on the respective edits of the plurality of layers and the text prompt without using the plurality of masks;

determining, using a contrastive language-image pre-training model, a first measure of similarity between the edited image and the text prompt and a second measure of similarity between the additional edited image and the text prompt; and

training the one or more machine learning models based on the first and second measures of similarity.

14 . The method of claim 12 , further comprising training the one or more machine learning models based on squared Euclidean norms of the latent edit vectors and a size of the local edit regions in the plurality of masks.

15 . A system, comprising:

a processing device; and

a computer-readable media storing instructions that, responsive to execution by the processing device, cause the processing device to perform operations including:

receiving a text prompt and an input image by a generator network, the generator network including a plurality of layers configured to perform respective edits for the text prompt at different resolutions;

generating, for a layer of the plurality of layers, a mask defining a local edit region of the input image based on the text prompt;

generating a feature of an edited image by editing, using the layer, the input image based on the text prompt and the mask; and

generating the edited image by incorporating the feature into the edited image.

16 . The system of claim 15 , wherein generating the mask includes:

segmenting the input image into semantic segments as part of processing the text prompt; and

selecting at least one semantic segment as the local edit region based on the text prompt and the respective edits of the layer.

17 . The system of claim 16 , wherein the generating the mask includes generating a matrix having columns that represent different layers of the generator network, rows that represent different semantic segments, and entries populated with confidence values indicating degrees of likelihood that respective layers affect corresponding semantic segments based on the text prompt and the respective edits of the respective layers.

18 . The system of claim 17 , wherein the selecting the at least one semantic segment includes selecting at least one entry from among the entries in a column associated with the layer, the at least one semantic segment having a confidence value that exceeds a threshold.

19 . A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

receiving a text prompt and an input image by a generator network, the generator network including a plurality of layers configured to perform respective edits for the text prompt at different resolution;

generating, for a layer of the plurality of layers, a mask defining a local edit region of the input image by conditioning a convolutional neural network associated with the layer on the text prompt and an unedited feature output using the layer;

generating a feature of an edited image by editing, using the layer, the input image based on the text prompt and the mask; and

generating the edited image by incorporating the feature into the edited image.

20 . The non-transitory computer-readable medium of claim 19 , wherein the generating the feature includes:

generating, using the layer, an unedited feature based on a latent vector that defines the input image;

determining a latent edit vector for the layer based on the text prompt;

generating a combined latent vector by combining the latent vector and the latent edit vector;

generating, using the layer, an edited feature based on the combined latent vector; and

generating the feature by blending the unedited feature and the edited feature based on the mask.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2023
From: REVANUR, AMBAREESH; BASU, DEBRAJ DEBASHISH; AGRAWAL, SHRADHA; AGARWAL, DHWANIT; PAI, DEEPAK
To: ADOBE INC.
Reel/Frame 063748/0602 →
Continuity (1)
Related Publication 20240386627A1 · Nov 21, 2024
References Cited (64)
US 20200242774A1 · Park · 2020 [cited by examiner]
US 20210383584A1 · Zhang · 2021 [cited by examiner]
US 20240169622A1 · Xie · 2024 [cited by examiner]
Morita, Ryugo, et al. “Interactive Image Manipulation with Complex Text Instructions.” arXiv preprint arXiv:2211.15352 (2022). (Year: 2022). [cited by examiner]
Couairon, G., Verbeek, J., Schwenk, H., & Cord, M. (2022). Diffedit: Diffusion-based semantic image editing with mask guidance. arXiv preprint arXiv:2210.11427. (Year: 2022). [cited by examiner]
Li, B., Qi, X., Lukasiewicz, T., & Torr, P. H. (2020). Manigan: Text-guided image manipulation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 7880-7889). (Year: 2020). [cited by examiner]
Avrahami, O., Lischinski, D., & Fried, O. (2022). Blended diffusion for text-driven editing of natural images. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 18208-18218). (Yea… [cited by examiner]
Xie, S., Zhang, Z., Lin, Z., Hinz, T., & Zhang, K. (2023). Smartbrush: Text and shape guided object inpainting with diffusion model. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (… [cited by examiner]
Watanabe, Y., Togo, R., Maeda, K., Ogawa, T., & Haseyama, M. (2023). Text-guided image manipulation via generative adversarial network with referring image segmentation-based guidance. IEEE Access, 11, 42534-42545. (Yea… [cited by examiner]
Abdal, Rameen , “CLIP2StyleGAN: Unsupervised Extraction of StyleGAN Edit Directions”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2112.05219.pdf>., D… [cited by applicant]
Abdal, Rameen et al., “Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?”, 2019 IEEE/CVF International Conference on Computer Vision (ICCV) [retrieved Feb. 20, 2023]. Retrieved from the Internet <https… [cited by applicant]
Abdal, Rameen et al., “Styleflow: Attribute-conditioned exploration of stylegan-generated images using conditional continuous normalizing flows”, ACM Transactions on Graphics, vol. 40, No. 3 [retrieved Feb. 20, 2023]. R… [cited by applicant]
Alaluf, Yuval et al., “HyperStyle: StyleGAN Inversion with HyperNetworks for Real Image Editing”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2111.15… [cited by applicant]
Alaluf, Yuval et al., “Only a Matter of Style: Age Transformation Using a Style-Based Regression Model”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/… [cited by applicant]
Alaluf, Yuval et al., “Third Time's the Charm? Image and Video Editing with StyleGAN3”, Advances in Image Manipulation Workshop—ECCV 2022 [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://yuval-alaluf.gith… [cited by applicant]
Avrahami, Omri et al., “Blended Diffusion for Text-driven Editing of Natural Images”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2111.14818.pdf>., N… [cited by applicant]
Bermano, Amit , “State-of-the-art in the architecture, methods and applications of stylegan”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2202.14020.… [cited by applicant]
Brock, Andrew et al., “Large Scale Gan Training for High Fidelity Natural Image Synthesis”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1809.11096.pd… [cited by applicant]
Collins, Edo et al., “Editing in Style: Uncovering the Local Semantics of GANs”, Cornell University arXiv, arXiv.org [retrieved Aug. 9, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2004.14367.pdf>., May 21,… [cited by applicant]
Deng, Jiankang et al., “ArcFace: Additive Angular Margin Loss for Deep Face Recognition”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1801.07698.pdf>… [cited by applicant]
Dong, Hao , “Semantic Image Synthesis via Adversarial Learning”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1707.06873.pdf>., Jul. 21, 2017, 9 Pages. [cited by applicant]
Gal, Rinon et al., “StyleGAN-NADA: CLIP-Guided Domain Adaptation of Image Generators”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2108.00946.pdf>., … [cited by applicant]
Goodfellow, Ian J. et al., “Generative Adversarial Nets”, In: Advances in neural information processing systems (2014) [retrieved Feb. 21, 2023]. Retrieved from the Internet <https://www.cs.utah.edu/˜zhe/teach/archived/… [cited by applicant]
Härkönen, Erik et al., “GANSpace: Discovering Interpretable GAN Controls”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2004.02546.pdf>., Dec. 14, 202… [cited by applicant]
Hou, Xianxu et al., “FEAT: Face Editing with Attention”, Cornell University arXiv, arXiv.org [retrieved Jun. 14, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2202.02713.pdf>., Feb. 6, 2022, 14 Pages. [cited by applicant]
Hyun Lee, Seung , “Sound-Guided Semantic Image Manipulation”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2112.00007.pdf>., Nov. 30, 2021, 11 Pages. [cited by applicant]
Kafri, Omer et al., “StyleFusion: Disentangling Spatial Segments in StyleGAN-Generated Images”, ACM Transactions on Graphics, vol. 41, No. 5 [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/… [cited by applicant]
Karras, Tero et al., “A Style-Based Generator Architecture for Generative Adversarial Networks”, Cornell University arXiv, arXiv.org [retrieved Jun. 28, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1812.049… [cited by applicant]
Karras, Tero et al., “Analyzing and Improving the Image Quality of StyleGAN”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. retrieved from the Internet <https://arxiv.org/pdf/1912.04958.pdf>., Dec. 2019… [cited by applicant]
Karras, Tero et al., “Progressive Growing of GANs for Improved Quality, Stability, and Variation”, Cornell University arXiv, arXiv.org [retrived Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1710.10… [cited by applicant]
Kim, Hyunsu et al., “Exploiting Spatial Dimensions of Latent in GAN for Real-Time Image Editing”, Cornell University arXiv, arXiv.org [retrieved Aug. 9, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2104.147… [cited by applicant]
Kingma, Diederik P. et al., “Adam: A Method for Stochastic Optimization”, Cornell University, arXiv Preprint, arXiv.org [retrieved Aug. 9, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1412.6980.pdf>., Jan. … [cited by applicant]
Kocasari, Umut et al., “StyleMC: Multi-Channel Based Fast Text-Guided Image Generation and Manipulation”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf… [cited by applicant]
Krause, Jonathan et al., “3D Object Representations for Fine-Grained Categorization”, 4th IEEE Workshop on 3D Representation and Recognition, at ICCV 2013 (3dRR-13) [retrieved Feb. 20, 2023]. Retrieved from the Internet… [cited by applicant]
Li, Bowen et al., “ManiGAN: Text-Guided Image Manipulation”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1912.06203.pdf>., Mar. 30, 2020, 16 Pages. [cited by applicant]
Ling, Huan et al., “EditGAN: High-Precision Semantic Image Editing”, Cornell University arXiv, arXiv.org [retrieved Aug. 9, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2111.03186.pdf>., Nov. 4, 2021, 38 Pa… [cited by applicant]
Liu, Yahui et al., “Describe What to Change: A Text-guided Unsupervised Image-to-Image Translation Approach”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org… [cited by applicant]
Nair, Vinod et al., “Rectified Linear Units Improve Restricted Boltzmann Machines”, In Proceedings of the 27th International Conference on Machine Learning (ICML-10) [retrieved Mar. 16, 2023]. Retrieved from the Interne… [cited by applicant]
Nam, Seonghyeon et al., “Text-Adaptive Generative Adversarial Networks: Manipulating Images with Natural Language”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arx… [cited by applicant]
Nichol, Alex et al., “Glide: Towards photorealistic image generation and editing with text-guided diffusion models”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://ar… [cited by applicant]
Nie, Weili et al., “Semi-Supervised StyleGAN for Disentanglement Learning”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://authors.library.caltech.edu/102273/2/2003.0… [cited by applicant]
Pakhomov, Daniil et al., “Segmentation in Style: Unsupervised Semantic Image Segmentation with Stylegan and CLIP”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxi… [cited by applicant]
Parmar, Gaurav et al., “On Aliased Resizing and Surprising Subtleties in GAN Evaluation”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2104.11222.pdf>… [cited by applicant]
Parmar, Gaurav et al., “Spatially-Adaptive Multilayer Selection for GAN Inversion and Editing”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2206.0835… [cited by applicant]
Patashnik, OR et al., “StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2103.17249.pdf>., Mar. 3… [cited by applicant]
Pinkney, Justin et al., “clip2latent: Text driven sampling of a pre-trained StyleGAN using denoising diffusion and CLIP”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https… [cited by applicant]
Radford, Alec et al., “Learning Transferable Visual Models From Natural Language Supervision”, Cornell University arXiv, arXiv.org [retrieved Jun. 29, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2103.00020… [cited by applicant]
Ramesh, Aditya et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, Cornell University arXiv, arXiv.org [retrieved Feb. 28, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2204.06125.pdf… [cited by applicant]
Ramesh, Aditya et al., “Zero-Shot Text-to-Image Generation”, Cornell University arXiv, arXiv.org [retrieved Feb. 21, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2102.12092.pdf>., Feb. 26, 2021, 20 Pages. [cited by applicant]
Reed, Scott et al., “Generative Adversarial Text to Image Synthesis”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1605.05396.pdf>., Jun. 5, 2016, 10 … [cited by applicant]
Richardson, Elad et al., “Encoding in Style: A StyleGAN Encoder for Image-to-Image Translation”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2008.009… [cited by applicant]
Shen, Yujun , “Closed-Form Factorization of Latent Semantics in GANs”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2007.06600.pdf>., Apr. 3, 2021, 9 … [cited by applicant]
Shen, Yujun et al., “Interpreting the Latent Space of GANs for Semantic Face Editing”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1907.10786.pdf>., … [cited by applicant]
Tewari, Ayush et al., “StyleRig: Rigging StyleGAN for 3D Control Over Portrait Images”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2004.00121.pdf>.,… [cited by applicant]
Wang, Zhou et al., “Multiscale structural similarity for image quality assessment”, The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers [retrieved Feb. 20, 2023]. Retrieved from the Internet <http://c… [cited by applicant]
Wu, Zongze et al., “StyleSpace Analysis: Disentangled Controls for StyleGAN Image Generation”, Cornell University arXiv, arXiv.org [retrieved Feb. 21, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2011.12799… [cited by applicant]
Xia, Weihao et al., “TediGAN: Text-Guided Diverse Face Image Generation and Manipulation”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2012.03308.pdf… [cited by applicant]
Xu, Bing et al., “Empirical Evaluation of Rectified Activations in Convolutional Network”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1505.00853.pdf… [cited by applicant]
Xu, Tao et al., “AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks”, Cornell University arXiv, arXiv.org [retrieved Feb. 21, 2023]. Retrieved from the Internet <https://arxi… [cited by applicant]
Zhang, Han et al., “StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://ar… [cited by applicant]
Zhang, Han , “StackGan++: Realistic Image Synthesis with Stacked Generative Adversarial Networks”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1710.1… [cited by applicant]
Zhang, Richard et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”, Cornell University arXiv, arXiv.org [retrieved Mar. 9, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1801.0392… [cited by applicant]
Zhang, Yuxuan et al., “DatasetGAN: Efficient Labeled Data Factory with Minimal Human Effort”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2104.06490.… [cited by applicant]
Zhou, Yufan et al., “TiGAN: Text-Based Interactive Image Generation and Manipulation”, AAAI Technical Track on Computer Vision III, vol. 36 No. 3 [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://ojs.aaai.… [cited by applicant]