IP Library › Granted Patent US 12,248,796
Granted Patent B2
US 12,248,796 · App. 17/384,109 · Granted Mar 11, 2025

Modifying digital images utilizing a language guided image editing model

Inventors: Ning Xu (Milpitas, CA); Zhe Lin (Fremont, CA)
Assignee: Adobe Inc.
G06F9/453G06F40/20G06N3/045G06T11/60G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,796
App. No.
17/384,109
Filed
Jul 23, 2021
Granted
Mar 11, 2025
Kind
B2
Art Unit
2635
USPC
382/156
Abstract

This disclosure describes one or more implementations of systems, non-transitory computer-readable media, and methods that perform language guided digital image editing utilizing a cycle-augmentation generative-adversarial neural network (CAGAN) that is augmented using a cross-modal cyclic mechanism. For example, the disclosed systems generate an editing description network that generates language embeddings which represent image transformations applied between a digital image and a modified digital image. The disclosed systems can further train a GAN to generate modified images by providing an input image and natural language embeddings generated by the editing description network (representing various modifications to the digital image from a ground truth modified image). In some instances, the disclosed systems also utilize an image request attention approach with the GAN to generate images that include adaptive edits in different spatial locations of the image.

Claims (56)

1. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to:

generate a natural language embedding representing a visual modification request from an input natural language text that describes the visual modification request for a digital image;

generate an attention matrix based on correlations between a visual feature map of the digital image and the natural language embedding, the attention matrix comprising an indication of a degree of editing at various locations of the digital image;

modify the visual feature map of the digital image utilizing an expanded natural language embedding that comprises reweighted elements based on the attention matrix to generate a modified visual feature map; and

generate, utilizing a generative adversarial neural network and based on the modified visual feature map, a modified digital image that comprises visual modifications from the visual modification request that vary across one or more spatial locations of the digital image.

2. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the computing device to generate the attention matrix by embedding the visual feature map and the natural language embedding within an embedded space.

3. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the reweighted elements of the expanded natural language embedding based on values from the attention matrix that indicate degrees of edits.

4. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the computing device to:

generate one or more modulation parameters utilizing the expanded natural language embedding; and

generate the modified visual feature map by scaling and shifting the visual feature map of the digital image utilizing the one or more modulation parameters.

5. The non-transitory computer-readable medium of claim 1 ,

wherein: the visual modification request comprises a request to modify at least one of a brightness, a contrast, a hue, a saturation, a tint, or a color within the digital image; and

further comprising instructions that, when executed by the at least one processor, cause the computing device to generate, utilizing the generative adversarial neural network and based on the modified visual feature map, the modified digital image to comprise:

a first modification of the at least one of the brightness, the contrast, the hue, the saturation, the tint, or the color within the digital image at a first spatial location of the digital image; and

a second modification of the at least one of the brightness, the contrast, the hue, the saturation, the tint, or the color within the digital image at a second spatial location of the digital image.

6. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the computing device to adjust one or more parameters corresponding to the attention matrix based on a loss between the modified digital image and a ground truth digital image comprising visual modifications from the visual modification request in different spatial locations of the digital image.

7. A computer-implemented method comprising:

generating, utilizing a generative adversarial neural network, a modified digital image reflecting modifications to a digital image indicated in a visual modification request from an input natural language embedding;

generating an additional natural language embedding that represents visual changes between the modified digital image and the digital image by utilizing an editing description network that outputs natural language embeddings that represent visual changes between digital images;

learning parameters of the editing description network from a comparison between the input natural language embedding and the additional natural language embedding; and

learning parameters of the generative adversarial neural network from multiple variations of the digital image and multiple natural language embeddings generated for the multiple variations of the digital image utilizing the editing description network.

8. The computer-implemented method of claim 7 , further comprising generating the input natural language embedding by utilizing a language encoder and an input natural language text that describes the visual modification request for the digital image.

9. The computer-implemented method of claim 7 , further comprising utilizing the editing description network to model editing operations between a feature map of the modified digital image and a feature map of the digital image to generate the additional natural language embedding.

10. The computer-implemented method of claim 9 , further comprising:

generating the feature map of the modified digital image and the feature map of the digital image utilizing an image encoder;

aligning features of the feature map of the modified digital image and the feature map of the digital image; and

generating the additional natural language embedding based on a merging of the aligned features.

11. The computer-implemented method of claim 7 , further comprising generating a variation for the multiple variations of the digital image by swapping the digital image and the modified digital image to:

generate, utilizing the editing description network with the digital image and the modified digital image, a variation-based natural language embedding that represents visual changes from the modified digital image to the digital image;

generate, utilizing the generative adversarial neural network, the variation-based natural language embedding, and the modified digital image, a variation-based modified digital image that mimics the digital image; and

learn the parameters of the generative adversarial neural network based on a comparison between the variation-based modified digital image and the digital image.

12. The computer-implemented method of claim 7 , further comprising generating a variation for the multiple variations of the digital image by randomly modifying at least one of a brightness, a contrast, a hue, a saturation, a tint, or a color of the digital image to generate a varied-digital image.

13. The computer-implemented method of claim 12 , further comprising:

generating, utilizing the editing description network with the digital image and the varied-digital image, a variation-based natural language embedding that represents visual changes from the digital image to the varied-digital image;

generating, utilizing the generative adversarial neural network, the variation-based natural language embedding, and the digital image, a variation-based modified digital image that mimics the varied-digital image; and

learning the parameters of the generative adversarial neural network based on a comparison between the variation-based modified digital image and the varied-digital image.

14. The computer-implemented method of claim 7 , further comprising learning parameters of the generative adversarial neural network utilizing an augmentation loss from a comparison of the multiple variations of the digital image and modified digital images generated by the generative adversarial neural network based on the multiple natural language embeddings generated for the multiple variations of the digital image.

15. A system comprising:

one or more memory devices comprising:

a digital image;

an input natural language text that describes a visual modification request for the digital image; and

a generative adversarial neural network; and

one or more processors configured to cause the system to:

generate a visual feature map for the digital image;

generate a modified visual feature map by modifying the visual feature map of the digital image utilizing an expanded natural language embedding from the input natural language text that comprises reweighted elements based on an attention matrix, the attention matrix comprising an indication of a degree of editing at various locations of the digital image; and

generate, utilizing the generative adversarial neural network and based on the modified visual feature map, a modified digital image comprising visual modifications from the visual modification request for the digital image.

16. The system of claim 15 , wherein the one or more processors are configured to cause the system to generate the modified digital image to comprise different visual modifications in different spatial locations of the digital image.

17. The system of claim 15 , wherein the generative adversarial neural network is trained utilizing a cyclic mechanism comprising multiple variations of a training digital image and an editing description network that outputs natural language embeddings that represent visual changes between digital images.

18. The system of claim 15 , wherein the one or more processors are configured to cause the system to receive the natural language text from a voice input.

19. The system of claim 15 , wherein the one or more processors are configured to:

display, within a graphical user interface of a client device, the digital image;

receive, from the client device, the natural language text that describes the visual modification request for the digital image; and

display, within the graphical user interface of the client device and in response to receiving the natural language text, the modified digital image comprising the visual modifications from the visual modification request for the digital image.

20. The system of claim 15 , wherein the visual modification request comprises:

at least one of a request to modify a brightness, a contrast, a hue, a saturation, a tint, or a color of the digital image; or

a request to remove an object depicted within the digital image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2021
From: XU, NING; LIN, ZHE
To: ADOBE INC.
Reel/Frame 056964/0791 →
Continuity (1)
Related Publication 20230042221A1 · Feb 9, 2023
References Cited (62)
US 20110288854A1 · Glass · 2011 [cited by examiner]
US 20120001934A1 · Bala · 2012 [cited by examiner]
US 20120032968A1 · Fan · 2012 [cited by examiner]
US 20140225899A1 · Bekmambetov · 2014 [cited by examiner]
US 20200053236A1 · Tsujii · 2020 [cited by examiner]
US 20200134090A1 · Mankovskii · 2020 [cited by examiner]
US 20200334486A1 · Joseph · 2020 [cited by examiner]
US 20210035554A1 · Iwase · 2021 [cited by examiner]
US 20210073191A1 · Hatami-Hanza · 2021 [cited by examiner]
US 20210341989A1 · Chen · 2021 [cited by examiner]
US 20220005235A1 · Gou · 2022 [cited by examiner]
US 20220245109A1 · Hatami-Hanza · 2022 [cited by examiner]
US 20220399017A1 · Xu · 2022 [cited by examiner]
WO WO2019073267A1 · 2019 [cited by examiner]
Bowen Li et al., “Controllable Text-to-Image Generation,” Dec. 14, 2019, Advances in Neural Information Processing Systems 32 (NeurIPS 2019),pp. 1-9. [cited by examiner]
Ramesh Manuvinakurike et al., “Edit me: A Corpus and a Framework for Understanding Natural Language Image Editing,” May 2018, Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LR… [cited by examiner]
Jing Shi et al., “A Benchmark and Baseline for Language-Driven Image Editing,” Nov. 2020, Proceedings of the Asian Conference on Computer Vision (ACCV), 2020, pp. 1-13. [cited by examiner]
Ramakrishna Vedantam et al., “CIDEr: Consensus-based Image Description Evaluation,” Jun. 2015, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 4566-4571. [cited by examiner]
Yuanming Hu et al., “Exposure: A White-Box Photo Post-Processing Framework,” May 2018, ACM Transactions on Graphics, vol. 37, No. 2, Article 26,pp. 26:1-26:16. [cited by examiner]
Hao Tan et al., “Expressing Visual Relationships via Language,” Jun. 19, 2019,mularXiv:1906.07689v2,pp. 1-9. [cited by examiner]
Tao Xu et al., “AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks,” Jun. 2018, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp… [cited by examiner]
Alaaeldin El-Nouby et al., “Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic Instruction,” Oct. 2018, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), … [cited by examiner]
Scott Reed, Zeynep Akata et al., “Generative Adversarial Text to Image Synthesis,” Jun. 11, 2016, Proceedings of the 33 rd International Conference on Machine Learning, New York, NY, USA, 2016. JMLR: W&CP vol. 48,pp. 1-… [cited by examiner]
Hai Wang et al.,“Learning to Globally Edit Images with Textual Description,” Oct. 13, 2018, arXiv:1810.05786v1,pp. 1-17. [cited by examiner]
Bowen Li et al., “Lightweight Generative Adversarial Networks for Text-Guided Image Manipulation,” 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada,pp. 1-9 [cited by examiner]
Tingting Qiao et al., “MirrorGAN: Learning Text-to-image Generation by Redescription,” Jun. 2019, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 1505-1510. [cited by examiner]
Seonghyeon Nam et al.,“Text-Adaptive Generative Adversarial Networks: Manipulating Images with Natural Language, ” 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montréal, Canada,pp. 1-8. [cited by examiner]
Wenbo Li et al., “Object-driven Text-to-Image Synthesis via Adversarial Training,” Jun. 2019, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 12174-12181. [cited by examiner]
Bowen Li et al., “ManiGAN: Text-Guided Image Manipulation,” Jun. 2020, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 7880-7887. [cited by examiner]
Satanjeev Banerjee and Alon Lavie. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for mach… [cited by applicant]
Madimir Bychkovsky, Sylvain Paris, Eric Chan, and Fredo Durand. Learning photographic global tonal adjustment with a database of input / output image pairs. In The Twenty-Fourth IEEE Conference on Computer Vision and Pa… [cited by applicant]
Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In Proceedings of the IEEE conference on co… [cited by applicant]
Hao Dong, Simiao Yu, Chao Wu, and Yike Guo. Semantic image synthesis via adversarial learning. In Proceedings of the IEEE International Conference on Computer Vision, pp. 5706-5714, 2017. [cited by applicant]
Alaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, De- von Hjelm, Layla El Asri, Samira Ebrahimi Kahou, Yoshua Bengio, and Graham W Taylor. Tell, draw, and repeat: Generating and modifying images based on continual ling… [cited by applicant]
Chen Gao, Yunpeng Chen, Si Liu, Zhenxiong Tan, and Shuicheng Yan. Adversarialnas: Adversarial neural architecture search for gans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp… [cited by applicant]
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in neural information processing systems, pp. 267… [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778, 2016. [cited by applicant]
Yuanming Hu, Hao He, Chenxi Xu, Baoyuan Wang, and Stephen Lin. Exposure: A white-box photo post-processing framework. ACM Transactions on Graphics (TOG), 37(2):1-17, 2018. [cited by applicant]
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1125-… [cited by applicant]
Bowen Li, Xiaojuan Qi, Thomas Lukasiewicz, and Philip HS Torr. Manigan: Text-guided image manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7880-7889, 2020. [cited by applicant]
Wenbo Li, Pengchuan Zhang, Lei Zhang, Qiuyuan Huang, Xiaodong He, Siwei Lyu, and Jianfeng Gao. Object-driven text-to-image synthesis via adversarial training. In Proceedings of the IEEE Conference on Computer Vision and… [cited by applicant]
Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. InText summarization branches out, pp. 74-81, 2004. [cited by applicant]
Ramesh Manuvinakurike, Jacqueline Brixey, Trung Bui, Walter Chang, Doo Soon Kim, Ron Artstein, and Kallirroi Georgila. Edit me: A corpus and a framework for understanding natural language image editing. In Proceedings o… [cited by applicant]
Mehdi Mirza and Simon Osindero. Conditional generative adversarial nets.arXiv preprint arXiv:1411.1784, 2014. [cited by applicant]
Seonghyeon Nam, Yunji Kim, and Seon Joo Kim. Text-adaptive generative adversarial networks: Manipulating images with natural language. In Advances in neural information processing systems, pp. 42-51, 2018. [cited by applicant]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp… [cited by applicant]
Jongchan Park, Joon-Young Lee, Donggeun Yoo, and In So Kweon. Distort-and-recover: Color enhancement using deep reinforcement learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p… [cited by applicant]
Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2337-2346… [cited by applicant]
Tingting Qiao, Jing Zhang, Duanqing Xu, and Dacheng Tao. Mirrorgan: Learning text-to-image generation by redescription. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1505-1514… [cited by applicant]
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. Generative adversarial text to image synthesis.arXiv preprint arXiv:1605.05396, 2016. [cited by applicant]
Jing Shi, Ning Xu, Trung Bui, Franck Dernoncourt, Zheng Wen, and Chenliang Xu. A benchmark and baseline for language-driven image editing.arXiv preprint arXiv:2010.02330, 2020 part 1. [cited by applicant]
Jing Shi, Ning Xu, Trung Bui, Franck Dernoncourt, Zheng Wen, and Chenliang Xu. A benchmark and baseline for language-driven image editing.arXiv preprint arXiv:2010.02330, 2020 part 2. [cited by applicant]
Hao Tan, Franck Dernoncourt, Zhe Lin, Trung Bui, and Mohit Bansal. Expressing visual relationships via language. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 1873-1883,… [cited by applicant]
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4566-4575, 2015. [cited by applicant]
Hai Wang, Jason D Williams, and SingBing Kang. Learning to globally edit images with textual description.arXiv preprint arXiv:1810.05786, 2018. [cited by applicant]
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In Proceedings of the IEEE con… [cited by applicant]
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiao-gang Wang, Xiaolei Huang, and Dimitris N Metaxas. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In Proceedings of t… [cited by applicant]
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiao-gang Wang, Xiaolei Huang, and Dimitris N Metaxas. Stack-gan++: Realistic image synthesis with stacked generative adversarial networks. IEEE transactions on pattern a… [cited by applicant]
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiao-gang Wang, Xiaolei Huang, and Dimitris N Metaxas. Stack-gan++: Realistic image synthesis with stacked generative adversarial networks. IEEE transactions on pattern a… [cited by applicant]
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiao-gang Wang, Xiaolei Huang, and Dimitris N Metaxas. Stack-gan++: Realistic image synthesis with stacked generative adversarial networks. IEEE transactions on pattern a… [cited by applicant]
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision, pp. … [cited by applicant]
Zhen Zhu, Tengteng Huang, Baoguang Shi, Miao Yu, Bofei Wang, and Xiang Bai. Progressive pose attention transfer for person image generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogniti… [cited by applicant]