IP Library › Granted Patent US 12,518,358
Granted Patent B2
US 12,518,358 · App. 18/178,212 · Granted Jan 6, 2026

Utilizing regularized forward diffusion for improved inversion of digital images

Inventors: Yijun Li (Seattle, WA); Richard Zhang (San Francisco, CA); Krishna Kumar Singh (San Jose, CA); Jingwan Lu (Santa Clara, CA); Gaurav Parmar (Pittsburg, PA); Jun-Yan Zhu (Cambridge, MA)
Assignee: Adobe Inc.
G06T5/70G06F40/126G06T5/50G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,358
App. No.
18/178,212
Granted
Jan 6, 2026
Kind
B2
Abstract

The present disclosure relates to systems, non-transitory computer-readable media, and methods for utilizing machine learning models to generate modified digital images. In particular, in some embodiments, the disclosed systems generate image editing directions between textual identifiers of two visual features utilizing a language prediction machine learning model and a text encoder. In some embodiments, the disclosed systems generated an inversion of a digital image utilizing a regularized inversion model to guide forward diffusion of the digital image. In some embodiments, the disclosed systems utilize cross-attention guidance to preserve structural details of a source digital image when generating a modified digital image with a diffusion neural network.

Claims (49)

1 . A computer-implemented method comprising:

generating, utilizing a diffusion layer of a diffusion neural network, a noise map from a source digital image;

generating a shifted noise map by shifting the noise map by an offset value;

determining a pairwise correlation loss by comparing one or more regions of the noise map and one or more regions of the shifted noise map;

generating a modified noise map based on the pairwise correlation loss; and

generating, from the modified noise map utilizing a denoising layer of the diffusion neural network conditioned with an image editing encoding representing an edit to the source digital image, a modified digital image including the edit to the source digital image while preserving structural details of the source digital image.

2 . The computer-implemented method of claim 1 , wherein generating the modified noise map based on the pairwise correlation loss comprises modifying the noise map to reduce a similarity metric between the one or more regions of the noise map and the one or more regions of the shifted noise map.

3 . The computer-implemented method of claim 1 , wherein determining the pairwise correlation loss comprises:

generating a pyramid of noise maps at different resolutions from the noise map;

generating a pyramid of shifted noise map at the different resolutions from the shifted noise map; and

determining the pairwise correlation loss by comparing the pyramid of noise maps and the pyramid of shifted noise maps.

4 . The computer-implemented method of claim 1 , further comprising:

generating, utilizing one or more subsequent additional diffusion layers of the diffusion neural network, an inversion of the source digital image from the modified noise map; and

generating the modified digital image from the modified noise map by utilizing a plurality of denoising layers of the diffusion neural network to generate the modified digital image from the inversion of the source digital image.

5 . The computer-implemented method of claim 1 , further comprising:

generating, utilizing a subsequent diffusion layer of the diffusion neural network, an additional noise map from the modified noise map;

determining an additional pairwise correlation loss by comparing one or more regions of the additional noise map with one or more regions of an additional shifted noise map; and

generating an additional modified noise map from the additional noise map based on the additional pairwise correlation loss.

6 . The computer-implemented method of claim 5 , further comprising generating the additional shifted noise map by shifting the additional noise map by an additional offset value different than the offset value of the shifted noise map.

7 . The computer-implemented method of claim 1 , wherein generating the modified noise map further comprises:

determining a divergence loss for the noise map relative to a standard distribution;

determining an auto-correlation regularization loss by combining the pairwise correlation loss and the divergence loss; and

generating the modified noise map based on the auto-correlation regularization loss.

8 . The computer-implemented method of claim 7 , wherein combining the pairwise correlation loss and the divergence loss comprises weighting the divergence loss by a first weight.

9 . A system comprising:

one or more memory devices; and

one or more processors coupled to the one or more memory devices that cause the system to perform operations comprising:

generating, utilizing a diffusion layer of a diffusion neural network, a noise map from a source digital image;

determining a pairwise correlation loss by comparing one or more regions of a noise map and one or more regions of a shifted noise map;

determining a divergence loss for the noise map relative to a standard distribution;

determining an auto-correlation regularization loss by combining the pairwise correlation loss and the divergence loss;

generating a modified noise map based on the auto-correlation regularization loss; and

generating, from the modified noise map utilizing a denoising layer of the diffusion neural network conditioned with an image editing encoding representing an edit to the source digital image, a modified digital image including the edit to the source digital image while preserving structural details of the source digital image.

10 . The system of claim 9 , wherein generating the noise map from the source digital image comprises generating, utilizing the diffusion layer of the diffusion neural network, the noise map from an initial latent vector corresponding to the source digital image.

11 . The system of claim 10 , wherein inverting the initial latent vector corresponding to the source digital image to generate the noise map comprises utilizing a deterministic forward diffusion model conditioned on a reference encoding of the source digital image.

12 . The system of claim 11 , wherein the operations further comprise generating, utilizing a text encoder, the reference encoding from an image caption describing the source digital image.

13 . The system of claim 9 , further comprising determining the divergence loss of the noise map relative to a reference mean and a unit variance.

14 . The system of claim 9 , wherein combining the pairwise correlation loss and the divergence loss comprises weighting the divergence loss by a first weight and weighting the pairwise correlation loss by a second weight.

15 . The system of claim 9 , wherein the operations further comprise generating an inversion of the source digital image from the modified noise map utilizing subsequent diffusion layers of the diffusion neural network.

16 . A non-transitory computer readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

generating, utilizing a diffusion layer of a diffusion neural network, a noise map from a source digital image;

generating a shifted noise map by shifting the noise map by an offset value;

determining a pairwise correlation loss by comparing one or more regions of the noise map and one or more regions of the shifted noise map;

generating a modified noise map based on the pairwise correlation loss; and

generating, from the modified noise map utilizing a denoising layer of the diffusion neural network conditioned with an image editing encoding representing an edit to the source digital image, a modified digital image including the edit to the source digital image while preserving structural details of the source digital image.

17 . The non-transitory computer readable medium of claim 16 , wherein generating the shifted noise map comprises randomly sampling the offset value.

18 . The non-transitory computer readable medium of claim 16 , wherein the operations further comprise generating an inversion of the source digital image from the modified noise map utilizing subsequent diffusion layers of the diffusion neural network conditioned by a reference encoding of the source digital image.

19 . The non-transitory computer readable medium of claim 18 , wherein generating the modified digital image from the modified noise map comprises generating, utilizing denoising layers of the diffusion neural network, the modified digital image from the inversion of the source digital image.

20 . The non-transitory computer readable medium of claim 19 , wherein generating the modified digital image from the inversion comprises conditioning the denoising layers of the diffusion neural network with the reference encoding of the source digital image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2023
From: LI, YIJUN; ZHANG, RICHARD; SINGH, KRISHNA KUMAR; LU, JINGWAN; PARMAR, GAURAV; ZHU, JUN-YAN
To: ADOBE INC.
Reel/Frame 062878/0432 →
Continuity (1)
Related Publication 20240338799A1 · Oct 10, 2024
References Cited (81)
US 10552968B1 · Wang et al. · 2020 [cited by applicant]
US 11693637B1 · Singh et al. · 2023 [cited by applicant]
US 11978141B2 · Saharia et al. · 2024 [cited by applicant]
US 11995803B1 · Karpman et al. · 2024 [cited by applicant]
US 20160232142A1 · Melnikov · 2016 [cited by applicant]
US 20210264203A1 · Fuxman et al. · 2021 [cited by applicant]
US 20230185540A1 · Srivastava et al. · 2023 [cited by applicant]
US 20230377226A1 · Saharia et al. · 2023 [cited by applicant]
US 20240119862A1 · Zheng et al. · 2024 [cited by applicant]
US 20240135509A1 · Liu · 2024 [cited by examiner]
US 20240161369A1 · Li et al. · 2024 [cited by applicant]
US 20240185035A1 · Yu et al. · 2024 [cited by applicant]
US 20240265690A1 · Anandkumar et al. · 2024 [cited by applicant]
US 20240282131A1 · Ren et al. · 2024 [cited by applicant]
US 20240338799A1 · Li et al. · 2024 [cited by applicant]
CN 113160084B · 2022 [cited by examiner]
Parmar et al., “Zero-shot Image-to-Image Translation”, arXiv:2302.03027v1 [cs.CV] Feb. 6, 2023. [cited by examiner]
Jeanneret, et al. , “Diffusion Models for Counterfactual Explanations”, arXiv:2203.15636v1 [cs.CV] Mar. 29, 2022. [cited by examiner]
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. arXiv preprint arXiv:2102.12092, 2021. [cited by applicant]
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. [cited by applicant]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language super-… [cited by applicant]
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv… [cited by applicant]
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:2208.01626, 2022. [cited by applicant]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. [cited by applicant]
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion models. arXiv preprint arXiv:2210.09276, 2022. [cited by applicant]
Chenlin Meng, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. Sdedit: Image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073, 2021. [cited by applicant]
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion mod… [cited by applicant]
Daniel Roich, Ron Mokady, Amit H Bermano, and Daniel Cohen-Or. Pivotal tuning for latent-based editing of real images. arXiv preprint arXiv:2106.05744, 2021. [cited by applicant]
David Bau, Hendrik Strobelt, William Peebles, Jonas Wulff, Bolei Zhou, Jun-Yan Zhu, and Antonio Torralba. Semantic photo manipulation with a generative image prior. 38(4):1- 11, 2019. [cited by applicant]
Diederik P Kingma and Max Welling. Auto-encoding variational bayes. In ICLR, 2014. [cited by applicant]
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014. [cited by applicant]
Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or. Encoding in style- a stylegan encoder for image-to-image translation. arXiv preprint arXiv:2008.00951, 2020. [cited by applicant]
Erik Harkonen, Aaron Hertzmann, Jaakko Lehtinen, and Sylvain Paris. Ganspace: Discovering interpretable gan controls. In NeurIPS, 2020. [cited by applicant]
Gaurav Parmar, Yijun Li, Jingwan Lu, Richard Zhang, Jun-948 Yan Zhu, and Krishna Kumar Sing. Spatially-adaptive multi-layer selection for gan inversion and editing. In Proceedings of the IEEE/CVF Conference on Computer … [cited by applicant]
Gwanghyun Kim, Taesung Kwon, and Jong Chul Ye. Diffusionclip: Text-guided diffusion models for robust image manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. … [cited by applicant]
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. CLIPScore: a reference-free evaluation metric for image captioning. In EMNLP, 2021. [cited by applicant]
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yin-fei Yang, Burcu Karagol Ayan, et al. Scaling autoregressive models for content- rich text-to-image generatio… [cited by applicant]
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In ICLR, 2021. [cited by applicant]
Jiapeng Zhu, Yujun Shen, Deli Zhao, and Bolei Zhou. In—domain gan inversion for real image editing. In ECCV, 2020. [cited by applicant]
Jinjin Gu, Yujun Shen, and Bolei Zhou. Image processing using multi-code gan prior. In CVPR, 2020. [cited by applicant]
John A Gubner. Probability and random processes for electrical and computer engineers. Cambridge University Press, 2006. [cited by applicant]
Jonas Wulff and Antonio Torralba. Improving inversion and generation diversity in stylegan using a gaussianized latent space. arXiv preprint arXiv:2009.06529, 2020. [cited by applicant]
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. [cited by applicant]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, 2020. [cited by applicant]
Jooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon, and Sungroh Yoon. Ilvr: Conditioning method for denoising diffusion probabilistic models. arXiv preprint arXiv:2108.02938, 2021. [cited by applicant]
Jun-Yan Zhu, Philipp Krähenbühl, Eli Shechtman, and Alexei A Efros. Generative visual manipulation on the natural image manifold. In ECCV, 2016. [cited by applicant]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, 2022. [cited by applicant]
Lakshmanan Nataraj, Tajuddin Manhar Mohammed, BS Manjunath, Shivkumar Chandrasekaran, Arjuna Flenner, Jawadul H Bappy, and Amit K Roy-Chowdhury. Detecting gan generated fake images using co-occurrence matrices. Electron… [cited by applicant]
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al. Cogview: Mastering text-to-image generation via transformers. Advances in Neural Information Proc… [cited by applicant]
Minyoung Huh, Richard Zhang, Jun-Yan Zhu, Sylvain Paris, and Aaron Hertzmann. Transforming and projecting images into class-conditional generative networks. In ECCV, 2020. [cited by applicant]
Omri Avrahami, Dani Lischinski, and Ohad Fried. Blended diffusion for text-driven editing of natural images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18208-18218, 2022. [cited by applicant]
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. Styleclip: Text-driven manipulation of stylegan imagery. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2… [cited by applicant]
Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman. Make-a-scene: Scene-based text-to-image generation with human priors. arXiv preprint arXiv:2203.13131, 2022. [cited by applicant]
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34:8780-8794, 2021. [cited by applicant]
Rameen Abdal, Peihao Zhu, John Femiani, Niloy Mitra, and Peter Wonka. Clip2stylegan: Unsupervised extraction of stylegan edit directions. In ACM SIGGRAPH 2022 Conference Proceedings, SIGGRAPH '22, New York, NY, USA, 202… [cited by applicant]
Rameen Abdal, Yipeng Qin, and Peter Wonka. age2stylegan: How to embed images into the stylegan latent space? In ICCV, 2019. [cited by applicant]
Rameen Abdal, Yipeng Qin, and Peter Wonka. age2stylegan++: How to edit the embedded images? CVPR, 2020. [cited by applicant]
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. [cited by applicant]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re… [cited by applicant]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. Stable Diffusion. https://github.com/CompVis/stable-diffusion, 2022. [cited by applicant]
Romain Beaumont. clip-retrieval. https://github.com/rom1504/clip-retrieval, 2022. [cited by applicant]
Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A Efros. Cnn-generated images are surprisingly easy to spot . . . for now. In CVPR, 2020. [cited by applicant]
Shengyu Zhao, Jonathan Cui, Yilun Sheng, Yue Dong, Xiao Liang, Eric I Chang, and Yan Xu. Large scale image completion via co-modulated generative adversarial networks. In ICLR, 2021. [cited by applicant]
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving he image quality of stylegan. In CVPR, 2020. [cited by applicant]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural … [cited by applicant]
Xingang Pan, Xiaohang Zhan, Bo Dai, Dahua Lin, Chen Change Loy, and Ping Luo. Exploiting deep generative prior for versatile image restoration and manipulation. IEEE TPAMI, 2021. [cited by applicant]
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krahenbuhl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. In ECCV, 2022. [cited by applicant]
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020. [cited by applicant]
Yujun Shen, Ceyuan Yang, Xiaoou Tang, and Bolei Zhou. Interfacegan: Interpreting the disentangled face representation learned by gans. IEEE TPAMI, 2020. [cited by applicant]
Zongze Wu, Dani Lischinski, and Eli Shechtman. Stylespace analysis: Disentangled controls for stylegan image generation. In CVPR, 2021. [cited by applicant]
Bo Zhang et al. “Spatial Transformation for Image Composition via Correspondence Learning.” arXiv preprint arXiv:2207.02398 (2022). [cited by applicant]
Chao Jai et al. “Scaling up visual and vision-language representation learning with noisy text supervision.” International conference on machine learning. PM LR, 2021. [cited by applicant]
Ethan Cooper. “Prompt-Free Image-to-Image and Synthesizing Time Lapse Videos of Paintings.” (Apr. 2022). [cited by applicant]
Guillaume Couairon, et al. “DiffEdit: Diffusion-based semantic image editing with mask guidance.” arXiv preprint arXiv:2210.11427 (2022). [cited by applicant]
Yogesh Balaji, et al. “eDiff-I: Text-to-image diffusion models with an ensemble of expert denoisers.” arXiv preprint arXiv:2211.01324 (2022). [cited by applicant]
U.S. Appl. No. 18/178,194, filed Dec. 19, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/178,194, filed Mar. 12, 2015, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/178,167, filed Feb. 26, 2025, Notice of Allowance. [cited by applicant]
Tim Brooks, Aleksander Holynski, and Alexei A Efros. “InstructPix2Pix: Learning to Follow Image Editing Instructions.” arXiv preprint arXiv:2211.09800 (Year: 2022). [cited by applicant]
Yu, Yingchen, et al. “Towards counterfactual image manipulation via clip.” Proceedings of the 3oth ACM International Conference on Multimedia. 2022 (Year: 2022). [cited by applicant]
U.S. Appl. No. 18/178,167, filed Jul. 15, 2025, Office Action. [cited by applicant]
Cited By (1)
US 12,749,293