IP Library › Granted Patent US 12,333,681
Granted Patent B1
US 12,333,681 · App. 18/067,598 · Granted Jun 17, 2025

Targeted generative visual editing of images

Inventors: Cheng-Yang Fu (San Francisco, CA); Tamara L. Berg (Laguna Beach, CA); Andrew Brown (New York, NY); Nicole Gallagher (Astoria, NY); Sen He (London, GB); Omkar Moreshwar Parkhi (Cupertino, CA); Antoine Toisoul (London, GB); Andrea Vedaldi (Oxford, GB); Tao Xiang (Ruislip, GB); Yanping Xie (London, GB)
Assignee: Meta Platforms Technologies, LLC
G06T5/50G06T9/00G06V10/25G06V10/82G06T2207/20081G06T2207/20132G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,681
App. No.
18/067,598
Granted
Jun 17, 2025
Kind
B1
Abstract

One embodiment of the present invention sets forth a technique for combining a source image and a driver image. The technique includes determining a first region of the source image to be blended with the driver image. The technique also includes inputting a second region of the source image that lies outside of the first region and the driver image into a neural network. The technique further includes generating, via the neural network, an output image that includes a third region corresponding to the first region of the source image and a fourth region corresponding to the second region of the source image, where the third region includes visual attributes of the driver image and a context associated with the source image and the fourth region includes visual attributes of the second region of the source image and the context associated with the source image.

Claims (47)

1. A method comprising:

determining, via a processor, a first region of a source image to be blended with a driver image;

inputting, via the processor, a second region of the source image that lies outside of the first region and the driver image into a neural network; and

generating, via the neural network, an output image that includes a third region corresponding to the first region of the source image and a fourth region corresponding to the second region of the source image, wherein the third region includes one or more visual attributes of the driver image and a context associated with the source image and the fourth region includes one or more visual attributes of the second region of the source image and the context associated with the source image.

2. The method of claim 1 , further comprising:

applying a transformation to a fifth region of a training output image to generate a training driver image;

inputting a sixth region of the training output image that lies outside of the fifth region and the training driver image into the neural network; and

training one or more components of the neural network based on a loss associated with the training output image and an output generated by the neural network based on the sixth region of the training output image and the training driver image.

3. The method of claim 2 , further comprising training one or more additional components of the neural network based on a first reconstruction of the training output image and a second reconstruction of the training driver image.

4. The method of claim 2 , wherein the transformation comprises a randomized crop of the fifth region.

5. The method of claim 2 , wherein the loss comprises a negative log-likelihood.

6. The method of claim 1 , wherein generating the output image comprises:

converting the second region of the source image into a first encoded representation; and

converting the driver image into a second encoded representation.

7. The method of claim 6 , wherein generating the output image further comprises:

converting the first encoded representation and the second encoded representation into a third encoded representation; and

decoding the third encoded representation to generate the output image.

8. The method of claim 1 , wherein inputting the second region of the source image into the neural network comprises:

combining the source image with a binary mask representing the first region to generate a masked source image; and

inputting the masked source image into the neural network.

9. The method of claim 1 , wherein the neural network comprises an autoregressive model that implements a conditional probability distribution associated with the output image, the source image, the driver image, and the first region of the source image.

10. The method of claim 1 , wherein the first region comprises at least one of a bounding box or a semantic segmentation.

11. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform:

determining a first region of a source image to be blended with a driver image;

inputting a second region of the source image that lies outside of the first region and the driver image into a neural network; and

generating, via the neural network, an output image that includes a third region corresponding to the first region of the source image and a fourth region corresponding to the second region of the source image, wherein the third region includes one or more visual attributes of the driver image and a context associated with the source image and the fourth region includes one or more visual attributes of the second region of the source image and the context associated with the source image.

12. The one or more non-transitory computer-readable media of claim 11 , wherein the instructions further cause the one or more processors to perform:

applying a transformation to a fifth region of a training output image to generate a training driver image;

inputting a sixth region of the training output image that lies outside of the fifth region and the training driver image into the neural network; and

training one or more components of the neural network based on a loss associated with the training output image and an output generated by the neural network based on the sixth region of the training output image and the training driver image.

13. The one or more non-transitory computer-readable media of claim 12 , wherein the instructions further cause the one or more processors to perform:

training one or more additional components of the neural network based on a first reconstruction of the training output image, a second reconstruction of the training driver image, and one or more perceptual losses.

14. The one or more non-transitory computer-readable media of claim 13 , wherein the one or more additional components comprise a set of encoders and a set of decoders.

15. The one or more non-transitory computer-readable media of claim 12 , wherein the one or more components comprise a transformer.

16. The one or more non-transitory computer-readable media of claim 12 , wherein the fifth region comprises at least one of a bounding box or a semantic segmentation.

17. The one or more non-transitory computer-readable media of claim 11 , wherein the generating the output image comprises:

converting the second region of the source image into a first encoded representation;

converting the driver image into a second encoded representation;

converting a token sequence that includes the first encoded representation and the second encoded representation into a third encoded representation; and

decoding the third encoded representation to generate the output image.

18. The one or more non-transitory computer-readable media of claim 17 , wherein at least one of the first encoded representation or the second encoded representation comprises a set of quantized tokens.

19. The one or more non-transitory computer-readable media of claim 11 , wherein determining the first region of the source image comprises receiving the first region from a user.

20. A system comprising:

one or more memories that store instructions; and

one or more processors that are coupled to the one or more memories, which when executing the instructions, are configured to perform: determine a first region of a source image to be blended with a driver image;

input a second region of the source image that lies outside of the first region and the driver image into a neural network; and

generate, via the neural network, an output image that includes a third region corresponding to the first region of the source image and a fourth region corresponding to the second region of the source image, wherein the third region includes one or more visual attributes of the driver image and a context associated with the source image and the fourth region includes one or more visual attributes of the second region of the source image and the context associated with the source image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2023
From: FU, CHENG-YANG; BERG, TAMARA L.; BROWN, ANDREW; GALLAGHER, NICOLE; HE, SEN; PARKHI, OMKAR MORESHWAR; TOISOUL, ANTOINE; VEDALDI, ANDREA; XIANG, TAO; XIE, YANPING
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 062317/0074 →
References Cited (83)
US 11127139B2 · Zhang et al. · 2021 [cited by examiner]
US 20220392255A1 · Moustafa · 2022 [cited by examiner]
Ramesh A., et al., “Zero-Shot Text-to-Image Generation,” International Conference on Machine Learning (ICML), Jul. 2021, 11 pages. [cited by applicant]
Razavi A., et al., “Generating Diverse High-Fidelity Images with VQ-VAE-2,” Advances in neural information processing systems 32, 2019, 11 pages. [cited by applicant]
Salimans T., et al., “Improved Techniques for Training Gans,” arXiv, arXiv: 1606.03498v1, Jun. 10, 2016, 10 pages. [cited by applicant]
Schaldenbrand P., et al., “StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Synthesis,” arXiv preprint arXiv:2111.03133v1, Nov. 4, 2021, 3 pages. [cited by applicant]
Schwettmann S., et al., “Toward a Visual Concept Vocabulary for Gan Latent Space,” Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, 9 pages. [cited by applicant]
Shen Y., et al., “Interpreting the Latent Space of Gans for Semantic Face Editing,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 9243-9252. [cited by applicant]
Shi J., et al., “SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Editing, ” arXiv:2112.00180v1, Nov. 30, 2021, 20 pages. [cited by applicant]
Shocher A., et al., “Semantic Pyramid for Image Generation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, 10 pages. [cited by applicant]
Szegedy C., et al., “Rethinking the Inception Architecture for Computer Vision,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2016, pp. 2818-2826. [cited by applicant]
Tov O., et al., “Designing an Encoder for Stylegan Image Manipulation,” arXiv:2102.02766v1, Feb. 4, 2021, 33 pages. [cited by applicant]
Tsai Y., et al., “Deep Image Harmonization,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 3789-3797. [cited by applicant]
Van Den Oord A., et al., “Neural Discrete Representation Learning,” In Advances in Neural Information Processing Systems, Dec. 2017, pp. 1-11. [cited by applicant]
Voynov A., et al., “Unsupervised Discovery of Interpretable Directions in the GAN Latent Space,” International conference on machine learning (PMLR), 2020, vol. 119, 11 pages. [cited by applicant]
Wang T-C., et al., “High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs,” IEEE Conference on Computer Vision and Pattern Recognition, Aug. 20, 2018, 14 pages. [cited by applicant]
Williams R.J., et al., “A Learning Algorithm for Continually Running Fully Recurrent Neural Networks,” Neural Computation, Jun. 1989, Vo. 1(2), pp. 270-280. [cited by applicant]
Wu C., et al., “NUWA: Visual Synthesis Pre-Training for Neural Visual World Creation,” arXiv:2111.12417v1, Nov. 24, 2021, 28 pages. [cited by applicant]
Wu Z., et al., “Stylespace Analysis: Disentangled Controls for Stylegan Image Generation,” arXiv:2011.12799v1, Nov. 25, 2020, 25 pages. [cited by applicant]
Xia W., et al., “TediGAN: Text-Guided Diverse Face Image Generation and Manipulation,” Computer Vision and Pattern Recognition (CVPR), Jun. 2021, pp. 2256-2265. [cited by applicant]
Xiao Z., et al., “Generative Latent Flow.,” arXiv:1905.10485, Sep. 22, 2019, 18 pages. [cited by applicant]
Xu Y., et al., “Generative Hierarchical Features from Synthesizing Images,” Computer Vision and Pattern Recognition (CVPR), Jun. 2021, pp. 4432-4442. [cited by applicant]
Yang C., et al., “Semantic Hierarchy Emerges in Deep Generative Representations for Scene Synthesis,” International Journal of Computer Vision, Feb. 11, 2020, 15 pages. [cited by applicant]
Yu F., et al., “LSUN: Construction of a Large-Scale Image Dataset using Deep Learning with Humans in the Loop.,” arXiv:1506.03365, Jun. 19, 2015, 9 pages. [cited by applicant]
Zhang R., et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, 10 pages. [cited by applicant]
Zhang Z., et al., “UFC-BERT: Unifying Multi-Modal Controls for Conditional Image Synthesis,” Neural Information Processing Systems (NeurIPS), Dec. 2021, 13 pages. [cited by applicant]
Zhao S., et al., “Large Scale Image Completion via Co-Modulated Generative Adversarial Networks,” International Conference on Learning Representations (ICLR), Mar. 2021, 25 pages. [cited by applicant]
Zhu J., et al., “In-Domain GAN Inversion for Real Image Editing,” European Conference on Computer Vision (ECCV), Aug. 23, 2020, pp. 592-608. [cited by applicant]
Zhu J-Y., et al., Generative Visual Manipulation on the Natural Image Manifold, European Conference on Computer Vision (ECCV), Sep. 2016, pp. 597-613. [cited by applicant]
Brown A., et al., “End-to-End Visual Editing with a Generatively Pre-Trained Artist,” arxiv.org, May 3, 2022, 33 pages. [cited by applicant]
European Search Report for European Patent Application No. 23207999.6, dated Mar. 27, 2024, 5 pages. [cited by applicant]
Zhang Z., et al., “M6-UFC: Unifying Multi-Modal Controls for Conditional Image Synthesis via Non-Autoregressive Generative Transformers,” arXiv:2105.14211v4, Feb. 19, 2022, 17 pages. [cited by applicant]
Zhang R., et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric,” In Proceedings of the EEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, 10 pages. [cited by applicant]
Issenhuth T., et al., “EdiBERT, A Generative Model for Image Editing,” arXiv:2111.15264, Nov. 30, 2021, 15 pages. [cited by applicant]
Jahanian A., et al., “On the “Steerability” of Generative Adversarial Networks,” International Conference on Learning Representations (ICLR), Apr. 2020, 31 pages. [cited by applicant]
Karras T., et al., “A Style-Based Generator Architecture for Generative Adversarial Networks,” In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 16-20, 2019, pp. 4401-4410. [cited by applicant]
Karras T., et al., “Analyzing and Improving the Image Quality of StyleGAN,” Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 8110-8119. [cited by applicant]
Karras T., et al., “Training Generative Adversarial Networks with Limited Data,” Neural Information Processing Systems (NeurIPS), Dec. 2020, 11 pages. [cited by applicant]
Kim H., et al., “Exploiting Spatial Dimensions of Latent in GAN for Real-Time Image Editing,” Computer Vision and Pattern Recognition (CVPR), Jun. 2021, pp. 852-861. [cited by applicant]
Lipton Z.C., et al., “Precise Recovery of Latent Vectors from Generative Adversarial Networks,” arXiv:1702.04782, Feb. 15, 2021, 4 pages. [cited by applicant]
Liu G., et al., “Image Inpainting for Irregular Holes using Partial Convolutions,” European Conference on Computer Vision (ECCV), Sep. 2018, 16 pages. [cited by applicant]
Liu X., et al., “More Control for Free! Image Synthesis with Semantic Diffusion Guidance.”, arXiv:2112.05744, Dec. 14, 2021, 16 pages. [cited by applicant]
Loshchilov I., et al., “Decoupled Weight Decay Regularization,” International Conference on Learning Representations (ICLR), Jan. 2019, 19 pages. [cited by applicant]
Mokady R., et al., “Mask Based Unsupervised Content Transfer,” arXiv preprint arXiv:1906.06558v1, Jun. 15, 2019, 31 pages. [cited by applicant]
Nichol A., et al., “Glide: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models,” arXiv preprint arXiv:2112.10741, Dec. 22, 2021, 20 pages. [cited by applicant]
Park T., et al., “Semantic Image Synthesis with Spatially-Adaptive Normalization,” In Conference on Computer Vision and Pattern Recognition, Mar. 18, 2019, 19 pages. [cited by applicant]
Patashnik O., et al., “Styleclip: Text-Driven Manipulation of Stylegan Imagery,” Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 2085-2094. [cited by applicant]
Peebles W., et al., “The Hessian Penalty: A Weak Prior for Unsupervised Disentanglement,” European Conference on Computer Vision, Nov. 7, 2020, vol. 12351, 23 pages. [cited by applicant]
Press O., et al., “Emerging Disentanglement in Auto-Encoder Based Unsupervised Image Content Transfer,” International Conference on Learning Representations (ICLR), 2019, 26 pages. [cited by applicant]
Dolhansky B., et al., “The Deepfake Detection Challenge Dataset,” arXiv preprint, arXiv:2006.07397v4, Oct. 28, 2020, 13 pages. [cited by applicant]
Esser P., et al., “A Disentangling Invertible Interpretation Network for Explaining Latent Representations,” Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 9223-9232. [cited by applicant]
Esser P., et al., “Imagebart: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis,” Neural Information Processing Systems 34 (NeurIPS 2021), Dec. 2021, 15 pages. [cited by applicant]
Esser P., et al., “Taming Transformers for High-Resolution Image Synthesis,” Computer Vision and Pattern Recognition (CVPR), Jun. 2021, pp. 12873-12883. [cited by applicant]
Fauw J.D., et al., “Hierarchical Autoregressive Image Models with Auxiliary Decoders,” arXiv:1903.04933, Mar. 6, 2019, 21 pages. [cited by applicant]
Gafni O., et al., “Make-a-Scene: Scene-based Text-to-Image Generation with Human Priors,” arXiv:2203.13131, Mar. 2022, 17 pages. [cited by applicant]
Galatolo F.A., et al., “Generating Images From Caption and Vice Versa via Clip-Guided Generative Latent Space Search,” International Conference on Image Processing and Vision Engineering (2021), Feb. 2021, 10 pages. [cited by applicant]
Ghosh P., et al., “InvGAN: Invertible GANs,” arXiv:2112.04598, Dec. 8, 2021, 15 pages. [cited by applicant]
Goyal A., et al., “Professor Forcing: A New Algorithm for Training Recurrent Networks,” Neural Information Processing Systems (NeurIPS), Dec. 2016, pp. 4608-4616. [cited by applicant]
Guan S., et al., “Collaborative Learning for Faster Stylegan Embedding,” arXiv:2007.01758, Jul. 3, 2020, pp. 4321-4330. [cited by applicant]
Harkonen E., et al., “GANSpace: Discovering Interpretable GAN Controls,” arXiv:2004.02546, Apr. 6, 2020, 14 pages. [cited by applicant]
Heusel M., et al., “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium,” Advances in Neural Information Processing Systems, 2017, 12 pages. [cited by applicant]
Holtzman A., et al., “The Curious Case of Neural Text Degeneration,” International Conference on Learning Representations (ICLR), Apr. 2020, 16 pages. [cited by applicant]
Iizuka S., et al., “Globally and Locally Consistent Image Completion,” ACM Transactions on Graphics, Jul. 2017, vol. 36, No. 4, pp. 107:1-107:14. [cited by applicant]
Isola P., et al., “Scene Collaging: Analysis and Synthesis of Natural Images with Semantic Layers,” International Conference on Computer Vision (ICCV), Dec. 2013, 8 pages. [cited by applicant]
Isola P., et al., “Image-to-Image Translation with Conditional Adversarial Networks,” IEEE Conference on Computer Vision and Pattern Recognition 2017, Nov. 26, 2018, 17 pages. [cited by applicant]
Radford A., et al., “Language Models are Unsupervised Multitask Learners,” 2019, 24 pages. [cited by applicant]
Radford A., et al., “Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks,” International Conference on Learning Representations, Jan. 2016, 16 pages. [cited by applicant]
Ramesh A., et al., “Hierarchical Text-Conditional Image Generation with Clip Latents,” arXiv preprint arXiv:2204.06125, Apr. 13, 2022, 22 pages. [cited by applicant]
Dai B., et al., “Diagnosing and Enhancing VAE Models,” International Conference on Learning Representations (ICLR), May 2019, 12 pages. [cited by applicant]
Deng J., et al., “ImageNet: A Large-Scale Hierarchical Image Database,” 2009 IEEE Conference on Computer Vision and Pattern Recognition, Aug. 18, 2009, 8 Pages. [cited by applicant]
Devlin J., et al., “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding,” In Proceedings of the 2019 Confer-ence of the North American Chapter of the Association for Computational Linguistic… [cited by applicant]
Ding M., et al., “CogView: Mastering Text-to-Image Generation via Transformers,” Neural Information Processing Systems 34 (NeurIPS 2021), 14 pages. [cited by applicant]
Abdal R., et al., “CLIP2StyleGAN: Unsupervised Extraction of StyleGAN Edit Directions,” arXiv:2112.05219, Dec. 9, 2021, 17 pages. [cited by applicant]
Abdal R., et al., “Image2stylegan++: How to Edit the Embedded Images?,” Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 8296-8305. [cited by applicant]
Abdal R., et al., “Image2StyleGAN: How to Embed Images into the StyleGAN Latent Space?,” In Proceedings of the IEEE International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, pp. 4432-4441. [cited by applicant]
Bau D., et al., “Semantic Photo Manipulation with a Generative Image Prior,” ACM Transactions on Graphics, Jul. 2019, vol. 38, No. 4, pp. 59:1-59:11. [cited by applicant]
Bau D., et al., “Inverting Layers of A Large Generator,” ICLR 2019 Debugging Machine Learning Models Workshop, May 2019, 4 pages. [cited by applicant]
Bau D., et al., “Paint by Word,” arXiv:2103.10951, Mar. 19, 2021, 10 pages. [cited by applicant]
Bau D., et al., “Understanding the Role of Individual Units in a Deep Neural Network,” Proceedings of the National Academy of Sciences (PNAS), Sep. 1, 2020, vol. 117, No. 48, pp. 30071-30078. [cited by applicant]
Chai L., et al., “Using Latent Space Regression to Analyze and Leverage Compositionality in GANs,” International Conference on Learning Representations (ICLR), May 2021, 30 pages. [cited by applicant]
Chen M., et al., “Generative Pretraining from Pixels,” International Conference on Machine Learning (ICML), Jul. 2020, 13 pages. [cited by applicant]
Choi Y., et al., “StarGAN: Unified Generative Adversarial Networks for Multi-Domain Image-to-Image Translation,” Computer Vision and Pattern Recognition (CVPR), Jun. 2018, pp. 8789-8797. [cited by applicant]
Crowson K: “VQGAN-CLIP,” Retrieved from the Internet: https://github.com/nerdyrodent/VQGAN-CLIP, 2021, 5 pages. [cited by applicant]
Cited By (1)
US 12,718,331