IP Library Granted Patent US 12,254,545
Granted Patent B2
US 12,254,545 · App. 18/298,138 · Granted Mar 18, 2025

Generating modified digital images incorporating scene layout utilizing a swapping autoencoder

Inventors: Taesung Park (Albany, CA); Alexei A Efros (Berkeley, CA); Elya Shechtman (Seattle, WA); Richard Zhang (San Francisco, CA); Junyan Zhu (Cambridge, MA)
Assignee: Adobe Inc.
G06T11/60G06N3/045G06N3/088G06T7/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,545
App. No.
18/298,138
Granted
Mar 18, 2025
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer readable media for accurately and flexibly generating modified digital images utilizing a novel swapping autoencoder that incorporates scene layout. In particular, the disclosed systems can receive a scene layout map that indicates or defines locations for displaying specific digital content within a digital image. In addition, the disclosed systems can utilize the scene layout map to guide combining portions of digital image latent code to generate a modified digital image with a particular textural appearance and a particular geometric structure defined by the scene layout map. Additionally, the disclosed systems can utilize a scene layout map that defines a portion of a digital image to modify by, for instance, adding new digital content to the digital image, and can generate a modified digital image depicting the new digital content.

Claims (50)

1. A method comprising:

extracting a texture code from a digital image utilizing an encoder of a swapping autoencoder that includes the encoder and a generator neural network;

extracting a structure code from the digital image utilizing the encoder of the swapping autoencoder;

receiving a scene layout map defining semantic regions that indicate boundaries for semantically labeled image content and indicating content of a semantic label not present within the digital image; and

generating, utilizing the generator neural network of the swapping autoencoder to combine the texture code and the structure code as guided by the scene layout map, a modified digital image by:

arranging content of the digital image according to the boundaries for the semantically labeled image content from the scene layout map; and

generating, utilizing the generator neural network, pixels depicting content of the semantic label not present within the digital image.

2. The method of claim 1 , wherein receiving the scene layout map comprises receiving indications for boundaries between pixels depicting content of different semantic labels.

3. The method of claim 1 , wherein receiving the scene layout map comprises receiving an indication for pixels depicting content of a semantic label not present within the digital image.

4. The method of claim 1 , wherein generating the pixels depicting the content of the semantic label not present within the digital image comprises utilizing the generator neural network to replace a portion of the structure code with replacement structure code for the semantic label.

5. The method of claim 1 , wherein generating the modified digital image comprises utilizing the generator neural network to modify the structure code by replacing a portion of the structure code with modified structure code corresponding to a semantic label introduced by the scene layout map and not present within the digital image.

6. The method of claim 5 , wherein modifying the structure code comprises:

generating an average-pooled structure code from a set of sample images depicting content corresponding to the semantic label introduced by the scene layout map; and

replacing the portion of the structure code with the average-pooled structure code.

7. The method of claim 1 , further comprising providing the modified digital image for display on a client device.

8. A system comprising:

a memory component; and

one or more processing devices coupled to the memory component, the one or more processing devices to perform operations comprising:

extracting, utilizing an encoder of a swapping autoencoder that includes the encoder and a generator neural network, a texture code from a digital image, the texture code comprising features corresponding to a textural appearance of the digital image;

extracting, utilizing the encoder of the swapping autoencoder, a structure code from the digital image, the structure code comprising features corresponding to a geometric structure of the digital image;

receiving a scene layout map defining semantic regions that indicate boundaries for semantically labeled image content, the semantic regions including a semantic region of pixels not depicted in the digital image; and

generating, utilizing a generator neural network and the scene layout map, a modified digital image by:

generating pixels depicting content from the digital image as defined by the texture code and the structure code and by arranging the pixels according to the semantic regions of the scene layout map; and

generating additional pixels depicting new content not in the digital image and corresponding to a semantic label indicated by the semantic region within the scene layout map.

9. The system of claim 8 , wherein receiving the scene layout map comprises:

receiving indications for boundaries between pixels depicting content of different semantic labels; and

receiving an indication for pixels depicting content of a semantic label not present within the digital image.

10. The system of claim 8 , wherein generating the modified digital image comprises:

generating a replacement structure code corresponding the semantic label of the semantic region different from pixels of the digital image by average pooling structure codes from a cluster of sample images depicting content corresponding to the semantic label of the semantic region; and

replacing a portion of the structure code of the digital image with the replacement structure code.

11. The system of claim 8 , wherein generating the modified digital image comprises utilizing the generator neural network to generate, from the texture code and the structure code, pixels depicting content arranged according to the boundaries defined by the semantic regions of the scene layout map.

12. The system of claim 8 , wherein receiving the scene layout map comprises receiving an indication of a user interaction defining the semantic label and a boundary for the semantic region of pixels different from pixels of the digital image.

13. The system of claim 8 , wherein generating the modified digital image comprises utilizing the generator neural network to generate pixels for the semantic region different from the pixels of the digital image.

14. The system of claim 8 , wherein the operations further comprise providing the modified digital image for display on a client device.

15. A non-transitory computer readable medium storing instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

extracting a texture code from a digital image utilizing an encoder of a swapping autoencoder that includes the encoder and a generator neural network;

extracting a structure code from the digital image utilizing the encoder of the swapping autoencoder;

receiving a scene layout map defining semantic regions that indicate boundaries for semantically labeled image content and indicating content of a semantic label not present within the digital image; and

generating, utilizing the generator neural network of the swapping autoencoder to combine the texture code and the structure code as guided by the scene layout map, a modified digital image by:

arranging content of the digital image according to the boundaries for the semantically labeled image content from the scene layout map; and

generating, utilizing the generator neural network, pixels depicting content of the semantic label not present within the digital image.

16. The non-transitory computer readable medium of claim 15 , wherein receiving the scene layout map comprises:

receiving indications for the boundaries between pixels depicting content of different semantic labels; and

receiving an indication for pixels depicting content of a semantic label not present within the digital image.

17. The non-transitory computer readable medium of claim 15 , wherein generating the modified digital image comprises:

generating a replacement structure code corresponding a semantic label for pixels not depicted in the digital image by average pooling structure codes from a cluster of sample images depicting content corresponding to the semantic label; and

replacing a portion of the structure code of the digital image with the replacement structure code.

18. The non-transitory computer readable medium of claim 15 , wherein generating the modified digital image comprises utilizing the generator neural network to generate, from the texture code and the structure code, pixels depicting content with boundaries defined by the semantic regions of the scene layout map.

19. The non-transitory computer readable medium of claim 15 , wherein generating the modified digital image comprises utilizing the generator neural network to generate, from the texture code and the structure code, pixels depicting different types of content separated by the semantic regions of the scene layout map.

20. The non-transitory computer readable medium of claim 15 , wherein the modified digital image depicts digital content corresponding to different labels fitted to different locations indicated by the scene layout map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2023
From: PARK, TAESUNG; EFROS, ALEXEI A; SHECHTMAN, ELYA; ZHANG, RICHARD; ZHU, JUNYAN
To: ADOBE INC.
Reel/Frame 063359/0540 →
Continuity (2)
Continuation 17091416 · Nov 6, 2020
Related Publication 20230245363A1 · Aug 3, 2023
References Cited (18)
US 4706243A · Noguchi · 1987 [cited by applicant]
US 10658005B1 · Bogan, III et al. · 2020 [cited by applicant]
US 20100202699A1 · Matsuzaka et al. · 2010 [cited by applicant]
US 20130265382A1 · Guleryuz et al. · 2013 [cited by applicant]
US 20160125572A1 · Yoo · 2016 [cited by examiner]
US 20160350930A1 · Lin et al. · 2016 [cited by applicant]
US 20190355103A1 · Baek · 2019 [cited by examiner]
H. Caesar, J. Uijlings, and V. Ferrari. Coco-stuff: Thing and stuff classes in context. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1209-1218, 2018. [cited by applicant]
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis… [cited by applicant]
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In Advances in neural information processing systems, pp. 2672-2680, 2014. [cited by applicant]
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. [cited by applicant]
T. Park, M.-Y. Liu, T.-C. Wang, and J.-Y. Zhu. Semantic image synthesis with spatially-adaptive normalization. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019. [cited by applicant]
T. Park, J.-Y. Zhu, O. Wang, J. Lu, E. Shechtman, A. A. Efros, and R. Zhang. Swapping autoencoder for deep image manipulation. In ArXiv, 2020. [cited by applicant]
T.-C.Wang, M.-Y. Liu, J.-Y. Zhu, A. Tao, J. Kautz, and B.Catanzaro. High-resolution image synthesis and semantic manipulation with conditional gans. In Proceedings of IEEE Conference on Computer Vision and Pattern Recog… [cited by applicant]
U.S. Appl. No. 17/091,416, filed Dec. 17, 2021, Office Action. [cited by applicant]
U.S. Appl. No. 17/091,416, filed Mar. 21, 2022, Office Action. [cited by applicant]
U.S. Appl. No. 17/091,416, filed Aug. 26, 2022, Office Action. [cited by applicant]
U.S. Appl. No. 17/091,416, filed Dec. 20, 2022, Notice of Allowance. [cited by applicant]