IP Library › Granted Patent US 12,223,623
Granted Patent B2
US 12,223,623 · App. 18/053,027 · Granted Feb 11, 2025

Harmonizing composite images utilizing a semantic-guided transformer neural network

Inventors: He Zhang (San Jose, CA); Hyun Joon Jung (Monte Sereno, CA)
Assignee: Adobe Inc.
G06T5/50G06T7/11G06T7/194G06V10/267G06V10/42G06V10/44G06V10/82G06T2200/24G06T2207/20084G06T2207/20092G06T2207/20132G06T2207/20212
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,623
App. No.
18/053,027
Granted
Feb 11, 2025
Kind
B2
Abstract

The present disclosure relates to systems, non-transitory computer-readable media, and methods that implement a multi-branch harmonization neural network architecture to harmonize composite images. For example, in one or more implementations, the semantic-guided transformer-based harmonization system uses a convolutional branch, a transformer branch, and a semantic branch to generate a harmonized composite image based on an input composite image and a corresponding segmentation mask. More particularly, the convolutional branch comprises a series of convolutional neural network layers followed by a style normalization layer to extract localized information from the input composite image. Further, the transformer branch comprises a series of transformer neural network layers to extract global information based on different resolutions of the input composite image. The semantic branch includes a visual neural network that generates semantic features that inform the harmonization of the composite images.

Claims (45)

1. A computer-implemented method comprising:

receiving a composite image and a segmentation mask of a foreground object portrayed against a background of the composite image;

utilizing a global neural network branch to extract global information of the background;

utilizing a semantic neural network branch to extract semantic information from the composite image; and

generating a harmonized composite image comprising the foreground object harmonized with the background by decoding the global information and the semantic information.

2. The computer-implemented method of claim 1 , wherein utilizing the global neural network branch to extract the global information of the background comprises extracting the global information utilizing a transformer neural network.

3. The computer-implemented method of claim 1 , wherein utilizing the semantic neural network branch to extract semantic information from the composite image comprises extracting semantic features utilizing a visual neural network.

4. The computer-implemented method of claim 1 , further comprising utilizing a local neural network branch to extract local information from the composite image.

5. The computer-implemented method of claim 4 , wherein generating the harmonized composite image comprises decoding the local information, the global information, and the semantic information.

6. The computer-implemented method of claim 5 , wherein utilizing the local neural network branch to extract local information comprises extracting local information of a background portion adjacent to the foreground object utilizing a convolutional neural network.

7. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

receiving a composite image and a segmentation mask of a foreground object portrayed against a background of the composite image;

utilizing a local neural network branch to extract local information from the composite image;

utilizing a global neural network branch to extract global information of the background;

utilizing a semantic neural network branch to extract semantic information from the composite image; and

generating a harmonized composite image comprising the foreground object harmonized with the background by decoding the local information, global information, and the semantic information.

8. The non-transitory computer-readable medium of claim 7 , wherein the operations further comprise:

generating a semantic aware image crop from the composite image;

generating a semantic aware segmentation mask crop from the segmentation mask; and

concatenating the semantic aware image crop and the semantic aware segmentation mask crop to generate a cropped four-channel semantic-aware input.

9. The non-transitory computer-readable medium of claim 8 , wherein utilizing the semantic neural network branch comprises extracting semantic features from the cropped four-channel semantic-aware input utilizing a visual neural network.

10. The non-transitory computer-readable medium of claim 9 , wherein extracting semantic features from the cropped four-channel semantic-aware input utilizing the visual neural network comprises extracting the semantic features from a convolutional layer just prior to average pooling and a classification layer of the visual neural network.

11. The non-transitory computer-readable medium of claim 9 , wherein utilizing the global neural network branch comprises utilizing one or more transformer neural network layers to extract the global information of the background from the cropped four-channel semantic-aware input.

12. The non-transitory computer-readable medium of claim 11 , wherein utilizing the one or more transformer neural network layers to extract the global information of the background comprises utilizing one or more layers comprising a self-attention neural network layer, a mix feed-forward neural network layer, and an overlap patch merging layer.

13. The non-transitory computer-readable medium of claim 9 , wherein utilizing the local neural network branch comprises utilizing one or more convolutional neural network layers to extract the local information of a background portion adjacent to the foreground object from the cropped four-channel semantic-aware input.

14. The non-transitory computer-readable medium of claim 13 , wherein utilizing the local neural network branch comprises utilizing one or more style normalization layers to extract style information from the background portion adjacent to the foreground object by determining a mean and standard deviation of pixel color values between pixels corresponding to the background portion adjacent to the foreground object.

15. The non-transitory computer-readable medium of claim 14 , wherein the operations further comprise utilizing the one or more style normalization layers to apply the style information to the foreground object.

16. The non-transitory computer-readable medium of claim 7 , wherein receiving the composite image and the segmentation mask comprises:

identifying, via a graphical user interface, a user selection of a background image for combining with the foreground object as the composite image;

identifying, via the graphical user interface, an additional user selection corresponding to a harmonization filter for harmonizing the foreground object and the background image; and

utilizing a segmentation neural network to generate the segmentation mask of the foreground object portrayed against the background of the composite image.

17. A system comprising:

a composite image comprising a foreground object portrayed against a background of the composite image;

a segmentation mask of the foreground object;

a local neural network branch to extract local information of a background portion adjacent to the foreground object;

a global neural network branch to extract global information of the background;

a semantic neural network branch to extract semantic information from the composite image;

a decoder to decode the local information, the global information, and the semantic information; and

one or more processors configured to cause the system to generate a harmonized composite image comprising the foreground object with modified pixel color values based on decoding of the local information, the global information, and the semantic information.

18. The system of claim 17 , wherein:

the local neural network branch comprises one or more convolutional neural network layers;

the global neural network branch comprises a transformer neural network; and

the semantic neural network branch comprises a visual neural network.

19. The system of claim 18 , wherein the global neural network branch comprises a series of transformer neural networks composed of at least one of a self-attention neural network layer, a mix feed-forward neural network layer, or an overlap patch merging layer.

20. The system of claim 18 , wherein the local neural network branch comprises a style normalization layer to generate style-normalized foreground feature vectors for the foreground object based on style information for the background portion adjacent to the foreground object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2022
From: ZHANG, HE; JUNG, HYUN JOON
To: ADOBE INC.
Reel/Frame 061673/0812 →
Continuity (1)
Related Publication 20240161240A1 · May 16, 2024
References Cited (34)
US 10186038B1 · Kluckner et al. · 2019 [cited by applicant]
US 10719742B2 · Shechtman et al. · 2020 [cited by applicant]
US 10867416B2 · Shen et al. · 2020 [cited by applicant]
US 11205073B2 · Hartman et al. · 2021 [cited by applicant]
US 11869125B2 · Bedi et al. · 2024 [cited by applicant]
US 11875510B2 · Wang et al. · 2024 [cited by applicant]
US 11935217B2 · Zhang et al. · 2024 [cited by applicant]
US 20040077393A1 · Kim et al. · 2004 [cited by applicant]
US 20180260668A1 · Shen et al. · 2018 [cited by applicant]
US 20210335480A1 · Johnsson et al. · 2021 [cited by applicant]
US 20230298148A1 · Zhang et al. · 2023 [cited by applicant]
US 20240161240A1 · Zhang et al. · 2024 [cited by applicant]
Guo, Zonghui, et al. “Transformer for image harmonization and beyond.” IEEE transactions on pattern analysis and machine intelligence 45.11 (2022): 12960-12977. (Year: 2022). [cited by examiner]
Li, Boyi, et al. “Language-driven semantic segmentation.” arXiv preprint arXiv:2201.03546 (2022). [cited by examiner]
Niu, Li, et al. “Making images real again: A comprehensive survey on deep image composition.” arXiv preprint arXiv:2106.14490 (2021). [cited by examiner]
Niu, Li, et al. “Making images real again: A comprehensive survey on deep image composition.” arXiv preprint arXiv:2106.14490 (2024). [cited by examiner]
Ling, Jun, et al. “Region-aware adaptive instance normalization for image harmonization.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2021. [cited by examiner]
Sofiiuk, Konstantin, Polina Popenova, and Anton Konushin. “Foreground-aware semantic representations for image harmonization.” Proceedings of the IEEE/CVF winter conference on applications of computer vision. 2021. [cited by examiner]
Tsai, Yi-Hsuan, et al. “Deep image harmonization.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2017. [cited by examiner]
Cun, Xiaodong, and Chi-Man Pun. “Improving the harmony of the composite image by spatial-separated attention module.” IEEE Transactions on Image Processing 29 (2020): 4759-4771. [cited by examiner]
Lin, Chen-Hsuan, et al. “St-gan: Spatial transformer generative adversarial networks for image compositing.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2018. [cited by examiner]
Wenyan Cong, Li Niu, Jianfu Zhang, Jing Liang, and Liqing Zhang. Bargainnet: Background-guided domain translation for image harmonization. In 2021 IEEE International Conference on Multimedia and Expo (ICME), pp. 1-6. IE… [cited by applicant]
Wenyan Cong, Jianfu Zhang, Li Niu, Liu Liu, Zhixin Ling, Weiyuan Li, and Liqing Zhang. Dovenet: Deep image harmonization via domain verification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern … [cited by applicant]
Xiaodong Cun and Chi-Man Pun. Improving the harmony of the composite image by spatial-separated attention module. IEEE Transactions on Image Processing, 29:4759-4771, 2020. [cited by applicant]
Zonghui Guo, Haiyong Zheng, Yufeng Jiang, Zhaorui Gu, and Bing Zheng. Intrinsic image harmonization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16367-16376, 2021. [cited by applicant]
Yifan Jiang, He Zhang, Jianming Zhang, Yilin Wang, Zhe Lin, Kalyan Sunkavalli, Simon Chen, Sohrab Amirghodsi, Sarah Kong, and Zhangyang Wang. Ssh: A selfsupervised framework for image harmonization. arXiv preprint arXiv… [cited by applicant]
Jun Ling, Han Xue, Li Song, Rong Xie, and Xiao Gu. Region-aware adaptive instance normalization for image harmonization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9361-937… [cited by applicant]
Patrick Perez, Michel Gangnet, and Andrew Blake. Poisson image editing. In ACM SIGGRAPH 2003 Papers, pp. 313-318. 2003. [cited by applicant]
Erik Reinhard, Michael Adhikhmin, Bruce Gooch, and Peter Shirley. Color transfer between images. IEEE Computer graphics and applications, 21(5):34-41, 2001. [cited by applicant]
Michael W Tao, Micah K Johnson, and Sylvain Paris. Error tolerant image compositing. In European Conference on Computer Vision, pp. 31-44. Springer, 2010. [cited by applicant]
Yi-Hsuan Tsai, Xiaohui Shen, Zhe Lin, Kalyan Sunkavalli, Xin Lu, and Ming-Hsuan Yang. Deep image harmonization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3789-3797, 2017. [cited by applicant]
Jun-Yan Zhu, Philipp Krahenbuhl, Eli Shechtman, and Alexei A Efros. Learning a discriminative model for the perception of realism in composite images. In Proceedings of the IEEE International Conference on Computer Visi… [cited by applicant]
U.S. Appl. No. 17/655,663, Jun. 10, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 17/655,663, Aug. 15, 2024, Notice of Allowance. [cited by applicant]
Cited By (1)
US 12,333,692