IP Library Granted Patent US 12,417,559
Granted Patent B2
US 12,417,559 · App. 17/744,995 · Granted Sep 16, 2025

Semantic image fill at high resolutions

Inventors: Tobias Hinz (Ulm, DE); Taesung Park (San Francisco, CA); Richard Zhang (San Francisco, CA); Matthew David Fisher (Burlingame, CA); Difan Liu (Amherst, MA); Evangelos Kalogerakis (Sunderland, MD)
Assignee: Adobe Inc.
G06T11/00G06T3/4046G06T7/11G06V10/235G06V10/44G06V10/513G06V10/7753G06V10/82G06V20/70G06T2207/20084G06T2207/20092
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,559
App. No.
17/744,995
Granted
Sep 16, 2025
Kind
B2
Abstract

Semantic fill techniques are described that support generating fill and editing images from semantic inputs. A user input, for example, is received by a semantic fill system that indicates a selection of a first region of a digital image and a corresponding semantic label. The user input is utilized by the semantic fill system to generate a guidance attention map of the digital image. The semantic fill system leverages the guidance attention map to generate a sparse attention map of a second region of the digital image. A semantic fill of pixels is generated for the first region based on the semantic label and the sparse attention map. The edited digital image is displayed in a user interface.

Claims (44)

1. A method comprising:

receiving, by a processing device, a user input indicating a selection of a first region of a digital image;

obtaining, by the processing device, a semantic label that corresponds to the first region;

generating, by the processing device, a guidance attention map of a second region of the digital image, the second region of the digital image being separate from the first region of the digital image;

generating, by the processing device, a sparse attention map of the second region of the digital image having a resolution greater than a resolution of the guidance attention map;

generating, by the processing device, pixels for the first region of the digital image based on the semantic label and the sparse attention map; and

displaying, by the processing device, the digital image with the generated pixels in the first region of the digital image in a user interface.

2. The method as recited in claim 1 , wherein the guidance attention map is generated based on a first model trained using machine learning, and the sparse attention map is generated based on a second model trained using machine learning.

3. The method as recited in claim 1 , wherein the sparse attention map is generated based on the guidance attention map.

4. The method as recited in claim 1 , further comprising splitting the digital image into first portions of the first region and second portions of the second region.

5. The method as recited in claim 4 , wherein the guidance attention map includes a plurality of guidance attention layers, each guidance attention layer corresponding to one of the first portions as a query portion.

6. The method as recited in claim 5 , further comprising generating a guidance attention layer by:

generating an initial attention layer of the second region for the query portion, each of the second portions having a corresponding attention weight; and

determining a guidance attention layer by selecting a subset of the second portions based on the corresponding attention weights.

7. The method as recited in claim 6 , further comprising generating a sparse attention layer based on the guidance attention layer for the query portion, each of the subset of the second portions having a corresponding sparse attention weight.

8. The method as recited in claim 1 , wherein the semantic label indicates a type of attention map.

9. A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

obtaining a digital image, a semantic label, and a first region of the digital image that corresponds to the semantic label;

generating a guidance attention map of a second region of the digital image that is outside of the first region;

generating a sparse attention map of the second region based on the guidance attention map, a resolution of the guidance attention map is less than a resolution of the sparse attention map; and

editing the digital image by generating pixels for the first region based on the semantic label and the sparse attention map.

10. The system as recited in claim 9 , the operations further comprising:

downsampling the digital image to the resolution of the guidance attention map; and

upsampling the guidance attention map to the resolution of the sparse attention map.

11. The system as recited in claim 9 , wherein the guidance attention map is generated based on a first model trained using machine learning, and the sparse attention map is generated based on a second model trained using machine learning.

12. The system as recited in claim 11 , wherein the second model trained using machine learning is trained based on the guidance attention map.

13. The system as recited in claim 12 , the operations further comprising identifying guidance portions of the digital image based on the guidance attention map, and wherein the sparse attention map is generated for the guidance portions.

14. A non-transitory computer-readable storage medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

receiving a semantic input that includes a selection of a first region of a digital image and a semantic label that corresponds to the first region;

generating a guidance attention map of a second region of the digital image that is separate from the first region of the digital image;

generating a sparse attention map of the second region of the digital image having a resolution greater than a resolution of the guidance attention map;

generating pixels for the first region of the digital image based on the semantic label and the sparse attention map; and

presenting the digital image with the generated pixels in the first region of the digital image.

15. The non-transitory computer-readable storage medium as described in claim 14 , wherein the sparse attention map is generated based on the guidance attention map.

16. The non-transitory computer-readable storage medium as described in claim 14 , wherein the guidance attention map is generated based on a guidance transformer model trained using machine learning.

17. The non-transitory computer-readable storage medium as described in claim 14 , wherein the sparse attention map is generated based on a sparse attention model trained using machine learning.

18. The non-transitory computer-readable storage medium as described in claim 14 , wherein the generating the guidance attention map includes downsampling the digital image to the resolution of the guidance attention map.

19. The non-transitory computer-readable storage medium as described in claim 14 , wherein the generating the sparse attention map includes upsampling the guidance attention map to the resolution of the sparse attention map.

20. The non-transitory computer-readable storage medium as described in claim 14 , the operations further comprising:

receiving an additional semantic input that includes an additional selection and an additional semantic label;

identifying one or more dependencies between the semantic input and the additional semantic input; and

determining a label order to process the semantic input and the additional semantic input based on the one or more dependencies.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2022
From: HINZ, TOBIAS; PARK, TAESUNG; ZHANG, RICHARD; FISHER, MATTHEW DAVID
To: ADOBE INC.
Reel/Frame 059917/0174 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2022
From: LIU, DIFAN; KALOGERAKIS, EVANGELOS
To: UNIVERSITY OF MASSACHUSETTS
Reel/Frame 059917/0254 →
Priority Claims (1)
GR 20220100358 · May 3, 2022 · national
Continuity (1)
Related Publication 20230360376A1 · Nov 9, 2023
References Cited (85)
US 20170371347A1 · Cohen · 2017 [cited by examiner]
US 20190333198A1 · Wang · 2019 [cited by examiner]
US 20210090289A1 · Karanam · 2021 [cited by examiner]
US 20220327657A1 · Zheng · 2022 [cited by examiner]
US 20250104291A1 · Yamada · 2025 [cited by examiner]
Bau, David et al., “Semantic Photo Manipulation with a Generative Image Prior”, Cornell University arXiv, arXiv.org [retrieved Jan. 31, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2005.07727.pdf>., Sep. 12… [cited by applicant]
Beltagy, IZ et al., “Longformer: The Long-Document Transformer”, Cornell University arXiv, arXiv.org [retrieved Jan. 31, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2004.05150.pdf>., Dec. 2, 2020, 17 Pages. [cited by applicant]
Caesar, Holger et al., “COCO-Stuff: Thing and Stuff Classes in Context”, Cornell University arXiv, arXiv.org [retrieved Jan. 31, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1612.03716.pdf>., Mar. 28, 2018,… [cited by applicant]
Cao, Chenjie et al., “The Image Local Autoregressive Transformer”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2106.02514.pdf>., Oct. 18, 2021, 20 Pag… [cited by applicant]
Chen, Jiawen et al., “Bilateral guided upsampling”, ACM Transactions on Graphics, vol. 35, No. 6 [retrieved Feb. 2, 2022]. Retrieved from the Internet <https://people.csail.mit.edu/hasinoff/pubs/ChenEtAl16-bgu.pdf>., De… [cited by applicant]
Chen, Liang-Chieh et al., “DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2022]. Retrieved from t… [cited by applicant]
Chen, Mark et al., “Generative Pretraining From Pixels”, Proceedings of the 37th International Conference on Machine Learning [retrieved Feb. 2, 2022]. Retrieved from the Internet <https://www.gwern.net/docs/ai/2020-che… [cited by applicant]
Chen, Qifeng et al., “Photographic Image Synthesis with Cascaded Refinement Networks”, Proceedings of the IEEE International Conference on Computer Vision [retrieved Feb. 2, 2022]. Retrieved from the Internet <https://c… [cited by applicant]
Cheng, Yu et al., “Sequential Attention GAN for Interactive Image Editing”, Proceedings of the 28th ACM International Conference on Multimedia [retrieved Feb. 2, 2022]. Retrieved from the Internet <https://arxiv.org/pdf… [cited by applicant]
Child, Rewon et al., “Generating Long Sequences with Sparse Transformers”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1904.10509.pdf>., Apr. 23, 2019… [cited by applicant]
Chu, Xiangxiang et al., “Conditional Positional Encodings for Vision Transformers”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2102.10882.pdf>., Mar.… [cited by applicant]
Chu, Xiangxiang et al., “Twins: Revisiting the Design of Spatial Attention in Vision Transformers”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2022]. Retrieved from the Internet <http://arxiv-export-lb.libra… [cited by applicant]
De Lutio, Riccardo et al., “Guided Super-Resolution As Pixel-to-Pixel Transformation”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <http://128.84.21.203/pdf/1904.01501>., Au… [cited by applicant]
Dhamo, Helisa et al., “Semantic Image Manipulation Using Scene Graphs”, Cornell University arXiv, arXiv.org [retrieved Feb. 2, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2004.03677.pdf>., Apr. 7, 2020, 15… [cited by applicant]
Dosovitskiy, Alexey et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale”, Cornell University arXiv, arXiv.org [retrieved Aug. 2, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/… [cited by applicant]
Esser, Patrick et al., “ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https:/… [cited by applicant]
Esser, Patrick et al., “Taming Transformers for High-Resolution Image Synthesis”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2012.09841.pdf>., Jun. 2… [cited by applicant]
Ferstl, David et al., “Image Guided Depth Upsampling Using Anisotropic Total Generalized Variation”, Proceedings of the 2013 IEEE International Conference on Computer Vision [retrieved Feb. 3, 2022]. Retrieved from the … [cited by applicant]
Gu, Shuyang et al., “Mask-Guided Portrait Editing With Conditional GANs”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1905.10346.pdf>., May 24, 2019, … [cited by applicant]
Heusel, Martin et al., “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium”, NIPS'17: Proceedings of the 31st International Conference on Neural Information Processing Systems [retrieved F… [cited by applicant]
Hinz, Tobias et al., “Generating Multiple Objects at Spatially Distinct Locations”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1901.00686.pdf>., Jan.… [cited by applicant]
Hinz, Tobias et al., “Improved Techniques for Training Single-Image GANs”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2003.11512.pdf>., Nov. 17, 2020… [cited by applicant]
Holtzman, Ari et al., “The Curious Case of Neural Text Degeneration”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1904.09751.pdf>., Feb. 14, 2020, 16 … [cited by applicant]
Hong, Seunghoon et al., “Learning Hierarchical Semantic Image Manipulation through Structured Representations”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.or… [cited by applicant]
Hui, Tak-Wai et al., “Depth Map Super-Resolution by Deep Multi-Scale Guidance”, European Conference on Computer Vision [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://doi.org/10.1007/978-3-319-46487-9_22>… [cited by applicant]
Isola, Phillip et al., “Image-to-Image Translation with Conditional Adversarial Networks”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) [retrieved Feb. 24, 2022]. Retrieved from the Internet <h… [cited by applicant]
Jiang, Yifan et al., “TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale Up”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/210… [cited by applicant]
Jo, Youngjoo et al., “SC-FEGAN: Face Editing Generative Adversarial Network With User's Sketch and Color”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf… [cited by applicant]
Kitaev, Nikita et al., “Reformer: The Efficient Transformer”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2001.04451.pdf>., Feb. 18, 2020, 12 Pages. [cited by applicant]
Kopf, Johannes et al., “Joint Bilateral Upsampling”, ACM Transactions on Graphics (Proceedings of SIGGRAPH) [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://johanneskopf.de/publications/jbu/paper/FinalPape… [cited by applicant]
Lee, Cheng-Han et al., “MaskGAN: Towards Diverse and Interactive Facial Image Manipulation”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1907.11922.pd… [cited by applicant]
Lewis, Mike et al., “BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension”, Cornell University arXiv, arXiv.org [retrieved Apr. 28, 2022]. Retrieved from the … [cited by applicant]
Ling, Huan et al., “EditGAN: High-Precision Semantic Image Editing”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2111.03186.pdf>., Nov. 4, 2021, 38 Pa… [cited by applicant]
Liu, Guilin et al., “Image Inpainting for Irregular Holes Using Partial Convolutions”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1804.07723.pdf>., D… [cited by applicant]
Liu, Hongyu et al., “DeFLOCNet: Deep Image Editing via Flexible Low-Level Controls”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2103.12723.pdf>., Mar… [cited by applicant]
Liu, Hongyu et al., “PD-GAN: Probabilistic Diverse GAN for Image Inpainting”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2105.02201.pdf>., May 5, 202… [cited by applicant]
Liu, Ming-Yu et al., “Joint Geodesic Upsampling of Depth Images”, Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://www.mer… [cited by applicant]
Liu, Ze et al., “Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2103.14030.pdf>… [cited by applicant]
Loshchilov, Ilya et al., “Decoupled Weight Decay Regularization”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1711.05101.pdf>., Jan. 4, 2019, 19 Pages. [cited by applicant]
Nam, Seonghyeon et al., “Text-Adaptive Generative Adversarial Networks: Manipulating Images with Natural Language”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxi… [cited by applicant]
Ntavelis, Evangelos et al., “SESAME: Semantic Editing of Scenes by Adding, Manipulating or Erasing Objects”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/p… [cited by applicant]
Park, Jaesik et al., “High quality depth map upsampling for 3D-TOF cameras”, Proceedings of the 2011 International Conference on Computer Vision [retrieved Feb. 3, 2022]. Retrieved from the Internet <http://www.cse.york… [cited by applicant]
Park, Taesung et al., “Semantic Image Synthesis With Spatially-Adaptive Normalization”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1903.07291.pdf>., … [cited by applicant]
Parmar, Niki et al., “Image Transformer”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1802.05751.pdf>., Jun. 15, 2018, 10 Pages. [cited by applicant]
Patashnik, Or et al., “StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2103.17249.pdf>., Mar. 31… [cited by applicant]
Ramesh, Aditya et al., “Zero-Shot Text-to-Image Generation”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2102.12092.pdf>., Feb. 26, 2021, 20 Pages. [cited by applicant]
Shaham, Tamar R. et al., “SinGAN: Learning a Generative Model From a Single Natural Image”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1905.01164.pdf… [cited by applicant]
Shaham, Tamar R. et al., “Spatially-Adaptive Pixelwise Networks for Fast Image Translation”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2012.02992.pd… [cited by applicant]
Shocher, Assaf et al., “”Zero-Shot“ Super-Resolution Using Deep Internal Learning”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1712.06087.pdf>., Dec.… [cited by applicant]
Su, Sitong et al., “Fully Functional Image Manipulation Using Scene Graphs in A Bounding-Box Free Way”, Proceedings of the 29th ACM International Conference on Multimedia [retrieved May 26, 2022]. Retrieved from the Int… [cited by applicant]
Suvorov, Roman et al., “Resolution-Robust Large Mask Inpainting With Fourier Convolutions”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2109.07161.pdf… [cited by applicant]
Tan, Zhentao et al., “Diverse Semantic Image Synthesis via Probability Distribution Modeling”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2103.06878.… [cited by applicant]
Tay, Yi et al., “Efficient Transformers: A Survey”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2009.06732.pdf>., Sep. 16, 2020, 28 Pages. [cited by applicant]
Tulsiani, Shubham et al., “PixelTransformer: Sample Conditioned Signal Generation”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2103.15813.pdf>., Mar.… [cited by applicant]
Vaswani, Ashish et al., “Attention Is All You Need”, Cornell University arXiv Preprint, arXiv.org [retrieved Feb. 17, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1706.03762.pdf>., Dec. 6, 2017, 15 pages. [cited by applicant]
Verdoliva, Luisa , “Media Forensics and DeepFakes: An Overview”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2001.06564.pdf>., Jan. 18, 2020, 24 Pages. [cited by applicant]
Mnker, Yael et al., “Image Shape Manipulation From a Single Augmented Training Sample”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://openaccess.thecvf.com/content/IC… [cited by applicant]
Wan, Ziyu et al., “High-Fidelity Pluralistic Image Completion with Transformers”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2103.14031.pdf>., Mar. 2… [cited by applicant]
Wang, Run et al., “FakeSpotter: A Simple yet Robust Baseline for Spotting AI-Synthesized Fake Faces”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1909… [cited by applicant]
Wang, Sheng-Yu et al., “CNN-Generated Images Are Surprisingly Easy to Spot . . . for Now”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1912.11035.pdf>… [cited by applicant]
Wang, Sinong et al., “Linformer: Self-Attention with Linear Complexity”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2006.04768.pdf>., Jun. 14, 2020, … [cited by applicant]
Wang, Wenhai et al., “Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction Without Convolutions”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv… [cited by applicant]
Wang, Xiaolong et al., “Non-local Neural Networks”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1711.07971.pdf>., Apr. 13, 2018, 10 pages. [cited by applicant]
Wang, Zhou et al., “Image Quality Assessment: From Error Visibility to Structural Similarity”, IEEE Transactions on Image Processing, vol. 13, No. 4 [retrieved Feb. 3, 2022]. Retrieved from the Internet <http://www.cns.… [cited by applicant]
Yang, Chao et al., “High-Resolution Image Inpainting Using Multi-Scale Neural Patch Synthesis”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1611.09969… [cited by applicant]
Yang, Jianwei et al., “Focal Self-attention for Local-Global Interactions in Vision Transformers”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2107.00… [cited by applicant]
Yang, Jingyu et al., “Color-Guided Depth Recovery From RGB-D Data Using an Adaptive Autoregressive Model”, IEEE Transactions on Image Processing [retrieved Feb. 3, 2022]. Retrieved from the Internet <http://citeseerx.is… [cited by applicant]
Yang, Qingxiong et al., “Spatial-Depth Super Resolution for Range Images”, IEEE Conference on Computer Vision and Pattern Recognition [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://web.archive.org/web/20… [cited by applicant]
Yu, Jiahui et al., “Free-Form Image Inpainting With Gated Convolution”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1806.03589.pdf>., Oct. 22, 2019, 1… [cited by applicant]
Yu, Jiahui et al., “Generative Image Inpainting With Contextual Attention”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1801.07892.pdf>., Mar. 21, 201… [cited by applicant]
Yu, Tao Yu et al., “Region Normalization for Image Inpainting”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1911.10375.pdf>., Nov. 23, 2019, 9 Pages. [cited by applicant]
Yu, Yingchen et al., “Diverse Image Inpainting with Bidirectional and Autoregressive Transformers”, Cornell University arXiv, arXiv.org [retrieved Feb. 3, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2104.1… [cited by applicant]
Zaheer, Manzil et al., “Big Bird: Transformers for Longer Sequences”, Cornell University arXiv, arXiv.org [retrieved Feb. 4, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2007.14062.pdf>., Jan. 8, 2021, 42 P… [cited by applicant]
Zhang, Pan et al., “Cross-Domain Correspondence Learning for Exemplar-Based Image Translation”, Cornell University arXiv, arXiv.org [retrieved Feb. 4, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2004.05571… [cited by applicant]
Zhang, Pengchuan et al., “Multi-Scale Vision Longformer: A New Vision Transformer for High- Resolution Image Encoding”, Cornell University arXiv, arXiv.org [retrieved Feb. 4, 2022]. Retrieved from the Internet <https://… [cited by applicant]
Zhang, Richard et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”, Cornell University arXiv, arXiv.org [retrieved Feb. 17, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1801.039… [cited by applicant]
Zheng, Haitian et al., “Semantic Layout Manipulation with High-Resolution Sparse Attention”, Cornell University arXiv, arXiv.org [retrieved Feb. 4, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2012.07288.pd… [cited by applicant]
Zhou, Bolei et al., “Scene Parsing Through ADE20K Dataset”, EEE Conference on Computer Vision and Pattern Recognition [retrieved Feb. 4, 2022]. Retrieved from the Internet <https://people.csail.mit.edu/xavierpuig/docume… [cited by applicant]
Zhou, Xingran et al., “CoCosNet v2: Full-Resolution Correspondence Learning for Image Translation”, IEEE/CVF Conference on Computer Vision and Pattern Recognition [retrieved Feb. 4, 2022]. Retrieved from the Internet <h… [cited by applicant]
Zhu, Peihao et al., “SEAN: Image Synthesis With Semantic Region-Adaptive Normalization”, Cornell University arXiv, arXiv.org [retrieved Feb. 4, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1911.12861.pdf>.,… [cited by applicant]
Cited By (1)
US 12,548,307