IP Library › Granted Patent US 12,743,755
Granted Patent B2
US 12,743,755 · App. 18/580,507 · Granted Sep 22, 2026

Guided contextual attention map for inpainting tasks

Inventors: Noritsugu Kanazawa (Campbell, CA); Neal Wadhwa (Cambridge, MA); Yael Pritch Knaan (Mountain View, CA); Kfir Aberman (San Mateo, CA)
Assignee: GOOGLE LLC
G06T5/77G06T5/60G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,755
App. No.
18/580,507
Granted
Sep 22, 2026
Kind
B2
Abstract

Systems and methods for augmenting data can leverage one or more machine-learned models and contextual attention data to provide more realistic and efficient data augmentation. For example, systems and methods for inpainting can leverage a machine-learned model to generate predicted contextual attention data and blend the predicted contextual attention data with obtained contextual attention data to determine replacement data for augmenting an image to replace one or more occlusions. The obtained contextual attention data can include user-guided contextual attention.

Claims (24)

1 . A computer-implemented method for training an inpainting model, the method comprising:

receiving, by a computing system comprising one or more processors, an input image and a ground truth image, wherein the ground truth image depicts a scene, and wherein the input image depicts the scene with one or more occlusions;

processing, by the computing system, the ground truth image with a contextual attention model to generate a contextual attention output;

processing, by the computing system, the input image and the contextual attention output with an augmentation model to generate a prediction image, wherein processing the input image and the contextual attention output with the augmentation model comprises:

processing the input image to generate predicted contextual attention data;

generating blended data based on the predicted contextual attention data and the contextual attention output;

generating the prediction image based on the blended data and the input image;

evaluating, by the computing system, a loss function that evaluates a difference between the prediction image and the ground truth image; and

adjusting, by the computing system, one or more parameters of the augmentation model based at least in part on the loss function.

2 . The computer-implemented method of claim 1 , wherein the augmentation model comprises a prediction model, a blend model, and an occlusion model, and wherein processing the input image and the contextual attention output with the augmentation model comprises:

processing, by the computing system, the input image with the prediction model to generate the predicted contextual attention data;

processing, by the computing system, the predicted contextual attention data and the contextual attention output with a blend model to generate the blended data;

processing, by the computing system, the blended data and the input image to generate the prediction image.

3 . The computer-implemented method of claim 2 , wherein the blend model is trained to randomly blend the predicted contextual attention data and the contextual attention output.

4 . The computer-implemented method of claim 1 , wherein the input image is generated by adding one or more occlusions to the ground truth image.

5 . The computer-implemented method of claim 1 , wherein the contextual attention model comprises a convolutional neural network and one or more contextual attention blocks.

6 . The computer-implemented method of claim 1 , wherein the contextual attention model is trained by:

processing, by the computing system, one or more training images with the contextual attention model to generate training contextual attention outputs;

processing, by the computing system, the training contextual attention outputs with an inpainting model to generate a training augmented image;

evaluating, by the computing system, a training loss function that evaluates a difference between the training augmented image and the ground truth image; and

adjusting, by the computing system, one or more contextual attention parameters of the contextual attention model based at least in part on the training loss function.

7 . The computer-implemented method of claim 1 , further comprising:

receiving, by the computing system, one or more inputs descriptive of a selection of a portion of the input image; and

wherein the prediction image is generated based at least in part on the one or more inputs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2024
From: KANAZAWA, NORITSUGU; ABERMAN, KFIR; PRITCH KNAAN, YAEL; WADHWA, NEAL
To: GOOGLE LLC
Reel/Frame 066226/0491 →
Continuity (1)
Related Publication 20250086760A1 · Mar 13, 2025
References Cited (22)
US 20180108137A1 · Price · 2018 [cited by examiner]
US 20180239867A1 · Kopylov · 2018 [cited by examiner]
US 20190238568A1 · Goswami · 2019 [cited by examiner]
US 20210217145A1 · El-Khamy et al. · 2021 [cited by applicant]
JP 2014142836 · 2014 [cited by applicant]
JP 2017058930 · 2017 [cited by applicant]
Liu et al., “Coherent Semantic Attention for Image Inpainting”, 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (Year: 2019). [cited by examiner]
Lizuka et al., “Globally and locally consistent image completion”, ACM Transactions on Graphics (TOG), vol. 36, Issue 4 (Aug. 2017) (Year: 2017). [cited by examiner]
Takahashi et al., “RICAP: Random Image Cropping and Patching Data Augmentation for Deep CNNs”, Proceedings of Machine Learning Research 95:786-798, 2018 (Year: 2018). [cited by examiner]
Cun et al, “Improving the Harmony of the Composite Image by Spatial-Separated Attention Module”, IEEE Transactions on Image Processing, vol. 29, 2020 (Year: 2020). [cited by examiner]
Zhao et al. “Guided Image Inpainting: Replacing an Image Region by Pulling Content from Another Image”, 2019 IEEE Winter Conference on Applications of Computer Vision (Year: 2019). [cited by examiner]
Hirahara et al., “Noise Removal and Defect Restoration for Sea Surface Temperature Images Using Adversarial Physics Model Loss”, The Institute of Electronics, Information and Communication Engineers, IEICE Technical Rep… [cited by applicant]
Yu et al., “Free-Form Image Inpainting with Gated Convolution”, IEEE/CVF International Conference on Computer Vision (ICCV), 2019, 10 pages. [cited by applicant]
International Preliminary Report on Patentability for Application No. PCT/US2021/042150, mailed Feb. 1, 2024, 12 pages. [cited by applicant]
github.com, “JiahuiYu/generative_inpainting”, https://github.com/Jiahui Yu/generative_inpainting, retrieved on Jan. 18, 2024, 3 pages. [cited by applicant]
International Search Report for PCT/US2021/042150, mailed on Jun. 13, 2022, 4 pages. [cited by applicant]
Mohite et al., “Image Inpainting with Contextual Attention and Partial Convolution”, 2020 International Conference on Artificial Intelligence and Signal Processing (AISP), Jan. 10-12, 2020, Amaravati, India, 6 pages. [cited by applicant]
Xiao et al., “Generative Image Inpainting by Hybrid Contextual Attention Network”, MultiMedia Modeling: 27th International Conference, Jun. 22-24, 2021, Prague, Czech Republic, pp. 162-173. [cited by applicant]
Xu et al., “Deep Flow-Guided Video Inpainting” arXiv:1905.02884v1, May 8, 2019, 10 pages. [cited by applicant]
Yu et al., “Generative Image Inpainting with Contextual Attention”, arXiv:1801.07892v2, Mar. 21, 2018, 15 pages. [cited by applicant]
Zhang et al., “Text-Guided Neural Image Inpainting”, arXiv:2004.03212v4, Mar. 22, 2021, 9 pages. [cited by applicant]
Zhao et al., “Guided Image Inpainting: Replacing an Image Region by Pulling Content from Another Image”, arXiv: 1803.08435v1, Mar. 22, 2018, 28 pages. [cited by applicant]