Guided contextual attention map for inpainting tasks
Systems and methods for augmenting data can leverage one or more machine-learned models and contextual attention data to provide more realistic and efficient data augmentation. For example, systems and methods for inpainting can leverage a machine-learned model to generate predicted contextual attention data and blend the predicted contextual attention data with obtained contextual attention data to determine replacement data for augmenting an image to replace one or more occlusions. The obtained contextual attention data can include user-guided contextual attention.
1 . A computer-implemented method for training an inpainting model, the method comprising:
receiving, by a computing system comprising one or more processors, an input image and a ground truth image, wherein the ground truth image depicts a scene, and wherein the input image depicts the scene with one or more occlusions;
processing, by the computing system, the ground truth image with a contextual attention model to generate a contextual attention output;
processing, by the computing system, the input image and the contextual attention output with an augmentation model to generate a prediction image, wherein processing the input image and the contextual attention output with the augmentation model comprises:
processing the input image to generate predicted contextual attention data;
generating blended data based on the predicted contextual attention data and the contextual attention output;
generating the prediction image based on the blended data and the input image;
evaluating, by the computing system, a loss function that evaluates a difference between the prediction image and the ground truth image; and
adjusting, by the computing system, one or more parameters of the augmentation model based at least in part on the loss function.
2 . The computer-implemented method of claim 1 , wherein the augmentation model comprises a prediction model, a blend model, and an occlusion model, and wherein processing the input image and the contextual attention output with the augmentation model comprises:
processing, by the computing system, the input image with the prediction model to generate the predicted contextual attention data;
processing, by the computing system, the predicted contextual attention data and the contextual attention output with a blend model to generate the blended data;
processing, by the computing system, the blended data and the input image to generate the prediction image.
3 . The computer-implemented method of claim 2 , wherein the blend model is trained to randomly blend the predicted contextual attention data and the contextual attention output.
4 . The computer-implemented method of claim 1 , wherein the input image is generated by adding one or more occlusions to the ground truth image.
5 . The computer-implemented method of claim 1 , wherein the contextual attention model comprises a convolutional neural network and one or more contextual attention blocks.
6 . The computer-implemented method of claim 1 , wherein the contextual attention model is trained by:
processing, by the computing system, one or more training images with the contextual attention model to generate training contextual attention outputs;
processing, by the computing system, the training contextual attention outputs with an inpainting model to generate a training augmented image;
evaluating, by the computing system, a training loss function that evaluates a difference between the training augmented image and the ground truth image; and
adjusting, by the computing system, one or more contextual attention parameters of the contextual attention model based at least in part on the training loss function.
7 . The computer-implemented method of claim 1 , further comprising:
receiving, by the computing system, one or more inputs descriptive of a selection of a portion of the input image; and
wherein the prediction image is generated based at least in part on the one or more inputs.