IP Library Granted Patent US 12,646,289
Granted Patent B2
US 12,646,289 · App. 18/336,423 · Granted Jun 2, 2026

Visual grounding of self-supervised representations for machine learning models utilizing difference attention

Inventors: Aishwarya Agarwal (Bengaluru, IN); Srikrishna Karanam (Bangalore, IN); Balaji Vasan Srinivasan (Bangalore, IN)
Assignee: Adobe Inc.
G06V10/751G06V10/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,289
App. No.
18/336,423
Granted
Jun 2, 2026
Kind
B2
Abstract

The present disclosure relates to systems, non-transitory computer-readable media, and methods for utilizing difference attention to evaluate and/or train machine learning models. In particular, in some embodiments, the disclosed systems generate, utilizing a machine learning model, a first feature vector from a digital image. In one or more implementations, the disclosed systems generate a masked digital image by masking a region from the digital image. Additionally, in some embodiments, the disclosed systems generate, utilizing the machine learning model, a second feature vector from the masked digital image. Moreover, in some implementations, the disclosed systems determine a difference feature vector between the first feature vector and the second feature vector. Furthermore, in some embodiments, the disclosed systems generate, from the difference feature vector, a difference attention map reflecting a visual grounding of the machine learning model relative to the region.

Claims (49)

1 . A computer-implemented method comprising:

generating, utilizing a machine-learning model, an image feature vector from a digital image;

generating a masked digital image by masking a region from the digital image;

generating, utilizing the machine-learning model, a masked image feature vector from the masked digital image;

determining a masked-image-difference feature vector by determining a difference between the image feature vector of the digital image and the masked image feature vector of the masked digital image; and

generating a difference attention map reflecting a visual grounding of the machine-learning model relative to the region by comparing activation maps of layers of the machine-learning model with the masked-image-difference feature vector, generated from the image feature vector of the digital image and the masked image feature vector of the masked digital image.

2 . The computer-implemented method of claim 1 , further comprising providing the difference attention map for display via a graphical user interface of a client device.

3 . The computer-implemented method of claim 1 , wherein generating the masked digital image comprises utilizing a saliency detection machine learning model to mask the region from the digital image.

4 . The computer-implemented method of claim 1 , wherein generating the difference attention map comprises determining a scalar signal from the masked-image-difference feature vector.

5 . The computer-implemented method of claim 4 , wherein generating the difference attention map further comprises generating a gradient matrix based on the scalar signal with respect to a feature map of the digital image.

6 . The computer-implemented method of claim 4 , wherein determining the scalar signal from the masked-image-difference feature vector comprises performing a thresholding operation on the masked-image-difference feature vector.

7 . The computer-implemented method of claim 4 , wherein determining the scalar signal from the masked-image-difference feature vector comprises determining an inner product of the masked-image-difference feature vector and the image feature vector.

8 . The computer-implemented method of claim 1 , further comprising modifying parameters of the machine-learning model by comparing the difference attention map and a saliency map of the digital image.

9 . A system comprising:

one or more memory devices comprising a digital image, a masked digital image comprising a portion of the digital image, and a machine-learning model; and

one or more processors configured to cause the system to:

generate, utilizing the machine-learning model, an image feature vector from the digital image;

generate, utilizing the machine-learning model, a masked image feature vector from the masked digital image;

generate a difference attention map by comparing activation maps of layers of the machine-learning model based on a difference between the image feature vector of the digital image and the masked image feature vector of the masked digital image; and

modify parameters of the machine-learning model by comparing the difference attention map and a saliency map of the digital image.

10 . The system of claim 9 , wherein the one or more processors are further configured to modify the parameters of the machine-learning model by:

determining a difference attention loss by comparing the difference attention map and the saliency map; and

training at least one of a classification machine-learning model, an object detection machine-learning model, or a segmentation machine-learning model based on the difference attention loss.

11 . The system of claim 9 , wherein the one or more processors are further configured to generate the difference attention map by:

determining a masked-image-difference feature vector between the image feature vector and the masked image feature vector; and

determining a differentiable scalar signal from the masked-image-difference feature vector.

12 . The system of claim 11 , wherein determining the differentiable scalar signal comprises combining the masked-image-difference feature vector and the image feature vector.

13 . The system of claim 11 , wherein the one or more processors are further configured to generate the difference attention map by generating a set of gradient matrices based on the differentiable scalar signal with respect to a set of feature maps of the digital image.

14 . The system of claim 9 , wherein comparing the difference attention map and the saliency map comprises determining an inner product of the difference attention map and the saliency map.

15 . A non-transitory computer-readable medium storing executable instructions that, when executed by a processing device, cause the processing device to perform operations comprising:

generating, utilizing a machine-learning model, an image feature vector from a digital image;

generating a masked digital image by masking a region from the digital image;

generating, utilizing the machine-learning model, a masked image feature vector from the masked digital image;

generating, based on a difference between the image feature vector of the digital image and the masked image feature vector of the masked digital image, a gradient matrix with respect to a feature map of the digital image; and

generating, based on the gradient matrix, a difference attention map that reflects a visual grounding of the machine-learning model relative to the region by comparing activation maps of layers of the machine-learning model with the difference between the image feature vector of the digital image and the masked image feature vector of the masked digital image.

16 . The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise providing the difference attention map for display with the digital image via a graphical user interface.

17 . The non-transitory computer-readable medium of claim 15 , wherein generating the gradient matrix with respect to the feature map of the digital image comprises:

determining a masked-image-difference feature vector between the image feature vector and the masked image feature vector; and

differentiating a scalar signal of the masked-image-difference feature vector with respect to the feature map.

18 . The non-transitory computer-readable medium of claim 15 , wherein generating the difference attention map comprises:

generating an additional gradient matrix with respect to an additional feature map having a second resolution different from a first resolution of the feature map; and

combining the gradient matrix and the additional gradient matrix.

19 . The non-transitory computer-readable medium of claim 18 , wherein combining the gradient matrix and the additional gradient matrix comprises:

generating a first weight by pooling the gradient matrix;

generating a second weight by pooling the additional gradient matrix; and

determining a feature sum by combining the feature map and the additional feature map utilizing the first weight and the second weight.

20 . The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

determining a difference attention loss by comparing the difference attention map and a saliency map; and

modifying parameters of the machine-learning model based on the difference attention loss.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2023
From: AGARWAL, AISHWARYA; KARANAM, SRIKRISHNA; SRINIVASAN, BALAJI VASAN
To: ADOBE INC.
Reel/Frame 063975/0509 →
Continuity (1)
Related Publication 20240420447A1 · Dec 19, 2024
References Cited (45)
US 20210272258A1 · Sharma · 2021 [cited by examiner]
US 20240255938A1 · Brikis · 2024 [cited by examiner]
US 20240259668A1 · Li · 2024 [cited by examiner]
Selvaraju, Ramprasaath R., et al. “CASTing Your Model: Learning to Localize Improves Self-Supervised Representations.” arXiv preprint arXiv:2012.04630 (2020).https://arxiv.org/abs/2012.04630 (Year: 2020). [cited by examiner]
Selvaraju, Ramprasaath R., et al. “Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization.” arXiv preprint arXiv:1610.02391 (2019).https://arxiv.org/abs/1610.02391 (Year: 2019). [cited by examiner]
Zheng, Meng, et al. “Visual similarity attention.” arXiv preprint arXiv: 1911.07381 (2019).https://arxiv.org/abs/1911.07381 (Year: 2022). [cited by examiner]
Adam Coates and Andrew Y Ng. Learning feature representations with k-means. In Neural networks: Tricks of the trade, pp. 561-580. Springer, 2012. [cited by applicant]
Aditya Chattopadhay, Anirban Sarkar, Prantik Howlader, and Vineeth N Balasubramanian. Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks. In 2018 IEEE winter conference on applica… [cited by applicant]
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp.… [cited by applicant]
Carsten Rother, Vladimir Kolmogorov, and Andrew Blake. “GrabCut”: interactive foreground extraction using iterated graph cuts. ACM transactions on graphics (TOG), 23(3):309-314, 2004. [cited by applicant]
Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked feature prediction for self-supervised visual pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision a… [cited by applicant]
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros. Context encoders: Feature learning by inpainting. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp… [cited by applicant]
Fei Tian, Bin Gao, Qing Cui, Enhong Chen, and Tie-Yan Liu. Learning deep representations for graph clustering. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 28, 2014. [cited by applicant]
Haofan Wang, Zifan Wang, Mengnan Du, Fan Yang, Zijian Zhang, Sirui Ding, Piotr Mardziel, and Xia Hu. Score-cam: Score-weighted visual explanations for convolutional neural networks. In Proceedings of the IEEE/CVF confer… [cited by applicant]
Jianwei Yang, Devi Parikh, and Dhruv Batra. Joint unsupervised learning of deep representations and image clusters. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5147-5156, 2016. [cited by applicant]
Kai Xiao, Logan Engstrom, Andrew Ilyas, and Aleksander Madry. Noise or signal: The role of image backgrounds in object recognition. arXiv preprint arXiv:2006.09994, 2020. [cited by applicant]
Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pp. 2961-2969, 2017. [cited by applicant]
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p… [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778, 2016. [cited by applicant]
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, … [cited by applicant]
Karan Desai and Justin Johnson. Virtex: Learning visual representations from textual annotations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11162-11173, 2021. [cited by applicant]
Lei Chen, Jianhui Chen, Hossein Hajimirsadeghi, and Greg Mori. Adapting grad-cam for embedding networks. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 2794-2803, 2020. [cited by applicant]
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions … [cited by applicant]
Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European conference on computer… [cited by applicant]
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results. http://www.pascal-network.org/challenges/VOC/voc2007/workshop/index.html. [cited by applicant]
Mark Everingham, SM Eslami, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes challenge: A retrospective. International journal of computer vision, 111(1):98-136, 2… [cited by applicant]
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. Advances in Neural Information Processing System… [cited by applicant]
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In Proceedings of the European conference on computer vision (ECCV), pp. 132-149, 2018. [cited by applicant]
Mathilde Caron, Piotr Bojanowski, Julien Mairal, and Armand Joulin. Unsupervised pre-training of image features on non-curated data. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2959-2… [cited by applicant]
Mehdi Noroozi and Paolo Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. In European conference on computer vision, pp. 69-84. Springer, 2016. [cited by applicant]
Meng Zheng, Srikrishna Karanam, Terrence Chen, Richard J. Radke, and Ziyan Wu. Visual similarity attention. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22, pp. 172… [cited by applicant]
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International j… [cited by applicant]
Peng-Tao Jiang, Chang-Bin Zhang, Qibin Hou, Ming-Ming Cheng, and Yunchao Wei. Layercam: Exploring hierarchical class activation maps for localization. IEEE Transactions on Image Processing, 30:5875-5888, 2021. [cited by applicant]
Raia Hadsell, Sumit Chopra, and Yann LeCun. Dimensionality reduction by learning an invariant mapping. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'06), vol. 2, pp. 1735-1742… [cited by applicant]
Ramprasaath R Selvaraju, Karan Desai, Justin Johnson, and Nikhil Naik. Casting your model: Learning to localize improves self-supervised representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and … [cited by applicant]
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE i… [cited by applicant]
Richard Zhang, Phillip Isola, and Alexei AEfros. Split-brain autoencoders: Unsupervised learning by cross-channel prediction. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1058-10… [cited by applicant]
S Ren, K He, R Girshick, and J Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6):1137-1149, 2016. [cited by applicant]
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28, 2015. [cited by applicant]
Tam Nguyen, Maximilian Dax, Chaithanya Kumar Mummadi, Nhung Ngo, Thi Hoai Phuong Nguyen, Zhongyu Lou, and Thomas Brox. Deepusps: Deep robust unsupervised saliency prediction via self-supervision. Advances in Neural Info… [cited by applicant]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pp. 1597-1607. PMLR, 2020. [cited by applicant]
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pp. 740-7… [cited by applicant]
Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recogniti… [cited by applicant]
Wenqian Liu, Runze Li, Meng Zheng, Srikrishna Karanam, Ziyan Wu, Bir Bhanu, Richard J Radke, and Octavia Camps. Towards visually explaining variational autoencoders. In Proceedings of the IEEE/CVF Conference on Computer… [cited by applicant]
Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15750-15758, 2021. [cited by applicant]