IP Library › Granted Patent US 12,230,016
Granted Patent B2
US 12,230,016 · App. 17/799,740 · Granted Feb 18, 2025

Explanation of machine-learned models using image translation

Inventors: Arunachalam Narayanaswamy (Sunnyvale, CA); Subhashini Venugopalan (Mountain View, CA); Avinash Vaidyanathan Varadarajan (Los Altos, CA)
Assignee: Google LLC
G06V10/776G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,230,016
App. No.
17/799,740
Granted
Feb 18, 2025
Kind
B2
Abstract

Systems and methods for identifying visual features that influence a predictive model are provided. The technology employs an image translation function to introduce a visual feature into an image to create a modified image that can be fed to a predictive model. When the predictive model generates a different prediction for a given image than it does for a modified version of that image, the image translation function can then be used to make further modified versions that exaggerate the introduced visual feature. The technology thus aids in identifying visual features that influence the predictive model so that the model's conclusions can be understood, and so that those visual features can be further studied and tested.

Claims (36)

1. A computer-implemented method for identifying visual features impacting model prediction, the method comprising:

generating, by one or more processors of a processing system, a first prediction based on a first image using a predictive model;

generating, by the one or more processors, a second prediction based on a second image using the predictive model, wherein the second image includes a visual feature created by modifying at least a portion of the first image using a translation function, and the second prediction is different than the first prediction; and

modifying, by the one or more processors, at least a portion of the second image using the translation function to create a third image in which the visual feature is amplified relative to the second image;

wherein the translation function is configured to translate a first class of imagery to appear more like a second class of imagery.

2. The method of claim 1 , further comprising:

modifying, by the one or more processors, at least a portion of the third image using the translation function to create a fourth image in which the visual feature is amplified relative to the third image.

3. The method of claim 1 , wherein the translation function is generated using a generative adversarial network.

4. The method of claim 1 , wherein the second image is generated using a generative adversarial network.

5. The method of claim 1 , wherein the predictive model is a neural network.

6. The method of claim 1 , wherein the visual feature included in the second image is created by modifying only a portion of the first image using the translation function.

7. The method of claim 6 , further comprising identifying, by the one or more processors, the portion of the first image using a spatial explanation model.

8. The method of claim 7 , wherein the spatial explanation model is a perturbation-based model.

9. The method of claim 7 , wherein the spatial explanation model is a backpropagation-based model.

10. The method of claim 6 , further comprising:

generating, by the one or more processors, a third prediction based on an ablated version of the first image using the predictive model, the third prediction being different than the first prediction; and

identifying, by the one or more processors, the portion of the first image based on the ablated version of the first image.

11. A processing system configured to identify visual features impacting model prediction, the processing system comprising:

a memory; and

one or more processors coupled to the memory and configured to:

generate a first prediction based on a first image using a predictive model;

generate a second prediction based on a second image using the predictive model, wherein the second image includes a visual feature created by modifying at least a portion of the first image using a translation function, and the second prediction is different than the first prediction; and

modify at least a portion of the second image using the translation function to create a third image in which the visual feature is amplified relative to the second image;

wherein the translation function is configured to translate a first class of imagery to appear more like a second class of imagery.

12. The system of claim 11 , wherein the one or more processors are further configured to:

modify at least a portion of the third image using the translation function to create a fourth image in which the visual feature is amplified relative to the third image.

13. The system of claim 11 , wherein the translation function is generated using a generative adversarial network.

14. The system of claim 11 , wherein the second image is generated using a generative adversarial network.

15. The system of claim 11 , wherein the predictive model is a neural network.

16. The system of claim 11 , wherein the visual feature included in the second image is created by modifying only a portion of the first image using the translation function.

17. The system of claim 16 , wherein the one or more processors are further configured to identify the portion of the first image using a spatial explanation model.

18. The system of claim 17 , wherein the spatial explanation model is a perturbation-based model.

19. The system of claim 17 , wherein the spatial explanation model is a backpropagation-based model.

20. The system of claim 16 , wherein the one or more processors are further configured to:

generate a third prediction based on an ablated version of the first image using the predictive model, the third prediction being different than the first prediction; and

identify the portion of the first image based on the ablated version of the first image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2022
From: NARAYANASWAMY, ARUNACHALAM; VENUGOPALAN, SUBHASHINI; VARADARAJAN, AVINASH VAIDYANATHAN
To: GOOGLE LLC
Reel/Frame 060805/0946 →
Continuity (1)
Related Publication 20230108319A1 · Apr 6, 2023
References Cited (78)
US 9271133B2 · Rodriguez · 2016 [cited by applicant]
US 10361802B1 · Hoffberg-Borghesani · 2019 [cited by examiner]
US 11188795B1 · Lee · 2021 [cited by examiner]
US 20110143811A1 · Rodriguez · 2011 [cited by examiner]
US 20140080428A1 · Rhoads · 2014 [cited by examiner]
US 20190180441A1 · Peng · 2019 [cited by examiner]
US 20190197357A1 · Anderson · 2019 [cited by examiner]
US 20200097858A1 · Baikalov · 2020 [cited by examiner]
US 20200184278A1 · Zadeh · 2020 [cited by examiner]
US 20200193075A1 · Allen · 2020 [cited by examiner]
US 20200302318A1 · Hetherington · 2020 [cited by examiner]
US 20210049503A1 · Nourian · 2021 [cited by examiner]
US 20210142161A1 · Huang · 2021 [cited by examiner]
US 20210166151A1 · Kennel · 2021 [cited by examiner]
US 20210182713A1 · Kar · 2021 [cited by examiner]
CA 3034644A1 · 2018 [cited by examiner]
CA 3079209A1 · 2019 [cited by examiner]
WO WO2010022185A1 · 2010 [cited by examiner]
International Search Report and Written Opinion for Application No. PCT/US20/20773 dated Nov. 9, 2020. [cited by applicant]
Weina Jin et al: “Applying Artificial Intelligence to Glioma Imaging: Advances and Challenges”, Arxiv.Org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov. 28, 2019, XP081541724. [cited by applicant]
Arrieta, Alejandro Barredo, et al., “Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI”, Arxiv.Org, Cornell University Library, 201 Olin Library Cornell … [cited by applicant]
Selvaraju, Ramprasaath R, et al., “Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization”, 2017 IEEE International Conference On Computer Vision (ICCV), IEEE, Mar. 21, 2017, pp. 618-626, XP033… [cited by applicant]
A.V. Varadarajan, et al. Predicting Optical Coherence Tomography-derived Diabetic Macular Edema Grades from Fundus Photographs Using Deep Learning. Nat Commun 11, 130 (2020). 8 Pages. [cited by applicant]
Alexander Mordvintsev, et al. DeepDream—a code example for visualizing Neural Networks. Google AI Blog, Jul. 1, 2015. Printed from: https://ai.googleblog.com/2015/07/deepdream-code-example-for-visualizing.html on Aug. 3… [cited by applicant]
Amit Dhurandhar, et al. Explanations Based on the Missing: Towards Contrastive Explanations with Pertinent Negatives. NIPS'18: Proceedings of the 32nd International Conference on Neural Information Processing Systems, D… [cited by applicant]
Andrei Kapishnikov, et al. XRAI: Better Attributions Through Regions. In Proceedings of the IEEE International Conference on Computer Vision, pp. 4948-4957, 2019. 10 Pages. [cited by applicant]
Andrew Miller, et al. Discriminative Regularization for Latent Variable Models with Applications to Electrocardiography. In Proceedings of the 36th Int'l Conf on Machine Learning, vol. 97, pp. 4585-4594, Jun. 9-15, 2019… [cited by applicant]
Aravindh Mahendran, et al. Understanding Deep Image Representations by Inverting Them. arXiv:1412.0035v1, Nov. 26, 2014. 9 Pages. [cited by applicant]
Avinash Varadarajan, et al. Predicting Optical Coherence Tomography-derived Diabetic Macular Edema Grades from Fundus Photographs Using Deep Learning. arXiv:1810.10342v1, Oct. 18, 2018. 30 Pages. [cited by applicant]
Avinash Varadarajan, et al. Predicting Optical Coherence Tomography-derived Diabetic Macular Edema Grades from Fundus Photographs Using Deep Learning. arXiv:1810.10342v2, Nov. 19, 2018. 31 Pages. [cited by applicant]
Avinash Varadarajan, et al. Predicting Optical Coherence Tomography-derived Diabetic Macular Edema Grades from Fundus Photographs Using Deep Learning. arXiv:1810.10342v3, Feb. 9, 2019. 35 Pages. [cited by applicant]
Avinash Varadarajan, et al. Predicting Optical Coherence Tomography-derived Diabetic Macular Edema Grades from Fundus Photographs Using Deep Learning. arXiv:1810.10342v4, Jul. 31, 2019. 38 Pages. [cited by applicant]
Casey Chu, et al. CycleGAN, A Master of Steganography. arXiv:1712.02950v1, Dec. 8, 2017. 5 Pages. [cited by applicant]
Casey Chu, et al. CycleGAN, A Master of Steganography. arXiv:1712.02950v2, Dec. 16, 2017. 6 Pages. [cited by applicant]
Chun-Hao Chang, et al. Explaining Image Classifiers by Counterfactual Generation. arXiv:1807.08024v1, Jul. 20, 2018. 11 Pages. [cited by applicant]
Chun-Hao Chang, et al. Explaining Image Classifiers by Counterfactual Generation. arXiv:1807.08024v2, Oct. 11, 2018. 15 Pages. [cited by applicant]
Chun-Hao Chang, et al. Explaining Image Classifiers by Counterfactual Generation. arXiv:1807.08024v3, Feb. 25, 2019. 19 Pages. [cited by applicant]
Daniel Smilkov, et al. SmoothGrad: removing noise by adding noise. arXiv:1706.03825v1, Jun. 12, 2017. 10 Pages. [cited by applicant]
David Bau, et al. Network Dissection: Quantifying Interpretability of Deep Visual Representations. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017): 3319-3327. [cited by applicant]
Ian J. Goodfellow, et al. Generative Adversarial Nets. NIPS'14: Proceedings of the 27th International Conference on Neural Information Processing Systems, vol. 2, Dec. 2014, pp. 2672-2680. 9 Pages. [cited by applicant]
Ishaan Gulrajani, et al. Improved Training of Wasserstein GANs. NIPS'17: Proceedings of the 31st International Conference on Neural Information Processing Systems, Dec. 2017, pp. 5769-5779. 11 Pages. [cited by applicant]
Jonathan Krause, et al. Grader Variability and the Importance of Reference Standards for Evaluating Machine earning Models for Diabetic Retinopathy. arXiv:1710.01711v1, Oct. 4, 2017. 18 Pages. [cited by applicant]
Jonathan Krause, et al. Grader Variability and the Importance of Reference Standards for Evaluating Machine earning Models for Diabetic Retinopathy. arXiv:1710.01711v2, May 30, 2018. 23 Pages. [cited by applicant]
Jonathan Krause, et al. Grader Variability and the Importance of Reference Standards for Evaluating Machine Learning Models for Diabetic Retinopathy. arXiv:1710.01711v3, Jul. 3, 2018. 23 Pages. [cited by applicant]
Joseph Paul Cohen, et al. Distribution Matching Losses Can Hallucinate Features in Medical Image Translation. arXiv:1805.08841v1, May 22, 2018. 11 Pages. [cited by applicant]
Joseph Paul Cohen, et al. Distribution Matching Losses Can Hallucinate Features in Medical Image Translation. arXiv:1805.08841v2, Oct. 3, 2018. 11 Pages. [cited by applicant]
Joseph Paul Cohen, et al. Distribution Matching Losses Can Hallucinate Features in Medical Image Translation. arXiv:1805.08841v3, Oct. 3, 2018. 11 Pages. [cited by applicant]
Jost Tobias Springenberg, et al. Striving for Simplicity: The All Convolutional Net. arXiv:1412.6806v1, Dec. 21, 2014. 12 Pages. [cited by applicant]
Jost Tobias Springenberg, et al. Striving for Simplicity: The All Convolutional Net. arXiv:1412.6806v2, Mar. 2, 2015. 14 Pages. [cited by applicant]
Jost Tobias Springenberg, et al. Striving for Simplicity: The All Convolutional Net. arXiv:1412.6806v3, Apr. 13, 2015. 14 Pages. [cited by applicant]
Jun-Yan Zhu, et al. Toward Multimodal Image-to-Image Translation. NIPS'17: Proceedings of the 31st International Conference on Neural Information Processing Systems, Dec. 2017, pp. 465-476. [cited by applicant]
Jun-Yan Zhu, et al. Unpaired Image-to-Image Translation Using Cycle-consistent Adversarial Networks. In Proceedings of the IEEE international conference on computer vision, pp. 2223-2232, 2017. 10 Pages. [cited by applicant]
Marco Tulio Ribeiro, et al. Why should I trust you ?: Explaining the Predictions of Any Classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, pp. 1135-1144. A… [cited by applicant]
Matthew D. Zeiler, et al. Visualizing and Understanding Convolutional Networks. Computer Vision—ECCV 2014. ECCV 2014. Lecture Notes in Computer Science, vol. 8689, pp. 818-833. 16 Pages. [cited by applicant]
Mukund Sundararajan, et al. Axiomatic Attribution for Deep Networks. ICML'17: Proceedings of the 34th International Conference on Machine Learning, vol. 70, pp. 3319-3328. Aug. 2017. 10 Pages. [cited by applicant]
Pouya Samangouei, et al. ExplainGAN: Model Explanation via Decision Boundary Crossing Transformations. Computer Vision—ECCV 2018. Lecture Notes in Computer Science, vol. 11214, pp. 681-696. 16 Pages. [cited by applicant]
Ruth C. Fong, et al. Interpretable Explanations of Black Boxes by Meaningful Perturbation. In Proceedings of the IEEE International Conference on Computer Vision, pp. 3429-3437, 2017. 9 Pages [cited by applicant]
Ruth Fong, et al. Understanding Deep Networks Via Extremal Perturbations and Smooth Masks. arXiv:1910.08485v1, Oct. 18, 2019. 10 Pages. [cited by applicant]
Ryan Lee, et al. Epidemiology of Diabetic Retinopathy, Diabetic Macular Edema and Related Vision Loss. Eye and vision, 2(1):17, 2015. 25 Pages. [cited by applicant]
Ryan Poplin, et al. Predicting Cardiovascular Risk Factors from Retinal Fundus Photographs using Deep Learning. arXiv: 1708.09843v1, Aug. 31, 2017. 21 Pages. [cited by applicant]
Ryan Poplin, et al. Predicting Cardiovascular Risk Factors from Retinal Fundus Photographs using Deep Learning. arXiv: 1708.09843v2, Sep. 21, 2017. 21 Pages. [cited by applicant]
Ryan Poplin, et al. Prediction of Cardiovascular Risk Factors from Retinal Fundus Photographs via Deep Learning. Nature Biomedical Engineering, 2(3):158, 2018. 9 Pages. [cited by applicant]
S. P. Harding, et al. Sensitivity and Specificity of Photography and Direct Ophthalmoscopy in Screening for Sight Threatening Eye Disease: The Liverpool Diabetic Eye Study. BMJ, 311(7013):1131-1135, 1995. 5 Pages. [cited by applicant]
Sarah Mackenzie, et al. SDOCT Imaging to Identify Macular Pathology in Patients Diagnosed with Diabetic Maculopathy by a Digital Photographic Retinal Screening Programme. PloS one, 6(5):e14811, 2011. 6 Pages. [cited by applicant]
Sarah Wolf. CycleGAN: Learning to Translate Images (Without Paired Training Data). Printed from: https://towardsdatascience.com/cyclegan-learning-to-translate-images-without-paired-training-data-5b4e93862c8d on Jan. 21,… [cited by applicant]
Shalmali Joshi, et al. Towards Realistic Individual Recourse and Actionable Explanations in Black-box Decision Making Systems. arXiv:1907.09615v1, Jul. 22, 2019. 19 Pages. [cited by applicant]
Shusen Liu, et al. Generative Counterfactual Introspection for Explainable Deep Learning. arXiv:1907.03077v1, Jul. 6, 2019. 8 Pages. [cited by applicant]
Sumedha Singla, et al. Explanation by Progressive Exaggeration. arXiv:1911.00483v1, Nov. 1, 2019. 19 Pages. [cited by applicant]
Sumedha Singla, et al. Explanation by Progressive Exaggeration. arXiv:1911.00483v2, Nov. 5, 2019. 19 Pages. [cited by applicant]
Sumedha Singla, et al. Explanation by Progressive Exaggeration. arXiv:1911.00483v3, Feb. 10, 2020. 20 Pages. [cited by applicant]
Taeksoo Kim, et al. Learning to Discover Cross-Domain Relations with Generative Adversarial Networks. In Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, PMLR 70, 2017. 9 Pages. [cited by applicant]
Varun Gulshan, et al. Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy In Retinal Fundus Photographs. Jama, 316(22):2402-2410, 2016. 9 Pages. [cited by applicant]
The First Examination Report for Indian Patent Application No. 202247050019, Jan. 13, 2023. [cited by applicant]
Vitali Petsiuk, et al. RISE: Randomized Input Sampling for Explanation of Black-box Models. arXiv:1806.07421v1, Jun. 19, 2018. 15 Pages. [cited by applicant]
Vitali Petsiuk, et al. RISE: Randomized Input Sampling for Explanation of Black-box Models. arXiv:1806.07421v2, Jul. 25, 2018. 17 Pages. [cited by applicant]
Vitali Petsiuk, et al. RISE: Randomized Input Sampling for Explanation of Black-box Models. arXiv:1806.07421v3, Sep. 25, 2018. 17 Pages. [cited by applicant]
Yu T. Wang, et al. Comparison of Prevalence of Diabetic Macular Edema Based on Monocular Fundus Photography vs Optical Coherence Tomography. JAMA ophthalmology, 134(2):222-228, 2016. 7 Pages. [cited by applicant]
Yunjey Choi, et al. Stargan: Unified Generative Adversarial Networks for Multi-domain Image-to-Image Translation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8789-8797, 2018. 9 … [cited by applicant]