IP Library › Granted Patent US 12,456,034
Granted Patent B2
US 12,456,034 · App. 17/229,275 · Granted Oct 28, 2025

Image classification explanation by generating boundary crossing examples with removed features via filter suppression

Inventor: Andres Mauricio Munoz Delgado (Weil Der Stadt, DE)
Assignee: ROBERT BOSCH GMBH
G06N3/045G06N3/0455G06N3/0464G06N3/0475G06N3/094G06N5/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,034
App. No.
17/229,275
Granted
Oct 28, 2025
Kind
B2
Abstract

A computer-implemented method for explaining a classification of one or more classifier inputs by a trained classifier. A generative model is used that generates inputs for the trained classifier. The generative model comprises multiple filters. Generator inputs corresponding to the one or more classifier inputs are obtained, where a generator input causes the generative model to approximately generate the corresponding classifier input. Filter suppression factors are determined for the multiple filters of the generative model. A filter suppression factor for a filter indicates a degree of suppression for a filter output of the filter. The filter suppression factors are determined based on an effect of adapting the classifier inputs according to the filter suppression factors on the classification by the trained classifier. The classification explanation is based on the filter suppression factors.

Claims (46)

1. A computer-implemented method of determining a classification explanation for a trained classifier, the classification explanation being for one or more classifier inputs classified by the trained classifier into a same class, the method comprising the following steps:

accessing model data defining the trained classifier and model data defining a generative model, the generative model being configured to generate a classifier input for the trained classifier from a generator input, the generative model including multiple filters, each filter of the generative model being configured to generate a filter output at an internal layer of the generative model;

obtaining generator inputs corresponding to the one or more classifier inputs, each generator input causing the generative model to approximately generate the corresponding classifier input;

determining filter suppression factors for the multiple filters of the generative model, each filter suppression factor for each filter indicating a degree of suppression for the filter output of the filter, the filter suppression factors being determined based on an effect of adapting the classifier inputs according to the filter suppression factors on the classification by the trained classifier, the determining including:

adapting each classifier input according to one or more of the filter suppression factors by applying the generative model to the generator input corresponding to the classifier input, while modulating the filter outputs of the filters of the generative model according to the one or more filter suppression factors, and

applying the trained classifier to each adapted classifier input to obtain a classifier output affected by the one or more filter suppression factors;

determining the classification explanation in terms of the filter suppression factors such that the classification explanation excludes any image representation and comprises a set of the filter suppression factors, and outputting the classification explanation, wherein:

the trained classifier is an image classifier, and

the classifier input includes an image of a product produced in a manufacturing process;

controlling the manufacturing process based on the classification of the classification explanation; and

determining the filter suppression factors by performing an optimization configured to: (i) minimize a difference between a target classifier output and affected classifier outputs of the trained classifier for the one or more classifier inputs affected by the filter suppression factors, and (ii) minimize an overall degree of suppression indicated by the filter suppression factors.

2. The method of claim 1 , the method further comprises classifying the classification explanation into a predefined set of possible anomalies.

3. The method of claim 1 , further comprising:

accessing a discriminative model configured to determine a degree to which an output of the generative model is synthetic, the optimization being further configured to minimize the degree for the one or more adapted classifier inputs.

4. The method of claim 1 , further comprising:

obtaining uniqueness scores indicating uniqueness of respective filters, wherein the minimization of the overall degree of suppression penalizes suppression of more unique filters less strongly than suppression of less unique filters.

5. The method of claim 1 , wherein the determining of the classification explanation includes determining a difference between each classifier input and a corresponding adapted classifier input.

6. The method of claim 5 , wherein the difference is a pixelwise difference, or a difference in color distribution, or a difference in entropy.

7. The method of claim 1 , wherein each filter output of the filters of the generative model is modulated according to a filter suppression factor by multiplying elements of the filter output with the filter suppression factor.

8. The method of claim 1 , further comprising obtaining a first classifier input, and determining a first generator input corresponding to the first classifier input.

9. The method of claim 1 , further comprising obtaining a class of the trained classifier and generating one or more generator inputs causing the generative model to generate classifier inputs from the class.

10. The method of claim 1 , further comprising outputting the classification explanation in a sensory perceptible manner to a user.

11. The method of claim 10 , wherein at least the adapted classifier input is output to the user, the method further comprising obtaining a desired classification of the adapted classifier input from the user for re-training the trained classifier using the adapted classifier input and the desired classification.

12. A system for determining a classification explanation for a trained classifier, the classification explanation being for one or more classifier inputs classified by the trained classifier into a same class, the system comprising:

a data interface configured to access model data defining the trained classifier and model data defining a generative model, the generative model being configured to generate a classifier input for the trained classifier from a generator input, the generative model including multiple filters, each filter of the generative model being configured to generate a filter output at an internal layer of the generative model;

a processor subsystem configured to:

obtain generator inputs corresponding to the one or more classifier inputs, each generator input causing the generative model to approximately generate the corresponding classifier input;

determine filter suppression factors for the multiple filters of the generative model, each filter suppression factor for each filter indicating a degree of suppression for the filter output of the filter, the filter suppression factors being determined based on an effect of adapting the classifier inputs according to the filter suppression factors on the classification by the trained classifier, the determining including:

adapting each classifier input according to one or more of the filter suppression factors by applying the generative model to the generator input corresponding to the classifier input, while modulating the filter outputs of the filters of the generative model according to the one or more of the filter suppression factors, and

applying the trained classifier to the adapted classifier input to obtain a classifier output affected by the one or more filter suppression factors; and

determine the classification explanation in terms of the filter suppression factors such that the classification explanation excludes any image representation and comprises a set of the filter suppression factors, and output the classification explanation, wherein:

the trained classifier is an image classifier,

the classifier input includes an image of a product produced in a manufacturing process, and

the manufacturing process is controlled based on the classification of the classification explanation; and

determine the filter suppression factors by performing an optimization configured to: (i) minimize a difference between a target classifier output and affected classifier outputs of the trained classifier for the one or more classifier inputs affected by the filter suppression factors, and (ii) minimize an overall degree of suppression indicated by the filter suppression factors.

13. A non-transitory computer-readable medium on which is stored a computer program for determining a classification explanation for a trained classifier, the classification explanation being for one or more classifier inputs classified by the trained classifier into a same class, the computer program, when executed by a processor system, causing the processor system to perform the following steps:

accessing model data defining the trained classifier and model data defining a generative model, the generative model being configured to generate a classifier input for the trained classifier from a generator input, the generative model including multiple filters, each filter of the generative model being configured to generate a filter output at an internal layer of the generative model;

obtaining generator inputs corresponding to the one or more classifier inputs, each generator input causing the generative model to approximately generate the corresponding classifier input;

determining filter suppression factors for the multiple filters of the generative model, each filter suppression factor for each filter indicating a degree of suppression for the filter output of the filter, the filter suppression factors being determined based on an effect of adapting the classifier inputs according to the filter suppression factors on the classification by the trained classifier, the determining including:

adapting each classifier input according to one or more of the filter suppression factors by applying the generative model to the generator input corresponding to the classifier input, while modulating the filter outputs of the filters of the generative model according to the one or more filter suppression factors, and

applying the trained classifier to each adapted classifier input to obtain a classifier output affected by the one or more filter suppression factors; and

determining the classification explanation in terms of the filter suppression factors such that the classification explanation excludes any image representation and comprises a set of the filter suppression factors, and outputting the classification explanation, wherein:

the trained classifier is an image classifier,

the classifier input includes an image of a product produced in a manufacturing process,

the manufacturing process is controlled based on the classification of the classification explanation; and

determining the filter suppression factors by performing an optimization configured to: (i) minimize a difference between a target classifier output and affected classifier outputs of the trained classifier for the one or more classifier inputs affected by the filter suppression factors, and (ii) minimize an overall degree of suppression indicated by the filter suppression factors.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2021
From: MUNOZ DELGADO, ANDRES MAURICIO
To: ROBERT BOSCH GMBH
Reel/Frame 058100/0294 →
Priority Claims (1)
EP 20170305 · Apr 20, 2020 · regional
Continuity (1)
Related Publication 20210326661A1 · Oct 21, 2021
References Cited (22)
Alzantot et al., “NeuroMask: Explaining Predictions of Deep Neural Networks through Mask Learning”, 2019, 2019 IEEE International Conference on Smart Computing (SMARTCOMP), vol. 2019, pp. 81-86 (Year: 2019). [cited by examiner]
Samangouei et al., “ExplainGAN: Model Explanation via Decision Boundary Crossing Transformations”, Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 666-681 (Year: 2018). [cited by examiner]
Fong et al., “Interpretable Explanations of Black Boxes by Meaningful Perturbation”, 2017, Proceedings of the IEEE International Conference on Computer Vision (ICCV), vol. 2017, pp. 3429-3437 (Year: 2017). [cited by examiner]
Kaneko et al., “Generative Attribute Controller with Conditional Filtered Generative Adversarial Networks”, 2017, Proceedings of the IEEE conference on computer vision and pattern recognition, vol. 2017, pp. 6089-6098 (… [cited by examiner]
Bau et al., “GAN Dissection: Visualizing and Understanding Generative Adversarial Networks”, 2018, arXiv, v2, pp. 1-18 (Year: 2018). [cited by examiner]
Yu et al., “Generative Image Inpainting with Contextual Attention”, 2018, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), vol. 2018, pp. 5505-5514 (Year: 2018). [cited by examiner]
Agarwal et al., “Removing input features via a generative model to explain their attributions to classifier's decisions”, 2019, arXiv, v1, pp. 1-35 (Year: 2019). [cited by examiner]
Mousa-Pasandi et al., “Convolutional Neural Network Pruning Using Filter Attenuation”, Feb. 9, 2020, arXiv, v1, pp. 1-5 (Year: 2020). [cited by examiner]
Abbasi-Asl et al., “Interpreting Convolutional Neural Networks Through Compression”, 2017, arXiv, v1, pp. 1-5 (Year: 2017). [cited by examiner]
Weimer et al., “Design of deep convolutional neural network architectures for automated feature extraction in industrial inspection”, 2016, CIRP Annals, vol. 65 No. 1, pp. 417-420 (Year: 2016). [cited by examiner]
Kaneko et al., “Class-Distinct and Class-Mutual Image Generation with GANs”, 2019, arXiv, v2, pp. 6089-6098 (Year: 2019). [cited by examiner]
Zhou et al., “Revisiting the Importance of Individual Units in CNNs via Ablation”, 2018, arXiv, v1, pp. 1-10 (Year: 2018). [cited by examiner]
Zhang et al., “Interpretable Convolutional Neural Networks”, 2018, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), vol. 2018, pp. 8827-8836 (Year: 2018). [cited by examiner]
Ruth Fong, et al., “Interpretable Explanations of Black Boxes By Meaningful Perturbation,” Cornell University, 2018, pp. 1-9. <https://arxiv.org/pdf/1704.03296.pdf> Downloaded Apr. 13, 2021. [cited by applicant]
Alec Radford et al., “Unsupervised Representation Learning With Deep Convolutional Generative Adversarial Networks,” Cornell University, 2016, pp. 1-16. <https://arxiv.org/pdf/1511.06434.pdf> Downloaded Apr. 13, 2021. [cited by applicant]
Olaf Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation,” Cornell University, 2015, pp. 1-8. <https://arxiv.org/pdf/1505.04597.pdf> Downloaded Apr. 13, 2021. [cited by applicant]
Diederik P. Kingma et al., “ADAM: A Method for Stochastic Optimization”, Cornell University, 2017, pp. 1-15. <https://arxiv.org/pdf/1412.6980.pdf> Downloaded Apr. 13, 2021. [cited by applicant]
Antonia Creswell et al., “Inverting the Generator of a Generative Adversarial Network,” Cornell University, 2018, pp. 1-8. <https://arxiv.org/pdf/1802.05701.pdf> Downloaded Apr. 13, 2021. [cited by applicant]
Chun-Hao Chang, et al., “Interpreting Neural Network Classifications With Variational Dropout Saliency Maps,” 2017, pp. 1-9. <http://www.cs.toronto.edu/˜kingsley/documents/interpreting-neural-network-camera-ready.pdf>. [cited by applicant]
Raghuram Mandyam Annasamy, et al., “Towards Better Interpretability in Deep Q-Networks,” Cornell University, 2018, pp. 1-16. <https://arxiv.org/pdf/1809.05630.pdf>. [cited by applicant]
Lukas Hoyer, et al., “Grid Saliency for Context Explanations of Semantic Segmentation,” Cornell University, 2019, pp. 1-24. <https://arxiv.org/pdf/1907.13054.pdf>. [cited by applicant]
Singh, et al.: “FCA-Net: Adversarial Learning for Skin Lesion Segmentation Based on Multi-Scale Features and Factorized Channel Attention,” IEEE Access, 7 (2019), pp. 130552-130656. [cited by applicant]