IP Library › Granted Patent US 12,694,286
Granted Patent B2
US 12,694,286 · App. 17/549,090 · Granted Jul 28, 2026

Extracting explanations from attention-based models

Inventors: Thai F. Le (West Palm Beach, FL); Supriyo Chakraborty (White Plains, NY); Mudhakar Srivatsa (White Plains, NY)
Assignee: International Business Machines Corporation
G06N3/08G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,694,286
App. No.
17/549,090
Granted
Jul 28, 2026
Kind
B2
Abstract

Providing an explanation for model outcome can include receiving input data, and passing the input data through an attention-based neural network, where the attention-based neural network learns attention weights associated with contextual embeddings corresponding to tokens of the input data and predicts an outcome corresponding to the input data. Based on an attention weight associated with a contextual embedding corresponding to a token of the input data, a signed relevance score can be determined to associate with the token for quantifying the token's relevance to the outcome. Based on the signed relevance score, an explanation of the token's contribution toward or against the outcome can be provided. The signed relevance score can be computed as a gradient of loss with respect to the attention weight.

Claims (34)

1 . A system comprising:

a processor;

a memory device coupled with the processor;

the processor configured to at least:

receive input data;

pass the input data through an attention-based neural network, wherein the attention-based neural network learns attention weights associated with contextual embeddings corresponding to tokens of the input data and predicts an outcome corresponding to the input data, the attention-based neural network generating the contextual embeddings through stacks of self-attention layers, the contextual embeddings being context-aware intermediate representations of the tokens of the input data;

based on an attention weight associated with a contextual embedding corresponding to a token of the input data, determine a signed relevance score to associate with the token for quantifying the token's relevance to the outcome, the signed relevance score providing a directional relevance, the signed relevance score computed as a negative value of a gradient of a loss with respect to the attention weight associated with the contextual embedding, a negative gradient indicating that the contextual embedding contributes toward the outcome, and a positive gradient indicating that the contextual embedding contributes against the outcome; and

based on the signed relevance score determined based on the attention weight associated with the contextual embedding corresponding to the token of the input data, provide a directional explanation of the token's contribution toward or against the outcome by explaining that the token's relevance to the outcome as quantified based on the signed relevance value is toward the outcome responsive to the signed relevance score being a positive value, and explaining that the token's relevance to the outcome as quantified based on the signed relevance value is against the outcome responsive to the signed relevance score being a negative value.

2 . The system of claim 1 , wherein the processor is further configured to determine the signed relevance score by multiplying the gradient with the attention weight.

3 . The system of claim 1 , wherein the processor is further configured to allocate the signed relevance score to the token by summing a relevance of all neurons in the contextual embedding, with each single neuron relevance being computed by back-propagation of a layer-wise neuron relevance.

4 . The system of claim 1 , wherein the processor is further configured to provide a test for accuracy of the explanation, the test testing for resiliency property.

5 . The system of claim 1 , wherein the processor is further configured to provide a test for accuracy of the explanation, the test testing for consistency property.

6 . The system of claim 1 , wherein the processor is further configured to provide a test for accuracy of the explanation testing for resiliency property, the test selecting a subset of tokens in the input data to replace, and further based on explanation performance resulting from the test, using the subset of tokens in an adversarial attack experiment of the attention-based neural network.

7 . The system of claim 1 , wherein the input data includes image data and the attention-based neural network is trained to generate captions associated with the image data.

8 . The system of claim 1 , wherein the input data includes a sequence of text in first language, wherein the attention-based neural network is trained to generate a translation of the sequence of text in second language.

9 . A method comprising:

receiving input data;

passing the input data through an attention-based neural network, wherein the attention-based neural network learns attention weights associated with contextual embeddings corresponding to tokens of the input data and predicts an outcome corresponding to the input data;

based on an attention weight associated with a contextual embedding corresponding to a token of the input data, determining a signed relevance score to associate with the token for quantifying the token's relevance to the outcome, the signed relevance score providing a directional relevance, the signed relevance score computed as a negative value of a gradient of a loss with respect to the attention weight associated with the contextual embedding, a negative gradient indicating that the contextual embedding contributes toward the outcome, and a positive gradient indicating that the contextual embedding contributes against the outcome; and

based on the signed relevance score determined based on the attention weight associated with the contextual embedding corresponding to the token of the input data, providing a directional explanation of the token's contribution to the outcome by explaining that the token's relevance to the outcome as quantified based on the signed relevance value is toward the outcome responsive to the signed relevance score being a positive value, and explaining that the token's relevance to the outcome as quantified based on the signed relevance value is against the outcome responsive to the signed relevance score being a negative value.

10 . The method of claim 9 , further including determining the signed relevance score by multiplying the gradient with the attention weight.

11 . The method of claim 9 , further including allocating the signed relevance score to the token by summing a relevance of all neurons in the contextual embedding, with each single neuron relevance being computed by back-propagation of a layer-wise neuron relevance.

12 . The method of claim 9 , further including providing a test for accuracy of the explanation, the test testing for resiliency property.

13 . The method of claim 9 , further including providing a test for accuracy of the explanation, the test testing for consistency property.

14 . The method of claim 9 , further including providing a test for accuracy of the explanation testing for resiliency property, the test selecting a subset of tokens in the input data to replace, and further based on explanation performance resulting from the test, using the subset of tokens in an adversarial attack experiment of the attention-based neural network.

15 . The method of claim 9 , wherein the input data includes image data and the attention-based neural network is trained to generate captions associated with the image data.

16 . The method of claim 9 , wherein the input data includes a sequence of text in first language, wherein the attention-based neural network is trained to generate a translation of the sequence of text in second language.

17 . A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions readable by a device to cause the device to:

receive input data;

pass the input data through an attention-based neural network, wherein the attention-based neural network learns attention weights associated with contextual embeddings corresponding to tokens of the input data and predicts an outcome corresponding to the input data;

based on an attention weight associated with a contextual embedding corresponding to a token of the input data, determine a signed relevance score to associate with the token for quantifying the token's relevance to the outcome, the signed relevance score providing a directional relevance, the signed relevance score computed as a negative value of a gradient of a loss with respect to the attention weight associated with the contextual embedding, a negative gradient indicating that the contextual embedding contributes toward the outcome, and a positive gradient indicating that the contextual embedding contributes against the outcome; and

based on the signed relevance score determined based on the attention weight associated with the contextual embedding corresponding to the token of the input data, provide a directional explanation of the token's contribution toward or against the outcome by explaining that the token's relevance to the outcome as quantified based on the signed relevance value is toward the outcome responsive to the signed relevance score being a positive value, and explaining that the token's relevance to the outcome as quantified based on the signed relevance value is against the outcome responsive to the signed relevance score being a negative value.

18 . The computer program product of claim 17 , wherein the device is further caused to determine the signed relevance score by multiplying the gradient with the attention weight.

19 . The computer program product of claim 17 , wherein the device is further caused to allocate the signed relevance score to the token by summing a relevance of all neurons in the contextual embedding, with each single neuron relevance being computed by back-propagation of a layer-wise neuron relevance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2021
From: LE, THAI F.; CHAKRABORTY, SUPRIYO; SRIVATSA, MUDHAKAR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058372/0780 →
Continuity (1)
Related Publication 20230186072A1 · Jun 15, 2023
References Cited (44)
US 10909401B2 · Burachas et al. · 2021 [cited by applicant]
US 11055616B2 · Dalli et al. · 2021 [cited by applicant]
US 20190108447A1 · Kounavis et al. · 2019 [cited by applicant]
US 20190156216A1 · Gupta et al. · 2019 [cited by applicant]
US 20190370647A1 · Doshi et al. · 2019 [cited by applicant]
US 20200167677A1 · Verma et al. · 2020 [cited by applicant]
US 20200193235A1 · Martinez-Canales et al. · 2020 [cited by applicant]
US 20200410355A1 · Zhu et al. · 2020 [cited by applicant]
US 20210232940A1 · Dalli et al. · 2021 [cited by applicant]
CN 113011192A · 2021 [cited by applicant]
CN 113761934A · 2021 [cited by applicant]
TW 201720118A · 2017 [cited by applicant]
TW I858384B · 2024 [cited by applicant]
WO 2023110182A1 · 2023 [cited by applicant]
Shrikumar, A., Greenside, P. Kundaje, A.. (2017). Learning Important Features Through Propagating Activation Differences. <i>Proceedings of the 34th International Conference on Machine Learning</i>, in <i>Proceedings of… [cited by examiner]
Sun, X., Yang, D., Li, X., Zhang, T., Meng, Y., Qiu, H., . . . Li, J. (2021). Interpreting Deep Learning Models in Natural Language Processing: A Review. arXiv [Cs.CL]. Retrieved from http://arxiv.org/abs/2110.10470 (Ye… [cited by examiner]
Serrano, Sofia, and Noah A. Smith. “Is Attention Interpretable?” Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, 2019. Crossref, https:… [cited by examiner]
Giulia Vilone, Luca Longo, Notions of explainability and evaluation approaches for explainable artificial intelligence, Information Fusion, vol. 76, 2021, pp. 89-106, ISSN 1566-2535, https://doi.org/10.1016/j.inffus.202… [cited by examiner]
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K. R., & Samek, W. (2015). On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation. PloS one, 10(7), e0130140. https:… [cited by examiner]
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and Klaus-Robert Müller. 2010. How to Explain Individual Classification Decisions. J. Mach. Learn. Res. 11 (Mar. 1, 2010), 1803-1831. (Y… [cited by examiner]
Mcinnes et al., (2018). UMAP: Uniform Manifold Approximation and Projection. Journal of Open Source Software, 3(29), 861, https://doi.org/10.21105/joss.00861 (Year: 2018). [cited by examiner]
Shrikumar, A., Greenside, P. &amp; Kundaje, A.. (2017). Learning Important Features Through Propagating Activation Differences. <i>Proceedings of the 34th International Conference on Machine Learning</i>, in <i>Proceedi… [cited by examiner]
NIST, “NIST Cloud Computing Program”, http://csrc.nist.gov/groups/SNS/cloud-computing/index.html, Created Dec. 1, 2016, Updated Oct. 6, 2017, 9 pages. [cited by applicant]
Ancona, M., et al., “Towards better understanding of gradient-based attribution methods for deep neural networks”, arXiv: 1711.06104v2, Feb. 6, 2018, 16 pages. [cited by applicant]
Simonyan, K., et al., “Deep inside convolutional networks: Visualising image classification models and saliency maps”, arXiv: 1312.6034v2, Apr. 19, 2014, 8 pages. [cited by applicant]
Selvaraju, R.R., et al., “Grad-cam: Why did you say that? visual explanations from deep networks via gradient-based localization”, arXiv:1610.02391v1, Oct. 7, 2016, 17 pages. [cited by applicant]
Sundararajan, M., et al., “Axiomatic Attribution for Deep Networks”, Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 2017, arXiv:1703.01365v2, Jun. 13, 2017, 11 pages. [cited by applicant]
Bahdanau, D., et al., “Neural Machine Translation by Jointly Learning to Align and Translate”, arXiv: 1409.0473v3, Oct. 7, 2014, 15 pages. [cited by applicant]
Luong, M.-T., et al., “Effective Approaches to Attention-based Neural Machine Translation”, Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Sep. 17-21, 2015, pp. 1412-1421. [cited by applicant]
Ba, J.L., et al., “Multiple Object Recognition with Visual Attention”, arXiv:1412.7755v2, Apr. 23, 2015, 10 pages. [cited by applicant]
Xu, K., et al., “Show, Attend and Tell: Neural Image Caption Generation with Visual Attention”, arXiv:1502.03044v3, Apr. 9, 2016, 22 pages. [cited by applicant]
Liu, Y, et al., “Learning Structured Text Representations”, Transactions of the Association for Computational Linguistics, Submission batch May 2017, Published Jan. 2018, pp. 63-75, vol. 6. [cited by applicant]
Mott, A., et al., “Towards interpretable reinforcement learning using attention augmented agents”, arXiv: 1906.02500v1, Jun. 6, 2019, 16 pages. [cited by applicant]
Jain, S., et al., “Attention is not explanation”, Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 2019, Jun. 2-7, 2019, pp. 3543-3556. [cited by applicant]
Vashishth, S., et al., “Attention Interpretability Across NLP Tasks”, arXiv:1909.11218v1, Sep. 24, 2019, 10 pages. [cited by applicant]
Wiegreffe, S., et al., “Attention is not not Explanation”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing… [cited by applicant]
Serrano, S., et al., “Is attention interpretable?”, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 28-Aug. 2, 2019, pp. 2931-2951. [cited by applicant]
Tensorflow, “Neural machine translation with attention”, https://www.tensorflow.org/text/tutorials/hmt_with_attention, Last updated Dec. 4, 2021, Accessed on Dec. 13, 2021, 55 pages. [cited by applicant]
International Search Report dated Jan. 16, 2023 issued in PCT/EP2022/077547, 14 pages. [cited by applicant]
Serrano, S., et al., “Is Attention Interpretable?”, arXiv:1906.03731v1, Jun. 9, 2019, 22 pages. [cited by applicant]
Vilone, G., et al., “Notions of explainability and evaluation approaches for explainable artificial intelligence”, Information Fusion, May 25, 2021, pp. 89-106, vol. 76. [cited by applicant]
Taiwanese Notice of Allowance dated Aug. 30, 2024 received in Taiwan Application No. 111132421, 9 pages. [cited by applicant]
Mcinnes, et al., UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction, arXiv: 1802.03426v3 [stat.ML], Sep. 18, 2020, 63 pages. [cited by applicant]
Shrikumar, et al. Not Just a Black Box: Learning Important Features Through Propagating Activation Differences, arXiv: 1605.01713v3 [cs.LG], Apr. 11, 2017, 6 pages. [cited by applicant]