IP Library › Granted Patent US 12,406,140
Granted Patent B2
US 12,406,140 · App. 17/807,160 · Granted Sep 2, 2025

System and method for generating contrastive explanations for text guided by attributes

Inventors: Saneem Ahmed Chemmengath (Bangalore, IN); Amar Prakash Azad (Bangalore, IN); Ronny Luss (Yorktown Heights, NY); Amit Dhurandhar (Yorktown Heights, NY)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F40/284G06F16/35G06F40/205G06F40/30G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,140
App. No.
17/807,160
Filed
Jun 16, 2022
Granted
Sep 2, 2025
Kind
B2
Art Unit
2656
USPC
704/260
Abstract

A method, computer program product and system are provided to generate perturbed text is provided. A processor receives a string of text from a user. A processor determines one or more classifications for at least one word in the string of text by a classification model. A processor determines a plurality of perturbations of the at least one word based on the one or more classifications, where the plurality of perturbations do not share the same one or more classifications as the at least one word in the string of text. A processor selects a perturbation of the string of text based on (i) an edit distance between the string of text and the plurality of perturbations, and (ii) a fluency metric for each of the plurality of perturbations. A processor provides the perturbation of the string of text to the user.

Claims (42)

1. A computer-implemented method comprising:

receiving a string of text from a user;

determining one or more classifications for at least one word in the string of text by a classification model;

determining a plurality of perturbations of the at least one word in the string of text such that the plurality of perturbations do not share a same classification with the one or more classifications of the at least one word in the string of text;

selecting, from the plurality of perturbations, a perturbation of the string of text that includes a respective perturbed classification that is different from the one or more classifications of the at least one word in the string of text based on (i) an edit distance between the string of text and the plurality of perturbations, and (ii) a fluency metric for each of the plurality of perturbations; and

providing, to the user, the perturbation of the string of text that includes the respective perturbed classification that is different from the one or more classifications of the at least one word in the string of text, wherein the providing identifies one or more attribute changes between the string of text and the perturbation of the string of text corresponding to the respective perturbed classification.

2. The computer-implemented method of claim 1 , wherein the perturbation of the string of text and the one or more attribute changes are provided to the classification model for reclassification of the perturbation.

3. The computer-implemented method of claim 2 , the method further comprising:

determining, by the one or more processors, one or more perturbed classifications for the perturbation of the string of text by the classification model; and

identifying, by the one or more processors, bias in the classification model based on the one or more perturbed classifications.

4. The computer-implemented method of claim 1 , wherein the fluency metric for each of the plurality of perturbations is based on output from a masked-language modeling bidirectional encoder representations from transformers (MLM-BERT) Language model.

5. The computer-implemented method of claim 1 , wherein at least one word in the plurality of perturbations is selected based on an integrated gradient neural network model of the string of text.

6. The computer-implemented method of claim 1 , wherein the edit distance indicates a number of words changed from the string of text to a perturbation of the string of text of the plurality of perturbations.

7. A computer program product comprising:

one or more computer-readable storage media and program instructions stored on the one or more computer-readable storage media, the program instructions comprising:

program instructions to receive a string of text from a user;

program instructions to determine one or more classifications for at least one word in the string of text by a classification model;

program instructions to determine a plurality of perturbations of the at least one word in the string of text such that the plurality of perturbations do not share a same classification with the one or more classifications of the at least one word in the string of text;

program instructions to select, from the plurality of perturbations, a perturbation of the string of text that includes a respective perturbed classification that is different from the one or more classifications of the at least one word in the string of text based on (i) an edit distance between the string of text and the plurality of perturbations, and (ii) a fluency metric for each of the plurality of perturbations; and

program instructions to provide, to the user, the perturbation of the string of text that includes the respective perturbed classification that is different from the one or more classifications of the at least one word in the string of text, wherein the program instructions to provide identifies one or more attribute changes between the string of text and the perturbation of the string of text corresponding to the respective perturbed classification.

8. The computer program product of claim 7 , wherein the perturbation of the string of text and the one or more attribute changes are provided to the classification model for reclassification of the perturbation.

9. The computer program product of claim 8 , the program instructions further comprising:

program instructions to determine one or more perturbed classifications for the perturbation of the string of text by the classification model; and

program instructions to identify bias in the classification model based on the one or more perturbed classifications.

10. The computer program product of claim 7 , wherein the fluency metric for each of the plurality of perturbations is based on output from a masked-language modeling bidirectional encoder representations from transformers (MLM-BERT) Language model.

11. The computer program product of claim 7 , wherein at least one word in the plurality of perturbations is selected based on an integrated gradient neural network model of the string of text.

12. The computer program product of claim 7 , wherein the edit distance indicates a number of words changed from the string of text to a perturbation of the string of text of the plurality of perturbations.

13. A computer system comprising:

one or more computer processors;

one or more computer readable storage media; and

program instructions stored on the computer readable storage media for execution by at least one of the one or more processors, the program instructions comprising:

program instructions to receive a string of text from a user;

program instructions to determine one or more classifications for at least one word in the string of text by a classification model;

program instructions to determine a plurality of perturbations of the at least one word in the string of text such that the plurality of perturbations do not share a same classification with the one or more classifications of the at least one word in the string of text;

program instructions to select, from the plurality of perturbations, a perturbation of the string of text that includes a respective perturbed classification that is different from the one or more classifications of the at least one word in the string of text based on (i) an edit distance between the string of text and the plurality of perturbations, and (ii) a fluency metric for each of the plurality of perturbations; and

program instructions to provide, to the user, the perturbation of the string of text that includes the respective perturbed classification that is different from the one or more classifications of the at least one word in the string of text, wherein the program instructions to provide identifies one or more attribute changes between the string of text and the perturbation of the string of text corresponding to the respective perturbed classification.

14. The computer system of claim 13 , wherein the perturbation of the string of text and the one or more attribute changes are provided to the classification model for reclassification of the perturbation.

15. The computer system of claim 14 , the program instructions further comprising:

program instructions to determine one or more perturbed classifications for the perturbation of the string of text by the classification model; and

program instructions to identify bias in the classification model based on the one or more perturbed classifications.

16. The computer system of claim 13 , wherein the fluency metric for each of the plurality of perturbations is based on output from a masked-language modeling bidirectional encoder representations from transformers (MLM-BERT) Language model.

17. The computer system of claim 13 , wherein at least one word in the plurality of perturbations is selected based on an integrated gradient neural network model of the string of text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2022
From: CHEMMENGATH, SANEEM AHMED; AZAD, AMAR PRAKASH; LUSS, RONNY; DHURANDHAR, AMIT
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 060222/0383 →
Continuity (1)
Related Publication 20230409832A1 · Dec 21, 2023
References Cited (31)
US 8264502B2 · Wang · 2012 [cited by examiner]
US 10977729B2 · Kamkar · 2021 [cited by examiner]
US 11222242B2 · Luss · 2022 [cited by examiner]
US 11450225B1 · Kilari · 2022 [cited by examiner]
US 11507787B2 · Dhurandhar · 2022 [cited by examiner]
US 11606389B2 · Chen · 2023 [cited by examiner]
US 11663404B2 · Wang · 2023 [cited by examiner]
US 11669687B1 · Joshi · 2023 [cited by examiner]
US 11720751B2 · Zohrevand · 2023 [cited by examiner]
US 11755948B2 · Kapishnikov · 2023 [cited by examiner]
US 20200193243A1 · Dhurandhar · 2020 [cited by applicant]
US 20200402658A1 · Tomsett · 2020 [cited by applicant]
US 20210067549A1 · Chen · 2021 [cited by examiner]
US 20210192382A1 · Kapishnikov · 2021 [cited by examiner]
US 20220129794A1 · McGrath · 2022 [cited by examiner]
US 20230016365A1 · Qiu · 2023 [cited by examiner]
US 20230120965A1 · Kilari · 2023 [cited by examiner]
US 20230376700A1 · Bista · 2023 [cited by examiner]
CN 111783443A · 2020 [cited by examiner]
CN 111931477B · 2021 [cited by examiner]
WO WO2022022163A1 · 2022 [cited by examiner]
Molnar, C., Casalicchio, G., & Bischl, B. (Sep. 2020). Interpretable machine learning: a brief history, state-of-the-art and challenges. In Joint European conference on machine learning and knowledge discovery in databa… [cited by examiner]
W. Zhang, Q. Chen and Y. Chen, “Deep Learning Based Robust Text Classification Method via Virtual Adversarial Training,” in IEEE Access, vol. 8, pp. 61174-61182, 2020. (Year: 2020). [cited by examiner]
W. Shi, M. Song and Y. Wang, “Perturbation-enhanced-based RoBERTa combined with BiLSTM model for Text classification,” ICETIS 2022; 7th International Conference on Electronic Technology and Information Science, Harbin, … [cited by examiner]
Chemmengath et al., “Let the CAT out of the bag: Contrastive Attributed explanations for Text”, Grace Period Disclosure, arXiv:2109.07983v1 [cs.CL] Sep. 16, 2021, 15 pages. [cited by applicant]
Jacovi et al., “Contrastive Explanations for Model Interpretability”, arXiv:2103.01378v3 [cs.CL] Sep. 14, 2021, 15 pages. [cited by applicant]
Luss et al., “Leveraging Latent Features for Local Explanations”, Research Track Paper, KDD '21, Aug. 14-18, 2021, Virtual Event, Singapore, 11 pages. [cited by applicant]
Madaan et al., “Generate Your Counterfactuals: Towards Controlled Counterfactual Generation for Text”, The Thirty-Fifth AAAI Conference on Artificial Intelligence (AAAI-21), 9 pages. [cited by applicant]
Paranjape et al., “Prompting Contrastive Explanations for Commonsense Reasoning Tasks”, Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, Aug. 1-6, 2021. © 2021 Association for Computational Li… [cited by applicant]
Ross et al. “Explaining NLP Models via Minimal Contrastive Editing (MICE)”, Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, Aug. 1-6, 2021, © 2021 Association for Computational Linguistics, 1… [cited by applicant]
Yin et al., “Interpreting Language Models with Contrastive Explanations”, arXiv:2202.10419v1 [cs.CL] Feb. 21, 2022, 13 pages. [cited by applicant]