IP Library Granted Patent US 11,521,065
Granted Patent B2
US 11,521,065 · App. 16/783,534 · Granted Dec 6, 2022

Generating explanations for context aware sequence-to-sequence models

Inventors: Rachamalla Anirudh Reddy (Warangal, IN); Pranay Kumar Lohia (Bhagalpur, IN); Samiulla Zakir Hussain Shaikh (Bangalore, IN); Diptikalyan Saha (Bangalore, IN); Sameep Mehta (Bangalore, IN)
Assignee: International Business Machines Corporation
G06N3/08G06F16/24575G06N3/0445G06N3/0454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,521,065
App. No.
16/783,534
Granted
Dec 6, 2022
Kind
B2
Abstract

Methods, systems, and computer program products for generating explanations for a semantic parser are provided herein. A computer-implemented method includes providing to a generative model (i) at least one query and (ii) a context of at least one dataset applicable to the at least one query, wherein the generative model generates a plurality of perturbations for the at least one input query based on the context; providing the plurality of perturbations as inputs to a context aware sequence-to-sequence model, thereby obtaining a plurality of outputs; and generating, for (i) an additional query provided as input to the context aware sequence-to-sequence model and (ii) a context applicable to the additional query, an explanation indicative of one or more parts of the additional query that contributes to an output corresponding to the additional query, based at least in part on the plurality of outputs corresponding to the perturbations.

Claims (43)

1. A computer-implemented method, comprising:

providing, to a first model, at least one query and a context of at least one first dataset applicable to the at least one query, wherein the first model generates a plurality of perturbations for the at least one input query based on the context;

providing the plurality of perturbations as inputs to a second model, thereby obtaining a plurality of outputs; and

generating, for an additional query provided as input to the second model and a context applicable to the additional query, an explanation indicative of one or more parts of the additional query that contributes to an output corresponding to the additional query, based at least in part on the plurality of outputs corresponding to the perturbations, wherein the explanation is generated at least in part by using a third model that is trained on a second dataset, wherein the second dataset is generated based on the plurality of outputs corresponding to the perturbations and each item in the second dataset indicates a change in one or more features of a given one of the perturbations relative to the at least one query;

wherein the method is carried out by at least one computing device.

2. The computer-implemented method of claim 1 , comprising:

training the third model, using the dataset, to classify a relative importance of the one or more features, wherein the second dataset comprises a binary features dataset.

3. The computer-implemented method of claim 2 , wherein the third model comprises a logistic regression classifier.

4. The computer-implemented method of claim 2 , wherein the second dataset is generated at least in part by identifying one or more n-grams of the at least one query.

5. The computer-implemented method of claim 1 , comprising:

debugging the second model based at least in part on the generated explanation.

6. The computer-implemented method of claim 1 , wherein the at least one query comprises a natural language query.

7. The computer-implemented method of claim 1 , wherein the context applicable to the at least one query corresponds to a relational database table.

8. The computer-implemented method of claim 1 , comprising:

encoding, using a long short-term memory model, one or more of the context applicable to the at least one query and the at least one query.

9. The computer-implemented method of claim 1 , wherein the first model comprises one or more of a generative adversarial network and a variational autoencoder.

10. The computer-implemented method of claim 1 , wherein the second model comprises a semantic parser.

11. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computing device to cause the computing device to:

provide, to a first model, at least one query and a context of at least one first dataset applicable to the at least one query, wherein the first model generates a plurality of perturbations for the at least one input query based on the context;

provide the plurality of perturbations as inputs to a second model, thereby obtaining a plurality of outputs; and

generate, for an additional query provided as input to the second model and (ii) a context applicable to the additional query, an explanation indicative of one or more parts of the additional query that contributes to an output corresponding to the additional query, based at least in part on the plurality of outputs corresponding to the perturbations, wherein the explanation is generated at least in part by using a third model that is trained on a second dataset, wherein the second dataset is generated based on the plurality of outputs corresponding to the perturbations and each item in the second dataset indicates a change in one or more features of a given one of the perturbations relative to the at least one query.

12. The computer program product of claim 11 , wherein the program instructions executable by a computing device further cause the computing device to:

train the third model, using the second dataset, to classify a relative importance of the one or more features, wherein the second dataset comprises a binary features dataset.

13. The computer program product of claim 12 , wherein the third model comprises a logistic regression classifier.

14. The computer program product of claim 12 , wherein the dataset is generated at least in part by identifying one or more n-grams of the at least one query.

15. The computer program product of claim 11 , wherein the program instructions executable by a computing device further cause the computing device to:

debug the second model based at least in part on the generated explanation.

16. The computer program product of claim 11 , wherein the program instructions executable by a computing device further cause the computing device to:

encode, using a long short-term memory model, one or more of the context applicable to the at least one query and the at least one query.

17. A system comprising:

a memory; and

at least one processor operably coupled to the memory and configured for:

providing, to a first model, at least one query and a context of at least one first dataset applicable to the at least one query, wherein the first model generates a plurality of perturbations for the at least one input query based on the context;

providing the plurality of perturbations as inputs to a second model, thereby obtaining a plurality of outputs; and

generating, for an additional query provided as input to the second model and a context applicable to the additional query, an explanation indicative of one or more parts of the additional query that contributes to an output corresponding to the additional query, based at least in part on the plurality of outputs corresponding to the perturbations, wherein the explanation is generated at least in part by using a third model that is trained on a second dataset, wherein the second dataset is generated based on the plurality of outputs corresponding to the perturbations and each item in the second dataset indicates a change in one or more features of a given one of the perturbations relative to the at least one query.

18. The system if claim 17 , wherein the first model comprises one or more of a generative adversarial network and a variational autoencoder.

19. The system if claim 17 , wherein the second model comprises a semantic parser.

20. A computer-implemented method, comprising:

providing, to a first model, at least one query and a context of at least one dataset applicable to the at least one query, wherein the first model generates a plurality of perturbations for the at least one input query based on the context;

providing the plurality of perturbations as inputs to a second model, thereby obtaining a plurality of outputs;

generating, for an additional query provided as input to the second model and a context applicable to the additional query, an explanation indicative of one or more parts of the additional query that contributes to an output corresponding to the additional query, based at least in part on the plurality of outputs corresponding to the perturbations; and

debugging the second model based at least in part on the generated explanation;

wherein the method is carried out by at least one computing device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2020
From: REDDY, RACHAMALLA ANIRUDH; LOHIA, PRANAY KUMAR; SHAIKH, SAMIULLA ZAKIR HUSSAIN; SAHA, DIPTIKALYAN; MEHTA, SAMEEP
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 051743/0096 →
Continuity (1)
Related Publication 20210248455A1 · Aug 12, 2021