IP Library Granted Patent US 10,572,601
Granted Patent B2
US 10,572,601 · App. 15/662,701 · Granted Feb 25, 2020

Unsupervised template extraction

Inventors: Eddy Hudson (Austin, TX); Joseph M. Kaufmann (Austin, TX); Niyati Parameswaran (Santa Clara, CA)
Assignee: International Business Machines Corporation
G06F17/279G06K9/6218G06K9/726H04L51/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,572,601
App. No.
15/662,701
Granted
Feb 25, 2020
Kind
B2
Abstract

An approach is provided that improves a question answering (QA) computer system by automatically generating relationship templates. Event patterns are extracted from data in a corpus utilized by the QA computer system. The extracted event patterns are analyzed with the analysis resulting in a number of clusters of related event patterns. Relationship templates are then created from the plurality of clusters of related event patterns and these relationship templates are then utilized to visually interact with the corpus.

Claims (57)

1. An information handling system comprising:

one or more processors;

one or more data stores accessible by at least one of the processors;

a memory coupled to at least one of the processors; and

a set of computer program instructions stored in the memory and executed by at least one of the processors in order to improve a question answering (QA) computer system by automatically generating relationship templates by performing actions of:

extracting a plurality of event patterns corresponding to a plurality of events from data in a corpus utilized by the QA computer system;

analyzing the extracted event patterns resulting in a plurality of clusters of related event patterns;

creating one or more relationship templates from the plurality of clusters of related event patterns, wherein a first one of the or more relationship templates comprises a first set of the plurality of events included in a first one of the plurality of clusters;

displaying the first relationship template as a first graphical representation on a display, wherein the first graphical representation displays the first set of events; and

in response to receiving an input selection that selects the first graphical representation, displaying a set of second graphical representations on the display that represent a set of roles within the first relationship template.

2. The information handling system of claim 1 wherein the actions further comprise:

expanding the corpus utilized by the QA system, wherein the expanding further comprises:

receiving a plurality of text data outside the corpus;

retrieving a plurality of sub-clusters, wherein each of the sub-clusters is based on the created one or more relationship templates; and

expanding a plurality of input arguments included in the event patterns by using a set of related event patterns found in the plurality of text data outside the corpus, wherein the expanding further comprises:

matching portions of the corpus with the cluster that includes the related event pattern corresponding to the plurality of input arguments.

3. The information handling system of claim 1 wherein the extracting further comprises:

clustering the event patterns using hierarchical agglomerative clustering techniques so that the event patterns that are closer together are clustered in the same event pattern, wherein each cluster of event patterns forms the basis of one of the one or more relationship templates.

4. The information handling system of claim 1 wherein the analyzing further comprises:

converting an argument in each of the extracted event patterns into vectors by using distributional semantics, wherein the converting results in a plurality of word values each corresponding to one of a plurality of words in the extracted event pattern;

calculating a similarity score between the plurality of words based on the word values pertaining to the respective words; and

identifying a plurality of sets of similar words based on a comparison of the calculated similarities.

5. The information handling system of claim 4 wherein the actions further comprise:

selecting each of the arguments and a successive argument to the selected argument; and

calculating the similarity score between the selected argument and the successive argument.

6. The information handling system of claim 5 wherein the actions further comprise:

performing the selecting and calculating on each pair of arguments and storing the similarity scores of all of the pairs;

comparing the similarity scores to a threshold;

in response to the threshold revealing a high similarity score between the arguments in one or more of the pairs of arguments:

clustering the event pattern corresponding with the selected arguments with the event pattern corresponding with the selected arguments' respective successive arguments.

7. The information handling system of claim 5 wherein the similarity is calculated using a cosine similarity algorithm.

8. A computer program product stored in a computer readable storage medium, comprising computer program code that, when executed by an information handling system, causes the information handling system to improve a question answering (QA) computer system by automatically generating relationship templates by performing actions comprising:

extracting a plurality of event patterns corresponding to a plurality of events from data in a corpus utilized by the QA computer system;

analyzing the extracted event patterns resulting in a plurality of clusters of related event patterns;

creating one or more relationship templates from the plurality of clusters of related event patterns, wherein a first one of the or more relationship templates comprises a first set of the plurality of events included in a first one of the plurality of clusters;

displaying the first relationship template as a first graphical representation on a display, wherein the first graphical representation displays the first set of events; and

in response to receiving an input selection that selects the first graphical representation, displaying a set of second graphical representations on the display that represent a set of roles within the first relationship template.

9. The computer program product of claim 8 wherein the actions further comprise:

expanding the corpus utilized by the QA system, wherein the expanding further comprises:

receiving a plurality of text data outside the corpus;

retrieving a plurality of sub-clusters, wherein each of the sub-clusters is based on the created one or more relationship templates; and

expanding a plurality of input arguments included in the event patterns by using a set of related event patterns found in the plurality of text data outside the corpus, wherein the obtaining further comprises:

matching portions of the corpus with the cluster that includes the related event pattern corresponding to the plurality of input arguments.

10. The computer program product of claim 8 wherein the extracting further comprises:

clustering the event patterns using hierarchical agglomerative clustering techniques so that the event patterns that are closer together are clustered in the same event pattern, wherein each cluster of event patterns forms the basis of one of the one or more relationship templates.

11. The computer program product of claim 8 wherein the analyzing further comprises:

converting an argument in each of the extracted event patterns into vectors by using distributional semantics, wherein the converting results in a plurality of word values each corresponding to one of a plurality of words in the extracted event pattern;

calculating a similarity score between the plurality of words based on the word values pertaining to the respective words; and

identifying a plurality of sets of similar words based on a comparison of the calculated similarities.

12. The computer program product of claim 11 wherein the actions further comprise:

selecting each of the arguments and a successive argument to the selected argument; and

calculating the similarity score between the selected argument and the successive argument.

13. The computer program product of claim 12 wherein the actions further comprise:

performing the selecting and calculating on each pair of arguments and storing the similarity scores of all of the pairs;

comparing the similarity scores to a threshold, wherein the similarity is calculated using a cosine similarity algorithm;

in response to the threshold revealing a high similarity score between the arguments in one or more of the pairs of arguments:

clustering the event pattern corresponding with the selected arguments with the event pattern corresponding with the selected arguments' respective successive arguments.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2017
From: HUDSON, EDDY; KAUFMANN, JOSEPH M.; PARAMESWARAN, NIYATI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 043126/0582 →
Continuity (1)
Related Publication 20190034408A1 · Jan 31, 2019