IP Library › Granted Patent US 8,738,359
Granted Patent B2
US 8,738,359 · App. 11/873,631 · Granted May 27, 2014

Scalable knowledge extraction

Inventors: Rakesh Gupta (Mountain View, CA); Quang Do (Champaign, IL)
Assignee: Honda Motor Co., Ltd.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,738,359
App. No.
11/873,631
Granted
May 27, 2014
Kind
B2
Abstract

The present invention provides a method for extracting relationships between words in textual data. Initially, a classifier is trained to identify text data having a specific format, such as situation-response or cause-effect, using a training corpus. The classifier receives input identifying components of the text data having the specified format and then extracts features from the text data having the specified format, such as the part of speech for words in the text data, the semantic role of words within the text data and sentence structure. These extracted features are then applied to text data to identify components of the text data which have the specified format. Rules are then extracted from the text data having the specified format.

Claims (62)

1. A method performed by a computer for using a classifier to find text data having a specific format, the method comprising the steps of:

training the classifier to identify the specific format using data stored in a training corpus, wherein training comprises identifying positive and negative instances of the specific format within components of the training corpus;

applying the trained classifier to a text corpus to identify components of the text corpus having a feature which corresponds to a characteristic of the specific format;

identifying a subset of the identified components of the text corpus with an above-threshold probability of having a feature which corresponds to the characteristic of the specific format;

gathering additional text data based on the identified subset of components of the text corpus;

applying the trained classifier to the additional text data to identify components of the additional text data having the feature which corresponds to a characteristic of the specified format;

generating by a computer rules from the identified subset of components of the text corpus and the identified components of the additional text data specifying the format of the identified subset of components of the text corpus and the identified components of the additional text data, wherein the rules correspond to the feature of the identified subset of components of the text corpus and the identified components of the additional text data; and

storing the generated rules in a computer storage medium.

2. The method of claim 1 , wherein the step of training the classifier to identify the specific format further comprises the steps of:

segmenting the training corpus into components;

receiving input which identifies one or more components of the training corpus having the specific format; and

extracting features from the one or more components of the training corpus having the specific format.

3. The method of claim 2 , wherein the step of extracting features from the one or more components of the training corpus comprises the steps of:

identifying a word within the training corpus and a part of speech associated with the identified word;

parsing the training corpus into one or more groups of words;

determining a syntactic relation between a plurality of words within the training corpus; and

identifying a verb-argument structure in a group of words within the training corpus.

4. The method of claim 1 , wherein the training corpus comprises text data describing a response to a situation.

5. The method of claim 1 , wherein generating a rule from a component of the identified subset of components of the text corpus or the identified components of the additional text data comprises the steps of:

examining a predicate argument of the component having the specific format;

determining an adjunct associated with the predicate argument; and

classifying the predicate argument based on the determined adjunct associated with the predicate argument.

6. The method of claim 1 , wherein the step of gathering additional text data comprises the steps of:

retrieving text data from a distributed data source; and

generating a subset of text data from the retrieved text data corresponding to text matching a parameter of a component of the identified subset of components of the text corpus.

7. The method of claim 6 , wherein the parameter comprises a maximum or minimum number of words included in the retrieved text.

8. The method of claim 6 , further comprising the step of:

resolving data defects in text included in the generated subset of text.

9. The method of claim 8 , wherein the step of resolving data defects in text comprises the steps of:

resolving a pronoun within the subset of text from the retrieved text data by replacing the pronoun with a corresponding noun; and

determining a discourse in which a component of the subset of text occurs.

10. The method of claim 1 , wherein the rules are generated by applying semantic role labeling to the components of the identified subset of components of the text corpus or the identified components of the additional text data.

11. A method performed by a computer for using a classifier to find text data having a specific format, the method comprising the steps of:

training the classifier to identify the specific format using data stored in a training corpus;

applying the trained classifier to a text corpus to identify components of the text corpus having a feature which corresponds to a characteristic of the specific format;

identifying a subset of the identified components of the text corpus with an above-threshold probability of having a feature which corresponds to the characteristic of the specific format;

gathering additional text data based on the identified subset of the components of the text corpus;

applying the trained classifier to the additional text data to identify components of the additional text data having the feature which corresponds to a characteristic of the specified format;

generating by a computer rules from the identified components of the additional text data specifying the format of the identified components of the additional text data, wherein the rules correspond to the feature of the identified components of the additional text data, and wherein the rules are generated by applying semantic role labeling to the identified components of the additional text data; and

storing the generated rules in a computer storage medium.

12. The method of claim 11 , wherein the step of training the classifier to identify the specific format comprises the steps of:

segmenting the training corpus into components;

receiving input which identifies one or more components of the training corpus having the specific format; and

extracting features from the one or more components of the training corpus having the specific format.

13. The method of claim 12 , wherein the step of extracting features from the one or more components of the training corpus comprises the steps of:

identifying a word within the training corpus and a part of speech associated with the identified word;

parsing the training corpus into one or more groups of words;

determining a syntactic relation between a plurality of words within the training corpus; and

identifying a verb-argument structure in a group of words within the training corpus.

14. The method of claim 11 , wherein the training corpus comprises text data describing a response to a situation.

15. The method of claim 11 , wherein generating a rule from a component of the identified components of the additional text data comprises the steps of:

examining a predicate argument in the component having the specific format;

determining an adjunct associated with the predicate argument; and

classifying the predicate argument based on the determined adjunct associated with the predicate argument.

16. The method of claim 11 , wherein the step of gathering additional text data comprises the steps of:

retrieving text data from a distributed data source; and

generating a subset of text data from the retrieved text data corresponding to text matching a parameter of a component of the identified subset of components of the text corpus.

17. The method of claim 16 , wherein the parameter comprises a maximum or minimum number of words included in the retrieved text.

18. The method of claim 16 , further comprising the step of:

resolving data defects in text included in the generated subset of text by:

resolving a pronoun within the subset of text from the retrieved text data by replacing the pronoun with a corresponding noun; and

determining a discourse in which a component of the subset of text occurs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2007
From: GUPTA, RAKESH; DO, QUANG XUAN
To: HONDA MOTOR CO., LTD.
Reel/Frame 019975/0350 →
Continuity (2)
Provisional Application 60852719 · Oct 18, 2006
Related Publication 20080097951A1 · Apr 24, 2008