IP Library › Granted Patent US 10,289,963
Granted Patent B2
US 10,289,963 · App. 15/444,051 · Granted May 14, 2019

Unified text analytics annotator development life cycle combining rule-based and machine learning based techniques

Inventors: Laura Chiticariu (San Jose, CA); Jeffrey Thomas Kreulen (San Jose, CA); Rajasekar Krishnamurthy (Campbell, CA); Prithviraj Sen (San Jose, CA); Shivakumar Vaithyanathan (San Jose, CA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N99/005G06F17/30705G06N5/025G06F17/241
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,289,963
App. No.
15/444,051
Granted
May 14, 2019
Kind
B2
Abstract

One embodiment provides a method for developing a text analytics program for extracting at least one target concept including: utilizing at least one processor to execute computer code that performs the steps of: initiating a development tool that accepts user input to develop rules for extraction of features of the at least one target concept within a dataset comprising textual information; developing, using the rules for feature extraction, an evaluation dataset comprising at least one document annotated with the at least one target concept to be extracted by the text analytics program; creating, using the rules for feature extraction, a rule-based annotator to extract the at least one target concept; training, using the evaluation dataset, a machine-learning annotator to extract the at least one target concept within the dataset; combining the rule-based annotator and the machine learning annotator to form a combined annotator; evaluating, using the evaluation dataset, extraction performance of the combined annotator against a predetermined threshold; and publishing, when the extraction performance of the combined annotator exceeds the predetermined threshold, the combined annotator for use in an application that extracts the at least one target concept from a plurality of datasets.

Claims (47)

1. A method for developing a text analytics program for extracting at least one target concept comprising:

utilizing at least one processor to execute computer code that performs the steps of:

initiating a development tool that accepts user input to develop rules for extraction of features of the at least one target concept within a dataset comprising textual information;

developing, using the rules for feature extraction, an evaluation dataset comprising at least one document annotated with the at least one target concept to be extracted by the text analytics program;

creating, using the rules for feature extraction, a rule-based annotator to extract the at least one target concept;

training, using the evaluation dataset, a machine-learning annotator to extract the at least one target concept within the dataset;

evaluating each of the rule-based annotator and the machine-learning annotator against the evaluation dataset and comparing the extraction results, of each of the rule-based annotator and the machine-learning annotator, from the evaluation against a threshold for accuracy;

combining, responsive to determining each of the rule-based annotator and the machine-learning annotator meet the threshold for accuracy, the rule-based annotator and the machine-learning annotator to form a combined annotator having features from both of the rule-based annotator and the machine-learning annotator;

evaluating, using the evaluation dataset, extraction performance of the combined annotator against a predetermined threshold; and

publishing, when the extraction performance of the combined annotator exceeds the predetermined threshold, the combined annotator for use in an application that extracts the at least one target concept from a plurality of datasets.

2. The method of claim 1 , wherein, in response to the extraction performance of the combined annotator being below the predetermined threshold, performing one of the following: iterating through one or more of the initiating, developing, creating, training, combining, and evaluating to form a refined, combined annotator.

3. The method of claim 1 , wherein one or more of the rule-based annotator and the machine learning annotator are based on one or more pre-existing annotators that is available from a catalog of annotators stored in memory.

4. The method of claim 1 , wherein the training dataset is annotated using one or more pre-existing annotators available from a catalog of annotators stored in memory.

5. The method of claim 1 , wherein the evaluation dataset comprises a testing dataset and a training dataset.

6. The method of claim 5 , wherein the training a machine-learning annotator comprises using the training dataset.

7. The method of claim 5 , wherein the testing dataset is reserved for testing an annotator selected from the group consisting of the rule-based annotator and the machine learning annotator.

8. The method of claim 1 , wherein the user input comprises at least one rule selected from the group consisting of: a regular expression, a dictionary of terms, parts of speech, and simple patterns.

9. The method of claim 1 , comprising providing a user interface for creating the combined annotator.

10. The method of claim 9 , wherein the user interface comprises graphical display elements facilitating formation of the combined annotator.

11. An apparatus for developing a text analytics program for extracting at least one target concept, comprising:

at least one processor, and

a computer readable storage medium having computer readable program code embodied therewith and executable by the at least one processor, the computer readable program code comprising:

computer readable program code that initiates a development tool that accepts user input to develop rules for extraction of features of the at least one target concept within a dataset comprising textual information;

computer readable program code that develops, using the rules for feature extraction, an evaluation dataset comprising at least one document annotated with the at least one target concept to be extracted by the text analytics program;

computer readable program code that creates, using the rules for feature extraction, a rule-based annotator to extract the at least one target concept;

computer readable program code that trains, using the evaluation dataset, a machine-learning annotator to extract the at least one target concept within the dataset;

computer readable program code that evaluates each of the rule-based annotator and the machine-learning annotator against the evaluation dataset and compares the extraction results, of each of the rule-based annotator and the machine-learning annotator, from the evaluation against a threshold for accuracy;

computer readable program code that combines, responsive to determining each of the rule-based annotator and the machine-learning annotator meet the threshold for accuracy, the rule-based annotator and the machine-learning annotator to form a combined annotator having features from both of the rule-based annotator and the machine-learning annotator;

computer readable program code that evaluates, using the evaluation dataset, extraction performance of the combined annotator against a predetermined threshold; and

computer readable program code that publishes, when the extraction performance of the combined annotator exceeds the predetermined threshold, the combined annotator for use in an application that extracts the at least one target concept from a plurality of datasets.

12. A computer program product comprising:

a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code executable by a processor and comprising:

computer readable program code that initiates a development tool that accepts user input to develop rules for extraction of features of the at least one target concept within a dataset comprising textual information;

computer readable program code that develops, using the rules for feature extraction, an evaluation dataset comprising at least one document annotated with the at least one target concept to be extracted by the text analytics program;

computer readable program code that creates, using the rules for feature extraction, a rule-based annotator to extract the at least one target concept;

computer readable program code that trains, using the evaluation dataset, a machine-learning annotator to extract the at least one target concept within the dataset;

computer readable program code that evaluates each of the rule-based annotator and the machine-learning annotator against the evaluation dataset and compares the extraction results, of each of the rule-based annotator and the machine-learning annotator, from the evaluation against a threshold for accuracy;

computer readable program code that combines, responsive to determining each of the rule-based annotator and the machine-learning annotator meet the threshold for accuracy, the rule-based annotator and the machine-learning annotator to form a combined annotator having features from both of the rule-based annotator and the machine-learning annotator;

computer readable program code that evaluates, using the evaluation dataset, extraction performance of the combined annotator against a predetermined threshold; and

computer readable program code that publishes, when the extraction performance of the combined annotator exceeds the predetermined threshold, the combined annotator for use in an application that extracts the at least one target concept from a plurality of datasets.

13. The computer program product of claim 12 , wherein, responsive to the extraction performance of the combined annotator being below the predetermined threshold, performing one of the following: iterating through one or more of the initiating, developing, creating, training, combining, and evaluating to form a refined, combined annotator.

14. The computer program product of claim 12 , wherein one or more of the rule-based annotator and the machine learning annotator are based on one or more pre-existing annotators available from a catalog of annotators stored in memory.

15. The computer program product of claim 12 , wherein the training dataset is annotated using one or more pre-existing annotators that is available from a catalog of annotators stored in memory.

16. The computer program product of claim 12 , wherein the evaluation dataset comprises a testing dataset and a training dataset and wherein the training a machine-learning annotator comprises using the training dataset.

17. The computer program product of claim 12 , wherein the evaluation dataset comprises a testing dataset and a training dataset and wherein the testing dataset is reserved for testing an annotator selected from the group consisting of the rule-based annotator and the machine learning annotator.

18. The computer program product of claim 12 , wherein the user input comprises at least one rule selected from the group consisting of: a regular expression, a dictionary of terms, parts of speech, and simple patterns.

19. The computer program product of claim 12 , comprising providing a user interface for creating the combined annotator.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2017
From: CHITICARIU, LAURA; KREULEN, JEFFREY THOMAS; KRISHNAMURTHY, RAJASEKAR; SEN, PRITHVIRAJ; VAITHYANATHAN, SHIVAKUMAR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 041387/0830 →
Continuity (1)
Related Publication 20180246867A1 · Aug 30, 2018
Cited By (2)
US 12,406,477 US 12,626,729