IP Library › Granted Patent US 10,387,560
Granted Patent B2
US 10,387,560 · App. 15/369,843 · Granted Aug 20, 2019

Automating table-based groundtruth generation

Inventors: Corville O. Allen (Morrisville, NC); Anne E. Gattiker (Austin, TX); Joseph N. Kozhaya (Morrisville, NC)
Assignee: International Business Machines Corporation
G06F17/248G06F17/245G06F17/278G06F17/2735G06F17/2775
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,387,560
App. No.
15/369,843
Filed
Dec 5, 2016
Granted
Aug 20, 2019
Kind
B2
Art Unit
2158
USPC
707/728
Abstract

A method, system and computer-usable medium are disclosed for automating the generation of table-based groundtruth, comprising: receiving a document comprising unstructured text and a table; generating questions by applying a template the contents of the table; performing QA pair generation operations on the table to generate QA pairs, each QA pair comprising a question generated by applying the template; and, assigning a score to each QA pair, the score providing an indicator of user interest to each QA pair, the score being based on a score generation methodology using the unstructured text and the table.

Claims (86)

1. A computer-implemented method for automating the generation of table-based groundtruth, comprising:

receiving a document comprising unstructured text and a table, the table comprising quantitative information in a tabular format, the table comprising repeated-structure content;

generating questions by applying a template to the contents of the table, the template comprising a direct statement template, the direct statement template comprising a statement containing data that is directly stated within the document;

performing QA pair generation operations on the table to generate QA pairs, each QA pair comprising a question generated by applying the template; and,

assigning a score to each QA pair, the score providing an indicator of user interest to each QA pair, the score being based on a score generation methodology using the unstructured text and the table, the score for each QA pair comprising a ranking score for each QA pair;

identifying groundtruth QA pairs based upon the ranking score of each QA pair; and,

training a QA system using the groundtruth QA pairs.

2. The method of claim 1 , further comprising:

parsing the table to generate label data associated with its columns and rows;

associating label data with a cell corresponding to a column and a row of the table;

determining whether the unstructured text includes a reference to the label data; and,

increasing the score for a QA pair associated with the label data.

3. The method of claim 2 , further comprising:

determining whether the unstructured text includes another reference to the label data; and,

further increasing the score for the QA pair associated with the label data if the reference and the another reference are proximate to one another.

4. The method of claim 1 , further comprising:

parsing the unstructured text in the document;

determining whether a parsed portion of the unstructured text references content within the table;

generating a first QA pair relating to the unstructured text referencing content within the table;

processing the first QA pair to identify entities and keywords;

creating a question template based on the entities and keywords;

using the question template to create a question for a second QA pair based upon the entities and keywords; and,

assigning scores to each of the QA pairs.

5. The method of claim 4 , further comprising:

assigning a higher score to a QA pair when a question of the QA pair is generated from the unstructured text.

6. The method of claim 5 , further comprising:

assigning a lower score to a QA pair when a question of the QA pair is generated using the question template created based on the entities and keywords.

7. A system comprising:

a processor;

a data bus coupled to the processor; and

a computer-usable medium embodying computer program code, the computer-usable medium being coupled to the data bus, the computer program code used for automating the generation of table-based groundtruths and comprising instructions executable by the processor and configured for:

receiving a document comprising unstructured text and a table, the table comprising quantitative information in a tabular format, the table comprising repeated-structure content;

generating questions by applying a template to the contents of the table, the template comprising a direct statement template, the direct statement template comprising a statement containing data that is directly stated within the document;

performing QA pair generation operations on the table to generate QA pairs, each QA pair comprising a question generated by applying the template; and,

assigning a score to each QA pair, the score providing an indicator of user interest to each QA pair, the score being based on a score generation methodology using the unstructured text and the table, the score for each QA pair comprising a ranking score for each QA pair;

identifying groundtruth QA pairs based upon the ranking score of each QA pair; and

training a QA system using the groundtruth QA pairs.

8. The system of claim 7 , wherein the instructions are further configured for:

parsing the table to generate label data associated with its columns and rows;

associating label data with a cell corresponding to a column and a row of the table;

determining whether the unstructured text includes a reference to the label data; and,

increasing the score for a QA pair associated with the label data.

9. The system of claim 8 , wherein the instructions are further configured for:

determining whether the unstructured text includes another reference to the label data; and,

further increasing the score for the QA pair associated with the label data if the reference and the another reference are proximate to one another.

10. The system of claim 7 , wherein the instructions are further configured for:

parsing the unstructured text in the document;

determining whether a parse portion of the unstructured text references content within the table;

generating a first QA pair relating to the unstructured text referencing content within the table;

processing the first QA pair to identify entities and keywords;

creating a question template based on the entities and keywords;

using the question template to create a question for a second QA pair based upon the entities and keywords; and,

assigning a score to the question for each of the QA pairs.

11. The system of claim 10 , wherein the instructions are further configured for:

assigning a higher score to a QA pair when a question of the QA pair is generated from the unstructured text.

12. The system of claim 11 , wherein the instructions are further configured for:

assigning a lower score to a QA pair when a question of the QA pair is generated using the question template created based on the entities and keywords.

13. A non-transitory, computer-readable storage medium embodying computer program code, the computer program code comprising computer executable instructions configured for:

receiving a document comprising unstructured text and a table, the table comprising quantitative information in a tabular format, the table comprising repeated-structure content;

generating questions by applying a template to the contents of the table, the template comprising a direct statement template, the direct statement template comprising a statement containing data that is directly stated within the document;

performing QA pair generation operations on the table to generate QA pairs, each QA pair comprising a question generated by applying the template; and,

assigning a score to each QA pair, the score providing an indicator of user interest to each QA pair, the score being based on a score generation methodology using the unstructured text and the table, the score for each QA pair comprising a ranking score for each QA pair;

identifying groundtruth QA pairs based upon the ranking score of each QA pair; and,

training a QA system using the groundtruth QA pairs.

14. The computer-readable storage medium of claim 13 , wherein the instructions are further configured for:

parsing the table to generate label data associated with its columns and rows;

associating label data with a cell corresponding to a column and a row of the table;

determining whether the unstructured text includes a reference to the label data; and,

increasing the score for a QA pair associated with the label data.

15. The computer-readable storage medium of claim 14 , wherein the instructions are further configured for:

determining whether the unstructured text includes another reference to the label data; and,

further increasing the score for the QA pair associated with the label data if the reference and the another reference are proximate to one another.

16. The computer-readable storage medium of claim 13 , wherein the instructions are further configured for:

parsing the unstructured text in the document;

determining whether a parse portion of the unstructured text references content within the table;

generating a first QA pair relating to the unstructured text referencing content within the table;

processing the first QA pair to identify entities and keywords;

creating a question template based on the entities and keywords;

using the question template to create a question for a second QA pair based upon the entities and keywords; and,

assigning a score to the question for each of the QA pairs.

17. The computer-readable storage medium of claim 16 , wherein the instructions are further configured for:

assigning a higher score to a QA pair when a question of the QA pair is generated from the unstructured text.

18. The computer-readable storage medium of claim 17 , wherein the instructions are further configured for:

assigning a lower score to a QA pair when a question of the QA pair is generated using the question template created based on the entities and keywords.

19. The non-transitory, computer-readable storage medium of claim 13 , wherein the computer executable instructions are deployable to a client system from a server system at a remote location.

20. The non-transitory, computer-readable storage medium of claim 13 , wherein the computer executable instructions are provided by a service provider to a user on an on-demand basis.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2016
From: ALLEN, CORVILLE O.; GATTIKER, ANNE E; KOZHAYA, JOSEPH N.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 040531/0413 →
Continuity (1)
Related Publication 20180157741A1 · Jun 7, 2018