IP Library Granted Patent US 10,169,337
Granted Patent B2
US 10,169,337 · App. 15/817,911 · Granted Jan 1, 2019

Converting data into natural language form

Inventors: John J. Bird (Rochester, MN); Doyle J. McCoy (Rochester, MN)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F17/2881G06F17/30011G06F17/30654
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,169,337
App. No.
15/817,911
Granted
Jan 1, 2019
Kind
B2
Abstract

Converting technical data from field oriented electronic data sources into natural language form is disclosed. An approach includes obtaining document data from an input document, wherein the document data is in a non-natural language form. The approach includes determining a data type of the document data from one of a plurality of data types defined in a detection and conversion database. The approach includes translating the document data to a natural language form based on the determined data type. The approach additionally includes outputting the translated document data in natural language form to an output data stream.

Claims (36)

1. A method implemented in a computer infrastructure, comprising:

obtaining document data from a document;

applying a first keyword translation to the document data;

translating, using a translation engine, a product of the first keyword translation to a natural language form;

applying a second keyword translation to a product of the natural language form translation, wherein the translation engine is provided with programming for determining how much document data is required for making a successful determination of data types prior to a translation to the natural language form; and

training a natural language engine by consuming the product of the second keyword translation.

2. The method of claim 1 , wherein the obtaining obtains the document data from an input data stream of the document.

3. The method of claim 1 , wherein the first keyword translation and the second keyword translation are applied using a keyword translation database.

4. The method of claim 1 , further comprising outputting the product of the second keyword translation to an output data stream.

5. The method of claim 4 , further comprising providing the output data stream to a question-answering system.

6. The method of claim 5 , further comprising saving the output data stream as an output document.

7. The method of claim 1 , wherein:

the document data comprises plural different portions of data; and

the translating is performed for each respective one of the plural different portions of data.

8. The method of claim 7 , wherein the plural different portions of data are arranged in an order in the document.

9. The method of claim 8 , further comprising storing metadata associated with each one of the plural different portions of data and determining a data type of the document data from a plurality of data types using a detection and conversion database.

10. The method of claim 9 , wherein the detection and conversion database comprises:

detection data used in identifying different types of data; and

translation data used in translating an identified data type into natural language form.

11. The method of claim 10 , wherein the translation data comprises rules, patterns and constructs.

12. The method of claim 11 , further comprising marking conversion information records from the detection and conversion database for use in the translation of the product of the first keyword translation to the natural language form.

13. The method of claim 12 , wherein the detection and conversion database further comprises plural different translation data corresponding to plural different data types.

14. The method of claim 13 , wherein the detection and conversion database further comprises data records for identifying the plural different data types.

15. The method of claim 14 , wherein each data record comprises first, second, and third elements.

16. The method of claim 15 , wherein the first element of each data record defines sections of the document in which predefined data types are found.

17. The method of claim 16 , wherein the predefined data types comprise document header data, document field data, table header data, table detail data, signature data and natural text.

18. The method of claim 17 , wherein the second element of each data record contains data, rules and logic for calculating respective confidence levels that a particular portion of data from the document corresponds to respective ones of the predefined data types.

19. The method of claim 18 , wherein the third element of each data record contains conversion rules for translating the product of the first keyword translation, the conversion rules comprise an input rule and an output rule, the input rule defines finding data of interest and the output rule defines how to present the data of interest in the natural language form.

20. A computer program product comprising a non-transitory computer usable tangible storage medium having readable program code embodied in the tangible storage medium, the computer program product includes at least one component configured to:

obtain portions of document data that are in a non-natural language form from a document;

for the portions, perform the steps of:

applying a first keyword translation to the portion;

translate the product of the first keyword translation to a natural language form, wherein a plurality of data types comprises: a document header data; a document field data; a table header data; a table detail data; and a signature data;

applying a second keyword translation to the product of the translation to the natural language form; and

determine whether an end of the document has been reached and concurrently place a product of the second keyword translation onto an output data stream; and

training a natural language engine by consuming the output data stream.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2021
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: KYNDRYL, INC.
Reel/Frame 057885/0644 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 20, 2017
From: BIRD, JOHN J.; MCCOY, DOYLE J.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 044181/0503 →
Continuity (4)
Continuation 15426717 · Feb 7, 2017
Continuation 14933624 · Nov 5, 2015
Continuation 13350107 · Jan 13, 2012
Related Publication 20180075025A1 · Mar 15, 2018
Cited By (2)
US 12,488,259 US 12,547,845