IP Library › Granted Patent US 11,182,545
Granted Patent B1
US 11,182,545 · App. 16/924,869 · Granted Nov 23, 2021

Machine learning on mixed data documents

Inventors: Yi-Chun Tsai (Taipei, TW); Ying-Chen Yu (Taipei, TW); June-Ray Lin (Taipei, TW); Pei-Hua Su (Kaohsiung, TW)
Assignee: International Business Machines Corporation
G06F40/177G06F40/169G06F40/295G06N3/0454G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,182,545
App. No.
16/924,869
Granted
Nov 23, 2021
Kind
B1
Abstract

A first natural language document is received. The document includes unstructured data and a first table structure that includes a plurality of first table entries. The first table structure is identified based on the document. The first table structure is extracted from the document in response to the identifying. A first machine learning output is generated based on a first machine learning model and from the document. A second machine learning output is generated based on a second machine learning model and from the first table structure. The first output of the document and the second output of the first table structure are combined.

Claims (65)

1. A method comprising:

receiving a first natural language document, wherein the first natural language document includes unstructured data and a first table structure, the first table structure includes a plurality of first table entries;

identifying, based on the first natural language document, the first table structure;

extracting, in response to the identifying, the first table structure from the first natural language document and;

generating, based on a first machine learning model and from the first natural language document after the first table structure was removed from the first natural language document, a first machine learning output;

generating, based on a second machine learning model and from the first table structure, a second machine learning output; and

combining the first machine learning output of the first natural language document and the second machine learning output of the first table structure of the first natural language document.

2. The method of claim 1 , wherein the plurality of first table entries is selected from the group consisting of bulleted items, numbered list items, tabular data, and structured data.

3. The method of claim 1 , wherein the generating the first machine learning output is based on a plurality of annotations for the first natural language document.

4. The method of claim 1 , wherein the generating the second machine learning output further comprises:

annotating, based on the first machine learning output of the natural language document and before the generating the second machine learning output, the table structure.

5. The method of claim 1 , wherein creating the first machine learning model comprises:

retrieving a plurality of natural language documents, wherein a subset of the plurality of natural language documents contains a table structure;

detecting, based on the plurality of natural language documents, the table structures in the subset;

extracting, from the subset, the table structures; and

training, based on the plurality of natural language documents other than the subset and based on the subset of the plurality after the extracting, the first machine learning model.

6. The method of claim 5 , wherein an output of training the first machine learning model includes entities and relationships, and wherein the method further comprises:

matching, based on the output of the first machine learning model and based on an annotation of the plurality of natural language documents, one or more entries of the extracted table structures with the entities and relationships;

annotating, based on the matching, the extracted table structures; and

training, based on the annotated extracted table structures, the second machine learning model.

7. The method of claim 1 , wherein the first machine learning model is selected from the group consisting of a neural network and a support vector machine.

8. The method of claim 1 , wherein the second machine learning model is selected from the group consisting of a neural network and a support vector machine.

9. A system, the system comprising:

a memory, the memory containing one or more instructions; and

a processor, the processor communicatively coupled to the memory, the processor, in response to reading the one or more instructions, configured to:

receive a first natural language document, wherein the first natural language document includes unstructured data and a first table structure, the first table structure includes a plurality of first table entries;

identify, based on the first natural language document, the first table structure;

extract, in response to the identifying, the first table structure from the first natural language document and;

generate, based on a first machine learning model and from the first natural language document after the first table structure was removed from the first natural language document, a first machine learning output;

generate, based on a second machine learning model and from the first table structure, a second machine learning output; and

combine the first machine learning output of the first natural language document and the second machine learning output of the first table structure of the first natural language document.

10. The system of claim 9 , wherein the plurality of first table entries is selected from the group consisting of bulleted items, numbered list items, tabular data, and structured data.

11. The system of claim 9 , wherein the generate the first machine learning output is based on a plurality of annotations for the first natural language document.

12. The system of claim 9 , wherein the generate the second machine learning output further comprises:

annotate, based on the first machine learning output of the natural language document and before the generating the second machine learning output, the table structure.

13. The system of claim 9 , wherein creating the first machine learning model comprises:

retrieve a plurality of natural language documents, wherein a subset of the plurality of natural language documents contains a table structure;

detect, based on the plurality of natural language documents, the table structures in the subset;

extract, from the subset, the table structures; and

train, based on the plurality of natural language documents other than the subset and based on the subset of the plurality after the extracting, the first machine learning model.

14. The system of claim 13 , wherein an output of training the first machine learning model includes entities and relationships, and wherein the processor is further configured to:

match, based on the output of the first machine learning model and based on an annotation of the plurality of natural language documents, one or more entries of the extracted table structures with the entities and relationships of the output;

annotate, based on the matching, the extracted table structures; and

train, based on the annotated extracted table structures, the second machine learning model.

15. A computer program product, the computer program product comprising:

one or more computer readable storage media; and

program instructions collectively stored on the one or more computer readable storage media, the program instructions configured to:

receive a first natural language document, wherein the first natural language document includes unstructured data and a first table structure, the first table structure includes a plurality of first table entries;

identify, based on the first natural language document, the first table structure;

extract, in response to the identifying, the first table structure from the first natural language document and;

generate, based on a first machine learning model and from the first natural language document after the first table structure was removed from the first natural language document, a first machine learning output;

generate, based on a second machine learning model and from the first table structure, a second machine learning output; and

combine the first machine learning output of the first natural language document and the second machine learning output of the first table structure of the first natural language document.

16. The computer program product of claim 15 , wherein the plurality of first table entries is selected from the group consisting of bulleted items, numbered list items, tabular data, and structured data.

17. The computer program product of claim 15 , wherein creating the first machine learning model comprises:

retrieve a plurality of natural language documents, wherein a subset of the plurality of natural language documents contains a table structure;

detect, based on the plurality of natural language documents, the table structures in the subset;

extract, from the subset, the table structures; and

train, based on the plurality of natural language documents other than the subset and based on the subset of the plurality after the extracting, the first machine learning model.

18. The computer program product of claim 17 , wherein an output of training the first machine learning model includes entities and relationships, and wherein the processor is further configured to:

match, based on the output of the first machine learning model and based on an annotation of the plurality of natural language documents, one or more entries of the extracted table structures with the entities and relationships of the output;

annotate, based on the matching, the extracted table structures; and

train, based on the annotated extracted table structures, the second machine learning model.

19. The computer program product of claim 15 , wherein the first machine learning model is selected from the group consisting of a neural network and a support vector machine.

20. The computer program product of claim 15 , wherein the second machine learning model is selected from the group consisting of a neural network and a support vector machine.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2020
From: TSAI, YI-CHUN; YU, YING-CHEN; LIN, JUNE-RAY; SU, PEI-HUA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 053166/0029 →
Cited By (3)
US 12,387,048 US 12,443,963 US 12,518,088