IP Library › Granted Patent US 11,308,268
Granted Patent B2
US 11,308,268 · App. 16/598,057 · Granted Apr 19, 2022

Semantic header detection using pre-trained embeddings

Inventors: Hassan Nadim (San Francisco, CA); Joshua S. Allen (Durham, NC); Kyle G. Christianson (Rochester, MN); Andrew R. Freed (Cary, NC)
Assignee: International Business Machines Corporation
G06F40/177G06F40/258G06V30/412G06V10/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,308,268
App. No.
16/598,057
Granted
Apr 19, 2022
Kind
B2
Abstract

A method, computer system, and a computer program product for detecting one or more semantic headers in one or more tabular structures by utilizing a custom pre-trained embeddings model is provided. The present invention may include receiving the custom pre-trained embeddings model. The present invention may also include computing one or more dot product values associated with the one or more tabular structures from the one or more documents based on the context of each cell associated with the one or more tabular structures in the one or more documents. The present invention may then include generating one or more similarity feature values based on the computed one or more dot product values. The present invention may further include detecting the one or more semantic headers associated with the one or more tabular structures based on the one or more similarity feature values.

Claims (85)

1. A computer-implemented method comprising:

receiving a custom pre-trained embeddings model,

wherein the received custom pre-trained embeddings model provides a context associated with each term included in each cell from a plurality of cells associated with one or more tabular structures in one or more documents;

computing one or more dot product values associated with the one or more tabular structures from the one or more documents based on the context of each cell from the plurality of cells associated with the one or more tabular structures in the one or more documents,

wherein the one or more tabular structures in the one or more documents is identified by parsing the one or more documents;

analyzing a plurality of cell contents in each table attribute associated with the one or more tabular structures;

dividing each cell from the plurality of cells in each tabular structure into two or more buckets,

wherein a first bucket is populated with a current word vector associated with each cell, and

wherein a second bucket is populated with one or more remaining word vectors associated with each cell;

generating one or more similarity feature values based on the computed one or more dot product values,

wherein the computed one or more dot product values are normalized; and

detecting one or more semantic headers associated with the one or more tabular structures from the one or more documents based on the one or more similarity feature values.

2. The method of claim 1 , further comprising:

adding the one or more remaining word vectors associated with each cell in each tabular structure on a table attribute-by-table attribute basis to compute a total second bucket vector; and

dividing the computed total second bucket vector with a number of cells.

3. The method of claim 2 , further comprising:

computing the one or more dot product values from the populated first bucket and the populated second bucket; and

computing a sum of the dot product values based on the plurality of cells in a same plane of the table attributes in each tabular structure to compute the one or more dot product values for each of the table attributes in each tabular structure.

4. The method of claim 1 , wherein generating the one or more similarity feature values based on the computed one or more dot product values, wherein the computed one or more dot product values are normalized, further comprises:

sorting the computed one or more dot product values based on a numerical value associated with each computed dot product value from the computed one or more dot product values; and

in response to determining a lowest dot product value, designating the table attribute associated with the lowest dot product value as the header associated with the tabular structure.

5. The method of claim 1 , further comprising:

combining the generated one or more similarity feature values associated with the one or more tabular structures;

transmitting, to a machine learning (ML) classifier, the combined one or more similarity feature values associated with the one or more tabular structures; and

classifying a plurality of contents associated with a plurality of cells from the one or more tabular structures.

6. The method of claim 1 , further comprising:

identifying one or more groups of similar records in a clustering model based on the generated one or more similarity feature values;

labeling the one or more tabular structures based on the identified one or more groups of similar records; and

storing the labeled one or more tabular structures in the clustering model, wherein the clustering model includes a database.

7. A computer system for detecting one or more semantic headers in one or more tabular structures by utilizing a custom pre-trained embeddings model, comprising:

one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage medium, and program instructions stored on at least one of the one or more tangible storage medium for execution by at least one of the one or more processors via at least one of the one or more memories, wherein the computer system is capable of performing a method comprising:

receiving the custom pre-trained embeddings model,

wherein the received custom pre-trained embeddings model provides a context associated with each term included in each cell from a plurality of cells associated with the one or more tabular structures in one or more documents;

computing one or more dot product values associated with the one or more tabular structures from the one or more documents based on the context of each cell from the plurality of cells associated with the one or more tabular structures in the one or more documents,

wherein the one or more tabular structures in the one or more documents is identified by parsing the one or more documents;

analyzing a plurality of cell contents in each table attribute associated with the one or more tabular structures;

dividing each cell from the plurality of cells in each tabular structure into two or more buckets,

wherein a first bucket is populated with a current word vector associated with each cell, and

wherein a second bucket is populated with one or more remaining word vectors associated with each cell;

generating one or more similarity feature values based on the computed one or more dot product values,

wherein the computed one or more dot product values are normalized; and

detecting the one or more semantic headers associated with the one or more tabular structures from the one or more documents based on the one or more similarity feature values.

8. The computer system of claim 7 , further comprising:

adding the one or more remaining word vectors associated with each cell in each tabular structure on a table attribute-by-table attribute basis to compute a total second bucket vector; and

dividing the computed total second bucket vector with a number of cells.

9. The computer system of claim 8 , further comprising:

computing one or more dot product values from the populated first bucket and the populated second bucket; and

computing a sum of the dot product values based on the plurality of cells in a same plane of the table attributes in each tabular structure to compute the one or more dot product values for each of the table attributes in each tabular structure.

10. The computer system of claim 7 , wherein generating the one or more similarity feature values based on the computed one or more dot product values, wherein the computed one or more dot product values are normalized, further comprises:

sorting the computed one or more dot product values based on a numerical value associated with each computed dot product value from the computed one or more dot product values; and

in response to determining a lowest dot product value, designating the table attribute associated with the lowest dot product value as the header associated with the tabular structure.

11. The computer system of claim 7 , further comprising:

combining the generated one or more similarity feature values associated with the one or more tabular structures;

transmitting, to a machine learning (ML) classifier, the combined one or more similarity feature values associated with the one or more tabular structures; and

classifying a plurality of contents associated with a plurality of cells from the one or more tabular structures.

12. The computer system of claim 7 , further comprising:

identifying one or more groups of similar records in a clustering model based on the generated one or more similarity feature values;

labeling the one or more tabular structures based on the identified one or more groups of similar records; and

storing the labeled one or more tabular structures in the clustering model, wherein the clustering model includes a database.

13. A computer program product for detecting one or more semantic headers in one or more tabular structures by utilizing a custom pre-trained embeddings model, comprising:

one or more computer-readable storage media and program instructions stored on at least one of the one or more tangible storage media, the program instructions executable by a processor to cause the processor to perform a method comprising:

receiving the custom pre-trained embeddings model,

wherein the received custom pre-trained embeddings model provides a context associated with each term included in each cell from a plurality of cells associated with the one or more tabular structures in one or more documents;

computing one or more dot product values associated with the one or more tabular structures from the one or more documents based on the context of each cell from the plurality of cells associated with the one or more tabular structures in the one or more documents,

wherein the one or more tabular structures in the one or more documents is identified by parsing the one or more documents;

analyzing a plurality of cell contents in each table attribute associated with the one or more tabular structures;

dividing each cell from the plurality of cells in each tabular structure into two or more buckets,

wherein a first bucket is populated with a current word vector associated with each cell, and

wherein a second bucket is populated with one or more remaining word vectors associated with each cell;

generating one or more similarity feature values based on the computed one or more dot product values,

wherein the computed one or more dot product values are normalized; and

detecting the one or more semantic headers associated with the one or more tabular structures from the one or more documents based on the one or more similarity feature values.

14. The computer program product of claim 13 , further comprising:

adding the one or more remaining word vectors associated with each cell in each tabular structure on a table attribute-by-table attribute basis to compute a total second bucket vector; and

dividing the computed total second bucket vector with a number of cells.

15. The computer program product of claim 14 , further comprising:

computing one or more dot product values from the populated first bucket and the populated second bucket; and

computing a sum of the dot product values based on the plurality of cells in a same plane of the table attributes in each tabular structure to compute the one or more dot product values for each of the table attributes in each tabular structure.

16. The computer program product of claim 13 , wherein generating the one or more similarity feature values based on the computed one or more dot product values, wherein the computed one or more dot product values are normalized; further comprises:

sorting the computed one or more dot product values based on a numerical value associated with each computed dot product value from the computed one or more dot product values; and

in response to determining a lowest dot product value, designating the table attribute associated with the lowest dot product value as the header associated with the tabular structure.

17. The computer program product of claim 13 , further comprising:

combining the generated one or more similarity feature values associated with the one or more tabular structures;

transmitting, to a machine learning (ML) classifier, the combined one or more similarity feature values associated with the one or more tabular structures; and

classifying a plurality of contents associated with a plurality of cells associated with the combined one or more similarity feature values from the one or more tabular structures.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 10, 2019
From: NADIM, HASSAN; ALLEN, JOSHUA S.; CHRISTIANSON, KYLE G.; FREED, ANDREW R.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 050676/0640 →
Continuity (1)
Related Publication 20210109993A1 · Apr 15, 2021