ADAPTIVE NATURAL LANGUAGE PROCESSING MODEL TRAINING WITH QUALITY ASSESSMENT
Systems and methods are disclosed for training a Natural Language Processing (NLP) model through iterative processing. A method involves performing multiple iterations to train the NLP model. Each iteration includes defining data points with assigned labels and preparing natural language queries for these points. A neural network processes real-world documents to generate extracted values from the real-world document to be used for querying. The extracted values are then validated by comparing them to the document data and identifying validated extracted values. The NLP model quality is evaluated using these validated values, and a quality score is determined. Upon identifying a suitable quality, the NLP model is defined as a fine-tuned model and configured to process new real-world documents for data point extraction.
1 . A system comprising:
one or more hardware processors of a machine; and
at least one memory storing instructions that, when executed by the one or more hardware processors, cause the machine to perform operations comprising:
performing a plurality of iterations to train a Natural Language Processing (NLP) model, each iteration comprising:
defining a list of data points and assigning labels to each data point;
preparing a query in natural language for each data point;
processing, via a neural network, a real-world document to generate an extracted value for the query;
validating the extracted value, the validating comprising comparing the extracted value to data in the real-world document to identify a validated extracted value;
evaluating, using the validated extracted value, a model quality of the NLP model; and
determining a quality score of the NLP model using the model quality;
defining a trained NLP model as a fine-tuned model; and
configuring the trained NLP model to process a new real-world document to extract a data point.
2 . The system of claim 1 , wherein processing the real-world document further comprises:
receiving, from the real-world document, at least one type of the data, the at least one type of the data comprising text data, layout data, or image data; and
processing the at least one type of the data using a text-image-layout transformer.
3 . The system of claim 2 , the operations further comprising:
generating contextualized image embeddings by processing the image data using a convolutional network;
adding the contextualized image embeddings to semantic embeddings of the text data to create combined embeddings; and
applying spatial bias augmentation to the combined embeddings.
4 . The system of claim 1 , wherein the neural network comprises:
an encoder-decoder model;
a spatial model; and
a multi-modal model.
5 . The system of claim 4 , wherein the spatial model further comprises:
a spatial-aware transformer configured to employ a self-attention equation; and
a word-centric mask configured to contextualize image data and text data.
6 . The system of claim 5 , wherein the operations further comprise:
extending a T5 transformer to enable consumption of multi-modal input, the extending comprising:
disregarding positional embeddings;
introducing relative bias by extending the self-attention equation with a sequential bias term;
calculating biases for a relative horizontal distance and a relative vertical distance between each pair of tokens in the text data; and
adding the calculated biases to the sequential bias term.
7 . The system of claim 1 , wherein the list of data points further comprises:
key information extraction from the real-world document;
document classification based on content and structure of the real-world document;
an answer to a natural language question about the content of the real-world document; and
extracted values corresponding to predefined data points within the real-world document.
8 . A method comprising:
performing, by at least one hardware processor, a plurality of iterations to train a Natural Language Processing (NLP) model, each iteration comprising:
defining a list of data points and assigning labels to each data point;
preparing a query in natural language for each data point;
processing, via a neural network, a real-world document to generate an extracted value for the query;
validating the extracted value, the validating comprising comparing the extracted value to data in the real-world document to identify a validated extracted value;
evaluating, using the validated extracted value, a model quality of the NLP model; and
determining a quality score of the NLP model using the model quality;
defining a trained NLP model as a fine-tuned model; and
configuring the trained NLP model to process a new real-world document to extract a data point.
9 . The method of claim 8 , wherein processing the real-world document further comprises:
receiving, from the real-world document, at least one type of the data, the at least one type of the data comprising text data, layout data, or image data; and
processing the at least one type of the data using a text-image-layout transformer.
10 . The method of claim 8 , further comprising:
generating contextualized image embeddings by processing image data using a convolutional network;
adding the contextualized image embeddings to semantic embeddings of text data to create combined embeddings; and
applying spatial bias augmentation to the combined embeddings.
11 . The method of claim 8 , wherein the neural network comprises:
an encoder-decoder model;
a spatial model; and
a multi-modal model.
12 . The method of claim 11 , wherein the spatial model further comprises:
a spatial-aware transformer configured to employ a self-attention equation; and
a word-centric mask configured to contextualize image data and text data.
13 . The method of claim 12 , further comprising:
extending a T5 transformer to enable consumption of multi-modal input, the extending comprising:
disregarding positional embeddings;
introducing relative bias by extending the self-attention equation with a sequential bias term;
calculating biases for a relative horizontal distance and a relative vertical distance between each pair of tokens in the text data; and
adding the calculated biases to the sequential bias term.
14 . The method of claim 8 , wherein the list of data points further comprises:
key information extraction from the real-world document;
document classification based on content and structure of the real-world document;
an answer to a natural language question about the content of the real-world document; and
extracted values corresponding to predefined data points within the real-world document.
15 . A non-transitory computer medium embodying instructions that, when executed by a machine, cause the computer medium to perform operations comprising:
performing, by at least one hardware processor, a plurality of iterations to train a Natural Language Processing (NLP) model, each iteration comprising:
defining a list of data points and assigning labels to each data point;
preparing a query in natural language for each data point;
processing, via a neural network, a real-world document to generate an extracted value for the query;
validating the extracted value, the validating comprising comparing the extracted value to data in the real-world document to identify a validated extracted value;
evaluating, using the validated extracted value, a model quality of the NLP model; and
determining a quality score of the NLP model using the model quality;
defining a trained NLP model as a fine-tuned model; and
configuring the trained NLP model to process a new real-world document to extract a data point.
16 . The non-transitory computer medium of claim 15 , wherein processing the real-world document further comprises:
receiving, from the real-world document, at least one type of the data, the at least one type of the data comprising text data, layout data, or image data; and
processing the at least one type of the data using a text-image-layout transformer.
17 . The non-transitory computer medium of claim 16 , further comprising:
generating contextualized image embeddings by processing the image data using a convolutional network;
adding the contextualized image embeddings to semantic embeddings of the text data to create combined embeddings; and
applying spatial bias augmentation to the combined embeddings.
18 . The non-transitory computer medium of claim 15 , wherein the neural network comprises:
an encoder-decoder model;
a spatial model; and
a multi-modal model.
19 . The non-transitory computer medium of claim 18 , wherein the spatial model further comprises:
a spatial-aware transformer configured to employ a self-attention equation; and
a word-centric mask configured to contextualize image data and text data.
20 . The non-transitory computer medium of claim 19 , wherein the operations further comprise:
extending a T5 transformer to enable consumption of multi-modal input, the extending comprising:
disregarding positional embeddings;
introducing relative bias by extending the self-attention equation with a sequential bias term;
calculating biases for a relative horizontal distance and a relative vertical distance between each pair of tokens in the text data; and
adding the calculated biases to the sequential bias term.