Facilitating identification of fillable regions in a form
Methods and systems are provided for facilitating identification of fillable regions and/or data associated therewith. In embodiments, a candidate fillable region indicating a region in a form that is a candidate for being fillable is obtained. Textual context indicating text from the form and spatial context indicating positions of the text within the form are also obtained. Fillable region data associated with the candidate fillable region is generated, via a machine learning model, using the candidate fillable region, the textual context, and the spatial context. Thereafter, a fillable form is generated using the fillable region data, the fillable form having one or more fillable regions for accepting input.
1 . A method comprising:
obtaining a candidate fillable region indicating a region in a form that is a candidate for being fillable, the candidate fillable region identified via a first machine learning model that analyzes images associated with the form;
obtaining textual context indicating text from the form and spatial context indicating positions of the text within the form;
generating a token sequence associated with the candidate fillable region using the textual context and the spatial context, wherein the token sequence comprises a region token representing the candidate fillable region and text tokens representing words from the form that surround the candidate fillable region, the region token and the text tokens being interleaved in the token sequence to represent a sequence of position data represented in the form, wherein an order of the region token and the text tokens for the token sequence is based on the position data associated with the text in the form and the candidate fillable region;
generating, via a second machine learning model, fillable region data associated with the candidate fillable region based on the token sequence provided as input to the second machine learning model, wherein the fillable region data comprises data indicating a type of content format for a fillable region, a sub-type of content related to the type of content format for the fillable region, a duplicative fillable region, or a combination thereof; and
using the fillable region data to automatically generate a fillable form that replicates content of the form and has one or more fillable regions for accepting input.
2 . The method of claim 1 , wherein the first machine learning model comprises a visual machine learning model.
3 . The method of claim 1 , wherein the images analyzed by the first machine learning model to identify the candidate fillable region include a raw image and a linguistic image that represents text within the form.
4 . The method of claim 1 , wherein the form is obtained in response to an indication of a desire to generate the fillable form from the form.
5 . The method of claim 1 further comprising:
obtaining image features associated with the form; and
predicting, via the first machine learning model, the candidate fillable region and probabilities associated with a set of types of candidate fillable regions.
6 . The method of claim 1 , wherein the textual context and the spatial context is identified from the form.
7 . The method of claim 1 , wherein the textual context comprises words from the form, and the spatial context comprises bounding boxes or coordinates associated with the words.
8 . The method of claim 1 , wherein the second machine learning model further uses candidate region features associated with the candidate fillable region to generate the fillable region data, the candidate region features being generated via the first machine learning model used to identify the candidate fillable region.
9 . The method of claim 1 , wherein the fillable region data further comprises data indicating a group of fillable regions.
10 . One or more non-transitory computer-readable storage media storing computer-executable instructions that, when executed by a processing device, cause the processing device to perform operations comprising:
obtaining a candidate fillable region indicating a position of a region in a form that is a candidate for being fillable, the candidate fillable region generated via a vision machine learning model that analyzes images associated with the form;
obtaining textual context indicating text from the form and spatial context indicating positions of the text within the form;
generating a token sequence associated with the candidate fillable region using the textual context and the spatial context, wherein the token sequence comprises a region token representing the candidate fillable region and text tokens representing words from the form that surround the candidate fillable region, the region token and the text tokens being interleaved in the token sequence, to represent a sequence represented in the form, based on position data associated with the text in the form and the candidate fillable region; and
inputting the token sequence into a layout machine learning model to generate fillable region data associated with the candidate fillable region using the token sequence and a type of content format for a fillable region identified via the vision machine learning model in association with the candidate fillable region.
11 . The media of claim 10 further comprising using the fillable region data to automatically generate a fillable form having one or more fillable regions for accepting input.
12 . The media of claim 10 , wherein the candidate fillable region and the spatial context comprise indications of coordinates or bounding boxes.
13 . The media of claim 10 , wherein the layout machine learning model is trained to generate fillable region data that indicates a particular type of content format for the fillable region comprising a signature, a checkbox, a text, or a non-fillable region.
14 . The media of claim 10 , wherein the layout machine learning model is trained to generate fillable region data that indicates a duplicate fillable region.
15 . The media of claim 10 , wherein the layout machine learning model is trained to generate fillable region data that indicates a sub-type of fillable region specifying a sub-type of the type of content format for the fillable region generated via the vision machine learning model.
16 . The media of claim 10 , wherein the layout machine learning model is trained to generate fillable region data that indicates a grouping associated with the candidate fillable region, wherein the grouping comprises at least one other candidate fillable region, a field, or a label.
17 . A system comprising one or more hardware processors and a memory component coupled to the one or more hardware processors, the one or more hardware processors to perform operations comprising:
obtaining a training dataset including a set of forms and training labels indicating positions of fillable regions, textual context indicating words in the set of forms, and spatial context indicating positions of the words in the set of forms; and
training a language machine learning model to generate fillable region data associated with candidate fillable regions, within a form, predicted by a visual machine learning model, wherein the language machine learning model is trained using the training dataset including the set of forms and the training labels indicating the positions of fillable regions, the textual context indicating words in the set of forms, and the spatial context indicating positions of the words in the set of forms, and wherein the trained language machine learning model is used to generate fillable region data that indicates a type of content format for a fillable region, a sub-type of content related to the type of content format for the fillable region, a redundant fillable region, or a combination thereof based on an input token sequence associated with the fillable region, the input token sequence including a region token representing the fillable region interleaved with text tokens representing the words that surround the fillable region, wherein an order of the region token and the text tokens is based on position data of the words and the fillable region.
18 . The system of claim 17 , wherein the language machine learning model is further trained using types of fillable regions generated by the visual machine learning model or candidate region features generated by the visual machine learning model.
19 . The system of claim 17 , wherein the trained language machine learning model is used to generate fillable region data that further indicates a grouping associated with a fillable region.