IP Library › Granted Patent US 12,614,405
Granted Patent B2
US 12,614,405 · App. 17/577,605 · Granted Apr 28, 2026

Facilitating identification of fillable regions in a form

Inventors: Ashutosh Mehra (Noida, IN); Christopher Alan Tensmeyer (Fulton, MD); Vlad Ion Morariu (Potomac, MD); Jiuxiang Gu (Baltimore, MD)
Assignee: Adobe Inc.
G06V30/412G06F40/174G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,405
App. No.
17/577,605
Granted
Apr 28, 2026
Kind
B2
Abstract

Methods and systems are provided for facilitating identification of fillable regions and/or data associated therewith. In embodiments, a candidate fillable region indicating a region in a form that is a candidate for being fillable is obtained. Textual context indicating text from the form and spatial context indicating positions of the text within the form are also obtained. Fillable region data associated with the candidate fillable region is generated, via a machine learning model, using the candidate fillable region, the textual context, and the spatial context. Thereafter, a fillable form is generated using the fillable region data, the fillable form having one or more fillable regions for accepting input.

Claims (32)

1 . A method comprising:

obtaining a candidate fillable region indicating a region in a form that is a candidate for being fillable, the candidate fillable region identified via a first machine learning model that analyzes images associated with the form;

obtaining textual context indicating text from the form and spatial context indicating positions of the text within the form;

generating a token sequence associated with the candidate fillable region using the textual context and the spatial context, wherein the token sequence comprises a region token representing the candidate fillable region and text tokens representing words from the form that surround the candidate fillable region, the region token and the text tokens being interleaved in the token sequence to represent a sequence of position data represented in the form, wherein an order of the region token and the text tokens for the token sequence is based on the position data associated with the text in the form and the candidate fillable region;

generating, via a second machine learning model, fillable region data associated with the candidate fillable region based on the token sequence provided as input to the second machine learning model, wherein the fillable region data comprises data indicating a type of content format for a fillable region, a sub-type of content related to the type of content format for the fillable region, a duplicative fillable region, or a combination thereof; and

using the fillable region data to automatically generate a fillable form that replicates content of the form and has one or more fillable regions for accepting input.

2 . The method of claim 1 , wherein the first machine learning model comprises a visual machine learning model.

3 . The method of claim 1 , wherein the images analyzed by the first machine learning model to identify the candidate fillable region include a raw image and a linguistic image that represents text within the form.

4 . The method of claim 1 , wherein the form is obtained in response to an indication of a desire to generate the fillable form from the form.

5 . The method of claim 1 further comprising:

obtaining image features associated with the form; and

predicting, via the first machine learning model, the candidate fillable region and probabilities associated with a set of types of candidate fillable regions.

6 . The method of claim 1 , wherein the textual context and the spatial context is identified from the form.

7 . The method of claim 1 , wherein the textual context comprises words from the form, and the spatial context comprises bounding boxes or coordinates associated with the words.

8 . The method of claim 1 , wherein the second machine learning model further uses candidate region features associated with the candidate fillable region to generate the fillable region data, the candidate region features being generated via the first machine learning model used to identify the candidate fillable region.

9 . The method of claim 1 , wherein the fillable region data further comprises data indicating a group of fillable regions.

10 . One or more non-transitory computer-readable storage media storing computer-executable instructions that, when executed by a processing device, cause the processing device to perform operations comprising:

obtaining a candidate fillable region indicating a position of a region in a form that is a candidate for being fillable, the candidate fillable region generated via a vision machine learning model that analyzes images associated with the form;

obtaining textual context indicating text from the form and spatial context indicating positions of the text within the form;

generating a token sequence associated with the candidate fillable region using the textual context and the spatial context, wherein the token sequence comprises a region token representing the candidate fillable region and text tokens representing words from the form that surround the candidate fillable region, the region token and the text tokens being interleaved in the token sequence, to represent a sequence represented in the form, based on position data associated with the text in the form and the candidate fillable region; and

inputting the token sequence into a layout machine learning model to generate fillable region data associated with the candidate fillable region using the token sequence and a type of content format for a fillable region identified via the vision machine learning model in association with the candidate fillable region.

11 . The media of claim 10 further comprising using the fillable region data to automatically generate a fillable form having one or more fillable regions for accepting input.

12 . The media of claim 10 , wherein the candidate fillable region and the spatial context comprise indications of coordinates or bounding boxes.

13 . The media of claim 10 , wherein the layout machine learning model is trained to generate fillable region data that indicates a particular type of content format for the fillable region comprising a signature, a checkbox, a text, or a non-fillable region.

14 . The media of claim 10 , wherein the layout machine learning model is trained to generate fillable region data that indicates a duplicate fillable region.

15 . The media of claim 10 , wherein the layout machine learning model is trained to generate fillable region data that indicates a sub-type of fillable region specifying a sub-type of the type of content format for the fillable region generated via the vision machine learning model.

16 . The media of claim 10 , wherein the layout machine learning model is trained to generate fillable region data that indicates a grouping associated with the candidate fillable region, wherein the grouping comprises at least one other candidate fillable region, a field, or a label.

17 . A system comprising one or more hardware processors and a memory component coupled to the one or more hardware processors, the one or more hardware processors to perform operations comprising:

obtaining a training dataset including a set of forms and training labels indicating positions of fillable regions, textual context indicating words in the set of forms, and spatial context indicating positions of the words in the set of forms; and

training a language machine learning model to generate fillable region data associated with candidate fillable regions, within a form, predicted by a visual machine learning model, wherein the language machine learning model is trained using the training dataset including the set of forms and the training labels indicating the positions of fillable regions, the textual context indicating words in the set of forms, and the spatial context indicating positions of the words in the set of forms, and wherein the trained language machine learning model is used to generate fillable region data that indicates a type of content format for a fillable region, a sub-type of content related to the type of content format for the fillable region, a redundant fillable region, or a combination thereof based on an input token sequence associated with the fillable region, the input token sequence including a region token representing the fillable region interleaved with text tokens representing the words that surround the fillable region, wherein an order of the region token and the text tokens is based on position data of the words and the fillable region.

18 . The system of claim 17 , wherein the language machine learning model is further trained using types of fillable regions generated by the visual machine learning model or candidate region features generated by the visual machine learning model.

19 . The system of claim 17 , wherein the trained language machine learning model is used to generate fillable region data that further indicates a grouping associated with a fillable region.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2022
From: MEHRA, ASHUTOSH; TENSMEYER, CHRISTOPHER ALAN; MORARIU, VLAD ION; GU, JIUXIANG
To: ADOBE INC.
Reel/Frame 058677/0327 →
Continuity (1)
Related Publication 20230230406A1 · Jul 20, 2023
References Cited (16)
US 10402640B1 · Becker et al. · 2019 [cited by applicant]
US 20190286691A1 · Sodhani · 2019 [cited by examiner]
US 20220292258A1 · Zeng · 2022 [cited by examiner]
US 20230084845A1 · Lu · 2023 [cited by examiner]
US 20230196748A1 · Hariharan · 2023 [cited by examiner]
US 20230351115A1 · Zeng · 2023 [cited by examiner]
EP 3882814A1 · 2021 [cited by applicant]
Soto, Carlos X. Visual detection with context fo rdocument layout analysis. No. BNL-212125-2019-COPA. Brookhaven National Lab.(BNL), Upton, NY (United States), 2019. (Year: 2019). [cited by examiner]
Ren, Shaoqing, et al. “Faster r-cnn: Towards real-time object detection with region proposal networks.” Advances in neural information processing systems 28 (2015). (Year: 2015). [cited by examiner]
Xu, Yiheng, et al. “Layoutlm: Pre-training of text and layout for document image understanding.” Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2020. [cited by applicant]
Xu, Yang, et al. “LayoutLMv2: Multi-modal pre-training for visually-rich document understanding.” arXiv preprint arXiv:2012.14740 (2020). [cited by applicant]
Li, C., Bi, B., Yan, M., Wang, W., Huang, S., Huang, F., & Si, L. (2021). StructuralLM: Structural Pre-training for Form Understanding. arXiv preprint arXiv:2105.11210. [cited by applicant]
Jaume, G., Ekenel, H. K., & Thiran, J. P. (Sep. 2019). Funsd: A dataset for form understanding in noisy scanned documents. In 2019 International Conference on Document Analysis and Recognition Workshops (ICDARW) (vol. 2… [cited by applicant]
Aggarwal, M., Sarkar, M., Gupta, H., & Krishnamurthy, B. (2020). Multi-modal association based grouping for form structure extraction. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision … [cited by applicant]
Bodla, N., Singh, B., Chellappa, R., & Davis, L. S. (2017). Soft-NMS—improving object detection with one line of code. In Proceedings of the IEEE international conference on computer vision (pp. 5561-5569). [cited by applicant]
Office Action received for German Patent Application No. 102022129588.5, mailed on Oct. 1, 2025, 12 pages of Original OA only. [cited by applicant]