IP Library › Granted Patent US 12,602,547
Granted Patent B2
US 12,602,547 · App. 18/240,480 · Granted Apr 14, 2026

Domain adapting graph networks for visually rich documents

Inventors: Amit Agarwal (Kolkata, IN); Srikant Panda (Bangalore, IN); Deepak Karmakar (Seraikella, IN); Kulbhushan Pachauri (Bangalore, IN)
Assignee: Oracle International Corporation
G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,547
App. No.
18/240,480
Filed
Aug 31, 2023
Granted
Apr 14, 2026
Kind
B2
Art Unit
2659
USPC
704/9
Abstract

In some implementations, techniques described herein may include identifying text in a visually rich document and determining a sequence for the identified text. The techniques may include selecting a language model based at least in part on the identified text and the determined sequence. Moreover, the techniques may include assigning each word of the identified text to a respective token to generate textual features corresponding to the identified text. The techniques may include extracting visual features corresponding to the identified text. The techniques may include determining positional features for each word of the identified text. The techniques may include generating a graph representing the visually rich document, each node in the graph representing each of the visual features, textual features, and positional features of a respective word of the identified text. The techniques may include training a classifier on the graph to classify each respective word of the identified text.

Claims (46)

1 . A computer-implemented method comprising:

identifying, by a computing system, text in a visually rich document;

determining, by the computing system, a sequence for the identified text, the sequence comprising a numerical order for each word of the identified text;

selecting, by the computing system, a pretrained language model based at least in part on the identified text, the determined sequence, and a domain of the identified text;

assigning, by the computing system, each word of the identified text to a respective token using the pretrained language model and the determined sequence to generate textual features corresponding to the identified text, each respective token comprising a string of one or more words;

extracting, by the computing system, visual features corresponding to the identified text, the visual features comprising information about a plurality of pixels representing each word of the identified text;

determining, by the computing system, positional features for each word of the identified text, the positional features comprising respective coordinates within the visually rich document for each word of the identified text, the respective coordinates for a word being coordinates corresponding to a position of the word or coordinates corresponding to a region of the visually rich document that contains the word;

fusing, by the computing system, the textual features and the visual features to generate fused features;

generating, by the computing system, a document model representing the visually rich document by assigning the positional features and the fused features to nodes in the document model, each node in the document model representing the fused features and the positional features of a respective word of the identified text;

training, by the computing system, a classifier to classify each respective word of the identified text, the classifier being trained on the document model representing the visually rich document; and

classifying, by the computing system, the respective word of the identified text with the classifier.

2 . The method of claim 1 , wherein the respective word is classified as a key or a value of a key value pair.

3 . The method of claim 1 , wherein the document model is a graph neural network.

4 . The method of claim 1 , wherein the domain comprises a language or a subject matter of the identified text.

5 . The method of claim 1 , wherein the visually rich document comprises at least one of an invoice, a receipt, an insurance form, a boarding pass, or an identification card.

6 . A non-transitory computer-readable medium storing a plurality of instructions that when executed control a computer system to perform operations comprising:

identifying, by a computing system, text in a visually rich document;

determining, by the computing system, a sequence for the identified text, the sequence comprising a numerical order for each word of the identified text;

selecting, by the computing system, a pretrained language model based at least in part on the identified text, the determined sequence, and a domain of the identified text;

assigning, by the computing system, each word of the identified text to a respective token using the pretrained language model and the determined sequence to generate textual features corresponding to the identified text, each respective token comprising a string of one or more words;

extracting, by the computing system, visual features corresponding to the identified text, the visual features comprising information about a plurality of pixels representing each word of the identified text;

determining, by the computing system, positional features for each word of the identified text, the positional features comprising respective coordinates within the visually rich document for each word of the identified text, the respective coordinates for a word being coordinates corresponding to a position of the word or coordinates corresponding to a region of the visually rich document that contains the word;

fusing, by the computing system, the textual features and the visual features to generate fused features;

generating, by the computing system, a document model representing the visually rich document by assigning the positional features and the fused features to nodes in the document model, each node in the document model representing the fused features and the positional features of a respective word of the identified text;

training, by the computing system, a classifier to classify each respective word of the identified text, the classifier being trained on the document model representing the visually rich document; and

classifying, by the computing system, the respective word of the identified text with the classifier.

7 . The non-transitory computer-readable medium of claim 6 , wherein the respective word is classified as a key or a value of a key value pair.

8 . The non-transitory computer-readable medium of claim 6 , wherein the document model is a graph neural network.

9 . The non-transitory computer-readable medium of claim 6 , wherein the domain comprises a language or a subject matter of the identified text.

10 . The non-transitory computer-readable medium of claim 6 , wherein the visually rich document comprises at least one of an invoice, a receipt, an insurance form, a boarding pass, or an identification card.

11 . A system comprising:

a computer-readable medium; and

one or more processors for executing instructions stored on the computer-readable medium to at least perform operations comprising:

identifying, by a computing system, text in a visually rich document;

determining, by the computing system, a sequence for the identified text, the sequence comprising a numerical order for each word of the identified text;

selecting, by the computing system, a pretrained language model based at least in part on the identified text, the determined sequence, and a domain of the identified text;

assigning, by the computing system, each word of the identified text to a respective token using the pretrained language model and the determined sequence to generate textual features corresponding to the identified text, each respective token comprising a string of one or more words;

extracting, by the computing system, visual features corresponding to the identified text, the visual features comprising information about a plurality of pixels representing each word of the identified text;

determining, by the computing system, positional features for each word of the identified text, the positional features comprising respective coordinates within the visually rich document for each word of the identified text, the respective coordinates for a word being coordinates corresponding to a position of the word or coordinates corresponding to a region of the visually rich document that contains the word;

fusing, by the computing system, the textual features and the visual features to generate fused features,

generating, by the computing system, a graph representing the visually rich document by assigning the positional features and the fused features to nodes in the document model, each node in the graph representing the fused features and the positional features of a respective word of the identified text;

training, by the computing system, a classifier to classify each respective word of the identified text, the classifier being trained on the graph representing the visually rich document; and

classifying, by the computing system, the respective word of the identified text with the classifier.

12 . The system of claim 11 , wherein the respective word is classified as a key or a value of a key value pair.

13 . The system of claim 11 , wherein the graph is a graph neural network.

14 . The system of claim 11 , wherein the domain comprises a language or a subject matter of the identified text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2023
From: AGARWAL, AMIT; PANDA, SRIKANT; KARMAKAR, DEEPAK; PACHAURI, KULBHUSHAN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 064799/0677 →
Continuity (1)
Related Publication 20240289551A1 · Aug 29, 2024
References Cited (81)
US 9424668B1 · Petrou et al. · 2016 [cited by applicant]
US 11087081B1 · Srivastava et al. · 2021 [cited by applicant]
US 11341367B1 · Barbosa et al. · 2022 [cited by applicant]
US 11823478B2 · Agarwal et al. · 2023 [cited by applicant]
US 11989964B2 · Agarwal et al. · 2024 [cited by applicant]
US 12106595B2 · Agarwal et al. · 2024 [cited by applicant]
US 12182498B1 · Sunkara · 2024 [cited by examiner]
US 20120062574A1 · Dhoolia et al. · 2012 [cited by applicant]
US 20160364608A1 · Sengupta · 2016 [cited by examiner]
US 20200125954A1 · Truong et al. · 2020 [cited by applicant]
US 20200380623A1 · Ranjan · 2020 [cited by examiner]
US 20200410231A1 · Chua et al. · 2020 [cited by applicant]
US 20210089587A1 · Gupta et al. · 2021 [cited by applicant]
US 20210133645A1 · Tazi et al. · 2021 [cited by applicant]
US 20210158093A1 · Kaynig-Fittkau et al. · 2021 [cited by applicant]
US 20210248323A1 · Maheshwari et al. · 2021 [cited by applicant]
US 20220092267A1 · Hou et al. · 2022 [cited by applicant]
US 20220171938A1 · Jalaluddin et al. · 2022 [cited by applicant]
US 20230146501A1 · Agarwal et al. · 2023 [cited by applicant]
US 20230153335A1 · Mcneill · 2023 [cited by applicant]
US 20230326224A1 · Agarwal et al. · 2023 [cited by applicant]
US 20230394235A1 · Rahman · 2023 [cited by examiner]
CN 107977345A · 2018 [cited by applicant]
CN 113936340A · 2022 [cited by applicant]
CN 114491010A · 2022 [cited by applicant]
WO 2022078922A1 · 2022 [cited by applicant]
Yu et al, (2020), “PICK: Processing Key Information Extraction from Documents using Improved Graph Learning-Convolutional Networks”, (Year: 2020). [cited by examiner]
U.S. Appl. No. 18/379,091, “Non-Final Office Action”, mailed Jun. 6, 2024, 19 pages. [cited by applicant]
U.S. Appl. No. 18/379,091, “Notice of Allowability”, mailed Aug. 14, 2024, 4 pages. [cited by applicant]
U.S. Appl. No. 18/379,091, “Notice of Allowance”, mailed Jul. 29, 2024, 7 pages. [cited by applicant]
International Application No. PCT/US2024/016876, “International Search Report and Written Opinion”, mailed Jun. 5, 2024, 16 pages. [cited by applicant]
Tang et al., “MatchVIE: Exploiting Match Relevancy between Entities for Visual Information Extraction”, Available online at: https://arxiv.org/pdf/2106.12940, Jun. 24, 2021, 7 pages. [cited by applicant]
Wei et al., “Robust Layout-Aware IE for Visually Rich Documents with Pre-Trained Language Models”, Cornell University Library, Available online at: https://arxiv.org/pdf/2005.11017, May 22, 2020, 10 pages. [cited by applicant]
Xu et al., “LayoutLMv2: Multi-Modal Pre-Training for Visually-Rich Document Understanding”, Available Online at: https://arxiv.org/pdf/2012.14740v4, Jan. 10, 2022, 13 pages. [cited by applicant]
“Augmentation Pipeline for Rendering Synthetic Paper Printing, Faxing, Scanning and Copy Machine Processes”, Available Online at: https://github.com/sparkfish/augraphy, Accessed from Internet on Apr. 4, 2022, pp. 1-13. [cited by applicant]
“BERT”, Available Online at: https://huggingface.co/docs/transformers/model_doc/bert, Accessed from Internet on Mar. 2, 2022, 114 pages. [cited by applicant]
“DataGen—CeDar”, Centre for Applied Data Analytics Research, Available Online at: https://old.ceadar.ie/wp-content/uploads/CeADAR_Flyer_DataGen_v2.pdf, 1 page. [cited by applicant]
“Datagen Synthetic Image Datasets for Computer Vision”, Available Online at: https://datagen.tech/, Accessed from Internet on Jun. 1, 2022, pp. 1-6. [cited by applicant]
“Deterministic Algorithm”, Wikipedia, Available Online at: https://en.wikipedia.org/wiki/Deterministic_algorithm, Accessed from Internet on Mar. 2, 2022, 4 pages. [cited by applicant]
“DistilBERT”, Available Online at: https://huggingface.co/docs/transformers/model_doc/distilbert, Accessed from Internet on Mar. 2, 2022, 62 pages. [cited by applicant]
“LayoutLMFT”, Available online at https://github.com/microsoft/unilm/tree/master/layoutlmft, Accessed from Internet on Aug. 19, 2021, 2 pages. [cited by applicant]
“Public Leader for SROIE”, Available online at https://rrc.cvc.uab.es/?ch=13&com=evaluation&task=3, Accessed from Internet on: Aug. 19, 2021, 5 pages. [cited by applicant]
“Scipy.Optimize.Linear_Sum_Assignment”, Available Online at: https://docs.scipy.org/doc/scipy-0.18.1/reference/generated/scipy.optimize.linear_sum_assignment.html, Sep. 19, 2016, 2 pages. [cited by applicant]
“Sklearn.Decomposition.PCA”, Available Online at: https://scikit-learn.org/stable/modules/generated/sklearn.decomposition.PCA.html, Accessed from Internet on Mar. 2, 2022, 6 pages. [cited by applicant]
“Spaczz: Fuzzy Matching and More for Spacy”, Available Online at: https://github.com/gandersen101/spaczz, Accessed from Internet on Mar. 2, 2022, 20 pages. [cited by applicant]
“Text Distance”, Available Online at: https://github.com/life4/textdistance, Accessed from Internet on Mar. 2, 2022, 9 pages. [cited by applicant]
“Welcome to Albumentations Documentation”, Available Online at: https://albumentations.ai/docs/, Accessed from Internet on Apr. 4, 2022, pp. 1-3. [cited by applicant]
“WordNet: A Lexical Database for English”, Princeton University, Available Online at: https://wordnet.princeton.edu/, Accessed from Internet on Mar. 2, 2022, 4 pages. [cited by applicant]
U.S. Appl. No. 17/714,806, “Notice of Allowance”, mailed Jun. 22, 2023, 14 pages. [cited by applicant]
U.S. Appl. No. 17/714,806, “Notice of Allowance”, mailed Jul. 26, 2023, 7 pages. [cited by applicant]
Abdelzad et al., “Detecting Out-of-Distribution Inputs in Deep Neural Networks Using an Early-Layer Output”, Available Online at: https://arxiv.org/pdf/1910.10307.pdf, Oct. 23, 2019, 15 pages. [cited by applicant]
Ba, “Meta-data Driven Key-Value Pairs Extraction with Azure Form Recognizer”, Available Online at: https://techcommunity.microsoft.com/t5/ai-cognitive-services-blog/meta-data-driven-key-value-pairs-extraction-with-azure… [cited by applicant]
Biswas et al., “DocSynth: A Layout Guided Approach for Controllable Document Image Synthesis”, Available Online at: https://arxiv.org/pdf/2107.02638.pdf, Jul. 6, 2021, pp. 1-15. [cited by applicant]
Brems, “A One-Stop Shop for Principal Component Analysis”, Towards Data Science, Available Online at: https://towardsdatascience.com/a-one-stop-shop-for-principal-component-analysis-5582fb7e0a9c, Apr. 18, 2017, 14 pages. [cited by applicant]
Chogovadze et al., “Controllable Data Augmentation Through Deep Relighting”, Available Online at: https://arxiv.org/pdf/2110.13996.pdf, Oct. 26, 2021, pp. 1-15. [cited by applicant]
Delalandre et al., “Generation of Synthetic Documents for Performance Evaluation of Symbol Recognition & Spotting Systems.”, International Journal on Document Analysis and Recognition, vol. 13, No. 3, Sep. 2010, pp. 187… [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, Available Online at: https://arxiv.org/pdf/1810.04805.pdf, May 24, 2019, 16 pages. [cited by applicant]
Gautam , “Form Data Augmentation: Repository for Augmenting Data in Forms, Invoices and Receipts for Document Image Understanding”, Available Online at: https://github.com/gautam-aayush/form-data-augmentation, Mar. 19, … [cited by applicant]
Ghosh , “Invoice Information Extraction Using OCR and Deep Learning”, Available Online at: https://medium.com/analytics-vidhya/invoice-information-extraction-using-ocr-and-deep-learning-b79464f54d69, Jan. 14, 2021, 31 p… [cited by applicant]
Huang et al., “Out-of-Distribution Detection for LiDAR-based 3D Object Detection”, Available Online at: https://arxiv.org/pdf/2209.14435v1.pdf, Sep. 28, 2022, 7 pages. [cited by applicant]
Jaadi , “A Step-by-Step Explanation of Principal Component Analysis (PCA)”, Builtin.com, Available Online at: https://builtin.com/data-science/step-step-explanation-principal-component-analysis, Apr. 1, 2021, 8 pages. [cited by applicant]
Journet et al., “DocCreator: A New Software for Creating Synthetic Ground-Truthed Document Images”, Journal of Imaging, vol. 3, Available Online at: https://hal.archives-ouvertes.fr/hal-01668915/file/jimaging.pdf, Dec. … [cited by applicant]
Lin et al., “An Efficient Data Augmentation Network for Out-of-Distribution Image Detection”, IEEE Access, vol. 9, Feb. 24, 2021, pp. 35313-35323. [cited by applicant]
Liu et al., “Self-Supervised Learning: Generative or Contrastive”, Available online at https://arxiv.org/pdf/2006.08218.pdf, Mar. 20, 2021, pp. 1-24. [cited by applicant]
Luan et al., “Out-Of-Distribution Detection for Deep Neural Networks with Isolation Forest and Local Outlier Factor”, Available Online at: https://www.researchgate.net/publication/354189192_Out-Of-Distribution_Detection… [cited by applicant]
Ma, “NLP Augmentation”, Available online at: https://github.com/makcedward/nlpaug, 2019, 4 pages. [cited by applicant]
Ma, “nlpaug: Data Augmentation for NLP”, Available Online at: https://github.com/makcedward/nlpaug, Accessed from Internet on Apr. 4, 2022, 21 pages. [cited by applicant]
Moore et al., “Hungarian Maximum Matching Algorithm”, Brilliant Math & Science Wiki, Available Online at: https://brilliant.org/wiki/hungarian-matching/, Accessed from Internet on Mar. 2, 2022, 7 pages. [cited by applicant]
Moore et al., “Matching (Graph Theory)”, Brilliant Math & Science Wiki, Available Online at: https://brilliant.org/wiki/matching/, Accessed from Internet on Mar. 2, 2022, 6 pages. [cited by applicant]
Moore et al., “Matching Algorithms (Graph Theory)”, Brilliant Math & Science Wiki, Available Online at: https://brilliant.org/wiki/matching-algorithms/, Accessed from Internet on Mar. 2, 2022, 5 pages. [cited by applicant]
Rawat et al., “PnPOOD : Out-Of-Distribution Detection for Text Classification via Plug and Play Data Augmentation”, Available Online at: https://arxiv.org/abs/2111.00506, Oct. 31, 2021, 9 pages. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, International Conference on Medical Image Computing and Computer-Assisted Intervention, Available Online at URL: https://arxiv.org/p… [cited by applicant]
Sebastianelli et al., “Automatic Dataset Builder for Machine Learning Applications to Satellite Imagery”, Software X, vol. 15, Jul. 2021, 7 pages. [cited by applicant]
Sun et al., “Spatial Dual-Modality Graph Reasoning for Key Information Extraction”, Journal of Latex Class Files, vol. 14, No. 8, Available online at https://arxiv.org/pdf/2103.14470.pdf, Aug. 2015, pp. 1-9. [cited by applicant]
Van Laer , “Recognition of Named Entities on Invoices for IxorDocs”, Available Online at: https://medium.com/ixorthink/recognition-of-named-entities-on-invoices-for-ixordocs-9bef38d24429, Aug. 2, 2018, 14 pages. [cited by applicant]
Veyseh et al., “Improving Keyphrase Extraction with Data Augmentation and Information Filtering”, Available Online at: https://www.researchgate.net/publication/363501653_Improving_Keyphrase_Extraction_with_Data_Augmenta… [cited by applicant]
White , “By 2024, 60% of the Data Used for the Development of AI and Analytics Projects Will Be Synthetically Generated”, Available Online at: https://blogs.gartner.com/andrew_white/2021/07/24/by-2024-60-of-the-data-use… [cited by applicant]
Xu et al., “LayoutLM: Pre-Training of Text and Layout for Document Image Understanding”, Available Online at: https://arxiv.org/pdf/1912.13318.pdf, Aug. 23-27, 2020, 9 pages. [cited by applicant]
You et al., “Graph Contrastive Learning with Augmentations”, 34th Conference on Neural Information Processing Systems, Available online at https://papers.nips.cc/paper/2020/file/3fe230348e9a12c13120749e3f9fa4cd-Paper.pd… [cited by applicant]
Yu et al., “PICK: Processing Key Information Extraction from Documents using Improved Graph Learning-Convolutional Networks”, arXiv:2004.07464, Available Online at: https://arxiv.org/pdf/2004.07464.pdf, Jul. 18, 2020, 8… [cited by applicant]
U.S. Appl. No. 17/524,157, “Notice of Allowance”, mailed Feb. 28, 2024, 24 pages. [cited by applicant]