IP Library › Granted Patent US 12,530,916
Granted Patent B2
US 12,530,916 · App. 17/589,370 · Granted Jan 20, 2026

Multimodal multitask machine learning system for document intelligence tasks

Inventors: Tharathorn Rimchala (San Francisco, CA); Peter Lee Frick (San Francisco, CA)
Assignee: Intuit Inc.
G06V30/413G06N3/084G06V30/10G06V30/412
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,916
App. No.
17/589,370
Granted
Jan 20, 2026
Kind
B2
Abstract

Multimodal multitask machine learning system for document intelligence tasks includes a feature extractor processing token values obtained from a document to obtain features, and a token extraction head classifying, using the features, the token values to obtain classified tokens. The classified tokens are aggregated into entities. A document classification model is executed on the features to classify the document and obtain a document label prediction. Further a confidence head model applying the document label prediction processes the entities to obtain a result.

Claims (46)

1 . A method comprising:

processing, by a feature extractor of an arrangement of machine-learning models, a plurality of token values obtained from a document, to obtain a plurality of features;

classifying, using the plurality of features, the plurality of token values by a token extraction head of the arrangement of machine-learning models, to obtain a plurality of classified tokens;

aggregating the plurality of classified tokens into a plurality of entities;

generating a document label prediction based on executing a document classification model on the plurality of features;

processing, by a confidence head of the arrangement of machine-learning models, based on applying the document label prediction, the plurality of entities to obtain a task consistent confidence loss wherein the task consistent confidence loss is a function of parameters comprising a binary cross entropy loss, a document label probability and an indicator function;

jointly optimizing, during training, the arrangement of machine-learning models comprising the feature extractor, the confidence head, and the token extraction head by processing the task consistent confidence loss through the feature extractor, the confidence head, and the token extraction head; and

processing a plurality of documents by the trained arrangement of machine-learning models to extract content comprising entity identifier, entity value pairs from the plurality of documents.

2 . The method of claim 1 , further comprising:

calculating a cross entropy loss by comparing a training document label with the document label prediction; and

backpropagating the cross entropy loss through the document classification model.

3 . The method of claim 1 , further comprising:

processing a document image through an optical character recognition (OCR) engine to obtain the plurality of token values, wherein the document comprises the document image.

4 . The method of claim 1 , further comprising:

obtaining the plurality of token values and a layout of the document; and

processing, by the feature extractor, the layout of the document, and the plurality of token values to obtain a plurality of multimodal features,

wherein the plurality of features comprises the plurality of multimodal features.

5 . The method of claim 1 , further comprising:

processing a document image through an image embedding model to obtain an image feature vector for the document image, wherein the document is the document image; and

processing, by the feature extractor, the plurality of token values, a layout of the document image, and the image feature vector to obtain a plurality of multimodal features,

wherein the plurality of features comprises the plurality of multimodal features.

6 . A system comprising:

an arrangement of machine-learning models, comprising:

a feature extractor executing on a computer processor for processing a plurality of token values obtained from a document to obtain a plurality of features,

a token extraction head executing on the computer processor for classifying, using the plurality of features, the plurality of token values to obtain a plurality of classified tokens,

a token aggregator executing on the computer processor for aggregating the plurality of classified tokens into a plurality of entities,

a document classification model executing on the computer processor for classifying, using the plurality of features, the document to obtain a document label prediction,

a confidence head model executing on the computer processor for processing, by applying the document label prediction, the plurality of entities to obtain a task consistent confidence loss, and

a task consistent loss function executing on the computer processor to calculate the task consistent confidence loss using parameters comprising a binary cross entropy loss, a document label probability, and an indicator function;

wherein the system is configured for:

jointly optimizing, during training, the arrangement of machine-learning models by processing the task consistent confidence loss through the feature extractor, the confidence head model, and the token extraction head; and

processing a plurality of documents by the trained arrangement of machine-learning models to extract content comprising entity identifier, entity value pairs from the plurality of documents.

7 . The system of claim 6 , further comprising:

a document classification loss function configured to

calculate a cross entropy loss by comparing a training document label with the document label prediction,

wherein the cross entropy loss is backpropagated through the document classification model.

8 . The system of claim 6 , further comprising:

an optical character recognition (OCR) engine for obtaining the plurality of token values from a document image, the document comprising the document image.

9 . The system of claim 6 , further comprising:

an optical character recognition (OCR) engine for obtaining the plurality of token values and a layout from a document image, the document comprising the document image,

wherein the feature extractor processes a layout of the document image, and the plurality of token values to obtain a plurality of multimodal features,

wherein the plurality of features comprises the plurality of multimodal features.

10 . The system of claim 6 , further comprising:

an image embedding model for processing a document image to obtain an image feature vector for the document image, wherein the document comprises the document image,

wherein the feature extractor processes the image feature vector, a layout of the document image, and the plurality of token values to obtain a plurality of multimodal features,

wherein the plurality of features comprises the plurality of multimodal features.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2022
From: RIMCHALA, THARATHORN; FRICK, PETER LEE
To: INTUIT INC.
Reel/Frame 060155/0550 →
Continuity (1)
Related Publication 20230245485A1 · Aug 3, 2023
References Cited (22)
US 11275934B2 · Reisswig · 2022 [cited by examiner]
US 20060047690A1 · Humphreys · 2006 [cited by examiner]
US 20110182500A1 · Esposito · 2011 [cited by examiner]
US 20200097768A1 · Freese · 2020 [cited by examiner]
US 20210073533A1 · Ast · 2021 [cited by examiner]
US 20210149993A1 · Torres · 2021 [cited by applicant]
US 20210303939A1 · Hu · 2021 [cited by examiner]
US 20220366168A1 · Kaur · 2022 [cited by examiner]
EP 3920044A1 · 2021 [cited by examiner]
EP 3923185A2 · 2021 [cited by examiner]
Pramanik et al., Towards a multi-modal, multi-task learning based pre-training framework for document representation learning, arXiv preprint arXiv:2009.14457 (Year: 2020). [cited by examiner]
Zheng, Xuanang, et al. “Span-based Joint Extracting Subjects and Objects and Classifying Relations with Multi-head Self-attention.” 2021 IEEE 9th International Conference on Information, Communication and Networks (ICIC… [cited by examiner]
Xu, Yiheng, et al. “Layoutlm: Pre-training of text and layout for document image understanding.” Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining. (Year: 2020). [cited by examiner]
Xu, Yang, et al. “Layoutlmv2: Multi-modal pre-training for visually-rich document understanding.” arXiv preprint arXiv:2012.14740 (Year: 2020). [cited by examiner]
Pramanik, S., et al., “Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning”, Sep. 30, 2020, 8 pages. [cited by applicant]
Xu, Y., et al, “LayoutLMV2: Mutli-Modal Pre-Training for Visually-Rich Document Understanding”, May 11, 2021, 16 pages. [cited by applicant]
Powalski, R., et al., “Going Full-Tilt Boogie on Document Understanding with Text-Image-Layout Transformer”, Jul. 12, 2021, 18 pages. [cited by applicant]
Appalaraju, S., et al., “DocFormer: End-to-End Transformer for Document Understanding”, Sep. 20, 2021, 22 pages. [cited by applicant]
Liu, X., et al., “Multi-Task Deep Neural Networks for Natural Language Understanding”, May 30, 2019, 10 pages. [cited by applicant]
Sanh, V., et al., “DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter”, Mar. 1, 2020, 5 pages. [cited by applicant]
Touvron, H., et al., “Training Data-Efficient Image Transformers and Distillation through Attention”, Jan. 15, 2021, 22 pages. [cited by applicant]
Torres, T., et al., “Document Understanding: Information Extraction from Financial Documents and Forms”, Mar. 25, 2020, 29 pages. [cited by applicant]
Cited By (1)
US 12,585,709