IP Library › Granted Patent US 12,394,233
Granted Patent B2
US 12,394,233 · App. 18/072,616 · Granted Aug 19, 2025

Entity extraction with encoder decoder machine learning model

Inventors: Tharathorn Rimchala (San Francisco, CA); Peter Frick (San Francisco, CA)
Assignee: Intuit Inc.
G06V30/19167G06V30/413
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,233
App. No.
18/072,616
Granted
Aug 19, 2025
Kind
B2
Abstract

A method includes executing an encoder machine learning model on multiple token values contained in a document to create an encoder hidden state vector. A decoder machine learning model executing on the encoder hidden state vector generates raw text comprising an entity value and an entity label for each of multiple entities. The method further includes generating a structural representation of the entities directly from the raw text and outputting the structural representation of the entities of the document.

Claims (57)

1. A method comprising:

executing an optical character recognition (OCR) engine on a document image of a document to obtain OCR output comprising a plurality of token values of the document image;

inputting the plurality of token values from the document image into an encoder machine learning model;

creating an encoder hidden state vector, the creating comprising executing the encoder machine learning model on the plurality of token values contained in the document;

generating, by a decoder machine learning model executing on the encoder hidden state vector, raw text comprising an entity value and an entity label for each entity of a plurality of entities, wherein the plurality of entities is in the document;

suppressing the plurality of entities in the raw text that are determined to have a less than threshold probability of being in the OCR output;

generating a structural representation of the plurality of entities directly from the raw text; and

outputting the structural representation of the plurality of entities.

2. The method of claim 1 , further comprising:

inputting formatting information generated from executing the OCR engine into the encoder machine learning model.

3. The method of claim 1 , further comprising:

inputting the document image into the encoder machine learning model, wherein the encoder machine learning model processes image features obtained from the document image.

4. The method of claim 1 , wherein the encoder machine learning model is executed directly on the OCR output without additional processing, the OCR output comprising the plurality of token values.

5. The method of claim 1 , further comprising:

comparing the plurality of entities with ground truth information to obtain a comparison result;

generating, by a loss function, a loss using the comparison result; and

backpropagating the loss through the decoder machine learning model and the encoder machine learning model.

6. The method of claim 1 , wherein executing the encoder machine learning model comprises:

executing a plurality of encoder blocks on the plurality of token values.

7. A method comprising:

creating an encoder hidden state vector, the creating comprising executing an encoder machine learning model on a plurality of token values contained in a document;

generating, by a decoder machine learning model executing on the encoder hidden state vector, raw text comprising an entity value and an entity label for each entity of a plurality of entities, wherein the plurality of entities is in the document;

generating a structural representation of the plurality of entities directly from the raw text;

executing a document classification model on the document to obtain a document type;

identifying a set of expected entity labels corresponding to the document type;

filtering, from the structural representation, an additional entity in the raw text generated by the decoder machine learning model based on the additional entity having the entity label failing to be in the set of expected entity labels; and

outputting the structural representation of the plurality of entities.

8. A system comprising:

an encoder machine learning model, executing on at least one computer processor, configured to process a plurality of token values extracted from a document to create an encoder hidden state vector;

a decoder machine learning model, executing on the at least one computer processor, configured to:

generate, by processing the encoder hidden state vector, raw text comprising an entity value and an entity label for each entity of a plurality of entities, wherein the plurality of entities is in the document,

a document classification model configured to classify the document to determine a document type of the document; and

a filter configured to:

identify a set of expected entity labels corresponding to the document type, and

filter, from a structural representation generated from the raw text, an additional entity in the raw text generated by the decoder machine learning model based on the additional entity having the entity label failing to be in the set of expected entity labels,

wherein the system is configured to output the plurality of entities of the document.

9. The system of claim 8 , further comprising:

an optical character recognition (OCR) engine processing a document image to generate a sequence of characters with bounding boxes representing a plurality of words of the document image,

wherein the plurality of words from the document image are used as the plurality of token values that are input into the encoder machine learning model.

10. The system of claim 9 , wherein the OCR engine is further configured to generate layout information that is input into the encoder machine learning model.

11. The system of claim 9 , wherein the encoder machine learning model receives, as input, the document image and is configured to generate the encoder hidden state vector using the document image.

12. The system of claim 8 , wherein the encoder machine learning model is executed directly on optical character recognition (OCR) output without additional processing, the OCR output comprising the plurality of token values.

13. The system of claim 8 , further comprising:

a loss function executing on the at least one computer processor and configured to:

compare the plurality of entities with ground truth information to obtain a comparison result, and

generate a loss using the comparison result,

wherein the loss is back-propagated through the decoder machine learning model and the encoder machine learning model.

14. The system of claim 8 , wherein the encoder machine learning model comprises:

a plurality of encoder blocks to execute on the plurality of token values.

15. A method comprising:

creating an encoder hidden state vector, the creating comprising executing an encoder machine learning model on a document image of a document;

generating, by a decoder machine learning model executing on the encoder hidden state vector, raw text comprising an entity value and an entity label for each entity of a plurality of entities, wherein the plurality of entities is in the document;

generating a structural representation of the plurality of entities directly from the raw text;

executing a document classification model on the document to obtain a document type;

identifying a set of expected entity labels corresponding to the document type;

filtering, from the structural representation, an additional entity in the raw text generated by the decoder machine learning model based on the additional entity having the entity label failing to be in the set of expected entity labels; and

outputting the plurality of entities of the document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2023
From: RIMCHALA, THARATHORN; FRICK, PETER
To: INTUIT INC.
Reel/Frame 062646/0381 →
Continuity (2)
Continuation 17829010 · May 31, 2022
Related Publication 20230386236A1 · Nov 30, 2023
References Cited (8)
US 20090208125A1 · Kajiwara · 2009 [cited by examiner]
US 20210012102A1 · Cristescu · 2021 [cited by examiner]
US 20210240932A1 · Erdemir · 2021 [cited by examiner]
US 20220067590A1 · Georgopoulos · 2022 [cited by examiner]
Raffel, C., et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”, Journal of Machine Learning Research 21 (2020), Jul. 28, 2020 (67 pages). [cited by applicant]
Powalski, R., et al., “Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer”, Jul. 12, 2021 (18 pages). [cited by applicant]
Rothe, S., et al., “Leveraging Pre-Trained Checkpoints for Sequence Generation Tasks”, Apr. 16, 2020 (17 pages). [cited by applicant]
von Platen, P., “Leveraging Pre-Trained Language Model Checkpoints for Encoder-Decoder Models”, Nov. 9, 2020 (49 pages). [cited by applicant]