IP Library › Granted Patent US 12,361,736
Granted Patent B2
US 12,361,736 · App. 18/149,795 · Granted Jul 15, 2025

Multi-stage machine learning model training for key-value extraction

Inventors: Yazhe Hu (Bellevue, WA); Jeaff Wang (Sammamish, WA); Mengqing Guo (Redmond, WA); Tao Sheng (Bellevue, WA); Jun Qian (Bellevue, WA)
G06V30/19147G06F40/284G06F40/30G06N3/08G06V30/1448
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,736
App. No.
18/149,795
Granted
Jul 15, 2025
Kind
B2
Abstract

Techniques for multi-stage training of a machine learning model to extract key-value pairs from documents are disclosed. A system trains a machine learning model using a set of training data including unlabeled documents of various document categories. The initial stage identifies relationships among tokens, or words, numbers, and punctuation, in documents. The system re-trains the machine learning model using a set of training data which includes a particular category of documents while excluding other categories of documents. The second training stage is a supervised machine learning stage in which the training data is labeled to identify key-value pairs in the documents. In the initial training stage, the system sets parameters of the machine learning model to an initial state. In the second stage, the system modifies the parameters of the machine learning model based on the characteristics of the training data set including the documents of the particular category.

Claims (67)

1. A non-transitory computer readable medium comprising instructions which, when executed by one or more hardware processors, causes performance of operations comprising:

executing a first training stage for training a machine learning model to generate first vectors encoding semantic information, and at least one of positional and visual information, of first textual content within documents at least by:

accessing a first plurality of training documents associated with a plurality of document categories, the first plurality of training documents including first textual information, first positional information, and first visual information;

training the machine learning model using the first plurality of training documents associated with the plurality of document categories to generate the first vectors encoding the semantic information, and the at least one of the positional and the visual information based on the first textual information, and at least one of the first positional information and the first visual information;

wherein the first training stage generates a first set of parameters for application of the machine learning model;

executing a second training stage for customizing the machine learning model to (a) generate second vectors to encode semantic information, and at least one of positional information and visual information of second textual content within documents of a particular document category and (b) extract key-value pairs at least by:

accessing a second plurality of training documents associated with the particular document category, the second plurality of training documents being tagged with key-value pairs and including second textual information, second positional information, and second visual information;

training the machine learning model using the second plurality of training documents associated with the particular document category to identify key-value pairs based at least in part on vector encodings of the second textual content that encode semantic relationships, and at least one of positional and visual relationships, between components of the key-value pairs,

wherein training the mahcine learning model using the second plurality of training documents results in a trained machien learning model, and

wherein the second training stage fine-tunes the first set of parameters to generate a second set of parameters for application of the machine learning model; and

applying the trained machine learning model with the second set of parameters to a first document of the particular document category to extract a first plurality of key-value pairs from the first document.

2. The non-transitory computer readable medium of claim 1 , wherein training the machine learning model using the second plurality of training documents comprises training the machine learning model to identify the key-value pairs based at least in part on analyzing vector encodings of the second textual content to determine each of: (a) semantic relationships, (b) positional relationships, and (c) visual relationships between components of the key-value pairs.

3. The non-transitory computer readable medium of claim 1 , wherein training the machine learning model using the first plurality of training documents includes generating, by the machine learning model for each document of the first plurality of training documents, a vector encoding the semantic relationships, positional relationships, and visual relationships among the respective textual content in each respective document of the first plurality of training documents.

4. The non-transitory computer readable medium of claim 1 , the operations further comprising:

accessing feedback data corresponding to the key-value pairs extracted from the second plurality of training documents using the trained machine learning model; and

updating the machine learning model based on the feedback data.

5. The non-transitory computer readable medium of claim 4 , wherein accessing the feedback data corresponding to the key-value pairs extracted from the second plurality of training documents using the trained machine learning model comprises:

applying a set of post-model rules to the key-value pairs to group tokens identified as separate values associated with a same key of a key-value pair into a single value of the key-value pair.

6. The non-transitory computer readable medium of claim 5 , wherein grouping the tokens identified as separate values associated with a same key of a key-value pair into the single value of the key-value pair comprises expanding a bounding box associated with the single value to encompass a first token identified by the machine learning model as a first value and a second token identified by the machine learning model as a second value.

7. The non-transitory computer readable medium of claim 4 , wherein obtaining the feedback data corresponding to the key-value pairs extracted from the second plurality of training documents using the trained machine learning model comprises:

applying a set of post-model rules to the key-value pairs to separate a single value of a key-value pair identified by the machine learning model into a first value corresponding to a first token and a second value corresponding to a second token.

8. The non-transitory computer readable medium of claim 1 , wherein the first plurality of training documents includes documents of the particular document category.

9. The non-transitory computer readable medium of claim 1 , wherein the particular document category comprises an invoice category.

10. The non-transitory computer readable medium of claim 1 , wherein the first plurality of training documents includes documents of a different category than the particular document category, and

wherein the second plurality of training documents excludes documents of any category different from the particular document category.

11. The non-transitory computer readable medium of claim 1 , wherein the second positional information includes bounding box location data associated with tokens in the second plurality of training documents, and

wherein the second visual information includes at least one of size data, shape data, and color data associated with the tokens in the second plurality of training documents.

12. The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

identifying tokens in each document of the first plurality of training documents; and

for each token identified in the first plurality of training documents: extracting text content, and at least one of (a) bounding box location data and (b) shape and color data associated with the token.

13. The non-transitory computer readable medium of claim 1 , wherein obtaining the second plurality of training documents associated with the particular document category comprises:

obtaining a first subset of training documents comprising historical documents of the particular document category; and

obtaining a second subset of training documents comprising synthetic documents of the particular document category.

14. The non-transitory computer readable medium of claim 13 , wherein obtaining the second subset of training documents comprises applying a set of document-generation rules to generate an image of a document of the particular document category and key-value pair labels associated with tokens in the image of the document.

15. The non-transitory computer readable medium of claim 1 , wherein the first plurality of training documents are unlabeled training documents,

wherein the second plurality of training documents are labeled training documents,

wherein training the machine learning model using the first plurality of training documents comprises applying a first algorithm implementing an unsupervised machine learning operation to the first plurality of training documents, and

wherein training the machine learning model using the second plurality of training documents comprises applying a second algorithm implementing a supervised machine learning operation to the second plurality of training documents.

16. A method comprising:

executing a first training stage for training a machine learning model to generate first vectors encoding semantic information, and at least one of positional and visual information, of first textual content within documents at least by:

accessing a first plurality of training documents associated with a plurality of document categories, the first plurality of training documents including first textual information, first positional information, and first visual information;

training the machine learning model using the first plurality of training documents associated with the plurality of document categories to generate the first vectors encoding the semantic information, and the at least one of the positional and the visual information based on the first textual information, and at least one of the first positional information and the first visual information;

wherein the first training stage generates a first set of parameters for application of the machine learning model;

executing a second training stage for customizing the machine learning model to (a) generate second vectors to encode semantic information, and at least one of positional information and visual information of second textual content within documents of a particular document category and (b) extract key-value pairs at least by:

accessing a second plurality of training documents associated with the particular document category, the second plurality of training documents being tagged with key-value pairs and including second textual information, second positional information, and second visual information;

training the machine learning model using the second plurality of training documents associated with the particular document category to identify key-value pairs based at least in part on analyzing vector encodings of the second textual content to determine semantic relationships, and at least one of positional and visual relationships, between components of the key-value pairs,

wherein training the machine learning model using the second plurality of training documents results in a trained machine learning model, and

wherein the second training stage fine-tunes the first set of parameters to generate a second set of parameters for application of the machine learning model; and

applying the trained machine learning model with the second set of parameters to a first document of the particular document category to extract a first plurality of key-value pairs from the first document.

17. The method of claim 16 , wherein training the machine learning model using the second plurality of training documents comprises training the machine learning model to identify the key-value pairs based at least in part on analyzing vector encodings of the second textual content to determine each of: (a) semantic relationships, (b) positional relationships, and (c) visual relationships between components of the key-value pairs.

18. The method of claim 16 , wherein training the machine learning model using the first plurality of training documents includes generating, by the machine learning model for each document of the first plurality of training documents, a vector encoding the semantic relationships, positional relationships, and visual relationships among the respective textual content in each respective document of the first plurality of training documents.

19. The method of claim 16 , further comprising:

accessing feedback data corresponding to the key-value pairs extracted from the second plurality of training documents using the trained machine learning model; and

updating the machine learning model based on the feedback data.

20. A system comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the system to perform operations comprising:

executing a first training stage for training a machine learning model to generate first vectors encoding semantic information, and at least one of positional and visual information, of first textual content within documents at least by:

accessing a first plurality of training documents associated with a plurality of document categories, the first plurality of training documents including first textual information, first positional information, and first visual information;

training the machine learning model using the first plurality of training documents associated with the plurality of document categories to generate the first vectors encoding the semantic information, and the at least one of the positional and the visual information based on the first textual information, and at least one of the first positional information and the first visual information;

wherein the first training stage generates a first set of parameters for application of the machine learning model;

executing a second training stage for customizing the machine learning model to (a) generate second vectors to encode semantic information, and at least one of positional information and visual information of second textual content within documents of a particular document category and (b) extract key-value pairs at least by:

accessing a second plurality of training documents associated with the particular document category, the second plurality of training documents being tagged with key-value pairs and including second textual information, second positional information, and second visual information;

training the machine learning model using the second plurality of training documents associated with the particular document category to identify key-value pairs based at least in part on vector encodings of the second textual content that encode semantic relationships, and at least one of positional and visual relationships, between components of the key-value pairs,

wherein training the machine learning model using the second plurality of training documents results in a trained machine learning model, and

wherein the second training stage fine-tunes the first set of parameters to generate a second set of parameters for application of the machine learning model; and

applying the trained machine learning model with the second set of parameters to a first document of the particular document category to extract a first plurality of key-value pairs from the first document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2023
From: HU, YAZHE; WANG, ZHENG; GUO, MENGQING; SHENG, TAO; QIAN, JUN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 062362/0232 →
Continuity (1)
Related Publication 20240221407A1 · Jul 4, 2024
References Cited (45)
US 5737442A · Alam · 1998 [cited by applicant]
US 5784487A · Cooperman · 1998 [cited by applicant]
US 6175844B1 · Stolin · 2001 [cited by applicant]
US 6336124B1 · Alam et al. · 2002 [cited by applicant]
US 6405175B1 · Ng · 2002 [cited by applicant]
US 7970213B1 · Ruzon et al. · 2011 [cited by applicant]
US 8249356B1 · Smith · 2012 [cited by applicant]
US 9495347B2 · Stadermann et al. · 2016 [cited by applicant]
US 9613267B2 · Dejean et al. · 2017 [cited by applicant]
US 9720896B1 · Wu et al. · 2017 [cited by applicant]
US 10241992B1 · Middendorf et al. · 2019 [cited by applicant]
US 10706322B1 · Yang et al. · 2020 [cited by applicant]
US 10713524B2 · Duta · 2020 [cited by applicant]
US 10878195B2 · Duta · 2020 [cited by applicant]
US 11288719B2 · Xu et al. · 2022 [cited by applicant]
US 12205395B1 · Lam · 2025 [cited by examiner]
US 20060271847A1 · Meunier · 2006 [cited by applicant]
US 20130318426A1 · Shu et al. · 2013 [cited by applicant]
US 20140064618A1 · Janssen, Jr. · 2014 [cited by applicant]
US 20160104077A1 · Jackson et al. · 2016 [cited by applicant]
US 20170017647A1 · Musuluri · 2017 [cited by examiner]
US 20170220859A1 · Grams · 2017 [cited by applicant]
US 20180373952A1 · Bui et al. · 2018 [cited by applicant]
US 20190005322A1 · Tripathi et al. · 2019 [cited by applicant]
US 20190171704A1 · Buisson et al. · 2019 [cited by applicant]
US 20200050845A1 · Foncubierta et al. · 2020 [cited by applicant]
US 20200117944A1 · Duta · 2020 [cited by applicant]
US 20200311410A1 · Prasad et al. · 2020 [cited by applicant]
US 20210019287A1 · Prasad et al. · 2021 [cited by applicant]
US 20230084845A1 · Lu · 2023 [cited by examiner]
CN 104881488B · 2017 [cited by applicant]
CN 106802884A · 2017 [cited by applicant]
EP 1739574B1 · 2007 [cited by applicant]
Cinnamon Ai, “Key-Value extraction using Key-Value Extraction”, Retrieved from https://cinnamonai.medium.com/key-value-extraction-using-graph-key-value-68718e1a4036, Sep. 30, 2020, pp. 5. [cited by applicant]
Clausner et al., “The Significance of Reading Order in Document Recognition and its Evaluation”, 12th International Conference on Document Analysis and Recognition, 2013, pp. 688-692. [cited by applicant]
Klampfl et al., “A Comparison of Two Unsupervised Table Recognition Methods from Digital Scientific Articles”, D-Lib Magazine, vol. 20, No. 11/12, 2014, 13 pages. [cited by applicant]
Liu et al., “TableSeer: Automatic Table Extraction, Search, and Understanding”, Department of Information Sciences and Technology, 2009, 171 pages. [cited by applicant]
Pan et al., “Document layout analysis and reading order determination for a reading robot”, TENCON 2010—2010 IEEE Region 10 Conference, 2010, 2 pages. [cited by applicant]
Patel et al., “Abstractive Information Extraction from Scanned Invoices (AIESI) using End-to-end Sequential Approach”, Sep. 12, 2020, pp. 6. [cited by applicant]
Perez-Arriaga et al., “TAO: System for Table Detection and Extraction from PDF Documents”, Proceedings of the Twenty-Ninth International Florida Artificial Intelligence Research Society Conference, 2016, pp. 591-596. [cited by applicant]
Qiao et al., “DavarOCR: A Toolbox for OCR and Multi-Modal Document Understanding”, Jul. 14, 2022, pp. 4. [cited by applicant]
Sassisegarane P., “How to extract structured data from invoices”, Retrieved from https://medium.com/nanonets/how-to-extract-structured-data-from-invoices-f7de539eb475, Apr. 30, 2021, pp. 26. [cited by applicant]
Shiraly K., “Automate Your Invoice Processing: Extract Data From Invoices in 6 Easy Steps”, Retrieved from https://www.width.ai/post/how-to-extract-data-from-invoices, Nov. 10, 2021, pp. 21. [cited by applicant]
Thomas M. Breuel, “Layout Analysis based on Text Line Segment Hypotheses”, available online at <https://pdfs.semanticscholar.org/e29b/5846e096fa3e858c8da73eb0dde2bf6f812b.pdf>, 4 pages. [cited by applicant]
Zhang et al., “TRIE: End-to-End Text Reading and Information Extraction for Document Understanding”, Oct. 25, 2021, pp. 10. [cited by applicant]
Cited By (1)
US 12,664,365