IP Library › Granted Patent US 12,242,809
Granted Patent B2
US 12,242,809 · App. 17/836,977 · Granted Mar 4, 2025

Techniques for pretraining document language models for example-based document classification

Inventors: Guoxin Wang (Bellevue, WA); Dinei Afonso Ferreira Florencio (Redmond, WA); Wenfeng Cheng (Bellevue, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F40/30G06F16/906G06F16/93G06F40/284G06F40/289G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,809
App. No.
17/836,977
Granted
Mar 4, 2025
Kind
B2
Abstract

A data processing system implements a method for training machine learning modes, including receiving a set of one or more unlabeled documents associated one or more first categories of documents to be used to train machine learning models to analyze the one or more unlabeled documents, and fine-tuning a first machine learning model and a second machine learning model based on the one or more unlabeled document to enable the first machine learning model to determine a semantic representation of the one or more first categories of document, and to enable the second machine learning model to classify the semantic representations according to the one or more first categories of documents, the first machine learning model and the second machine learning model having been trained using first unlabeled training data including a second plurality of categories of documents that do not include the one or more first categories of documents.

Claims (67)

1. A data processing system comprising:

a processor; and

a machine-readable medium storing executable instructions that, when executed, cause the processor to perform operations comprising:

receiving a set of one or more unlabeled documents associated with one or more first categories of documents to be used to train machine learning models to analyze the set of one or more unlabeled documents; and

fine-tuning a first machine learning model and a second machine learning model based on the set of one or more unlabeled documents to enable the first machine learning model to determine a semantic representation of the one or more first categories of document, and to enable the second machine learning model to classify semantic representations according to the one or more first categories of documents, the first machine learning model and the second machine learning model having been trained using first unlabeled training data including a second plurality of categories of documents, the second plurality of categories of documents not including the one or more first categories of documents.

2. The data processing system of claim 1 , wherein the machine-readable medium includes instructions configured to cause the processor to perform operations of:

pretraining, using an unsupervised training method using first unlabeled training data, an instance of the first machine learning model for determining a document representation of a document received as an input, the document representation comprising the semantic representation of the document, and an instance of the second machine learning model for determining a category of the document based on the semantic representation output by the first machine learning model, the first unlabeled training data including the second plurality of categories of documents.

3. The data processing system of claim 2 , wherein the second machine learning model is a distance-based classifier trained to predict a nearest class of document based on a distance between the semantic representation of the document and a predicted document class for the document, and wherein the machine-readable medium includes instructions configured to cause the processor to perform operations of:

analyzing documents of the first unlabeled training data to determine key values for each of the second plurality of categories of documents; and

pretraining the second machine learning model using a heuristic process in which pairs of documents from the first unlabeled training data are compared by comparing the key values associated with a first document of a pair of documents of the pairs of documents with key values associated with a second document of the pair to determine a distance between the first document and the second document.

4. The data processing system of claim 1 , wherein the first machine learning model is configured to perform operations of:

receiving a first document as an input;

tokenizing the first document into a plurality of tokens;

segmenting the plurality of tokens into a plurality of chunks comprising sequential subset of the plurality of tokens;

analyzing each of the plurality of chunks using a transformer encoder layer of the first machine learning model to generate a token representation for each of the plurality of chunks; and

combining the token representation for each of the plurality of chunks to generate a document representation for the document comprising a semantic representation of the first document.

5. The data processing system of claim 4 , wherein combining the token representation for each of the plurality of chunks to generate the document representation further comprises:

combining the token representation for each of the plurality of chunks to generate the document representation for the document using a self-attention fusion module of the first machine learning model.

6. The data processing system of claim 1 , wherein the machine-readable medium includes instructions configured to cause the processor to perform operations of:

receiving an input file to the first machine learning model that includes a plurality of documents;

segmenting the input file into a plurality of representations of document pages;

providing the plurality of representations of document pages to a classifier module configured to output classification predictions including a predicted category for each document page;

providing the plurality of representations of document pages to a splitter module configured to output splitting predictions including a prediction whether each document page is a single-page document or a page of a multipage document; and

combining the classification predictions and the splitting predictions to obtain segmentation results that identify each document predicted to be included in the input file, a number of pages associated with each document, and a predicted category for the document.

7. The data processing system of claim 6 , wherein combining the classification predictions and the splitting predictions further comprises:

analyzing the classification predictions and the splitting predictions using a Viterbi decoder.

8. The data processing system of claim 6 , wherein segmenting the input file into a plurality of document pages further comprises:

providing the input file to an instance of a pretrained LayoutLM; and

obtaining the plurality of representations of document pages as an output of the pretrained LayoutLM.

9. A method implemented in a data processing system for classifying a document, the method comprising:

receiving a set of one or more unlabeled documents associated with one or more first categories of documents to be used to train machine learning models to analyze the set of one or more unlabeled documents; and

fine-tuning a first machine learning model and a second machine learning model based on the set of one or more unlabeled documents to enable the first machine learning model to determine a semantic representation of the one or more first categories of document, and to enable the second machine learning model to classify semantic representations according to the one or more first categories of documents, the first machine learning model and the second machine learning model having been trained using first unlabeled training data including a second plurality of categories of documents, the second plurality of categories of documents not including the one or more first categories of documents.

10. The method of claim 9 , further comprising:

pretraining, using an unsupervised training method using first unlabeled training data, an instance of the first machine learning model for determining a document representation of a document received as an input, the document representation comprising the semantic representation of the document, and an instance of the second machine learning model for determining a category of the document based on the semantic representation output by the first machine learning model, the first unlabeled training data including the second plurality of categories of documents.

11. The method of claim 10 , wherein the second machine learning model is a distance-based classifier trained to predict a nearest class of document based on a distance between the semantic representation of the document and a predicted document class for the document, and the method further comprising:

analyzing documents of the first unlabeled training data to determine key values for each of the second plurality of categories of documents; and

pretraining the second machine learning model using a heuristic process in which pairs of documents from the first unlabeled training data are compared by comparing the key values associated with a first document of a pair of documents of the pairs of documents with key values associated with a second document of the pair to determine a distance between the first document and the second document.

12. The method of claim 9 , further comprising performing, with the first machine learning model, operations of;

receiving a first document as an input;

tokenizing the first document into a plurality of tokens;

segmenting the plurality of tokens into a plurality of chunks comprising sequential subset of the plurality of tokens;

analyzing each of the plurality of chunks using a transformer encoder layer of the first machine learning model to generate a token representation for each of the plurality of chunks; and

combining the token representation for each of the plurality of chunks to generate a document representation for the document comprising a semantic representation of the first document.

13. The method of claim 12 , wherein combining the token representation for each of the plurality of chunks to generate the document representation further comprises:

combining the token representation for each of the plurality of chunks to generate the document representation for the document using a self-attention fusion module of the first machine learning model.

14. The method of claim 9 , further comprising:

receiving an input file to the first machine learning model that includes a plurality of documents;

segmenting the input file into a plurality of representations of document pages;

providing the plurality of representations of document pages to a classifier module configured to output classification predictions including a predicted category for each document page;

providing the plurality of representations of document pages to a splitter module configured to output splitting predictions including a prediction whether each document page is a single-page document or a page of a multipage document; and

combining the classification predictions and the splitting predictions to obtain segmentation results that identify each document predicted to be included in the input file, a number of pages associated with each document, and a predicted category for the document.

15. The method of claim 14 , wherein combining the classification predictions and the splitting predictions further comprises:

analyzing the classification predictions and the splitting predictions using a Viterbi decoder.

16. The method of claim 14 , wherein segmenting the input file into a plurality of document pages further comprises:

providing the input file to an instance of a pretrained LayoutLM; and

obtaining the plurality of representations of document pages as an output of the pretrained LayoutLM.

17. A machine-readable medium on which are stored instructions that, when executed, cause a processor of a programmable device to perform operations of:

receiving a set of one or more unlabeled documents associated with one or more first categories of documents to be used to train machine learning models to analyze the set of one or more unlabeled documents; and

fine-tuning a first machine learning model and a second machine learning model based on the set of one or more unlabeled documents to enable the first machine learning model to determine a semantic representation of the one or more first categories of document, and to enable the second machine learning model to classify semantic representations according to the one or more first categories of documents, the first machine learning model and the second machine learning model having been trained using first unlabeled training data including a second plurality of categories of documents, the second plurality of categories of documents not including the one or more first categories of documents.

18. The machine-readable medium of claim 17 , further comprising instructions configured to cause the processor to perform operations of:

pretraining, using an unsupervised training method using first unlabeled training data, an instance of the first machine learning model for determining a document representation of a document received as an input, the document representation comprising the semantic representation of the document, and an instance of the second machine learning model for determining a category of the document based on the semantic representation output by the first machine learning model, the first unlabeled training data including the second plurality of categories of documents.

19. The machine-readable medium of claim 18 , wherein the first machine learning model is configured to perform operations of:

receiving a first document as an input;

tokenizing the first document into a plurality of tokens;

segmenting the plurality of tokens into a plurality of chunks comprising sequential subset of the plurality of tokens;

analyzing each of the plurality of chunks using a transformer encoder layer of the first machine learning model to generate a token representation for each of the plurality of chunks; and

combining the token representation for each of the plurality of chunks to generate the document representation for the document.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE THIRD INVENTOR'S EXECUTION DATE ON THE COVER SHEET PREVIOUSLY RECORDED AT REEL: 060270 FRAME: 0898. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 21, 2023
From: WANG, GUOXIN; FLORENCIO, DINEI AFONSO FERREIRA; CHENG, WENFENG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 063876/0827 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2022
From: WANG, GUOXIN; FLORENCIO, DINEI AFONSO FERREIRA; CHENG, WENFENG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 060270/0898 →
Continuity (1)
Related Publication 20230401386A1 · Dec 14, 2023
References Cited (14)
US 11074412B1 · Leeman-munk et al. · 2021 [cited by applicant]
US 20200279105A1 · Muffat et al. · 2020 [cited by applicant]
US 20210142181A1 · Liu · 2021 [cited by examiner]
US 20210286989A1 · Zhong · 2021 [cited by examiner]
US 20220092101A1 · Yun et al. · 2022 [cited by applicant]
CN 110442684B · 2020 [cited by applicant]
JP 4994199B2 · 2012 [cited by applicant]
JP 5346841B2 · 2013 [cited by applicant]
Pramanik et al., “Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning”, public paper, Jan. 5, 2022. [cited by examiner]
Bahdanau, et al., “Neural Machine Translation by Jointly Learning to Align and Translate”, In Proceedings of 3rd International Conference on Learning Representations, May 7, 2015, 15 Pages. [cited by applicant]
Ferrando, et al., “Improving Accuracy and Speeding up Document Image Classification Through Parallel Systems”, In Proceedings of 20th International Conference on Computational Science, Jun. 3, 2020, pp. 387-400. [cited by applicant]
Khandve, et al., “Hierarchical Neural Network Approaches for Long Document Classification”, In Repository of arXiv:2201.06774v1, Jan. 18, 2022, pp. 1-11. [cited by applicant]
Xu, et al., “LayoutLM: Pre-training of Text and Layout for Document Image Understanding”, In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Aug. 23, 2020, pp. 1192-1200. [cited by applicant]
Xu, et al., “LayoutLMv2: Multi-Modal Pre-training for Visually-Rich Document Understanding”, In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Co… [cited by applicant]