IP Library › Granted Patent US 12,619,828
Granted Patent B2
US 12,619,828 · App. 18/563,002 · Granted May 5, 2026

Reading order detection in a document

Inventors: Lei Cui (Beijing, CN); Yiheng Xu (Beijing, CN); Yang Xu (Harbin, CN); Furu Wei (Beijing, CN); Zilong Wang (Shanghai, CN)
Assignee: Microsoft Technology Licensing, LLC
G06F40/30G06V30/414G06F40/106
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,828
App. No.
18/563,002
Granted
May 5, 2026
Kind
B2
Abstract

According to embodiments of the present disclosure, there is provided a solution for reading order detection in a document. In the solution, a computer-implemented method includes: determining a text sequence and layout information presented in a document, the text sequence comprising a plurality of text elements, the layout information indicating a spatial layout of the plurality of text elements in the document; generating a plurality of semantic feature representations corresponding to the plurality of text elements based at least on the text sequence and the layout information; and determining a reading order of the plurality of text elements in the document based on the plurality of semantic feature representations. According to the solution, the introduction of the layout information can better characterize a spatial layout manner of the text elements in a specific document, thereby determining the reading order more effectively and accurately.

Claims (61)

1 . A computer-implemented method comprising:

executing a feature extraction neural network model that has been trained using example documents to determine a text sequence and layout information presented in a document, the text sequence comprising a plurality of text elements, the layout information indicating a spatial layout of the plurality of text elements in the document;

generating a plurality of semantic feature representations corresponding to the plurality of text elements based at least on the text sequence and the layout information; wherein using the trained feature extraction neural network model to generate the plurality of semantic feature representations comprises:

converting the text sequence and the layout information into a first embedding representation and a second embedding representation, respectively;

concatenating the first embedding representation and the second embedding representation, to obtain a concatenated embedding representation; and

applying the concatenated embedding representation into the trained feature extraction neural network model to generate the plurality of semantic feature representations; and

determining a reading order of the plurality of text elements in the document based on the plurality of semantic feature representations provided by the trained feature extraction neural network model.

2 . The method of claim 1 , wherein generating the plurality of semantic feature representations comprises:

for a first text element in the text sequence,

determining an attention weight for the first text element with respect to a second text element in the text sequence based on at least one of the following: a relative spatial positioning of the first text element with respect to the second text element in the document, and a relative ranking position of the first text element with respect to the second text element in the text sequence, the attention weight indicating an importance degree of the second text element to the first text element; and

determining a semantic feature representation of the first text element by weighting an embedding representation of the second text element with the determined attention weight.

3 . The method of claim 1 , wherein generating the plurality of semantic feature representations comprises:

determining an image-format file corresponding to the document;

determining visual information from the image-format file, the visual information indicating visual appearances of the plurality of text elements presented in the document; and

generating the plurality of semantic feature representations further based on the visual information.

4 . A computer-implemented method comprising:

executing a feature extraction neural network model that has been trained using example documents to determine a text sequence, layout information and order labeling information presented in a first sample document, the text sequence comprising a first set of text elements, the layout information indicating a spatial layout of the first set of text elements in the first sample document, the order labeling information indicating a ground-truth reading order of the first set of text elements in the first sample document;

generating, using the feature extraction neural network model, respective semantic feature representations of the first set of text elements based at least on the text sequence and the layout information;

executing an order determination neural network model to determine a predicted reading order of the first set of text elements in the first sample document based on the semantic feature representations; and

training the feature extraction neural network model and order determination neural network model based on a difference between the predicted reading order and the ground-truth reading order.

5 . The method of claim 4 , wherein the first sample document comprises an editable text document, and determining the order labeling information comprises:

determining format information corresponding to the editable text document, the format information at least indicating the ground-truth reading order of the first set of text elements.

6 . The method of claim 4 , wherein determining the layout information comprises:

determining a vector file corresponding to the first sample document; and

determining the layout information of the first set of text elements from the vector file.

7 . The method of claim 6 , wherein a plurality of text elements that occur at different positions in the first sample document and represent a same text are assigned with different indices, and the plurality of text elements are labeled with different colors in the vector file, the color with which each text element is labeled being determined based on the index assigned to the text element; and

wherein determining the layout information of the first set of text elements from the vector file comprises:

assigning, based on the indices and the colors assigned to the plurality of text elements, layout information determined from the vector file to the plurality of text elements extracted from the first sample document.

8 . The method of claim 4 , wherein generating the semantic feature representations comprises:

determining visual information from a first image-format file corresponding to the first sample document, the visual information representing visual appearances of the first set of text elements presented in the first sample document; and

generating, using the feature extraction neural network model, the semantic feature representations further based on the visual information.

9 . The method of claim 4 , further comprising obtaining the pre-trained feature extraction neural network model by:

determining a second image-format file corresponding to a second sample document, the second sample document comprising a second set of text elements;

generating, by masking at least one text element of the second set of text elements in the second image-format file, first masking information to indicate that the at least one text element is masked and other text elements of the second set of text elements are not masked;

determining, using the feature extraction neural network model, respective semantic feature representations of the second set of text elements;

determining second masking information based on the respective semantic feature representations of the second set of text elements, the second masking information indicating whether respective text elements of the second set of text elements are masked; and

pre-training the feature extraction model based on a difference between the first masking information and the second masking information.

10 . The method of claim 8 , wherein the feature extraction neural network model is further configured to generate a first visual feature representation of the first image-format file based on the text sequence, the layout information and the visual information,

wherein the method further comprises obtaining the pre-trained feature extraction neural network model by:

determining a third sample document, a third image-format file, and match labeling information, the match labeling information indicating whether the third image-format file matches with the third sample document;

generating, using the feature extraction neural network model, a second visual feature representation of the third image-format file based on the third sample document and the third image-format file;

determining, based on the second visual feature representation, a match result indicating whether the third image-format file matches with the third sample document; and

pre-training the feature extraction neural network model based on a difference between the match result and the match labeling information.

11 . An electronic device, comprising:

a processor, and

a memory coupled to the processor and having instructions stored thereon, the instructions, when executed by the processor, causing the device to perform acts comprising:

executing a feature extraction neural network model that has been trained using example documents to determine a text sequence and layout information presented in a document, the text sequence comprising a plurality of text elements, the layout information indicating a spatial layout of the plurality of text elements in the document;

generating a plurality of semantic feature representations corresponding to the plurality of text elements based at least on the text sequence and the layout information; wherein using the trained feature extraction neural network model to generate the plurality of semantic feature representations comprises:

for a first text element in the text sequence,

determining an attention weight for the first text element with respect to a second text element in the text sequence based on at least one of the following: a relative spatial positioning of the first text element with respect to the second text element in the document, and a relative ranking position of the first text element with respect to the second text element in the text sequence, the attention weight indicating to the trained feature extraction neural network model an importance degree of the second text element to the first text element; and

determining a semantic feature representation of the first text element by weighting an embedding representation of the second text element with the determined attention weight; and

determining a reading order of the plurality of text elements in the document based on the plurality of semantic feature representations provided by the trained feature extraction neural network model.

12 . An electronic device, comprising:

a processor; and

a memory coupled to the processor and having instructions stored thereon, the instructions, when executed by the processor, causing the device to perform acts comprising:

executing a feature extraction neural network model that has been trained using example documents to determine a text sequence, layout information and order labeling information presented in a first sample document, the text sequence comprising a first set of text elements, the layout information indicating a spatial layout of the first set of text elements in the first sample document, the order labeling information indicating a ground-truth reading order of the first set of text elements in the first sample document;

generating, using the feature extraction neural network model, respective semantic feature representations of the first set of text elements based at least on the text sequence and the layout information;

executing an order determination neural network model to determine a predicted reading order of the first set of text elements in the first sample document based on the semantic feature representations; and

training the feature extraction neural network model and order determination neural network model based on a difference between the predicted reading order and the ground-truth reading order.

13 . The device of claim 12 , wherein the first sample document comprises an editable text document, and determining the order labeling information comprises:

determining format information corresponding to the editable text document, the format information at least indicating the ground-truth reading order of the first set of text elements.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2023
From: CUI, LEI; XU, YIHENG; XU, YANG; WEI, FURU; WANG, ZILONG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065638/0095 →
Priority Claims (1)
CN 202110739466.3 · Jun 30, 2021 · national
Continuity (1)
Related Publication 20240265206A1 · Aug 8, 2024
References Cited (56)
US 8539342B1 · Lewis · 2013 [cited by applicant]
US 8649600B2 · Saund · 2014 [cited by applicant]
US 9014478B2 · Itoko · 2015 [cited by applicant]
US 10796145B2 · Anisimovskiy · 2020 [cited by applicant]
US 20180121393A1 · Masalovitch · 2018 [cited by applicant]
US 20200311185A1 · Agrawal · 2020 [cited by examiner]
US 20200320329A1 · Bui · 2020 [cited by applicant]
US 20210081729A1 · Huang et al. · 2021 [cited by applicant]
US 20210272599A1 · Patterson · 2021 [cited by examiner]
CN 105760507A · 2016 [cited by applicant]
CN 109933780A · 2019 [cited by applicant]
CN 117912045A · 2024 [cited by examiner]
Pan, et al. “Document Layout Analysis and Reading Order Determination for a Reading Robot,” IEEE, 2010. (Year: 2010). [cited by examiner]
“Duplicate Word Finder”, accessed on link https://web.archive.org/web/20201018010405/https:/codepen.io/finnhvman/pen/oPwXRa, Oct. 18, 2020, 1 page. [cited by applicant]
“PDF3: Ensuring correct tab and reading order in PDF documents”, accessed on link https://www.w3.org/TR/WCAG20-TECHS/PDF3.html#:˜:text=The%20reading%20order%20of%20a,PDF%20document's%20content%20tree%20structure, Oct. 7… [cited by applicant]
Aiello, et al., “Bidimensional Relations for Reading Order Detection”, In Publication of University of Groningen, Johann Bernoulli Institute for Mathematics and Computer Science, 2003, 5 pages. [cited by applicant]
Bello, et al., “Attention Augmented Convolutional Networks”, arXiv:1904.09925, Sep. 9, 2020, 13 pages. [cited by applicant]
Ceci, et al., “A Data Mining Approach to Reading Order Detection”, In Proceedings of Ninth International Conference on Document Analysis and Recognition, vol. 2, Sep. 23, 2007, 5 pages. [cited by applicant]
Clausner, et al., “The Significance of Reading Order in Document Recognition and its Evaluation”, In 12th International Conference on Document Analysis and Recognition, Aug. 25, 2013, pp. 688-692. [cited by applicant]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: H… [cited by applicant]
Dong, et al., “Unified Language Model Pre-training for Natural Language Understanding and Generation”, In Proceedings of Annual Conference on Neural Information Processing Systems, Dec. 8, 2019, 13 pages. [cited by applicant]
Ferilli, et al., “Abstract Argumentation for Reading Order Detection”, In Proceedings of the ACM Symposium on Document Engineering, Sep. 16, 2014, pp. 45-48. [cited by applicant]
Gralinski, et al., “Kleister: A novel task for Information Extraction involving Long Documents with Complex Layout”, arXiv:2003.02356, Mar. 4, 2020, 14 pages. [cited by applicant]
International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/030466, Sep. 14, 2022, 13 pages. [cited by applicant]
Jaume, et al., “Funsd: A Dataset for Form Understanding in Noisy Scanned Documents”, In Proceedings of International Conference on Document Analysis and Recognition Workshops, Sep. 22, 2019, 6 pages. [cited by applicant]
Kiram, Ravi., “Find and Remove Repeated Words Using GREP”, accessed on link https://web.archive.org/web/20201127033936/https:/creativepro.com/find-remove-repeated-words-grep/, Nov. 27, 2020, 7 pages. [cited by applicant]
Kruk, et al., “Integrating Text and Image: Determining Multimodal Document Intent in Instagram Posts”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, Apr. 19, 2019, pp. 4622-4632. [cited by applicant]
Li, et al., “An End-to-End OCR Text Re-organization Sequence Learning for Rich-text Detail Image Comprehension”, In Proceedings of 16th European Conference of Computer Vision, Aug. 23, 2020, 16 pages. [cited by applicant]
Li, et al., “DocBank: A Benchmark Dataset for Document Layout Analysis”, In Proceedings of the 28th International Conference on Computational Linguistics, Dec. 8, 2020, pp. 949-960. [cited by applicant]
Li, et al., “TableBank: Table Benchmark for Image-based Table Detection and Recognition”, In Proceedings of the 12th Language Resources and Evaluation Conference, May 11, 2020, pp. 1918-1925. [cited by applicant]
Liu, et al., “Graph Convolution for Multimodal Information Extraction from Visually Rich Documents”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human… [cited by applicant]
Lockard, et al., “ZeroShotCeres: Zero-Shot Relation Extraction from Semi-Structured Webpages”, In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 2020, pp. 8105-8117. [cited by applicant]
Majumder, et al., “Representation Learning for Information Extraction from Form-like Documents”, In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 2020, pp. 6495-6507. [cited by applicant]
Malerba, et al., “Learning to Order: A Relational Approach”, In Proceedings of International Workshop on Mining Complex Data, Springer, Sep. 2007, pp. 209-223. [cited by applicant]
Malerba, et al., “Machine Learning for Reading Order Detection in Document Image Understanding”, In Machine Learning in Document Analysis and Recognition, vol. 90, 2008, pp. 45-69. [cited by applicant]
Malerba, et al., “Machine Learning for Reading Order Detection in Document Image Understanding”, Machine Learning in Document Analysis and Recognition, Dec. 2007, vol. 90, pp. 49-72. [cited by applicant]
Mathew, et al., “DocVQA: A Dataset for VQA on Document Images”, arXiv:2007.00398v2, Jul. 2020, 23 pages. [cited by applicant]
Papineni, et al., “BLEU: a Method for Automatic Evaluation of Machine Translation”, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL), Jul. 6, 2002, pp. 311-318. [cited by applicant]
Park, et al., “CORD: A Consolidated Receipt Dataset for Post-OCR Parsing”, In Proceedings of 33rd Conference on Neural Information Processing Systems, 2019, 4 pages. [cited by applicant]
Pramanik, et al., “Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning”, arXiv:2009.14457, Sep. 30, 2020, 8 pages. [cited by applicant]
Sarkhel, et al., “Deterministic Routing between Layout Abstractions for Multi-Scale classification of Visually Rich Documents”, In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligen… [cited by applicant]
Shaw, et al., “Self-Attention with Relative Position Representations”, In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, v… [cited by applicant]
Siegel, et al., “Extracting Scientific Figures with Distantly Supervised Neural Networks”, In Proceedings of the 18th ACM/IEEE on Joint Conference on Digital Libraries, Jun. 3, 2018, pp. 223-232. [cited by applicant]
Stanislawek, et al., “Kleister: Key Information Extraction Datasets Involving Long Documents with Complex Layouts”, arXiv:2105.05796, May 12, 2021, 16 pages. [cited by applicant]
Wang, et al., “Layout Reader: Pre-training of Text and Layout for Reading Order Detection”, arXiv:2108.11591, Aug. 27, 2021, 10 pages. [cited by applicant]
Wei, et al., “Robust Layout-aware IE for Visually Rich Documents with Pre-trained Language Models”, In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, Jul… [cited by applicant]
Wolf, et al., “Huggingface's Transformers: State-of-the-Art Natural Language Processing”, arXiv:1910.03771, Oct. 16, 2019, 11 pages. [cited by applicant]
Xu, et al., “LayoutLM: Pre-training of Text and Layout for Document Image Understanding”, In Proceedings of the 26thACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Aug. 23, 2020, pp. 1192-1200. [cited by applicant]
Xu, et al., “LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding”, arXiv:2012.14740v1, Dec. 29, 2020, 16 pages. [cited by applicant]
Xu, et al., “Layoutlmv2: Multi-Modal Pre-Training for Visually-Rich Document Understanding”, In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, 2021, pp. 2579-2591. [cited by applicant]
Yang, et al., “Learning to Extract Semantic Structure from Documents Using Multimodal Fully Convolutional Neural Networks”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jul. 21, 2017, pp… [cited by applicant]
Yu, et al., “Pick: Processing Key Information Extraction from Documents Using Improved Graph Learning-Convolutional Networks”, arXiv:2004.07464, Jul. 18, 2020, 8 pages. [cited by applicant]
Zhang, et al., “TRIE: End-to-End Text Reading and Information Extraction for Document Understanding”, arXiv:2005.13118, May 2020, 9pages. [cited by applicant]
Zhao, Herin., “Microsoft ImageBERT | Cross-modal Pretraining with Large-scale Image-Text Data”, accessed on link https://medium.com/syncedreview/microsoft-imagebert-cross-modal-pretraining-with-large-scale-image-text-da… [cited by applicant]
Zhong, et al., “PubLayNet: largest dataset ever for document layout analysis”, In Proceedings of International Conference on Document Analysis and Recognition, Sep. 20, 2019, pp. 1015-1022. [cited by applicant]
First Office Action Received for Chinese Application No. 202110739466.3, mailed on Jan. 13, 2026, 23 pages. (English translation Provided). [cited by applicant]