IP Library › Granted Patent US 12,333,844
Granted Patent B2
US 12,333,844 · App. 18/055,752 · Granted Jun 17, 2025

Extracting document hierarchy using a multimodal, layer-wise link prediction neural network

Inventors: Vlad Morariu (Potomac, MD); Puneet Mathur (College Park, MD); Rajiv Jain (Vienna, VA); Ashutosh Mehra (Noida, IN); Jiuxiang Gu (College Park, MD); Franck Dernoncourt (Sunnyvale, CA); Anandhavelu N (Kangayam, IN); Quan Tran (San Jose, CA); Verena Kaynig-Fittkau (Cambridge, MA); Nedim Lipka (Santa Clara, CA); Ani Nenkova (Philadelphia, PA)
Assignee: Adobe Inc.
G06V30/413G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,844
App. No.
18/055,752
Granted
Jun 17, 2025
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media that generate a digital document hierarchy comprising layers of parent-child element relationships from the visual elements. For example, for a layer of the layers, the disclosed systems determine, from the visual elements, candidate parent visual elements and child visual elements. In addition, for the layer of the layers, the disclosed systems generate, from the feature embeddings utilizing a neural network, element classifications for the candidate parent visual elements and parent-child element link probabilities for the candidate parent visual elements and the child visual elements. Moreover, for the layer, the disclosed systems select parent visual elements from the candidate parent visual elements based on the parent-child element link probabilities. Further, the disclosed systems utilize the digital document hierarchy to generate an interactive digital document from the digital document image.

Claims (72)

1. A method comprising:

generating feature embeddings from visual elements of a digital document image;

generating a digital document hierarchy comprising layers of parent-child element relationships from the visual elements by, for a layer of the layers:

determining, from the visual elements, child visual elements and candidate parent visual elements for the child visual elements;

generating, from the feature embeddings utilizing a neural network, element classifications for the candidate parent visual elements and parent-child element link probabilities for the candidate parent visual elements and the child visual elements; and

selecting, from the candidate parent visual elements and based on the parent-child element link probabilities, parent visual elements for the child visual elements; and

utilizing the digital document hierarchy to generate an interactive digital document from the digital document image.

2. The method of claim 1 , wherein generating the feature embeddings from the visual elements of the digital document image comprises:

generating spatial feature embeddings from positional coordinates of the visual elements within the digital document image; and

generating the element classifications for the candidate parent visual elements and the parent-child element link probabilities utilizing the spatial feature embeddings.

3. The method of claim 1 , wherein generating the feature embeddings from the visual elements of the digital document image comprises:

extracting text from a visual text element of the visual elements portrayed in the digital document image;

generating, utilizing a trained language machine learning model, semantic feature embeddings from the text of the visual text element; and

generating the element classifications for the candidate parent visual elements and the parent-child element link probabilities utilizing the semantic feature embeddings.

4. The method of claim 1 , wherein generating the digital document hierarchy comprising the layers of parent-child element relationships from the visual elements comprises, for an additional layer of the layers:

determining, from the parent visual elements of the layer, an additional set of candidate parent visual elements and an additional set of child visual elements for the additional layer;

generating, utilizing the neural network for the additional layer, additional parent-child element link probabilities for the additional set of candidate parent visual elements and the additional set of child visual elements; and

selecting additional parent elements for the additional layer from the additional set of candidate parent visual elements based on the additional parent-child element link probabilities.

5. The method of claim 4 , wherein generating the feature embeddings comprises:

generating structural embeddings from the element classifications for the layer; and

generating the additional parent-child element link probabilities for the additional layer utilizing the structural embeddings.

6. The method of claim 1 , further comprising generating, utilizing the neural network, the parent-child element link probabilities by:

generating, utilizing a multimodal encoder of the neural network, parent vector representations from feature embeddings of the candidate parent visual elements;

generating, utilizing the multimodal encoder of the neural network, child vector representations from feature embeddings of the child visual elements;

utilizing a first neural network layer of the neural network to generate the parent-child element link probabilities from the parent vector representations and the child vector representations; and

utilizing a second neural network layer of the neural network to generate the element classifications for the candidate parent visual elements from the parent vector representations.

7. The method of claim 1 , further comprising selecting the parent visual elements from the candidate parent visual elements based on the parent-child element link probabilities by:

determining cost metrics from the parent-child element link probabilities, parent visual element sizes for the candidate parent visual elements, and child visual element sizes for the child visual elements; and

selecting the parent visual elements from the candidate parent visual elements by comparing the cost metrics.

8. The method of claim 1 , further comprising:

determining, from the visual elements, the candidate parent visual elements by applying spatial constraints to identify particular visual elements that encompass one or more other visual elements; and

utilizing the digital document hierarchy to determine a text order for the digital document image.

9. A system comprising:

at memory component; and

one or more processing devices coupled to the memory component, the one or more processing devices to perform operations comprising:

receiving a parent-child element training dataset comprising a digital document image comprising a plurality of visual elements, ground truth parent-child element relationships and ground truth element classifications; and

training a neural network utilizing the parent-child element training dataset to generate, for candidate parent visual elements predicted from the plurality of visual elements, element classifications and parent-child element links for digital document images.

10. The system of claim 9 , wherein training the neural network comprises:

generating, for the candidate parent visual elements from the plurality of visual elements using the neural network, predicted element classifications and predicted parent-child element links; and

modifying parameters of the neural network by comparing the predicted element classifications with the ground truth element classifications and comparing the parent-child element links with the ground truth parent-child element relationships.

11. The system of claim 10 , wherein generating the predicted parent-child element links comprises:

generating, utilizing a multimodal encoder of the neural network, parent vector representations from the candidate parent visual elements; and

generating, utilizing the multimodal encoder of the neural network, child vector representations from child visual elements of the plurality of visual elements.

12. The system of claim 11 , wherein generating the predicted parent-child element links comprises:

combining the parent vector representations with the child vector representations to generate combined parent-child vector representations; and

utilizing a first neural network layer of the neural network to generate the predicted parent-child element links from the combined parent-child vector representations.

13. The system of claim 11 , wherein generating the element classifications comprises utilizing a second neural network layer of the neural network to generate the element classifications from the parent vector representations.

14. The system of claim 10 , wherein:

comparing the predicted element classifications with the ground truth element classifications comprises determining an element classification loss utilizing a first loss function,

comparing the predicted parent-child element links with the ground truth parent-child element relationships comprises determining a parent-child link loss utilizing a second loss function, and

modifying the parameters of the neural network comprises utilizing back propagation to modify the parameters based on the element classification loss and the parent-child link loss.

15. A non-transitory computer-readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

generating feature embeddings from visual elements of a digital document image;

generating a digital document hierarchy comprising layers of parent-child element relationships from the visual elements by, for a layer of the layers:

determining, from the visual elements, child visual elements and candidate parent visual elements for the child visual elements;

generating, from the feature embeddings utilizing a neural network, element classifications for the candidate parent visual elements and parent-child element link probabilities for the candidate parent visual elements and the child visual elements; and

selecting, from the candidate parent visual elements and based on the parent-child element link probabilities, parent visual elements for the child visual elements; and

utilize the digital document hierarchy to determine a text order for the digital document image.

16. The non-transitory computer-readable medium of claim 15 , wherein generating the feature embeddings from the visual elements of the digital document image comprises:

generating spatial feature embeddings from positional coordinates of the visual elements within the digital document image;

generating, utilizing a trained language machine learning model, semantic feature embeddings from text of the visual elements; and

generating the element classifications for the candidate parent visual elements and the parent-child element link probabilities utilizing the spatial feature embeddings and the semantic feature embeddings.

17. The non-transitory computer-readable medium of claim 15 , further comprising utilizing the digital document hierarchy to generate an interactive digital document from the digital document image.

18. The non-transitory computer-readable medium of claim 15 , further comprising:

determining, from the parent visual elements of the layer, an additional set of candidate parent visual elements and an additional set of child visual elements for an additional layer;

generating, utilizing the neural network for the additional layer, additional element classifications and additional parent-child element link probabilities; and

selecting additional parent elements for the additional layer from the additional set of candidate parent visual elements based on the additional parent-child element link probabilities.

19. The non-transitory computer-readable medium of claim 15 , wherein generating, utilizing the neural network, the parent-child element link probabilities comprises:

generating, utilizing a multimodal encoder of the neural network, parent vector representations from feature embeddings of the candidate parent visual elements;

generating, utilizing the multimodal encoder of the neural network, child vector representations from feature embeddings of the child visual elements; and

generating the parent-child element link probabilities utilizing the parent vector representations and the child vector representations.

20. The non-transitory computer-readable medium of claim 15 , wherein selecting the parent visual elements from the candidate parent visual elements based on the parent-child element link probabilities comprises utilizing an optimization model to select the parent visual elements from the candidate parent visual elements based on the parent-child element link probabilities, parent visual element sizes for the candidate parent visual elements, child visual element sizes for the child visual elements, and boundaries of the visual elements.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2022
From: MORARIU, VLAD; MATHUR, PUNEET; JAIN, RAJIV; MEHRA, ASHUTOSH; GU, JIUXIANG; DERNONCOURT, FRANCK; N, ANANDHAVELU; TRAN, QUAN; KAYNIG-FITTKAU, VERENA; LIPKA, NEDIM; NENKOVA, ANI
To: ADOBE INC.
Reel/Frame 061782/0356 →
Continuity (1)
Related Publication 20240161529A1 · May 16, 2024
References Cited (56)
US 20230079343A1 · Roy · 2023 [cited by examiner]
US 20230351115A1 · Zeng · 2023 [cited by examiner]
Carbonell et al., “Named Entity Recognition and Relation Extraction with Graph Neural Networks in Semi Structured Documents”, 2021, by Manual Carbonell, Pau Riba, Mauricio Villegas, Alicia Fornes, and Josep Llados, in 2… [cited by examiner]
Milan Aggarwal, Hiresh Gupta, Mausoom Sarkar, and Balaji Krishnamurthy. 2020. Form2Seq : A Framework for Higher-Order Form Structure Extraction. In Proceedings of the 2020 Conference on Empirical Methods in Natural Lang… [cited by applicant]
Milan Aggarwal, Mausoom Sarkar, Hiresh Gupta, and Balaji Krishnamurthy. 2020. Multi-Modal Association based Grouping for Form Structure Extraction. 2020 IEEE Winter Conference on Applications of Computer Vision (WACV) (… [cited by applicant]
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R Manmatha. 2021. DocFormer: End-to-End Transformer for Document Understanding. arXiv preprint arXiv:2106.11539 (2021). [cited by applicant]
Manuel Carbonell, Pau Riba, Mauricio Villegas, Alicia Fornés, and Josep Lladós. 2021. Named Entity Recognition and Relation Extraction with Graph Neural Networks in Semi Structured Documents. In 2020 25th International … [cited by applicant]
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. In European conference on computer vision. Springer, 213… [cited by applicant]
Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen, Xiwei Xu, Liming Zhu, and Guoqiang Li. 2020. Object detection for graphical user interface: old fashioned or deep learning or a combination? Proceedings of the 28… [cited by applicant]
Christian Clausner, Stefan Pletschacher, and Apostolos Antonacopoulos. 2013. The significance of reading order in document recognition and its evaluation. In 2013 12th International Conference on Document Analysis and R… [cited by applicant]
Tuan Anh Nguyen Dang, Duc Thanh Hoang, Quang Bach Tran, Chih-Wei Pan, and Thanh Dat Nguyen. 2021. End-to-End Hierarchical Relation Extraction for Generic Form Understanding. In 2020 25th International Conference on Patt… [cited by applicant]
Brian Davis, Bryan Morse, Brian Price, Chris Tensmeyer, and Curtis Wiginton. 2021. Visual FUDGE: Form Understanding via Dynamic Graph Editing. arXiv preprint arXiv:2105.08194 (2021). [cited by applicant]
Biplab Deka, Zifeng Huang, Chad Franzen, Joshua Hibschman, Daniel Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar. 2017. Rico: A Mobile App Dataset for Building Data-Driven Design Applications. In Proceedings of t… [cited by applicant]
Dan Deng, Haifeng Liu, Xuelong Li, and Deng Cai. 2018. PixelLink: Detecting Scene Text via Instance Segmentation. ArXiv abs/1801.01315 (2018). [cited by applicant]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL. [cited by applicant]
Lukasz Garncarek, Rafal Powalski, Tomasz Stanislawek, Bartosz Topolski, Piotr Halama, Michal P. Turski, and Filip Grali'nski. 2021. LAMBERT: Layout-Aware Language Modeling for Information Extraction. In ICDAR. [cited by applicant]
Aditya Gupta, Anuj Kumar, Mayank, Vishwa Nath Tripathi, and Sashikala Tapaswi. 2007. Mobile web: web manipulation for small displays using multi-level hierarchy page segmentation. In Mobility '07. [cited by applicant]
Jaekyu Ha, Robert M. Haralick, and Ihsin T. Phillips. 1995. Document page decomposition by the bounding-box project. Proceedings of 3rd International Conference on Document Analysis and Recognition 2 (1995), 1119-1122 v… [cited by applicant]
Leipeng Hao, Liangcai Gao, Xiaohan Yi, and Zhi Tang. 2016. A Table Detection Method for PDF Documents Based on Convolutional Neural Networks. 2016 12th IAPR Workshop on Document Analysis Systems (DAS) (2016), 287-292. [cited by applicant]
Adam W Harley, Alex Ufkes, and Konstantinos G Derpanis. 2015. Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval. In International Conference on Document Analysis and Recognition (ICDA… [cited by applicant]
Dafang He, Scott D. Cohen, Brian L. Price, Daniel Kifer, and C. Lee Giles. 2017. Multi-Scale Multi-Task FCN for Semantic Page Segmentation and Table Detection. 2017 14th IAPR International Conference on Document Analysi… [cited by applicant]
Teakgyu Hong, DongHyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park. 2020. Bros: A Pre-trained Language Model for Understanding Texts in Document. (2020). [cited by applicant]
Teakgyu Hong, Donghyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park. 2021. BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents. arXiv prepri… [cited by applicant]
Wonseok Hwang, Jinyeong Yim, Seunghyun Park, Sohee Yang, and Minjoon Seo. 2020. Spatial Dependency Parsing for Semi-Structured Document Information Extraction. arXiv preprint arXiv:2005.00642 (2020). [cited by applicant]
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019. Funsd: A dataset for form understanding in noisy scanned documents. In 2019 International Conference on Document Analysis and Recognition Workshops (I… [cited by applicant]
Mohamed Khemakhem, Axel Herold, and Laurent Romary. 2018. Enhancing Usability for Automatically Structuring Digitised Dictionaries. [cited by applicant]
Franck Lebourgeois, Zbigniew Bublinski, and Hubert Emptoz. 1992. A fast and efficient method for extracting text paragraphs and graphics from unconstrained documents. Proceedings., 11th IAPR International Conference on … [cited by applicant]
Chen-Yu Lee, Chun-Liang Li, Chu Wang, Renshen Wang, Yasuhisa Fujii, Siyang Qin, Ashok Popat, and Tomas Pfister. 2021. ROPE: Reading Order Equivariant Positional Encoding for Graph-based Document Information Extraction. … [cited by applicant]
Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020. On the Sentence Embeddings from Pre-trained Language Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language P… [cited by applicant]
Kai Li, Curtis Wigington, Chris Tensmeyer, Handong Zhao, Nikolaos Barmpalios, Vlad I Morariu, Varun Manjunatha, Tong Sun, and Yun Fu. 2020. Cross-domain document object detection: Benchmark suite and method. In Proceedi… [cited by applicant]
Minghao Li, Yiheng Xu, Lei Cui, Shaohan Huang, Furu Wei, Zhoujun Li, and Ming Zhou. 2020. DocBank: A Benchmark Dataset for Document Layout Analysis. In Proceedings of the 28th International Conference on Computational L… [cited by applicant]
Yulin Li, Yuxi Qian, Yuechen Yu, Xiameng Qin, Chengquan Zhang, Yan Liu, Kun Yao, Junyu Han, Jingtuo Liu, and Errui Ding. 2021. StrucTexT: Structured Text Understanding with Multi-Modal Transformers. In Proceedings of th… [cited by applicant]
Minghui Liao, Baoguang Shi, Xiang Bai, Xinggang Wang, and Wenyu Liu. 2017. TextBoxes: A Fast Text Detector with a Single Deep Neural Network. In AAAl. [cited by applicant]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Con… [cited by applicant]
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems. 3111-3119. [cited by applicant]
Tom Murray. 1999. Authoring Intelligent Tutoring Systems: An analysis of the state of the art. [cited by applicant]
Lawrence O'Gorman. 1993. The document spectrum for page layout analysis. IEEE Transactions on pattern analysis and machine intelligence 15, 11 (1993), 1162-1173. [cited by applicant]
Xiaobao Peng, Zhijian Yin, and Zhen Yang. 2020. Deeplab_v3_plus-net for Image Semantic Segmentation with Channel Compression. 2020 IEEE 20th International Conference on Communication Technology (ICCT) (2020), 1320-1324. [cited by applicant]
Rafal Powalski, Łukasz Borchmann, Dawid Jurkiewicz, Tomasz Dwojak, Michal Pietruszka, and Gabriela Pałka. 2021. Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer. In ICDAR. [cited by applicant]
W. M. Rand. 1971. Objective Criteria for the Evaluation of Clustering Methods. J. Amer. Statist. Assoc. 66 (1971), 846-850. [cited by applicant]
Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084 (2019). [cited by applicant]
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015. Faster RCNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 39 (2015), 11… [cited by applicant]
Anikó Simon, Jean-Christophe Pret, and A. Peter Johnson. 1997. A Fast Algorithm for Bottom-Up Document Layout Analysis. IEEE Trans. Pattern Anal. Mach. Intell. 19 (1997), 273-277. [cited by applicant]
Ray Smith. 2007. An overview of the Tesseract OCR engine. In Ninth international conference on document analysis and recognition (ICDAR 2007), vol. 2. IEEE, 629-633. [cited by applicant]
Robert Endre Tarjan and Anthony E Trojanowski. 1977. Finding a maximum independent set. SIAM J. Comput. 6, 3 (1977), 537-546. [cited by applicant]
Zilong Wang, Yiheng Xu, Lei Cui, Jingbo Shang, and Furu Wei. 2021. LayoutReader: Pre-training of Text and Layout for Reading Order Detection. In Proceedings of the 2021 Conference on Empirical Methods in Natural Languag… [cited by applicant]
Zilong Wang, Mingjie Zhan, Xuebo Liu, and Ding Liang. 2020. DocStruct: A Multimodal Method to Extract Hierarchy Structure in Document for General Form Understanding. In Findings of the Association for Computational Ling… [cited by applicant]
Zilong Wang, Mingjie Zhan, Houxing Ren, Zhaohui Hou, Yuwei Wu, Xingyan Zhang, and Ding Liang. 2021. GroupLink: An End-to-end Multitask Method for Word Grouping and Relation Extraction in Form Understanding. ArXiv abs/21… [cited by applicant]
Ronald J. Williams and David Zipser. 1989. A Learning Algorithm for Continually Running Fully Recurrent Neural Networks. Neural Comput. 1, 2 (Jun. 1989), 270-280. https://doi.org/10.1162/neco.1989.1.2.270. [cited by applicant]
Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. 2019. Detectron2. https://github.com/facebookresearch/ detectron2. [cited by applicant]
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020. Layoutlm: Pre-training of text and layout for document image understanding. In Proceedings of the 26th ACM SIGKDD International Conference on… [cited by applicant]
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei A. F. Florencio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou. 2021. LayoutLMv2: Multi-modal Pre-training for Visually-Rich Docume… [cited by applicant]
Xiao Yang, Ersin Yumer, Paul Asente, Mike Kraley, Daniel Kifer, and C. Lee Giles. 2017. Learning to Extract Semantic Structure from Documents Using Multimodal Fully Convolutional Neural Networks. 2017 IEEE Conference on… [cited by applicant]
Peng Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu, Liang Qiao, Yi Niu, and Fei Wu. 2020. Trie: End-to-end text reading and information extraction for document understanding. In Proceedings of the 28th ACM Inter… [cited by applicant]
Xiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White, Kyle Murray, Lisa Yu, Qi Shan, Jeffrey Nichols, Jason Wu, Chris Fleizach, et al. 2021. Screen Recognition: Creating Accessibility Metadata for Mobile Applic… [cited by applicant]
Yue Zhang, Bo Zhang, Rui Wang, Junjie Cao, Chen Li, and Zuyi Bao. 2021. Entity Relation Extraction as Dependency Parsing in Visually Rich Documents. arXiv preprint arXiv:2110.09915 (2021). [cited by applicant]