IP Library › Granted Patent US 12,333,845
Granted Patent B2
US 12,333,845 · App. 17/528,061 · Granted Jun 17, 2025

Unified pretraining framework for document understanding

Inventors: Jiuxiang Gu (Greenbelt City, MD); Ani Nenkova Nenkova (Philadelphia, PA); Nikolaos Barmpalios (Palo Alto, CA); Vlad Ion Morariu (Potomac, MD); Tong Sun (San Ramon, CA); Rajiv Bhawanji Jain (Falls Church, VA); Jason wen yong Kuen (Santa Clara, CA); Handong Zhao (San Jose, CA)
Assignee: Adobe Inc.
G06V30/414G06F18/214G06F18/253G06F40/30G06N3/02G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,845
App. No.
17/528,061
Granted
Jun 17, 2025
Kind
B2
Abstract

The technology described includes methods for pretraining a document encoder model based on multimodal self cross-attention. One method includes receiving image data that encodes a set of pretraining documents. A set of sentences is extracted from the image data. A bounding box for each sentence is generated. For each sentence, a set of predicted features is generated by using an encoder machine-learning model. The encoder model performs cross-attention between a set of masked-textual features for the sentence and a set of masked-visual features for the sentence. The set of masked-textual features is based on a masking function and the sentence. The set of masked-visual features is based on the masking function and the corresponding bounding box. A document-encoder model is pretrained based on the set of predicted features for each sentence and pretraining tasks. The pretraining tasks includes masked sentence modeling, visual contrastive learning, or visual-language alignment.

Claims (46)

1. A non-transitory computer-readable storage medium having instructions stored thereon, which, when executed by a processor of a computing device cause the processor to perform actions comprising:

receiving image data that encodes a document;

extracting, from the image data, a set of sentences and a corresponding bounding box for each sentence of the set of sentences;

for each sentence of the set of sentences, generating a set of predicted features using an encoder machine learning (ML) model that performs cross-attention between a set of masked-textual features for the sentence and a set of masked-visual features for the sentence, wherein the set of masked-textual features are based on a masking function and the sentence and the set of masked-visual features are based on the masking function and the corresponding bounding box for the sentence; and

pretraining a document-encoder ML model based on the set of predicted features for each sentence of the set of sentences and one or more pretraining tasks including visual-language alignment to enforce alignment between text and image regions and jointly pretraining, in association with pretraining the document-encoder ML, an image encoder that derives visual features for semantic regions, wherein at least one visual feature comprises a table, a font size, a style, or a figure.

2. The computer-readable storage medium of claim 1 , wherein the actions further comprise:

for each sentence of the set of sentences, generating a textual embedding by using a sentence encoder model and a corresponding visual embedding by using a convolution model and a portion of the image data associated with the corresponding bounding box; and

for each sentence of the set of sentences, generating the set of predicted features further based on the textual embedding for the sentence and the corresponding visual embedding.

3. The computer-readable storage medium of claim 2 , wherein the actions further comprise:

for each sentence of the set of sentences, generating the set of masked-textual features and the set of masked-visual features by using the masking function, the textual embedding for the sentence, and the corresponding visual embedding.

4. The computer-readable storage medium of claim 2 , wherein generating a textual embedding for a sentence of the set of sentences comprises:

generating a sentence embedding for the sentence by using the sentence encoding model and a multiset of tokens included in the sentence;

generating a position embedding for the corresponding bounding box based on a position, within the document, of the corresponding bounding box; and

generating the textual embedding for the sentence by using a combination of the sentence embedding and the position embedding for the bounding box.

5. The computer-readable storage medium of claim 2 , wherein generating a corresponding visual embedding for a sentence of the set of sentences comprises:

generating a position embedding for the corresponding bounding box based on a position, within the document, of the corresponding bounding box;

generating a region-of-interest (ROI) embedding for the corresponding bounding box by using the convolution model and the portion of the image data associated with the corresponding bounding box; and

generating the corresponding visual embedding for the sentence based on a combination of the ROI embedding and the position embedding for the bounding box.

6. The computer-readable storage medium of claim 5 , wherein the actions further comprise:

for each sentence of the set of sentences, generating the set of predicted features based further on the position embedding for the bounding box.

7. The computer-readable storage medium of claim 2 , wherein the actions further comprise:

for each sentence of the set of sentences, generating a corresponding set of visual representations by using a vector quantization method to discretize the corresponding visual embedding; and

for each sentence of the set of sentences, generating the set of masked-visual features by applying the visual mask on the corresponding set of visual representations.

8. The one or more computer-readable storage media of claim 2 , wherein generating the set of masked-textual features and the set of masked-visual features is further based on the masking function stochastically masking the textual embedding for the sentence and the corresponding visual embedding.

9. The one or more computer-readable storage media of claim 1 , wherein the one or more pretraining tasks includes at least one of masked sentence modeling or a visual contrastive learning.

10. A method comprising:

receiving image data that encodes a document;

extracting, from the image data, a set of sentences and a corresponding bounding box for each sentence of the set of sentences;

for each sentence of the set of sentences, generating a textual embedding by using a sentence encoder model and a corresponding visual embedding by using a convolution model and a portion of the image data associated with the corresponding bounding box;

for each sentence of the set of sentences, generating a set of masked-textual features and a set of masked-visual features by using a masking function, the textual embedding for the sentence, and the corresponding visual embedding;

for each sentence of the set of sentences, generating a set of predicted features by using an encoder machine learning (ML) model that performs cross-attention between the set of masked-textual features and the set of masked-visual features for the sentence; and

pretraining a document-encoder ML model based on the set of predicted features for each sentence of the set of sentences and one or more pretraining tasks including masked sentence modeling, visual contrastive learning, and visual-language alignment to enforce alignment between text and image regions, and jointly pretraining, in association with pretraining the document-encoder ML, the convolution model that generates visual embeddings for semantic regions.

11. The method of claim 10 , wherein generating a textual embedding for a sentence of the set of sentences comprises:

generating a sentence embedding for the sentence by using the sentence encoding model and a multiset of tokens included in the sentence;

generating a position embedding for the corresponding bounding box based on a position, within the document, of the corresponding bounding box; and

generating the textual embedding for the sentence based on a combination of the sentence embedding and the position embedding for the bounding box.

12. The method of claim 10 , wherein generating a corresponding visual embedding for a sentence of the set of sentences comprises:

generating a position embedding for the corresponding bounding box based on a position, within the document, of the corresponding bounding box;

generating a region-of-interest (ROI) embedding for the corresponding bounding box by using the convolution model and the portion of the image data associated with the corresponding bounding box; and

generating the corresponding visual embedding for the sentence based on a combination of the ROI embedding and the position embedding for the bounding box.

13. The method of claim 12 , wherein the actions further comprise:

for each sentence of the set of sentences, generating the set of predicted features based further on the position embedding for the bounding box.

14. The method of claim 10 , wherein the actions further comprise:

for each sentence of the set of sentences, generating a corresponding set of visual representations by using a vector quantization method to discretize the corresponding visual embedding; and

for each sentence of the set of sentences, generating the set of masked-visual features by applying the visual mask on the corresponding set of visual representations.

15. The method of claim 10 , wherein generating the set of masked-textual features and the set of masked-visual features is further based on the masking function stochastically masking the textual embedding for the sentence and the corresponding visual embedding.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2021
From: MORARIU, VLAD ION; SUN, TONG; JAIN, RAJIV; BARMPALIOS, NIKOLAOS; GU, JIUXIANG; KUEN, JASON WEN YONG; ZHAO, HANDONG; NENKOVA, ANI NENKOVA
To: ADOBE INC.
Reel/Frame 058130/0573 →
Continuity (1)
Related Publication 20230154221A1 · May 18, 2023
References Cited (53)
US 12182524B2 · Zhang · 2024 [cited by examiner]
US 20210027157A1 · Chen · 2021 [cited by examiner]
US 20220284321A1 · Yuan · 2022 [cited by examiner]
US 20230367977A1 · Nagata · 2023 [cited by examiner]
Wu L, Li S, Hsieh CJ, Sharpnack JL. Stochastic shared embeddings: Data-driven regularization of embedding layers. Advances in Neural Information Processing Systems. 2019;32. (Year: 2019). [cited by examiner]
Stefanini M, Cornia M, Baraldi L, Cascianelli S, Fiameni G, Cucchiara R. From show to tell: A survey on deep learning-based image captioning. IEEE transactions on pattern analysis and machine intelligence. Feb. 7, 2022;… [cited by examiner]
Xu Y, Li M, Cui L, Huang S, Wei F, Zhou M. Layoutlm: Pre-training of text and layout for document image understanding. InProceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining A… [cited by examiner]
Gu J, Nenkova AN, Barmpalios N, Morariu VI, Sun T, Jain RB, Kuen JW, Zhao H, inventors; Adobe Inc, assignee. Unified pretraining framework for document understanding. U.S. Appl. No. 17/528,061, filed May 18, 2023. (Year… [cited by examiner]
Zhang J, Xie Y, Ding W, Wang Z. Cross on cross attention: Deep fusion transformer for image captioning. IEEE Transactions on Circuits and Systems for Video Technology. Feb. 9, 2023;33(8):4257-68. (Year: 2023). [cited by examiner]
Han K, Wang Y, Chen H, Chen X, Guo J, Liu Z, Tang Y, Xiao A, Xu C, Xu Y, Yang Z. A survey on visual transformer. arXiv preprint arXiv:2012.12556. Dec. 23, 2020. (Year: 2020). [cited by examiner]
Zhang C, Yang Z, He X, Deng L. Multimodal intelligence: Representation learning, information fusion, and applications. IEEE Journal of Selected Topics in Signal Processing. Mar. 2020;14(3):478-93. (Year: 2020). [cited by examiner]
Li P, Gu J, Kuen J, Morariu VI, Zhao H, Jain R, Manjunatha V, Liu H. Selfdoc: Self-supervised document representation learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2021 (p… [cited by examiner]
Khalifa M, Vyas Y, Wang S, Horwood G, Mallya S, Ballesteros M. Contrastive training improves zero-shot classification of semi-structured documents. arXiv preprint arXiv:2210.05613. Oct. 11, 2022. (Year: 2022). [cited by examiner]
Wu, Y., et al., “Detectron2”, Retrieved from Internet URL : https://github.com/facebookresearch/detectron2, accessed on Feb. 25, 2022, pp. 4. [cited by applicant]
“EasyOCR”, Retrieved from Internet URL : https://github.com/JaidedAI/EasyOCR, accessed on Feb. 25, 2022, pp. 9. [cited by applicant]
Ba, J. L., et al., “Layer Normalization”, arXiv:1607.06450v1, pp. 1-14 (Jul. 21, 2016). [cited by applicant]
Baevski, A., et al., “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations”, 34th Conference on Neural Information Processing Systems, pp. 1-12 (2020). [cited by applicant]
Cer, D., et al., “SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Cross-lingual Focused Evaluation”, Proceedings of the 11th International Workshop on Semantic Evaluations (SemEval-2017), pp. 1-14 (201… [cited by applicant]
Dai, Z., et al., “Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context”, arXiv:1901.02860v3, pp. 1-20 (Jun. 2, 2019). [cited by applicant]
Deng, J., et al., “ImageNet: A Large-Scale Hierarchical Image Database”, 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1-8 (2009). [cited by applicant]
Goldstein, J., et al., “Multi-Document Summarization by Sentence Extraction”, In NAACL-ANLP Workshop, pp. 40-48 (2000). [cited by applicant]
Girshick, R., “Fast R-CNN”, CVPR, pp. 1440-1448 (2015). [cited by applicant]
Gu, J., et al., “Recent Advances in Convolutional Neural Networks”, arXiv:1512.07108v6, pp. 1-38 (2017). [cited by applicant]
Gu, J., et al., “Self-Supervised Relationship Probing”, 34th Conference on Neural Information Processing Systems, pp. 1-13 (2020). [cited by applicant]
Harley, A. D., et al., “Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval”, 2015 13th International Conference on Document Analysis and Recognition (ICDAR), pp. 1-5 (2015). [cited by applicant]
He, K., et al., “Mask R-CNN”, In ICCV, pp. 2961-2969 (2017). [cited by applicant]
Hu, J., et al., “Squeeze-and-Excitation Networks”, CVPR, pp. 7132-7141 (2018). [cited by applicant]
Jegou, H., et al., “Product Quantization for Nearest Neighbor Search”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 33, No. 1, pp. 1-14 (2010). [cited by applicant]
Jang, E., et al., “Categorical Reparameterization With Gumbel-Softmax”, Published as a conference paper at ICLR, pp. 1-13 (Aug. 5, 2017). [cited by applicant]
Jaume, G., et al., “Funsd: A Dataset for Form Understanding in Noisy Scanned Documents”, ICDARW, pp. 1-6 (Oct. 29, 2019). [cited by applicant]
Kay, A., “Tesseract: An open-source optical character recognition engine”, Linux Journal, pp. 1-12 (Jul. 2007). [cited by applicant]
Kingma, D. P., and Ba, J. L., “Adam: A Method for Stochastic Optimization”, ICLR, arXiv:1412.6980v1, pp. 1-9 (Dec. 22, 2014). [cited by applicant]
Lewis, D., et al., “Building a Test Collection for Complex Document Information Processing”, In SIGIR, pp. 1-2 (Aug. 6-9, 2006). [cited by applicant]
Lu, J., et al., “ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks”, 33rd Conference on Neural Information Processing Systems, pp. 1-11 (2019). [cited by applicant]
Li, L. H., et al., “Visualbert: A Simple and Performant Baseline for Vision and Language”, arXiv:1908.03557v1, pp. 1-14 (Aug. 9, 2019). [cited by applicant]
Liu, Y., et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach”, arXiv:1907.11692, pp. 1-13 (Jul. 26, 2019). [cited by applicant]
Li, G., et al., “Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training”, arXiv:1908.06066v3, pp. 1-8 (Dec. 2, 2019). [cited by applicant]
Li, P., et al., “SelfDoc: Self-Supervised Document Representation Learning”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5652-5660 (2021). [cited by applicant]
Oord, A. V. D., et al., “Neural Discrete Representation Learning”, 31st Conference on Neural Information Processing Systems, pp. 1-10 (2017). [cited by applicant]
Park, S., et al., “CORD: A Consolidated Receipt Dataset for Post-OCR Parsing”, 33rd Conference on Neural Information Processing Systems, pp. 1-4 (2019). [cited by applicant]
Patil, A. G., et al., “READ: Recursive Autoencoders for Document Layout Generation”, CVPR, pp. 1-10 (2020). [cited by applicant]
Pramanik, S., et al., “Towards a multi-modal, multi-task learning based pre-training framework for document representation learning”, arXiv:2009.14457v1, pp. 1-8 (Sep. 30, 2020). [cited by applicant]
Powalski, R., et al., “Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer”, arXiv:2102.09550v3, pp. 1-18 (Jul. 12, 2021). [cited by applicant]
Reimers, N., and Gurevych, I., “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Con… [cited by applicant]
Su, W., et al., “VL-BERT: Pre-Training of Generic Visual-Linguistic Representations”, arXiv:1908.08530v3, pp. 1-15 (Nov. 22, 2019). [cited by applicant]
Tan, H., and Bansal, M., “LXMERT: Learning Cross-Modality Encoder Representations from Transformers”, In EMNLP, pp. 1-14 (Dec. 3, 2019). [cited by applicant]
Tung, F., and Mori, G., “Similarity-Preserving Knowledge Distillation”, ICCV, pp. 1365-1374 (2019). [cited by applicant]
Williams, A., et al., “A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference”, NAACL, pp. 1-11 (Feb. 19, 2018). [cited by applicant]
Wu, T-L., et al., “LAMPRET: Layout-Aware Multimodal PreTraining for Document Understanding”, arXiv:2104.08405v1, pp. 1-14 (2021). [cited by applicant]
Xu, Y., et al., “LayoutLM: Pre-training of Text and Layout for Document Image Understanding”, KDD '20: Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 1192-1200 (202… [cited by applicant]
Xu, Y., et al., “LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding”, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, pp. 2579-2591 (Aug. 1-6, 2021). [cited by applicant]
Yang, Z., et al., “Hierarchical Attention Networks for Document Classification”, In NAACL, Association for Computational Linguistics, pp. 1480-1489 (2016). [cited by applicant]
Zhong, X., et al., “Publaynet: largest dataset ever for document layout analysis”, arXiv:1908.07836v1, pp. 1-8 (Aug. 16, 2019). [cited by applicant]