IP Library › Granted Patent US 12,314,668
Granted Patent B2
US 12,314,668 · App. 18/362,886 · Granted May 27, 2025

Natural language processing text-image-layout transformer

Inventors: Lukasz Konrad Borchmann (Warsaw, PL); Dawid Andrzej Jurkiewicz (Poznan, PL); Tomasz Dwojak (Poznan, PL); Michal Waldemar Pietruszka (Cracow, PL); Gabriela Klaudia Palka (Poznan, PL)
Assignee: Snowflake Inc.
G06F40/295G06F40/106G06F40/30G06N3/08G06T11/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,668
App. No.
18/362,886
Granted
May 27, 2025
Kind
B2
Abstract

Disclosed herein is a system, method, and storage medium for Natural Language Processing (NLP) of real-world documents via a cloud data platform. The system combines three NLP models, including an encoder-decoder model, a spatial model, and a multi-modal model not previously combined. A text-image-layout transfer NLP system receives multi-modal input data and trains the multi-modal input data using the combination of the three NLP models.

Claims (88)

1. A method comprising:

providing access to a machine learning model for iterative training on Natural Language Processing (NLP) of real-world documents, the providing access to the machine learning model comprising:

receiving, at a text-image-layout transformer (TILT) NLP system of a cloud data platform, multi-modal input data comprising text data, layout data, and image data;

executing multiple NLP models on the multi-modal input data, the multiple NLP models comprising:

an encoder-decoder model configured to generate text-based features not present in the text data;

a spatial model configured to implement spatial relationship features in the layout data; and

a multi-modal model configured to add visual context features to process the image data;

receiving additional data associated with the multi-modal input data, the additional data comprising semantic data associated with the text-based features, sequential distance data associated with the spatial relationship features, and spatial relationship data associated with the visual context features;

maintaining a distinction between the semantic data, the sequential distance data, and the spatial relationship data in the real-world documents, the maintaining the distinction comprising separating the semantic data from the sequential distance data and the spatial relationship data;

providing regularization augmentation to each of the text-based features, the spatial relationship features, and the visual context features while enabling cross-modal learning among the encoder-decoder model, the spatial model, and the multi-modal model; and

training the machine learning model on the multi-modal input data, the text-based features, the spatial relationship features, and the visual context features.

2. The method of claim 1 , further comprising:

processing the multi-modal input data, the processing comprising maintaining distinct processing paths for the text-based features extracted from the text data, the spatial relationship features derived from the layout data, and the visual context features extracted from the image data; and

unifying the distinct processing paths through an end-to-end neural architecture to preserve independence of the text-based features, the spatial relationship features, and the visual context features.

3. The method of claim 1 , wherein the training on the NLP of the real-world documents comprises:

analyzing the multi-modal input data; and

receiving at least one question regarding the multi-modal input data.

4. The method of claim 3 , further comprising:

generating output including at least one of answers to the at least one question, key information, and document classification.

5. The method of claim 1 , wherein the TILT NLP system of the cloud data platform performs operations comprising:

receiving the real-world documents;

extending biases with spatial relationships that include relative attention biases; and

providing additional image semantics to the received real-world documents.

6. The method of claim 1 , further comprising:

employing spatial bias augmentation, wherein biases are extended with spatial relationships; and

generating contextualized image embeddings, wherein additional image semantics are provided with the multi-modal input data.

7. The method of claim 6 , further comprising:

embedding distributional and contextualized semantics of a text token into a multi-dimensional vector space;

adding text embeddings to visual features of the multi-modal input data; and

assigning, to the text token, the visual features relative to a position and surrounding in the multi-dimensional vector space.

8. A system comprising:

one or more hardware processors of a machine; and

at least one memory storing instructions that, when executed by the one or more hardware processors, cause the machine to perform operation comprising:

providing access to a machine learning model for iterative training on Natural Language Processing (NLP) of real-world documents, the providing access to the machine learning model comprising:

receiving, at a text-image-layout transformer (TILT) NLP system of a cloud data platform, multi-modal input data comprising text data, layout data, and image data;

executing multiple NLP models on the multi-modal input data, the multiple NLP models comprising:

an encoder-decoder model configured to generate text-based features not present in the text data;

a spatial model configured to implement spatial relationship features in the layout data; and

a multi-modal model configured to add visual context features to process the image data;

receiving additional data associated with the multi-modal input data, the additional data comprising semantic data associated with the text-based features, sequential distance data associated with the spatial relationship features, and spatial relationship data associated with the visual context features;

maintaining a distinction between the semantic data, the sequential distance data, and the spatial relationship data in the real-world documents, the maintaining the distinction comprising separating the semantic data from the sequential distance data and the spatial relationship data;

providing regularization augmentation to each of the text-based features, the spatial relationship features, and the visual context features while enabling cross-modal learning among the encoder-decoder model, the spatial model, and the multi-modal model; and

training the machine learning model on the multi-modal input data, the text-based features, the spatial relationship features, and the visual context features.

9. The system of claim 8 , the operations further comprising:

processing the multi-modal input data, the processing comprising maintaining distinct processing paths for the text-based features extracted from the text data, the spatial relationship features derived from the layout data, and the visual context features extracted from the image data; and

unifying the distinct processing paths through an end-to-end neural architecture to preserve independence of the text-based features, the spatial relationship features, and the visual context features.

10. The system of claim 8 , wherein the training on the NLP of the real-world documents further comprises:

analyzing the multi-modal input data; and

receiving at least one question regarding the multi-modal input data.

11. The system of claim 10 , the operations further comprising:

generating output including at least one of answers to the at least one question, key information, and document classification.

12. The system of claim 8 , wherein the TILT NLP system of the cloud data platform performs operations further comprising:

receiving the real-world documents;

extending biases with spatial relationships that include relative attention biases; and

providing additional image semantics to the received real-world documents.

13. The system of claim 8 , the operations further comprising:

employing spatial bias augmentation, wherein biases are extended with spatial relationships; and

generating contextualized image embeddings, wherein additional image semantics are provided with the multi-modal input data.

14. The system of claim 13 , the operations further comprising:

embedding distributional and contextualized semantics of a text token into a multi-dimensional vector space;

adding text embeddings to visual features of the multi-modal input data; and

assigning, to the text token, the visual features relative to a position and surrounding in the multi-dimensional vector space.

15. A non-transitory computer readable storage medium embodying instructions that, when executed by a machine, cause the computer to perform operations comprising:

providing access to a machine learning model for iterative training on Natural Language Processing (NLP) of real-world documents, the providing access to the machine learning model comprising:

receiving, at a text-image-layout transformer (TILT) NLP system of a cloud data platform, multi-modal input data comprising text data, layout data, and image data;

executing multiple NLP models on the multi-modal input data, the multiple NLP models comprising:

an encoder-decoder model configured to generate text-based features not present in the text data;

a spatial model configured to implement spatial relationship features in the layout data; and

a multi-modal model configured to add visual context features to process the image data;

receiving additional data associated with the multi-modal input data, the additional data comprising semantic data associated with the text-based features, sequential distance data associated with the spatial relationship features, and spatial relationship data associated with the visual context features;

maintaining a distinction between the semantic data, the sequential distance data, and the spatial relationship data in the real-world documents, the maintaining the distinction comprising separating the semantic data from the sequential distance data and the spatial relationship data;

providing regularization augmentation to each of the text-based features, the spatial relationship features, and the visual context features while enabling cross-modal learning among the encoder-decoder model, the spatial model, and the multi-modal model; and

training the machine learning model on the multi-modal input data, the text-based features, the spatial relationship features, and the visual context features.

16. The non-transitory computer readable storage medium of claim 15 , the operations further comprising:

processing the multi-modal input data, the processing comprising maintaining distinct processing paths for the text-based features extracted from the text data, the spatial relationship features derived from the layout data, and the visual context features extracted from the image data; and

unifying the distinct processing paths through an end-to-end neural architecture to preserve independence of the text-based features, the spatial relationship features, and the visual context features.

17. The non-transitory computer readable storage medium of claim 15 , wherein the training on the NLP of the real-world documents comprises:

analyzing the multi-modal input data; and

receiving at least one question regarding the multi-modal input data.

18. The non-transitory computer readable storage medium of claim 17 , the operations further comprising:

generating output including at least one of answers to the at least one question, key information, and document classification.

19. The non-transitory computer readable storage medium of claim 15 , wherein the TILT NLP system of the cloud data platform performs operations further comprising:

receiving the real-world documents;

extending biases with spatial relationships that include relative attention biases; and

providing additional image semantics to the received real-world documents.

20. The non-transitory computer readable storage medium of claim 15 , the operations further comprising:

employing spatial bias augmentation, wherein biases are extended with spatial relationships; and

generating contextualized image embeddings, wherein additional image semantics are provided with the multi-modal input data.

Assignments (2)
CONFIRMATORY ASSIGNMENT Recorded May 19, 2025
From: APPLICA SP. Z O.O.
To: SNOWFLAKE INTERNATIONAL HOLDINGS INC.
Reel/Frame 071296/0386 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2023
From: BORCHMANN, LUKASZ KONRAD; JURKIEWICZ, DAWID ANDRZEJ; DWOJAK, TOMASZ; PIETRUSZKA, MICHAL WALDEMAR; PALKA, GABRIELA KLAUDIA
To: APPLICA SP. Z O.O.
Reel/Frame 065185/0005 →
Continuity (3)
Continuation 17651311 · Feb 16, 2022
Provisional Application 63150271 · Feb 17, 2021
Related Publication 20240028832A1 · Jan 25, 2024
References Cited (140)
US 9953008B2 · Zaric et al. · 2018 [cited by applicant]
US 10636074B1 · Bentley et al. · 2020 [cited by applicant]
US 10956673B1 · Ramezani et al. · 2021 [cited by applicant]
US 10990645B1 · Shi · 2021 [cited by applicant]
US 11455468B2 · Dancewicz et al. · 2022 [cited by applicant]
US 11620451B2 · Dancewicz et al. · 2023 [cited by applicant]
US 11645712B2 · Wellmann et al. · 2023 [cited by applicant]
US 11704090B2 · Li et al. · 2023 [cited by applicant]
US 11763087B2 · Borchmann et al. · 2023 [cited by applicant]
US 11842391B2 · Wellmann et al. · 2023 [cited by applicant]
US 20190294874A1 · Orlov et al. · 2019 [cited by applicant]
US 20200176098A1 · Lucas et al. · 2020 [cited by applicant]
US 20200349178A1 · Raju · 2020 [cited by applicant]
US 20200349415A1 · Raju · 2020 [cited by applicant]
US 20210081613A1 · Begun et al. · 2021 [cited by applicant]
US 20210081729A1 · Huang et al. · 2021 [cited by applicant]
US 20210271707A1 · Lin et al. · 2021 [cited by applicant]
US 20210286989A1 · Zhong et al. · 2021 [cited by applicant]
US 20210342785A1 · Mann et al. · 2021 [cited by applicant]
US 20220036063A1 · Bhuyan et al. · 2022 [cited by applicant]
US 20220076109A1 · Srivastava et al. · 2022 [cited by applicant]
US 20220157341A1 · Adato et al. · 2022 [cited by applicant]
US 20220229983A1 · Zohrevand et al. · 2022 [cited by applicant]
US 20220261547A1 · Dancewicz et al. · 2022 [cited by applicant]
US 20220270311A1 · Borchmann et al. · 2022 [cited by applicant]
US 20220327286A1 · Dancewicz et al. · 2022 [cited by applicant]
US 20220335518A1 · Wellmann et al. · 2022 [cited by applicant]
US 20230259709A1 · Dancewicz et al. · 2023 [cited by applicant]
US 20240211691A1 · Dancewicz et al. · 2024 [cited by applicant]
CN 109840287A · 2019 [cited by applicant]
CN 111767732A · 2020 [cited by applicant]
CN 117043783 · 2023 [cited by applicant]
CN 117083605 · 2023 [cited by applicant]
WO 2022054079 · 2022 [cited by applicant]
WO WO2022175847A1 · 2022 [cited by applicant]
WO WO2022175849A1 · 2022 [cited by applicant]
Xu, Yang, et al. “Layoutlmv2: Multi-modal pre-training for visually-rich document understanding.” arXiv preprint arXiv:2012.14740 (2020). (Year: 2020). [cited by examiner]
Raffel, Colin, et al. (“Exploring the limits of transfer learning with a unified text-to-text transformer.” arXiv preprint arXiv:1910.10683 (2019).) (Year: 2019). [cited by examiner]
Daniel Hewlett, et al. 2016. WikiReading: A Novel Large-scale Language Understanding Task over Wikipedia. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistic (vol. 1: Long Papers), … [cited by examiner]
Hermann, Karl Moritz, et al. “Teaching machines to read and comprehend.” Advances in neural information processing systems 28 (2015). (Year: 2015). [cited by examiner]
Sarkhel, Ritesh, and Arnab Nandi. “Deterministic routing between layout abstractions for multi-scale classification of visually rich documents.” 28th International Joint Conference on Artificial Intelligence (IJCAI), 20… [cited by examiner]
Powalski, Rafał, et al. “Going full-tilt boogie on document understanding with text-image-layout transformer.” Document Analysis and Recognition-ICDAR 2021: 16th International Conference, Lausanne, Switzerland, Sep. 5-1… [cited by examiner]
Xu, Canwen, Zhenzhong Chen, and Chenliang Li. (“Obj-glove: Scene-based contextual object embedding.” arXiv preprint arXiv:1907.01478 (2019).) (Year: 2019). [cited by examiner]
“European Application Serial No. 22706921.8, Voluntary Amendment filed Apr. 26, 2024”, 12 pages. [cited by applicant]
“Chinese Application Serial No. 2022800156830, Voluntary Amendment filed Jul. 2, 2024”, with English claims, 32 pages. [cited by applicant]
“European Application Serial No. 22709037.0, Response to Communication Pursuant to Rules 161 and 162 EPC filed Apr. 26, 2024”, 20 pages. [cited by applicant]
“U.S. Appl. No. 18/127,458, Notice of Allowance mailed Dec. 5, 2023”, 9 pages. [cited by applicant]
“U.S. Appl. No. 18/127,458, Non Final Office Action mailed Aug. 16, 2023”, 10 pgs. [cited by applicant]
“U.S. Appl. No. 18/127,458, Response filed Nov. 16, 2023 to Non Final Office Action mailed Aug. 16, 2023”, 9 pgs. [cited by applicant]
“U.S. Appl. No. 17/651,311, Examiner Interview Summary mailed May 9, 2023”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 17/651,311, Examiner Interview Summary mailed Nov. 22, 2022”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 17/651,311, Final Office Action mailed Sep. 21, 2022”, 23 pgs. [cited by applicant]
“U.S. Appl. No. 17/651,311, Non Final Office Action mailed Feb. 7, 2023”, 27 pgs. [cited by applicant]
“U.S. Appl. No. 17/651,311, Non Final Office Action mailed Jun. 7, 2022”, 22 pgs. [cited by applicant]
“U.S. Appl. No. 17/651,311, Notice of Allowance mailed May 25, 2023”, 9 pgs. [cited by applicant]
“U.S. Appl. No. 17/651,311, Response filed May 8, 2023 to Non Final Office Action mailed Feb. 7, 2023”, 11 pgs. [cited by applicant]
“U.S. Appl. No. 17/651,311, Response filed Sep. 6, 2022 to Non Final Office Action mailed Jun. 7, 2022”, 9 pgs. [cited by applicant]
“U.S. Appl. No. 17/651,311, Response filed Dec. 20, 2022 to Final Office Action mailed Sep. 21, 2022”, 16 pgs. [cited by applicant]
“U.S. Appl. No. 17/651,313, Corrected Notice of Allowability mailed Jul. 19, 2022”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 17/651,313, Notice of Allowance mailed May 17, 2022”, 9 pgs. [cited by applicant]
“U.S. Appl. No. 17/807,313, Corrected Notice of Allowability mailed Dec. 21, 2022”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 17/807,313, Non Final Office Action mailed Aug. 19, 2022”, 12 pgs. [cited by applicant]
“U.S. Appl. No. 17/807,313, Notice of Allowance mailed Sep. 28, 2022”, 9 pgs. [cited by applicant]
“U.S. Appl. No. 17/807,313, Notice of Allowance mailed Dec. 2, 2022”, 9 pgs. [cited by applicant]
“U.S. Appl. No. 17/807,313, Response filed Sep. 6, 2022 to Non Final Office Action mailed Aug. 19, 2022”, 1 pg. [cited by applicant]
“U.S. Appl. No. 18/127,458, Preliminary Amendment Filed Mar. 28, 2023”, 67 pgs. [cited by applicant]
“International Application Serial No. PCT/IB2022/051392, International Search Report mailed May 31, 2022”, 3 pgs. [cited by applicant]
“International Application Serial No. PCT/IB2022/051392, Written Opinion mailed May 31, 2022”, 7 pgs. [cited by applicant]
“International Application Serial No. PCT/IB2022/051394, International Search Report mailed May 25, 2022”, 3 pgs. [cited by applicant]
“International Application Serial No. PCT/IB2022/051394, Written Opinion mailed May 25, 2022”, 8 pgs. [cited by applicant]
Cho, Minseok, et al., “Adversarial TableQA: Attention supervision for question answering on tables”, Proceedings of Machine Leaning Research 95, (2018), 391-406. [cited by applicant]
Choi, Eunsol, et al., “QuAC: Question Answering in Context”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, (2018), 2174-2184. [cited by applicant]
Chuang, Yung-Sung, et al., “SpeechBERT: An audio-and-text jointly learned language model for end-to-end spoken question answering”, Interspeech 2020, Oct. 25-29, 2020, Shanghai, China, (2020), 5 pgs. [cited by applicant]
Clark, Jonathan H., et al., “TyDi Qa: A benchmark for information-seeking question answering in typologically diverse languages TACL (2020)”, Transactions of the Association for Computational Linguistics, vol. 8, (2020)… [cited by applicant]
Dai, Jifeng, et al., “R-FCN: Object detection via region-based fully convolutional networks. In: NeurIPS (2016)”, Advances in Neural Information Processing Systems 29 (NIPS 2016), (2016), 1-9. [cited by applicant]
Daniel, Hewlett, “Wiki Reading: A Novel Large-scale Language Understanding Task over Wikipedia”, In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistic, (Aug. 7-12, 2016), 1535-1545. [cited by applicant]
Denk, Timo I., et al., “BERTgrid: Contextualized Embedding for 2d Document Representation and UnderstandinD”, arXiv preprint, arXiv:1909.04948v2 [cs.CL] Oct. 4, 2019, (2019), 4 pgs. [cited by applicant]
Dodge, Jesse, et al., “Finetuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping”, ArXiv preprint, ARXIV:2002.06305v1 [cs.CL] Feb. 15, 2020, (2020), 11 pgs. [cited by applicant]
Dua, Dheeru, et al., “DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs”, Proceedings of NAACL-HLT 2019, (2019), 2368-2378. [cited by applicant]
Dwojak, Tomasz, et al., “From Dataset Recycling to Multi-Property Extraction and Beyond”, Proceedings of the 24th Conference on Computational Natural Language Learing, (2020), 641-651. [cited by applicant]
Ethayarajh, Kawin, et al., “How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embedding”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language… [cited by applicant]
Garncarek, Lukasz, et al., “LAMBERT: Layout-Aware Language Modeling for Information Extraction”, accepted to ICDAR 2021, (221021), 1-16. [cited by applicant]
Guu, Kevin, et al., “Retrieval augmented language model pre-training”, Proceedings of the 37th International Conference on Machine Learning, PMLR 119, (2020), 10 pgs. [cited by applicant]
Han, Kai, et al., “A Survey on Visual Transformer”, ArXiv preprint, arXiv:2012.12556v3 [cs.CV] Jan. 30, 2021, (2021), 26 pgs. [cited by applicant]
Harley, Adam D., et al., “Evaluation of deep convolutional nets for document image classification and retrieval”, 2015 International Conference on Document Analysis and Recognition (ICDAR), (2015), 991-995. [cited by applicant]
Hermann, Karl Moritz, et al., “Teaching machines to read and comprehend”, Advances in neural information processing systems 28, (2015). [cited by applicant]
Herzig, Jonathan, et al., “TaPas: Weakly supervised table parsing via pre-training”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, (2020), 4320-4333. [cited by applicant]
Hewlitt, Daniel, et al., “WikiReading: A novel large-scale language understanding task over Wikipedia”, Proceeding of the 54th Annual Meeting of the Association for Computational Linguistics (vol. 1, Long Papers), (2016… [cited by applicant]
Ho, Jonathan, et al., “Axial attention in multidimensional transformers”, arXiv preprint, arXiv:1912.12180v1 [cs.CV] Dec. 20, 2019, (2019), 11 pgs. [cited by applicant]
Hong, Teakgyu, et al., “BROS: A pre-trained language model for understanding texts in document openreview.net preprint”, openreview.net preprint, ICLR, (2021), 17 pgs. [cited by applicant]
Huang, Zheng, et al., “ICDAR2019 Competition on Scanned Receipt OCR and information Extraction”, 2019 International Conference on Document Analysis and Recognition (ICDAR), (2019), 1516-1520. [cited by applicant]
Hwang, Wonseok, et al., “Spatial dependency parsing for semi-structured document information extraction”, ArXiv preprint, arXiv:2005.006442v3 [cs.CL] Jul. 1, 2021, (2020), 14 pgs. [cited by applicant]
Jaume, Guillaume, et al., “Funsd: A dataset for form understanding in noisy scanned documents”, 2019 International Conference on Document Analysis Workshops (ICDARW), (2019), 1-6. [cited by applicant]
Kae, K., et al., “DVQA: understanding data visualizations via question answering”, InCVPR, (2018), 5648-5656. [cited by applicant]
Kafle, Kushal, et al., “DVQA: understanding data visualizations via question answering”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2018), 5648-5656. [cited by applicant]
Kahou, Samira E., et al., “FigureQA: An annotated figure dataset for visual reasoning”, Workshop track, ArXiv preprint, arXiv:1710.07300v2 [cs.CV] Feb. 22, 2018, (2018), 20 pgs. [cited by applicant]
Kasai, Jungo, et al., “Deep encoder, shallow decoder: Reevaluating the speed-quality tradeoff in machine translation”, Published as a conference paper at ICLR 2021, (2021), 1-16. [cited by applicant]
Keskar, Nitish, et al., “Unifying question answering and text classification via span extraction”, ArXiv preprint, arXiv:1904.09286v2 [cs.CL] Sep. 20, 2019, (2019), 10 pgs. [cited by applicant]
Khashabi, Daniel, et al., “UnifiedQA: Crossing format boundaries with a single QA system”, Findings of the Association for Computational Linguistics (EMNLP 2020), (2020), 1896-1907. [cited by applicant]
Knot, Tushar, et al., “QASC: A dataset for question answering via sentence composition.”, The Thirty-Fourth AAAI Conference on Artificial Intelligence (AAAI-20), (2020). [cited by applicant]
Kudo, T.Aku, et al., “Subword regularization: Improving neural network translation models with multiple subword candidates.”, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol.… [cited by applicant]
Kumar, Ankit, et al., “Ask me anything: Dynamic memory networks for natural language processing.”, Proceedings of the 33rd International Conference on Machine Learning, vol. 48, (2016), 10 pgs. [cited by applicant]
Kwiatkowski, Tom, et al., “Natural questions: A benchmark for question answering research. TACL (2019)”, Transactions of the Association for Computational Linguistics, vol. 7, (2019), 453-466. [cited by applicant]
Lai, Guokun., et al., “RACE: Large-scale Reading comprehension dataset from examinations”, Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, (2017), 785-794. [cited by applicant]
Le, Hung, et al., “Multimodal Transformer Networks for End-to-End Video-Grounded Dialogue System”, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, (2019), 5612-5623. [cited by applicant]
Lee, Kuang-H H., et al., “Stacked Cross Attention for Image-Text Matching”, ECCV 2018, (2018). [cited by applicant]
Lewis, Mike, et al., “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension”, Proceedings of the 58th Annual Meeting of the Association for Computational Lingu… [cited by applicant]
Li, Liunian H., et al., “VisualBERT: A simple and performant baseline for vision and language”, arXiv preprint, arXiv:1908.03557v1 [cs.CV] 9Aug2019, (2019), 14 pgs. [cited by applicant]
Liu, Xiaojing, et al., “Graph convolution for multimodal information extraction from visually rich documents. In:”, Proceedings of NAACL-HLT 2019, (2019), 32-39. [cited by applicant]
Ma, Junteng, et al., “Fusion of image-text attention for transformer-based multimodal machine translation.”, 2019 International Conference on Asian Language Processing (IALP), (2019), 199-204. [cited by applicant]
Mathew, Minesh, et al., “DocVQA: A dataset for VQA on document images”, IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), (2021), 2200-2209. [cited by applicant]
McCann, Bryan, et al., “The Natural Language Decathlon: Multitask Learning as Question answering”, arXiv preprint, arXiv:1806.087230v1 [cs.CL] Jun. 20, 2018, (2018), 23 pgs. [cited by applicant]
Palm, Rasmus B., et al., “CloudScan—a configuration-free invoice analysis system using recurrent neural networks”, Proceedings of 2017 14th IAPR International Conference on Document Analysis and Recognition, (2017), 8 p… [cited by applicant]
Park, Seunghyun, et al., “Cord: A consolidated receipt dataset for post-ocr parsing”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada, (2019), 1-4. [cited by applicant]
Powalski, R., et al., “Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer”, arXiv:2102.09550v3 [cs.CL] Jul. 12, 2021, (2021), 18 pgs. [cited by applicant]
Powalski, Rafal, et al., “UniCase { rethinking casing in language models”, ArXiv preprint, arXiv:2010.11936v1 [cs.CL] 122Oct. 2020, (2020), 5 pgs. [cited by applicant]
Radford, Alec, et al., “Language models are unsupervised multitask learners”, Technical Report, OpenAI, (2019), 24 pgs. [cited by applicant]
Raffel, Colin, et al., “Exploring the limits of transfer learning with a unified text-to-text transformer”, Journal of Machine Learning Research 21, (2019), 1-67. [cited by applicant]
Raffel, Colin, et al., “Exploring the Limits of Transfer Learning With a Unified Text-to-Text Transformet”, Journal of Machine Learning Research 21, (2020), 1-67. [cited by applicant]
Rajpurkar, Pranav, et al., “SQUAD: 100,000+ questions for machine comprehension of text In: EMNLP (2016)”, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, (2016), 2383-2392. [cited by applicant]
Reddy, Siva, et al., “CoQA: A Conversational Question Answering Challenge”, Transactions of the Association for Computational Linguistic, vol. 7, (2019), 249-266. [cited by applicant]
Ren, Yi, et al., “A study of non-autoregressive model for sequence generation.”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, (2020), 149-159. [cited by applicant]
Ronneberger, O., et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation.”, Medical Image Computing and Computer Invention (MICCAI) (LNCS 9351), (2015), 234-241. [cited by applicant]
Sarkhel, Ritesh, et al., “Deterministic routing between layout abstractions for multi-scale classification of visually rich documents”, 28th International Joint Conference on Artificial Intelligence (IJCAI), 2019, (2019… [cited by applicant]
Sennrich, Rico, et al., “Neural Machine Translation of Rare Words with Subword Units”, Proceedings of the 54th Annual Meeting of the Association for Linguistics, Computational(vol. 1: Long Papers), (2016), 1715-1725. [cited by applicant]
Sidorov, Oleksii, et al., “TextCaps: A dataset for image captioning with reading comprehension”, ArXiv preprint, arXiv:2003.12462v2 [cs.CV] Aug. 4, 2020, (2020), 26 pgs. [cited by applicant]
Singh, Amanpreet, et al., “Towards VQA models that can read”, CVPR, (2019), 10 pgs. [cited by applicant]
Stanislawek, Tomasz, et al., “Kleister: Key information extraction datasets involving long documents with complex layouts”, ArXiv preprint, arXiv: submit/3741295 [cs.CL] May 12, 2021, (2021), 16 pgs. [cited by applicant]
Su, Weijie, et al., “VL-BERT: pre-training of generic visual-linguistic representations”, Published as a conference paper at ICLR 2020, (2020), 16 pgs. [cited by applicant]
Vaswani, Ashish, et al., “Attention Is All You Need”, Proceedings, 31st Conference on Neural Processing Systems (NIPS 2017), (2017), 1-11. [cited by applicant]
Xu, et al., “Layoutimv2: Multi-Modal Pre-Training for Visually-Rich Document Understanding”, arXiv prepringrXiv, (Dec. 29, 2020). [cited by applicant]
Xu, et al., “LayoutLM: Pre-training of Text for Layout for Document Image Understanding”, Proceedings of the 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 20),, (2020), 9 pgs. [cited by applicant]
Xu, Canwen, et al., “Obj-glove: Scene-based contextual object embedding”, arXiv:1907.01478v1 [cs.CV] Jul. 2, 2019, (2019), 14 pgs. [cited by applicant]
Xu, Yang, et al., “LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding”, arXiv: 2012.14740v4 [cs.CL] Jan. 10, 2022, (2022), 13 pgs. [cited by applicant]
Xu, Yang, et al., “LayoutLMv2: Multi-modal pre-training for visually-rich document understanding”, Arxiv preprint, arXiv:2012.14740v1 [cs.CL] Dec. 29, 2020, (Dec. 29, 2020), 16 pgs. [cited by applicant]
Xu, Yang, et al., “LayoutLMv2: Multi-modal pre-training for visually-rich document understanding”, ArXiv preprint, arXiv:2012.14740v4 [cs.CL] Jan. 10, 2022, (2020), 13 pgs. [cited by applicant]
Xu, Yiheng, et al., “LayoutLM: Pre-training of text and layout for document image understanding”, KDD' 20: The 26th AC, SIGKDD Conference on Knowledge Discovery and Dat Mining, (2020), 14 pgs. [cited by applicant]
Yin, P., et al., “TaBERT: Pretraining for joint understanding of textual and tabular data”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, (2020), 8413-8426. [cited by applicant]
“U.S. Appl. No. 18/428,859, Notice of Allowance mailed Sep. 11, 2024”, 9 pgs. [cited by applicant]
“Chinese Application Serial No. 2022800156830, Office Action mailed Oct. 19, 2024”, W/O English Translation, 7 pgs. [cited by applicant]