Text-image-layout transformer [TILT]
Disclosed herein is a system and method for Natural Language Processing (NLP) of real world documents. The system and method combine various models not previously combined and overcome the challenges of this combination. Models include an encoder-decoder model, a spatial model, and a multi-modal model.
1. A system for Natural Language Processing (NLP) of real-world documents, the system comprising:
a text-image-layout transformer (TILT) NLP system that is executed on one or more processors, the TILT NLP system comprises executable instructions that when executed by the one or more processors, cause the TILT NLP system to perform operations comprising:
receiving multi-modal input data, the received multi-modal input data including at least text data, layout data, and image data;
executing an encoder-decoder model, a spatial model, and a multi-modal model, the encoder-decoder model including values absent from the text data, the spatial model including a spatial-aware transformer, and the multi-modal model including adding visual features to word embeddings that are contextualized on multiple resolution levels of an image; and
operating on the received multi-modal input data to generate a useful output that relates to analysis of the received data, the operating on the received multi-modal input data comprising:
maintaining a distinction between semantics and sequential distances;
extending biases with spatial relationships that include relative attention biases; and
providing additional image semantics to the received multi-modal input data, such that contextualized image embeddings cover image region semantics in a context of an entire visual neighborhood, the additional image semantics including text embeddings obtained from the image data.
2. The system of claim 1 , wherein executing the models comprises name entity recognition (NER)-based extraction, and disconnecting name entity recognition (NER)-based extraction from the useful output.
3. The system of claim 1 , wherein the operations further comprising:
receiving one or more questions regarding the received multi-modal input data.
4. The system of claim 3 , wherein the useful output comprises answers to the one or more questions.
5. The system of claim 3 , wherein the useful output comprises key information.
6. The system of claim 3 , wherein the useful output comprises document classification.
7. The system of claim 1 , wherein the spatial-aware transformer employs self-attention and a word-centric masking method that concerns both images and text.
8. The system of claim 1 , wherein the operations further comprising:
extending a T 5 transformer to enable consumption of the multi-modal input data.
9. The system of claim 1 , wherein the multi-modal model further comprises reliance on the relative attention biases.
10. The system of claim 1 , wherein the operations further comprising:
extending a T 5 architectural approach to spatial dimensions.
11. The system of claim 1 , wherein the operations further comprising:
generating contextualized image embeddings.
12. The system of claim 1 , wherein the operations further comprising:
employing spatial bias augmentation.
13. A method for natural language processing (NLP), the method comprising:
receiving, by at least one hardware processor, a real-world document that includes multi-modal input data, the multi-modal input data including at least one of text data, layout data, and image data;
executing an encoder-decoder model, a spatial model, and a multi-modal model, the encoder-decoder model including values absent from the text data, the spatial model including a spatial-aware transformer, and the multi-modal model including adding visual features to word embeddings that are contextualized on multiple resolution levels of an image; and
operating on the real-world document to generate a useful output that relates to analysis of the received data, the operating on the received real-world document comprising:
maintaining a distinction between semantics and sequential distances;
extending biases with spatial relationships that include relative attention biases; and
providing additional image semantics to the received real-world document, such that contextualized image embeddings cover image region semantics in a context of an entire visual neighborhood, the additional image semantics including text embeddings obtained from the image data.
14. The method of claim 13 , wherein executing the models comprises name entity recognition (NER)-based extraction, and disconnecting name entity recognition (NER)-based extraction from the useful output.
15. The method of claim 13 , wherein the method further comprises receiving one or more questions regarding the received real-world document.
16. The method of claim 15 , wherein the useful output comprises answers to the one or more questions.
17. The method of claim 15 , wherein the useful output comprises key information.
18. The method of claim 15 , wherein the useful output comprises document classification.
19. The method of claim 13 , wherein the spatial-aware transformer employs self-attention and a word-centric masking method that concerns both images and text.
20. The method of claim 13 , wherein the method further comprises extending a T 5 transformer to enable consumption of the multi-modal input data.
21. The method of claim 13 , wherein the multi-modal model further comprises reliance on the relative attention biases.
22. The method of claim 13 , wherein the method further comprises extending a T 5 architectural approach to spatial dimensions.
23. The method of claim 13 , wherein the method further comprises generating contextualized image embeddings.
24. The method of claim 13 , wherein the method further comprises employing spatial bias augmentation.
25. A non-transitory computer storage medium embodying instructions that, when executed by a machine, cause the machine to perform operations comprising:
receiving, by at least one hardware processor, a real-world document that includes at least one of text data, layout data, and image data;
executing an encoder-decoder model, a spatial model, and a multi-modal model, the encoder-decoder model including values absent from the text data, the spatial model including a spatial-aware transformer, and the multi-modal model including adding visual features to word embeddings that are contextualized on multiple resolution levels of an image; and
operating on the real-world document to generate a useful output that relates to analysis of the received data, the operating on the received real-world document comprising:
maintaining a distinction between semantics and sequential distances;
extending biases with spatial relationships that include relative attention biases; and
providing additional image semantics to the received real-world document, such that contextualized image embeddings cover image region semantics in a context of an entire visual neighborhood, the additional image semantics including text embeddings obtained from the image data.
26. The non-transitory computer storage medium of claim 25 , wherein executing the models comprises name entity recognition (NER)-based extraction, and disconnecting name entity recognition (NER)-based extraction from the useful output.
27. The non-transitory computer storage medium of claim 25 , the operations further comprising:
receiving one or more questions regarding the received data.
28. The non-transitory computer storage medium of claim 27 , wherein the useful output comprises answers to the one or more questions.