IP Library › Granted Patent US 11,763,087
Granted Patent B2
US 11,763,087 · App. 17/651,311 · Granted Sep 19, 2023

Text-image-layout transformer [TILT]

Inventors: Lukasz Konrad Borchmann (Poznan, PL); Dawid Andrzej Jurkiewicz (Poznan, PL); Tomasz Dwojak (Poznan, PL); Michal Waldemar Pietruszka (Cracow, PL); Gabriela Klaudia Palka (Poznan, PL)
Assignee: APPLICA SP. Z.O.O.
G06F40/295G06F40/106G06F40/30G06N3/08G06T11/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,763,087
App. No.
17/651,311
Granted
Sep 19, 2023
Kind
B2
Abstract

Disclosed herein is a system and method for Natural Language Processing (NLP) of real world documents. The system and method combine various models not previously combined and overcome the challenges of this combination. Models include an encoder-decoder model, a spatial model, and a multi-modal model.

Claims (53)

1. A system for Natural Language Processing (NLP) of real-world documents, the system comprising:

a text-image-layout transformer (TILT) NLP system that is executed on one or more processors, the TILT NLP system comprises executable instructions that when executed by the one or more processors, cause the TILT NLP system to perform operations comprising:

receiving multi-modal input data, the received multi-modal input data including at least text data, layout data, and image data;

executing an encoder-decoder model, a spatial model, and a multi-modal model, the encoder-decoder model including values absent from the text data, the spatial model including a spatial-aware transformer, and the multi-modal model including adding visual features to word embeddings that are contextualized on multiple resolution levels of an image; and

operating on the received multi-modal input data to generate a useful output that relates to analysis of the received data, the operating on the received multi-modal input data comprising:

maintaining a distinction between semantics and sequential distances;

extending biases with spatial relationships that include relative attention biases; and

providing additional image semantics to the received multi-modal input data, such that contextualized image embeddings cover image region semantics in a context of an entire visual neighborhood, the additional image semantics including text embeddings obtained from the image data.

2. The system of claim 1 , wherein executing the models comprises name entity recognition (NER)-based extraction, and disconnecting name entity recognition (NER)-based extraction from the useful output.

3. The system of claim 1 , wherein the operations further comprising:

receiving one or more questions regarding the received multi-modal input data.

4. The system of claim 3 , wherein the useful output comprises answers to the one or more questions.

5. The system of claim 3 , wherein the useful output comprises key information.

6. The system of claim 3 , wherein the useful output comprises document classification.

7. The system of claim 1 , wherein the spatial-aware transformer employs self-attention and a word-centric masking method that concerns both images and text.

8. The system of claim 1 , wherein the operations further comprising:

extending a T 5 transformer to enable consumption of the multi-modal input data.

9. The system of claim 1 , wherein the multi-modal model further comprises reliance on the relative attention biases.

10. The system of claim 1 , wherein the operations further comprising:

extending a T 5 architectural approach to spatial dimensions.

11. The system of claim 1 , wherein the operations further comprising:

generating contextualized image embeddings.

12. The system of claim 1 , wherein the operations further comprising:

employing spatial bias augmentation.

13. A method for natural language processing (NLP), the method comprising:

receiving, by at least one hardware processor, a real-world document that includes multi-modal input data, the multi-modal input data including at least one of text data, layout data, and image data;

executing an encoder-decoder model, a spatial model, and a multi-modal model, the encoder-decoder model including values absent from the text data, the spatial model including a spatial-aware transformer, and the multi-modal model including adding visual features to word embeddings that are contextualized on multiple resolution levels of an image; and

operating on the real-world document to generate a useful output that relates to analysis of the received data, the operating on the received real-world document comprising:

maintaining a distinction between semantics and sequential distances;

extending biases with spatial relationships that include relative attention biases; and

providing additional image semantics to the received real-world document, such that contextualized image embeddings cover image region semantics in a context of an entire visual neighborhood, the additional image semantics including text embeddings obtained from the image data.

14. The method of claim 13 , wherein executing the models comprises name entity recognition (NER)-based extraction, and disconnecting name entity recognition (NER)-based extraction from the useful output.

15. The method of claim 13 , wherein the method further comprises receiving one or more questions regarding the received real-world document.

16. The method of claim 15 , wherein the useful output comprises answers to the one or more questions.

17. The method of claim 15 , wherein the useful output comprises key information.

18. The method of claim 15 , wherein the useful output comprises document classification.

19. The method of claim 13 , wherein the spatial-aware transformer employs self-attention and a word-centric masking method that concerns both images and text.

20. The method of claim 13 , wherein the method further comprises extending a T 5 transformer to enable consumption of the multi-modal input data.

21. The method of claim 13 , wherein the multi-modal model further comprises reliance on the relative attention biases.

22. The method of claim 13 , wherein the method further comprises extending a T 5 architectural approach to spatial dimensions.

23. The method of claim 13 , wherein the method further comprises generating contextualized image embeddings.

24. The method of claim 13 , wherein the method further comprises employing spatial bias augmentation.

25. A non-transitory computer storage medium embodying instructions that, when executed by a machine, cause the machine to perform operations comprising:

receiving, by at least one hardware processor, a real-world document that includes at least one of text data, layout data, and image data;

executing an encoder-decoder model, a spatial model, and a multi-modal model, the encoder-decoder model including values absent from the text data, the spatial model including a spatial-aware transformer, and the multi-modal model including adding visual features to word embeddings that are contextualized on multiple resolution levels of an image; and

operating on the real-world document to generate a useful output that relates to analysis of the received data, the operating on the received real-world document comprising:

maintaining a distinction between semantics and sequential distances;

extending biases with spatial relationships that include relative attention biases; and

providing additional image semantics to the received real-world document, such that contextualized image embeddings cover image region semantics in a context of an entire visual neighborhood, the additional image semantics including text embeddings obtained from the image data.

26. The non-transitory computer storage medium of claim 25 , wherein executing the models comprises name entity recognition (NER)-based extraction, and disconnecting name entity recognition (NER)-based extraction from the useful output.

27. The non-transitory computer storage medium of claim 25 , the operations further comprising:

receiving one or more questions regarding the received data.

28. The non-transitory computer storage medium of claim 27 , wherein the useful output comprises answers to the one or more questions.

Assignments (2)
CONFIRMATORY ASSIGNMENT Recorded May 19, 2025
From: APPLICA SP. Z O.O.
To: SNOWFLAKE INTERNATIONAL HOLDINGS INC.
Reel/Frame 071296/0386 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 16, 2022
From: BORCHMANN, LUKASZ KONRAD; JURKIEWICZ, DAWID ANDRZEJ; DWOJAK, TOMASZ; PIETRUSZKA, MICHAL WALDEMAR; PALKA, GABRIELA KLAUDIA
To: APPLICA SP. Z O.O.
Reel/Frame 059142/0680 →
Continuity (2)
Provisional Application 63150271 · Feb 17, 2021
Related Publication 20220270311A1 · Aug 25, 2022
Cited By (4)
US 12,314,668 US 12,346,658 US 12,737,934 US 12,743,728