IP Library › Granted Patent US 11,455,468
Granted Patent B2
US 11,455,468 · App. 17/651,313 · Granted Sep 27, 2022

Iterative training for text-image-layout transformer

Inventors: Adam Dancewicz (Warsaw, PL); Filip Gralinski (Warsaw, PL); Lukasz Konrad Borchmann (Poznan, PL)
Assignee: APPLICA SP. Z O.O.
G06F40/295G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,455,468
App. No.
17/651,313
Granted
Sep 27, 2022
Kind
B2
Abstract

Disclosed herein is a system and method for Natural Language Processing (NLP) of real world documents. the system and method combines various models not previously combined and overcomes the challenges of this combination. Models include an encoder-decoder model, a spatial model, and a multi-modal model. An iterative training process receives documents and generates outputs, wherein the iterative training process comprises enabling information retrieval from documents without training data.

Claims (53)

1. A system for Natural Language Processing (NLP) of real world documents, the system comprising:

a text-image-layout transformer (TILT) natural language processing (NLP) system that is executed on one or more processors, the TILT system comprises executable instructions that when executed by the processor, perform a method, the method comprising,

executing one or more models selected from a group comprising,

an encoder-decoder model;

a spatial model, and;

a multi-modal model;

receiving data comprising at least text data, layout data, and image data; and

operating on the received data to generate a useful output that relates to analysis of the received data; and

the system further executing an iterative training process that receives documents and generates outputs, wherein the iterative training process comprises enabling information retrieval from documents without training data.

2. The system of claim 1 , wherein executing the one or more models comprises name entity recognition (NER)-based extraction, and disconnecting name entity recognition (NER)-based extraction from the useful output.

3. The system of claim 1 , wherein the method further comprises receiving one or more questions regarding the received data.

4. The system of claim 3 , wherein the useful output comprises answers to the one or more questions.

5. The system of claim 3 , wherein the useful output comprises key information.

6. The system of claim 3 , wherein the useful output comprises document classification.

7. The system of claim 1 , wherein the spatial model comprises a spatial-aware transformer that employs self-attention and a word-centric masking method that concerns both images and text.

8. The system of claim 1 , wherein the method further comprises extending a T5 transformer to enable consumption of multi-modal input.

9. The system of claim 1 , wherein the multi-modal model comprises adding visual features to word embeddings that are contextualized on multiple resolution levels of an image.

10. The system of claim 9 , wherein the multi-modal model further comprises reliance on relative attention biases.

11. The system of claim 1 , wherein the method further comprises extending a T5 architectural approach to spatial dimensions.

12. The system of claim 1 , wherein the method further comprises generating contextualized image embeddings.

13. The system of claim 1 , wherein the method further comprises spatial bias augmentation.

14. A method for natural language processing (NLP), the method comprising:

one or more processors executing instructions to perform the method, the instructions comprising,

receiving a real-world document that comprises at least text data, layout data, and image data;

executing one or more models selected from a group comprising,

an encoder-decoder model;

a spatial model, and;

a multi-modal model;

operating on the real-world document to generate a useful output; and

further executing an iterative training process that receives documents and generates outputs, wherein the iterative training process comprises enabling information retrieval from documents without training data.

15. The method of claim 14 , wherein executing the one or more models comprises name entity recognition (NER)-based extraction, and disconnecting name entity recognition (NER)-based extraction from the useful output.

16. The method of claim 14 , wherein the method further comprises receiving one or more questions regarding the received data.

17. The method of claim 16 , wherein the useful output comprises answers to the one or more questions.

18. The method of claim 16 , wherein the useful output comprises key information.

19. The method of claim 16 , wherein the useful output comprises document classification.

20. The method of claim 14 , wherein the spatial model comprises a spatial-aware transformer that employs self-attention and a word-centric masking method that concerns both images and text.

21. The method of claim 14 , wherein the method further comprises extending a T5 transformer to enable consumption of multi-modal input.

22. The method of claim 14 , wherein the multi-modal model comprises adding visual features to word embeddings that are contextualized on multiple resolution levels of an image.

23. The method of claim 22 , wherein the multi-modal model further comprises reliance on relative attention biases.

24. The method of claim 14 , wherein the method further comprises extending a T5 architectural approach to spatial dimensions.

25. The method of claim 14 , wherein the method further comprises generating contextualized image embeddings.

26. The method of claim 14 , wherein the method further comprises spatial bias augmentation.

27. A non-transient computer medium having stored therein instructions, which when executed by a processor perform a method, the method comprising:

receiving a real-world document that comprises at least text data, layout data, and image data;

executing one or more models selected from a group comprising,

an encoder-decoder model;

a spatial model, and;

a multi-modal model; and

operating on the real-world document to generate a useful output; and

further executing an iterative training process that receives documents and generates outputs, wherein the iterative training process comprises enabling information retrieval from documents without training data.

28. The method of claim 27 , wherein executing the one or more models comprises name entity recognition (NER)-based extraction, and disconnecting name entity recognition (NER)-based extraction from the useful output.

29. The method of claim 27 , wherein the method further comprises receiving one or more questions regarding the received data.

30. The method of claim 29 , wherein the useful output comprises answers to the one or more questions.

Assignments (2)
CONFIRMATORY ASSIGNMENT Recorded May 19, 2025
From: APPLICA SP. Z O.O.
To: SNOWFLAKE INTERNATIONAL HOLDINGS INC.
Reel/Frame 071296/0386 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 16, 2022
From: DANCEWICZ, ADAM; GRALINSKI, FILIP; BORCHMANN, LUKASZ KONRAD
To: APPLICA SP. Z O.O.
Reel/Frame 059143/0083 →
Continuity (2)
Provisional Application 63150271 · Feb 17, 2021
Related Publication 20220261547A1 · Aug 18, 2022
Cited By (2)
US 12,314,668 US 12,346,658