IP Library › Granted Patent US 11,934,786
Granted Patent B2
US 11,934,786 · App. 18/127,458 · Granted Mar 19, 2024

Iterative training for text-image-layout data in natural language processing

Inventors: Adam Dancewicz (Warsaw, PL); Filip Gralinkski (Warsaw, PL); Lukasz Konrad Borchmann (Warsaw, PL)
Assignee: APPLICA SP. Z O.O.
G06F40/295G06F40/106G06F40/30G06N3/08G06T11/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,934,786
App. No.
18/127,458
Granted
Mar 19, 2024
Kind
B2
Abstract

Methods, systems, and computer programs are presented for providing access to a cloud data platform including a machine learning model for performing a plurality of iterations, by at least one hardware processor, to generate a Natural Language Processing (NLP) model. The cloud data platform performs each iteration by receiving real-world documents and enabling information retrieval from the real-world documents without annotated training data. Each iteration includes receiving data comprising text data, layout data, and image data and analyzing the text data, the layout data, and the image data. The cloud data platform generates one or more outputs from the machine learning model by applying the iterative training on new data, based at least in part on the analyzing of the text data, the layout data, and the image data.

Claims (89)

1. A system comprising:

one or more processors of a machine; and

at least one memory storing instructions that, when executed by the one or more processors, cause the machine to perform operations comprising:

providing access to a cloud data platform including a machine learning model for performing a plurality of iterations to generate a Natural Language Processing (NLP) model, each iteration comprising:

receiving real-world documents;

enabling information retrieval from the real-world documents without annotated training data;

receiving data comprising text data, layout data, and image data;

analyzing the text data, the layout data, and the image data; and

generating one or more outputs from the machine learning model by applying the plurality of iterations on new data, based at least in part on the analyzing of the text data, the layout data, and the image data.

2. The system of claim 1 , the operations further comprising:

generating a prediction model for extracting a set of data points from the real-world documents; and

returning a value for each data point in the set of data points, wherein the value for each data point is included in the one or more outputs.

3. The system of claim 2 , the operations further comprising:

enabling verification of correctness of one or more predictions provided by the prediction model, wherein the verification of correctness includes comparing extracted values to information in the real-world documents.

4. The system of claim 3 , wherein performing the plurality of iterations further comprises:

evaluating a quality of the prediction model at an end of each iteration, wherein the quality includes a percentage of correct predictions; and

stopping the plurality of iterations after the quality of the prediction model is satisfactory to a user.

5. The system of claim 1 , the operations further comprising:

executing one or more models selected from:

an encoder-decoder model;

a spatial model, and;

a multi-modal model.

6. The system of claim 5 , the operations further comprising:

executing a layout-aware model formulated within the encoder-decoder model; and

applying the encoder-decoder model to information extraction and question answering.

7. The system of claim 5 , wherein executing the encoder-decoder model includes generating one or more values not included in the text data.

8. The system of claim 5 , wherein executing the spatial model includes considering layout information directly as being positional embeddings or indirectly as being contextualized on spatial neighborhoods.

9. The system of claim 5 , wherein the spatial model includes a spatial-aware transformer comprising:

adding bias to self-attention; and

augmenting one or more spatial biases by multiplying a horizontal and a vertical distance between one or more tokens.

10. The system of claim 5 , wherein the multi-modal model includes a multi-modal transformer to add one or more visual features to word embeddings.

11. A method comprising:

providing access to a cloud data platform including a machine learning model for performing a plurality of iterations, by at least one hardware processor, to generate a Natural Language Processing (NLP) model, each iteration comprising:

receiving real-world documents;

enabling information retrieval from the real-world documents without annotated training data;

receiving data comprising text data, layout data, and image data;

analyzing the text data, the layout data, and the image data; and

generating one or more outputs from the machine learning model by applying the iterative training on new data, based at least in part on the analyzing of the text data, the layout data, and the image data.

12. The method of claim 11 , further comprising:

generating a prediction model for extracting a set of data points from the real-world documents; and

returning a value for each data point in the set of data points, wherein the value for each data point is included in the one or more outputs.

13. The method of claim 12 , further comprising:

enabling verification of correctness of one or more predictions provided by the prediction model, wherein the verification of correctness includes comparing extracted values to information in the real-world document.

14. The method of claim 13 , wherein performing the plurality of iterations further comprises:

evaluating a quality of the prediction model at an end of each iteration, wherein the quality includes a percentage of correct predictions; and

stopping the iterations after the quality of the prediction model is satisfactory to a user.

15. The method of claim 11 , further comprising:

executing one or more models selected from:

an encoder-decoder model;

a spatial model, and;

a multi-modal model.

16. The method of claim 15 , further comprising:

executing a layout-aware model formulated within the encoder-decoder model; and

applying the encoder-decoder model to information extraction and question answering.

17. The method of claim 15 , wherein executing the encoder-decoder model includes generating one or more values not included in the text data.

18. The method of claim 15 , wherein executing the spatial model includes considering layout information directly as being positional embeddings or indirectly as being contextualized on spatial neighborhoods.

19. The method of claim 15 , wherein the spatial model includes a spatial-aware transformer comprising:

adding bias to self-attention; and

augmenting one or more spatial biases by multiplying a horizontal and a vertical distance between one or more tokens.

20. The method of claim 15 , wherein the multi-modal model includes a multi-modal transformer to add one or more visual features to word embeddings.

21. A non-transitory computer medium embodying instructions that, when executed by a machine, cause the computer medium to perform operations comprising:

providing access to a cloud data platform including a machine learning model for performing a plurality of iterations, by at least one hardware processor, to generate an NLP model, each iteration comprising:

receiving real-world documents;

enabling information retrieval from the real-world documents without annotated training data;

receiving data comprising text data, layout data, and image data;

analyzing the text data, the layout data, and the image data; and

generating one or more outputs from the machine learning model by applying the iterative training on new data, based at least in part on the analyzing of the text data, the layout data, and the image data.

22. The non-transitory computer medium of claim 21 , further comprising:

generating a prediction model for extracting a set of data points from the real-world documents; and

returning a value for each data point in the set of data points, wherein the value for each data point is included in the one or more outputs.

23. The non-transitory computer medium of claim 22 , further comprising:

enabling verification of correctness of one or more predictions provided by the prediction model, wherein the verification of correctness includes comparing extracted values to information in the real-world document.

24. The non-transitory computer medium of claim 23 , wherein performing the plurality of iterations further comprises:

evaluating a quality of the prediction model at an end of each iteration, wherein the quality includes a percentage of correct predictions; and

stopping the iterations after the quality of the prediction model is satisfactory to a user.

25. The non-transitory computer medium of claim 21 , further comprising:

executing one or more models selected from:

an encoder-decoder model;

a spatial model, and;

a multi-modal model.

26. The non-transitory computer medium of claim 25 , further comprising:

executing a layout-aware model formulated within the encoder-decoder model; and

applying the encoder-decoder model to information extraction and question answering.

27. The non-transitory computer medium of claim 25 , wherein executing the encoder-decoder model includes generating one or more values not included in the text data.

28. The non-transitory computer medium of claim 25 , wherein executing the spatial model includes considering layout information directly as being positional embeddings or indirectly as being contextualized on spatial neighborhoods.

29. The non-transitory computer medium of claim 25 , wherein the spatial model includes a spatial-aware transformer comprising:

adding bias to self-attention; and

augmenting one or more spatial biases by multiplying a horizontal and a vertical distance between one or more tokens.

30. The non-transitory computer medium of claim 25 , wherein the multi-modal model includes a multi-modal transformer to add one or more visual features to word embeddings.

Assignments (2)
CONFIRMATORY ASSIGNMENT Recorded May 19, 2025
From: APPLICA SP. Z O.O.
To: SNOWFLAKE INTERNATIONAL HOLDINGS INC.
Reel/Frame 071296/0386 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: DANCEWICZ, ADAM; GRALINSKI, FILIP; BORCHMANN, LUKASZ KONRAD
To: APPLICA SP. Z.O.O.
Reel/Frame 063135/0794 →
Continuity (4)
Continuation 17807313 · Jun 16, 2022
Continuation 17651313 · Feb 16, 2022
Provisional Application 63150271 · Feb 17, 2021
Related Publication 20230259709A1 · Aug 17, 2023