IP Library › Granted Patent US 12,158,874
Granted Patent B2
US 12,158,874 · App. 18/508,895 · Granted Dec 3, 2024

Encoder-decoder transformer for table generation

Inventors: Lukasz Konrad Borchmann (Warsaw, PL); Tomasz Dwojak (Poznan, PL); Lukasz Slawomir Garncarek (Warsaw, PL); Dawid Andrzej Jurkiewicz (Poznan, PL); Michal Waldemar Pietruszka (Cracow, PL); Gabriela Klaudia Palka (Poznan, PL); Karolina Szyndler (Szczecin, PL); Michal Turski (Warsaw, PL)
Assignee: APPLICA SP. Z O.O.
G06F16/2282G06F16/211
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,158,874
App. No.
18/508,895
Granted
Dec 3, 2024
Kind
B2
Abstract

Systems and methods for generating tables are provided. The systems and methods perform operations comprising accessing a text document comprising a plurality of strings; processing the text document by a machine learning model to generate a table comprising a plurality of entries that organizes the plurality of strings into rows and columns over a plurality of iterations; and at each of the plurality of iterations, estimating by the machine learning model a first value of a first entry of the plurality of entries based on a second value of a second entry of the plurality of entries that has been determined in a prior iteration.

Claims (68)

1. A system comprising:

at least one hardware processor; and

at least one memory storing instructions that cause the at least one hardware processor to execute operations comprising:

accessing a text document comprising a plurality of strings that are unstructured, the plurality of strings forming one or more sentences;

processing the text document by a machine learning model to generate a table that organizes words of the one or more sentences of the plurality of strings;

and

populating a second cell that is at a position that is non-adjacent relative to a position of a first cell in the table with a second word of the words of the one or more sentences rather than populating the second word in a third cell that is adjacent next to or underneath the first cell, the third cell being left empty in response to populating the second cell; and

at each of a plurality of iterations, predicting, by the machine learning model, a first value of the first cell based on a second value of the second cell that has been determined in a prior iteration.

2. The system of claim 1 , the operations comprising:

generating a first table instance comprising a first set of cells based on the plurality of strings;

at a first iteration, generating, by the machine learning model, a first plurality of confidence scores for values in each cell of the first set of cells of the first table instance; and

selecting a first subset of the first set of cells of the first table instance based on the first plurality of confidence scores.

3. The system of claim 2 , the operations comprising:

retrieving a confidence threshold;

comparing the first plurality of confidence scores for each cell of the first set of cells of the first table instance to the confidence threshold; and

identifying the first subset of the first set of cells that are associated with respective confidence scores that transgress the confidence threshold.

4. The system of claim 3 , the operations comprising generating a second table instance comprising a second set of entries by:

after selecting the first subset of the first set of cells, resetting the values associated with a remaining set of cells that are excluded from the first subset of the first set of cells; and

retaining the values associated with the first subset of the first set of cells in the second table instance.

5. The system of claim 4 , the operations comprising:

at a second iteration, processing the second table instance by the machine learning model, to generate a second plurality of confidence scores for each cell of the second set of cells of the second table instance; and

selecting a second subset of the second set of cells based on the second plurality of confidence scores.

6. The system of claim 5 , the operations comprising:

comparing the second plurality of confidence scores to the confidence threshold; and

identifying the second subset of the second set of cells that are associated with respective confidence scores that transgress the confidence threshold.

7. The system of claim 6 , the operations comprising:

repeating generation of additional table instances until all cells of a given table instance are associated with confidence scores that transgress the confidence threshold.

8. The system of claim 1 , the operations comprising:

training the machine learning model based on a corpus of training documents to maximize an expected log likelihood for a training table across all random permutations of a factorization order.

9. The system of claim 1 , wherein the machine learning model is trained to populate a plurality of cells of the table in any order in a way that maximizes confidence in each individual entry and reduces error accumulation.

10. The system of claim 1 , the operations comprising:

training the machine learning model based on a corpus of training documents by populating a plurality of training tables representing different permutations of strings in the training documents to maximize an expected log likelihood.

11. The system of claim 1 , the machine learning model generatively predicting text to include in an empty cell of the table by performing operations comprising:

identifying a category or column header corresponding to the empty cell;

processing data from a row corresponding to the empty cell and corresponding headers of each populated entry from the row corresponding to the empty cell; and

using information from other cells in the row corresponding to the empty cell to estimate or predict the text to include in the empty cell.

12. The system of claim 1 , wherein the machine learning model estimates column headers of the table.

13. The system of claim 1 , the operations comprising:

inferring one or more values for individual cells of a plurality of cells with words excluded from the plurality of strings based on values of other cells of the plurality of cells; and

populating an individual cell of the plurality of cells of the table using the inferred one or more values without retrieving values from the text document.

14. A method comprising:

accessing, by one or more processors, a text document comprising a plurality of strings that are unstructured, the plurality of strings forming one or more sentences;

processing the text document by a machine learning model to generate a table that organizes words of the one or more sentences of the plurality of strings;

and

populating a second cell that is at a position that is non-adjacent relative to a position of a first cell in the table with a second word of the words of the one or more sentences rather than populating the second word in a third cell that is adjacent next to or underneath the first cell, the third cell being left empty in response to populating the second cell; and

at each of a plurality of iterations, predicting, by the machine learning model, a first value of the first cell based on a second value of the second cell that has been determined in a prior iteration.

15. The method of claim 14 , further comprising:

receiving a query that indicates values for columns, wherein the table is generated based on the values indicated for the columns.

16. The method of claim 14 , further comprising:

generating a first table instance comprising a first set of cells based on the plurality of strings;

at a first iteration, generating, by the machine learning model, a first plurality of confidence scores for values in each cell of the first set of cells of the first table instance; and

selecting a first subset of the first set of cells of the first table instance based on the first plurality of confidence scores.

17. The method of claim 16 , further comprising:

retrieving a confidence threshold;

comparing the first plurality of confidence scores for each cell of the first set of cells of the first table instance to the confidence threshold; and

identifying the first subset of the first set of cells that are associated with respective confidence scores that transgress the confidence threshold.

18. The method of claim 17 , further comprising generating a second table instance comprising a second set of cells by:

after selecting the first subset of the first set of cells, resetting the values associated with a remaining set of cells that are excluded from the first subset of the first set of cells; and

retaining the values associated with the first subset of the first set of cells in the second table instance.

19. A non-transitory computer-storage medium comprising instructions that, when executed by a processor of a machine, configure the machine to perform operations comprising:

accessing a text document comprising a plurality of strings that are unstructured, the plurality of strings forming one or more sentences;

processing the text document by a machine learning model to generate a table that organizes words of the one or more sentences of the plurality of strings;

and

populating a second cell that is at a position that is non-adjacent relative to a position of a first cell in the table with a second word of the words of the one or more sentences rather than populating the second word in a third cell that is adjacent next to or underneath the first cell, the third cell being left empty in response to populating the second cell; and

at each of a plurality of iterations, predicting, by the machine learning model, a first value of the first cell based on a second value of the second cell that has been determined in a prior iteration.

20. The non-transitory computer-storage medium of claim 19 , the operations comprising:

inferring one or more values for individual cells of a plurality of cells with words excluded from the plurality of strings based on values of other cells of the plurality of cells; and

populating an individual cell of the plurality of cells of the table using the inferred one or more values without retrieving values from the text document.

Assignments (2)
CONFIRMATORY ASSIGNMENT Recorded May 19, 2025
From: APPLICA SP. Z O.O.
To: SNOWFLAKE INTERNATIONAL HOLDINGS INC.
Reel/Frame 071296/0386 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: PIETRUSZKA, MICHAL WALDEMAR; TURSKI, MICHAL; BORCHMANN, LUKASZ KONRAD; DWOJAK, TOMASZ; PALKA, GABRIELA KLAUDIA; SZYNDLER, KAROLINA; JURKIEWICZ, DAWID ANDRZEJ; GARNCAREK, LUKASZ SLAWOMIR
To: APPLICA SP. Z O.O.
Reel/Frame 065560/0313 →
Continuity (3)
Continuation 18152083 · Jan 9, 2023
Provisional Application 63267174 · Jan 26, 2022
Related Publication 20240086388A1 · Mar 14, 2024
Cited By (1)
US 12,626,186