IP Library › Granted Patent US 12,645,662
Granted Patent B2
US 12,645,662 · App. 19/061,638 · Granted Jun 2, 2026

Text-based machine learning extraction of table data from a read-only document

Inventors: Hongyang Yu (Wentworth Point, AU); Hanieh Borhanazad (Sydney, AU); Sandip Mandlecha (Pune, IN)
Assignee: Coupa Software Incorporated
G06F16/2282G06F16/93G06N3/045G06V30/412G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,662
App. No.
19/061,638
Granted
Jun 2, 2026
Kind
B2
Abstract

Embodiments of the disclosed technologies provide solutions for automatically reading digital electronic documents that contain tables and correctly extracting table data, rows and columns from the documents with high accuracy and high throughput. Embodiments are capable of converting a table portion of a read-only document to a searchable, editable data record using text rectangle (TR)-level numerical data that indicates probabilities of TRs belonging to canonicals and at least one convolutional neural network (CNN) that processes the TR-level numerical data to produce table-level numerical data.

Claims (36)

1 . A computer-implemented method, comprising:

receiving a read-only digital document having one or more text rectangles;

determining text rectangle (TR)-level numerical data for each of the one or more text rectangles, the TR-level numerical data being associated with a feature probability vector;

generating a stretched digital document by stretching one or more coordinates of the one or more text rectangles in a first direction and a second direction;

partitioning the stretched digital document into a grid having a plurality of cells;

projecting the TR-level numerical data onto the grid;

determining that a first text rectangle of the one or more text rectangles occupies a length of the grid formed by multiple cells of the plurality of cells, wherein the first text rectangle is associated with a feature weight of the feature probability vector;

de-biasing the feature probability vector by reducing the feature weight; and

processing the TR-level numerical data using a convolutional neural network (CNN) to classify the one or more text rectangles as belonging to a particular canonical.

2 . The computer-implemented method of claim 1 , wherein the determining of the TR-level numerical data is performed by a first neural network that divides the text rectangles into a first class representing text rectangles that are a label and a second class representing text rectangles that are not a label.

3 . The computer-implemented method of claim 2 , wherein the first class of text rectangles is further classified by a second neural network configured to output a label feature probability vector of the text rectangle belonging to a label canonical and the second class of text rectangles is further classified by a third neural network configured to output a value feature probability vector of the text rectangle belonging to a value canonical.

4 . The computer-implemented method of claim 3 , wherein the label canonical is quantity, price, or description.

5 . The computer-implemented method of claim 3 , wherein the value canonical is numeric, text, currency, or date.

6 . The computer-implemented method of claim 3 , wherein the TR-level numerical data is projected onto the grid by a feature map generator configured to map the label feature probability vector and the value feature probability vector to corresponding locations on the grid.

7 . The computer-implemented method of claim 1 , wherein the grid is of a grid size of about g cells by g cells, where each cell has dimensions of c pixels by c pixels, where c corresponds to a minimum font size of the text and g is a multiple of c.

8 . The computer-implemented method of claim 7 , wherein the minimum font size is larger than a predetermined minimum limit.

9 . The computer-implemented method of claim 7 , wherein the coordinates of the text rectangles of the stretched document are stretched to occupy a feature map of size p pixels by p pixels and p=g multiplied by c.

10 . The computer-implemented method of claim 1 , wherein the feature weight is reduced by applying a fading process such as an exponential smoothing or a weight decay function to the feature probability vector.

11 . One or more non-transitory computer-readable storage media, storing instructions which, when executed, cause one or more processors to execute:

receiving a read-only digital document having one or more text rectangles;

determining text rectangle (TR)-level numerical data for each of the one or more text rectangles, the TR-level numerical data being associated with a feature probability vector;

generating a stretched digital document by stretching one or more coordinates of the one or more text rectangles in a first direction and a second direction;

partitioning the stretched digital document into a grid having a plurality of cells;

projecting the TR-level numerical data onto the grid;

determining that a first text rectangle of the one or more text rectangles occupies a length of the grid formed by multiple cells of the plurality of cells, wherein the first text rectangle is associated with a feature weight of the feature probability vector;

de-biasing the feature probability vector by reducing the feature weight; and

processing the TR-level numerical data using a convolutional neural network (CNN) to classify the one or more text rectangles as belonging to a particular canonical.

12 . The one or more non-transitory computer-readable storage media of claim 11 , wherein the determining of the TR-level numerical data is performed by a first neural network that divides the text rectangles into a first class representing text rectangles that are a label and a second class representing text rectangles that are not a label.

13 . The one or more non-transitory computer-readable storage media of claim 12 , wherein the first class of text rectangles is further classified by a second neural network configured to output a label feature probability vector of the text rectangle belonging to a label canonical and the second class of text rectangles is further classified by a third neural network configured to output a value feature probability vector of the text rectangle belonging to a value canonical.

14 . The one or more non-transitory computer-readable storage media of claim 13 , wherein the label canonical is quantity, price, or description.

15 . The one or more non-transitory computer-readable storage media of claim 13 , wherein the value canonical is numeric, text, currency, or date.

16 . The one or more non-transitory computer-readable storage media of claim 13 , wherein the TR-level numerical data is projected onto the grid by a feature map generator configured to map the label feature probability vector and the value feature probability vector to corresponding locations on the grid.

17 . The one or more non-transitory computer-readable storage media of claim 11 , wherein the grid is of a grid size of about g cells by g cells, where each cell has dimensions of c pixels by c pixels, where c corresponds to a minimum font size of the text and g is a multiple of c.

18 . The one or more non-transitory computer-readable storage media of claim 17 , wherein the minimum font size is larger than a predetermined minimum limit.

19 . The one or more non-transitory computer-readable storage media of claim 17 , wherein the coordinates of the text rectangles of the stretched document are stretched to occupy a feature map of size p pixels by p pixels and p=g multiplied by c.

20 . The one or more non-transitory computer-readable storage media of claim 11 , wherein the feature weight is reduced by applying a fading process such as an exponential smoothing or a weight decay function to the feature probability vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 24, 2025
From: YU, HONGYANG; BORHANAZAD, HANIEH; MANDLECHA, SANDIP
To: COUPA SOFTWARE INCORPORATED
Reel/Frame 070311/0167 →
Priority Claims (1)
IN 202011037847 · Sep 2, 2020 · national
Continuity (4)
Continuation 18415062 · Jan 17, 2024
Continuation 17973511 · Oct 25, 2022
Continuation 17074957 · Oct 20, 2020
Related Publication 20250190418A1 · Jun 12, 2025
References Cited (15)
US 10878173B2 · Morariu · 2020 [cited by applicant]
US 11500843B2 · Yu · 2022 [cited by examiner]
US 20030028503A1 · Giuffrida · 2003 [cited by applicant]
US 20030028801A1 · Liberman · 2003 [cited by examiner]
US 20080243512A1 · Breebaart · 2008 [cited by examiner]
US 20100131841A1 · Saito · 2010 [cited by examiner]
US 20130033390A1 · Hoshikawa et al. · 2013 [cited by applicant]
US 20180203984A1 · Agrawal et al. · 2018 [cited by applicant]
US 20190087444A1 · Arakawa · 2019 [cited by applicant]
US 20200265224A1 · Gurav et al. · 2020 [cited by applicant]
US 20200285878A1 · Wang · 2020 [cited by examiner]
US 20200410231A1 · Chua et al. · 2020 [cited by applicant]
US 20210141781A1 · Chadha et al. · 2021 [cited by applicant]
US 20210241331A1 · Katzenelson et al. · 2021 [cited by applicant]
US 20230252231A1 · Kotwal et al. · 2023 [cited by applicant]