IP Library Granted Patent US 12,205,391
Granted Patent B2
US 12,205,391 · App. 17/562,471 · Granted Jan 21, 2025

Extracting structured information from document images

Inventors: Mikhail Lanin (Moscow, RU); Stanislav Semenov (Moscow, RU)
Assignee: ABBYY Development Inc.
G06V30/412G06V30/19107G06V30/19173G06V30/413G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,391
App. No.
17/562,471
Granted
Jan 21, 2025
Kind
B2
Abstract

An example method of extracting structured information from document images comprises: receiving a document image; detecting a tabular structure within the document image; identifying a plurality of rows of the tabular structure, wherein each row of the plurality of rows comprises one or more lines; for each row of the plurality of rows, identifying a set of field types of one or more fields comprised by each line of the one or more lines comprised by the respective row; detecting, in each line of the one or more lines, a set of fields corresponding to a respective set of field types; and extracting information from the set of fields.

Claims (67)

1. A method, comprising:

receiving, by a processing device, a document image;

detecting a tabular structure within the document image;

identifying a plurality of rows of the tabular structure, wherein each row of the plurality of rows comprises plurality of lines;

for each row of the plurality of rows, identifying a respective plurality of lines comprised by the row;

detecting, in each line of the plurality of lines, a set of fields corresponding to a respective set of field types; and

extracting information from the set of fields.

2. The method of claim 1 , wherein the tabular structure is provided by a table.

3. The method of claim 1 , further comprising:

identifying vertical boundaries of the tabular structure.

4. The method of claim 1 , wherein identifying the plurality of rows further comprises:

identifying a plurality of lines of the tabular structure; and

classifying each line of the plurality of lines.

5. The method of claim 1 , wherein identifying the plurality of rows further comprises:

identifying a plurality of lines of the tabular structure; and

clustering the plurality of lines into a plurality of clusters.

6. The method of claim 1 , wherein identifying the set of field types further comprises:

determining, for each line of the one or more lines, a corresponding line type derived from a corresponding set of field types of the one or more fields comprised by the line.

7. The method of claim 1 , further comprising:

detecting a multi-line field by grouping two or more single-line field portions.

8. A system comprising:

a memory; and

a processing device operatively coupled to the memory, the processing device configured to:

receive a document image;

detect a tabular structure within the document image;

identify a plurality of rows of the tabular structure, wherein each row of the plurality of rows comprises plurality of lines;

for each row of the plurality of rows, identify a respective plurality of lines comprised by the row;

identify a set of field types of one or more fields comprised by each line of the one or more lines comprised by the respective row;

detect, in each line of the plurality of lines, a set of fields corresponding to a respective set of field types; and

extract information from the set of fields.

9. The system of claim 8 , wherein the processing device is further configured to:

identify vertical boundaries of the tabular structure.

10. The system of claim 8 , wherein identifying the plurality of rows further comprises:

identifying a plurality of lines of the tabular structure; and

classifying each line of the plurality of lines.

11. The system of claim 8 , wherein identifying the plurality of rows further comprises:

identifying a plurality of lines of the tabular structure; and

clustering the plurality of lines into a plurality of clusters.

12. The system of claim 8 , wherein identifying the set of field types further comprises:

determining, for each line of the one or more lines, a corresponding line type derived from a corresponding set of field types of the one or more fields comprised by the line.

13. The system of claim 8 , wherein the processing device is further configured to:

detecting a multi-line field by grouping two or more single-line field portions.

14. The system of claim 8 , wherein the processing device is further configured to:

train a table detection classifier for detecting one or more tabular structures within a document;

train a vertical boundary detection classifier for identifying vertical boundaries of the one or more tabular structures;

train a row detection classifier for determining a row layout of the one or more tabular structures;

train a line classifier for extracting information from the fields of the one or more tabular structures;

train a field detection module for detecting fields of the one or more tabular structures.

15. A non-transitory computer-readable storage medium including executable instructions that, when executed by a computing system, cause the computing system to:

receive a document image;

detect a tabular structure within the document image;

identify a plurality of rows of the tabular structure, wherein each row of the plurality of rows comprises plurality of lines;

for each row of the plurality of rows, identify a respective plurality of lines comprised by the row;

identify a set of field types of one or more fields comprised by each line of the one or more plurality of lines comprised by the respective row;

detect, in each line of the plurality of lines, a set of fields corresponding to a respective set of field types; and

extract information from the set of fields.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the tabular structure is provided by a table.

17. The non-transitory computer-readable storage medium of claim 15 , further comprising executable instructions that, when executed by the computing system, cause the computing system to:

identifying vertical boundaries of the tabular structure.

18. The non-transitory computer-readable storage medium of claim 15 , wherein identifying the plurality of rows further comprises:

identifying a plurality of lines of the tabular structure; and

classifying each line of the plurality of lines.

19. The non-transitory computer-readable storage medium of claim 15 , wherein identifying the plurality of rows further comprises:

identifying a plurality of lines of the tabular structure; and

clustering the plurality of lines into a plurality of clusters.

20. The non-transitory computer-readable storage medium of claim 15 , wherein identifying the set of field types further comprises:

determining, for each line of the one or more lines, a corresponding line type derived from a corresponding set of field types of the one or more fields comprised by the line.

Assignments (2)
SECURITY INTEREST Recorded Aug 14, 2023
From: ABBYY INC.; ABBYY USA SOFTWARE HOUSE INC.; ABBYY DEVELOPMENT INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 064730/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 11, 2022
From: LANIN, MIKHAIL; SEMENOV, STANISLAV
To: ABBYY DEVELOPMENT INC.
Reel/Frame 058618/0426 →
Cited By (1)
US 12,639,971