IP Library Granted Patent US 8,547,589
Granted Patent B2
US 8,547,589 · App. 12/470,425 · Granted Oct 1, 2013

Data capture from multi-page documents

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,547,589
App. No.
12/470,425
Granted
Oct 1, 2013
Kind
B2
Abstract

A method for processing a batch of scanned images is provided. The method comprises processing the scanned images into documents; for documents comprising multiple pages maintaining a page-based coordinate system to specify a location of structures within a page and joining the pages to form a multi-page sheet having a sheet-based coordinate system to specify a location of structures within the multi-page sheet; performing a data extraction operation to extract data from each document, said data extraction operation comprising a page mode wherein structures are detected on individual pages using the page-based coordinate system and a document mode wherein structures are detected within the entire document using the sheet-based coordinate system.

Claims (56)

1. A method for processing a batch of document images, the method comprising:

processing, by a computing device, the document images into one or more documents, wherein a document of the one or more documents includes multiple pages;

maintaining, by the computing device, a page-based coordinate system to specify a location of structures within individual pages of the document;

combining, by the computing device, the multiple pages to form a multi-page sheet, wherein a sheet-based coordinate system specifies a location of structures within the multi-page sheet; and

performing, by the computing device, a data extraction operation to extract data from the document, said data extraction operation including:

detecting the structures on individual pages using the page-based coordinate system;

defining a repeating group of fields, wherein the repeating group of fields is capable of flowing over from one page onto another page;

detecting whether all fields of an instance of the repeating group of fields are found on consecutive pages; and

depending on whether all fields of the instance of the repeating group of fields are found on consecutive pages, detecting structures using the sheet-based coordinate system, detecting structures within the document using the sheet-based coordinate system.

2. The method of claim 1 , wherein processing the document images into one or more documents includes identifying the one or more documents based on header and footer information in predefined flexible structure descriptions.

3. The method of claim 2 , wherein a flexible structure description for the document includes descriptions of data fields to be detected and descriptions of anchor elements and their relationships within documents of a given type.

4. The method of claim 1 , wherein detecting the structures on the individual pages includes detecting header information that spans more than one page as a repetitive field.

5. The method of claim 4 , wherein the header information is excluded from the data extraction operation.

6. The method of claim 4 , wherein the header information includes table header information associated with a table, and a running title associated with a page.

7. The method of claim 1 , wherein detecting the structures within the entire document using the sheet-based coordinate system includes:

detecting data in a table row as a repetitive field that spans several rows.

8. The method of claim 1 , wherein the sheet-based coordinate system includes a global system of coordinates for the document.

9. The method of claim 8 , wherein the sheet-based coordinate system supports parallel shifts, wherein each page includes its own shift or shifts.

10. A non-transitory computer-readable medium embodying a set of instructions which, when executed by a computer, cause the computer to:

process document images into one or more documents, wherein a document of the one or more documents includes multiple pages;

maintain a page-based coordinate system to specify a location of structures within individual pages of the document;

combine the multiple pages to form a multi-page sheet, wherein a sheet-based coordinate system specifies a location of structures within the multi-page sheet; and

perform, by the computing device, a data extraction operation to extract data from the document, said data extraction operation including:

detecting the structures on the individual pages using the page-based coordinate system;

defining a repeating group of fields, wherein the repeating group of fields is capable of flowing over from one page onto another page;

detecting whether all fields of an instance of the repeating group of fields are found on consecutive pages of the document; and

depending on whether all fields of the instance of the repeating group of fields are found on consecutive pages, detecting structures within the document using the sheet-based coordinate system.

11. The non-transitory computer-readable medium of claim 10 , wherein the instructions further cause the computer to:

identify the one or more documents based on header and footer information in predefined flexible structure descriptions.

12. The non-transitory computer-readable medium of claim 10 , wherein the instructions further cause the computer to:

detect header information that spans more than one page as a repetitive field.

13. The non-transitory computer-readable medium of claim 12 , wherein the header information includes table header information associated with a table or a running title associated with a page.

14. The non-transitory computer-readable medium of claim 10 , wherein the instructions further cause the computer to:

detect data in a table row as a repetitive field that spans several rows.

15. The non-transitory computer-readable medium of claim 10 , wherein a flexible structure description for the document includes descriptions of data fields to be detected and descriptions of anchor elements and their relationships within documents of a given type.

16. The non-transitory computer-readable medium of claim 10 , wherein the sheet-based coordinate system includes a global system of coordinates for the document.

17. The non-transitory computer-readable medium of claim 10 , wherein the sheet-based coordinate system supports parallel shifts, wherein each page includes its own shift or shifts.

18. A system for capturing data from a document image, the system comprising:

an imaging component capable of capturing the document image of a document;

a processor; and

a memory coupled to the processor and in electronic communication with the imaging component, the memory configured with instructions for causing the processor to:

process document images into one or more documents, wherein a document of the one or more documents includes multiple pages;

maintain a page-based coordinate system to specify a location of structures within individual pages of the document;

combine the multiple pages to form a multi-page sheet, wherein a sheet- based coordinate system specifies a location of structures within the multi-page sheet; and

perform a data extraction operation to extract data from the document, said data extraction operation including:

detecting the structures on the individual pages using the page- based coordinate system;

defining a repeating group of fields, wherein the repeating group of fields is capable of flowing over from one page onto another page;

detecting whether all fields of an instance of the repeating group of fields are found on consecutive pages of the document; and

depending on whether all fields of the instance of the repeating group of fields are found on consecutive pages, detecting structures within the document using the sheet-based coordinate system.

19. The system of claim 18 , which is operable to identify the one or more documents based on header and footer information in predefined flexible structure descriptions.

20. The system of claim 18 , which is operable to detect header information that spans more than one page as a repetitive field.

21. The system of claim 18 , wherein the header information includes table header information associated with a table or a running title associated with a page.

22. The system of claim 18 , which is operable to detect data in a table row as a repetitive field that spans several rows.

23. The system of claim 18 , wherein a flexible structure description for the document includes descriptions of data fields to be detected and descriptions of anchor elements and their relationships within documents of a given type.

24. The system of claim 18 , wherein the sheet-based coordinate system includes a global system of coordinates for the document.

25. The system of claim 18 , wherein the sheet-based coordinate system supports parallel shifts, wherein each page includes its own shift or shifts.

Assignments (5)
SECURITY INTEREST Recorded Aug 14, 2023
From: ABBYY INC.; ABBYY USA SOFTWARE HOUSE INC.; ABBYY DEVELOPMENT INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 064730/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: ABBYY PRODUCTION LLC
To: ABBYY DEVELOPMENT INC.
Reel/Frame 059249/0873 →
MERGER Recorded May 2, 2019
From: ABBYY PRODUCTION LLC; ABBYY DEVELOPMENT LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 049079/0942 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2013
From: ABBYY SOFTWARE LTD.
To: ABBYY DEVELOPMENT LLC
Reel/Frame 031085/0834 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 22, 2009
From: TUGANBAEV, DIAR; ZLOBIN, SERGEY; FILIMONOVA, IRINA
To: ABBYY SOFTWARE LTD
Reel/Frame 022728/0073 →