IP Library Granted Patent US 8,260,049
Granted Patent B2
US 8,260,049 · App. 12/236,054 · Granted Sep 4, 2012

Model-based method of document logical structure recognition in OCR systems

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,260,049
App. No.
12/236,054
Granted
Sep 4, 2012
Kind
B2
Abstract

In one embodiment, the invention provides a method for determining a logical structure of a document. The method comprises generating at least one document hypothesis for the whole document; for each document hypothesis, verifying said document hypothesis including (a) generating at least one block hypothesis for each block in the document based on the document hypothesis; and (b) selecting a best block hypothesis for each block; selecting as a best document hypothesis the document hypothesis that has the best degree of correspondence with the selected best block hypotheses for the document; and forming the document based on the best document hypothesis.

Claims (43)

1. A method for determining a logical structure of a document, the method comprising:

acquiring an image of the document;

identifying one or more blocks in the image of the document;

generating a hypothesis for at least one of the identified blocks in the image of the document (a “block hypothesis”);

generating at least one document hypothesis for the image of the document, wherein said generating included referencing a plurality of document models, wherein each document model describes one or more possible logical structures, and wherein such logical structures are based on the presence of one or more blocks;

selecting a document hypothesis based on its degree of correspondence with at least one block hypothesis; and

forming a representation of the document based on the selected document hypothesis.

2. The method of claim 1 , wherein the generating the at least one document hypothesis for the image of the document includes generating a plurality of document hypotheses in order of differing probabilities.

3. The method of claim 1 ,

wherein indentifying the one or more blocks includes performing a physical structure analysis to identify each said one or more blocks.

4. The method of claim 1 , the method further comprising:

saving the representation of the document in an extended format in a memory.

5. The method of claim 1 , wherein the generating the at least one document hypothesis for the image of the document is based on information about a possible arrangement of blocks in the image of the document.

6. The method of claim 1 , wherein said generating the hypothesis for at least one of the identified blocks in the document is based on the at least one document hypothesis.

7. The method of claim 1 , wherein generating the hypothesis for at least one of the identified blocks includes generating a hypothesis for each identified block in the image of the document.

8. The method of claim 1 , wherein a block comprises one or more form elements.

9. A system comprising:

a processor; and

a memory coupled to the processor, the memory storing instructions which when executed by the processor, cause the system to perform a method for determining a logical structure for a document, comprising:

identifying one or more blocks in an image of the document;

generating a hypothesis for at least one of the identified blocks in the image of the document;

generating at least one document hypothesis for the image of the document, wherein said generating includes referencing a plurality of document models, wherein each document model describes one or more possible logical structures, and wherein such logical structures are based on the presence of one or more blocks;

selecting a document hypothesis based on its degree of correspondence with at least one block hypothesis; and

forming a representation of the document based on the selected document hypothesis.

10. The system of claim 1 , wherein the generating the at least one document hypothesis for the document includes generating a plurality of document hypotheses in order of differing probabilities.

11. The system of claim 1 , wherein wherein identifying the one or more blocks includes performing a physical structure analysis to identify each said one or more blocks.

12. The system of claim 1 , wherein the method further comprises saving the representation of the document in an extended format in a memory.

13. The system of claim 1 , wherein the document hypothesis is selected automatically.

14. The system of claim 1 , wherein the document hypothesis is selected by receiving a selection through a user interface element.

15. The system of claim 1 , wherein the generating the at least one document hypothesis for the image of the document is based on information about a possible arrangement of blocks in the image of the document.

16. A non-transitory computer-readable medium having stored thereon instructions which when executed by a computer, cause the computer to perform a method for determining a logical structure for a document, comprising:

identifying one or more blocks in an image of the document;

generating a hypothesis for at least one of the identified blocks in the image of the document;

generating at least one document hypothesis for the image of the document, wherein said generating includes referencing a plurality of document models, wherein each document model describes one or more possible logical structures, and wherein such logical structures are based on the presence of one or more blocks;

selecting a document hypothesis based on its degree of correspondence with at least one block hypothesis; and

forming a representation of the document based on the selected document hypothesis.

17. The non-transitory computer-readable medium of claim 16 , wherein the generating includes generating a plurality of document hypotheses in order of differing probabilities.

18. The non-transitory computer-readable medium of claim 16 , wherein wherein identifying the one or more blocks includes performing a physical structure analysis to identify each said one or more blocks.

19. The non-transitory computer-readable medium of claim 16 , wherein the method further comprises saving a representation of the document in an extended format in a memory.

20. The non-transitory computer-readable medium of claim 16 , wherein the document hypothesis is selected automatically.

21. The non-transitory computer-readable medium of claim 16 , wherein the document hypothesis is selected by receiving a selection through a user interface element.

22. The method of claim 1 , wherein generating said at least one document hypothesis includes comparing the hypothesis for said at least one of the identified blocks against a block model representing a possible logical structure for a corresponding block.

23. The method of claim 1 , wherein generating said at least one document hypothesis includes comparing the hypothesis for said at least one of the identified blocks on the basis of a degree of correspondence between the hypothesis for said at least one of the identified blocks and each of a plurality of block models.

Assignments (5)
SECURITY INTEREST Recorded Aug 14, 2023
From: ABBYY INC.; ABBYY USA SOFTWARE HOUSE INC.; ABBYY DEVELOPMENT INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 064730/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: ABBYY PRODUCTION LLC
To: ABBYY DEVELOPMENT INC.
Reel/Frame 059249/0873 →
MERGER Recorded May 2, 2019
From: ABBYY PRODUCTION LLC; ABBYY DEVELOPMENT LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 049079/0942 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2013
From: ABBYY SOFTWARE LTD.
To: ABBYY DEVELOPMENT LLC
Reel/Frame 031085/0834 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2009
From: DERYAGIN, DMITRY; ANISIMOVICH, KONSTANTIN
To: ABBYY SOFTWARE LTD
Reel/Frame 022403/0553 →