IP Library Granted Patent US 7,305,612
Granted Patent B2
US 7,305,612 · App. 10/403,106 · Granted Dec 4, 2007

Systems and methods for automatic form segmentation for raster-based passive electronic documents

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,305,612
App. No.
10/403,106
Granted
Dec 4, 2007
Kind
B2
Abstract

Systems and methods for automatically extracting form information (document structure, elements, format, etc.) from electronic documents such as raster-based passive documents, and storing such form information in a file in accordance with a predetermined DTD (document type definition).

Claims (54)

1. A method for processing electronic documents, comprising the steps of:

receiving as input an electronic document, wherein the electronic document is a PDF (portable document format) file and wherein at least a portion of the electronic document is raster-based;

extracting form information from text portions and non-text portions of the electronic document, the form information including form lines and table boxes extracted from raster-based data; and

generating a structured document for the electronic document, wherein the structured document represents the extracted form information of the text portions and the extracted form information of the non-text portions in a well-defined, hierarchical structure based on a predefined document type definition,

wherein the step of extracting form information comprises the steps of:

segmenting text portions and non-text portions;

separately processing the text portions and non-text portions to extract associated form information; and

combining the extracted form information of the text portions and non-text portions, and

wherein the step of separately processing the non-text portions comprises the steps of:

extracting images and location information for the images;

converting grayscale and color extracted images to black and white images;

determining horizontal and vertical lines in the extracted images; and

determining table entries and form fields in the extracted images using the horizontal and vertical lines.

2. The method of claim 1 , wherein the step of processing the text portions of the electronic document comprises:

extracting text segments and corresponding location information of the extracted text segments; and

determining a context of each text segment using the location information.

3. The method of claim 2 , wherein the context of a text segment comprises of a paragraph, heading title and subheading.

4. The method of claim 1 , wherein the step of determining horizontal and vertical lines comprises the steps of:

performing a Hough Transformation to extract line information from an image;

determining lines in the image using the extracted line information; and

discarding lines having a slope that exceeds a predetermined threshold.

5. The method of claim 1 , further comprising the step of removing skew from the horizontal and vertical lines.

6. The method of claim 1 , wherein the step of determining table entries and form fields using the horizontal and vertical lines, comprises the steps of:

determining intersections between horizontal and vertical line pairs;

determining table entries from the determined intersections; and

associating horizontal lines that do not intersect other lines as part of form fields.

7. The method of claim 1 , wherein the structured document is based on SGML, XML or extensions thereof.

8. A program storage device readable by machine, tangibly embodying a program of instructions executable by the machine to perform method steps for processing electronic documents, the method steps comprising:

receiving as input an electronic document, wherein the electronic document is a PDF (portable document format) file and wherein at least a portion of the electronic document is raster-based;

extracting form information from text portions and non-text portions of the electronic document, the form information including form lines and table boxes extracted from raster-based data; and

generating a structured document for the electronic document, wherein the structured document represents the extracted form information of the text portions and the extracted form information of the non-text portions in a well-defined, hierarchical structure based on a predefined document type definition,

wherein the instruction for extracting form information comprise instructions for performing the method steps of:

segmenting text portions and non-text portions;

separately processing the text portions and non-text portions to extract associated form information; and

combining the extracted form information of the text portions and non-text portions, and

wherein the instructions for separately processing the non-text portions comprise instructions for performing the method steps of:

extracting images and location information for the images;

converting grayscale and color extracted images to black and white images;

determining horizontal and vertical lines in the extracted images; and

determining table entries and form fields in the extracted images using the horizontal and vertical lines.

9. The program storage device of claim 8 , wherein the instructions for processing the text portions of the electronic document comprise instructions for performing the method steps of:

extracting text segments and corresponding location information of the extracted text segments; and

determining a context of each text segment using the location information.

10. The program storage device of claim 9 , wherein the context of a text segment comprises one of a paragraph, heading, title and subheading.

11. The program storage device of claim 8 , wherein the instructions for determining horizontal and vertical lines comprise instructions for performing the method steps of:

performing a Hough Transformation to extract line information from an image;

determining lines in the image using the extracted line information; and

discarding lines having a slope that exceeds a predetermined threshold.

12. The program storage device of claim 8 , further comprising instructions for performing the step of removing skew from the horizontal and vertical lines.

13. The program storage device of claim 8 , wherein the instructions for determining table entries and form fields using the horizontal and vertical lines comprise instructions for performing the steps of:

determining intersections between horizontal and vertical line pairs;

determining table entries from the determined intersections; and

associating horizontal lines that do not intersect other lines as part of form fields.

14. The program storage device of claim 8 , wherein the structured document is based on SGML, XML or extensions thereof.

Assignments (1)
MERGER Recorded Apr 5, 2010
From: SIEMENS CORPORATE RESEARCH, INC.
To: SIEMENS CORPORATION
Reel/Frame 024185/0042 →