IP Library › Granted Patent US 7,676,089
Granted Patent B2
US 7,676,089 · App. 11/362,755 · Granted Mar 9, 2010

Document layout analysis with control of non-character area

Assignee: Ricoh Company, Ltd.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,676,089
App. No.
11/362,755
Granted
Mar 9, 2010
Kind
B2
Abstract

An apparatus, method, system, computer program and product, each capable of applying document layout analysis to a document image with control of a non-character area. A non-character area is extracted from a document image to be processed. A character image is generated from the document image by removing the non-character area from the document image. The character image is segmented into a plurality of sections to generate a segmented image. The segmented image is adjusted using a selected component of the non-character image to generate an adjusted segmented image. A segmentation result is output, which is generated based on the adjusted segmented image.

Claims (50)

1. An image processing apparatus, comprising:

means for inputting a document image to be processed;

means for extracting a non-character area from the document image, wherein the non-character area comprises a line;

means for generating a character image by removing the non-character area from the document image;

means for segmenting the character image into a plurality of sections to generate a segmented image;

means for adjusting the segmented image using an area previously occupied by the previously removed line of the non-character area as a separator to separate a merged section of the plurality of sections of the segmented image into different sections to generate an adjusted segmented image in which the merged section is separated into the different sections; and

means for outputting a segmentation result generated based on the adjusted segmented image.

2. The apparatus of claim 1 , wherein the character image is a binary image comprising a plurality of white pixels and a plurality of black pixels.

3. The apparatus of claim 2 , wherein the means for segmenting comprises:

means for extracting one or more maximal white rectangles covering the plurality of white pixels,

wherein the character image is segmented using the maximal white rectangles as a separator.

4. The apparatus of claim 1 , wherein the line of the non-character area comprises a rule line.

5. The apparatus of claim 4 , wherein the line of the non-character area further comprises a table line.

6. The apparatus of claim 1 , wherein the line of the non-character area comprises a table line.

7. An image processing method, comprising the steps of:

inputting a multivalue document image to be processed;

generating a binary document image from the multivalue document image;

extracting a non-character area from the binary document image, wherein the non-character area comprises a line;

removing the non-character area from the binary document image to generate a character image;

generating a segmented image by segmenting the character image into a plurality of sections;

adjusting the segmented image using an area previously occupied by the previously removed line of the non-character area as a separator that separates a merged section of the plurality of sections of the segmented image into different sections to generate an adjusted segmented image in which the merged section is separated into the different sections, wherein the merged section originally included the line of the non-character area that was previously removed from the binary document image; and

outputting a segmentation result generated based on the adjusted segmented image.

8. An image processing method, comprising the steps of:

inputting a multivalue document image to be processed;

extracting a non-character area from the multivalue document image, wherein the non-character area comprises a line;

generating a binary document image from the multivalue document image;

removing the non-character area from the binary document image to generate a character image;

generating a segmented image by segmenting the character image into a plurality of sections;

adjusting the segmented image using an area previously occupied by the previously removed line of the non-character area as a separator that separates a merged section of the plurality of sections of the segmented image into different sections to generate an adjusted segmented image in which the merged section is separated into the different sections, wherein the merged section originally included the line of the non-character area that was previously removed from the binary document image; and

outputting a segmentation result generated based on the adjusted segmented image.

9. An image processing system, comprising:

a processor; and

a storage device configured to store a plurality of instructions which, when executed by the processor, causes the processor to perform at least one function of the plurality of functions, the plurality of functions comprising:

inputting a document image to be processed;

extracting a non-character area from the document image, wherein the non-character area comprises a selected component;

generating a character image by removing the non-character area from the document image;

segmenting the character image into a plurality of sections to generate a segmented image;

adjusting the segmented image using the area previously occupied by the selected component of the previously removed non-character area as a separator that separates a merged section of the plurality of sections of the segmented image into different sections to generate an adjusted segmented image in which the merged section is separated into the different sections, wherein the merged section originally included the selected component of the non-character area that was previously removed from the document image; and

outputting a segmentation result generated based on the adjusted segmented image.

10. The system of claim 9 , further comprising:

a scanner configured to scan an original document image into the document image to be processed.

11. The system of claim 10 , further comprising:

an output device configured to output the segmentation result in a form visible to a user.

12. A computer readable medium storing computer instructions for performing an image processing method comprising the steps of:

inputting a document image to be processed;

extracting a non-character area from the document image, wherein the non-character area comprises a line;

generating a character image by removing the non-character area from the document image;

segmenting the character image into a plurality of sections to generate a segmented image;

adjusting the segmented image using the line of the non-character area that is previously removed from the document image as a separator that separates a merged section of the plurality of sections of the segmented image into different sections to generate an adjusted segmented image in which the merged section is separated into the different sections, wherein the merged section originally has the line of the non-character area that is previously removed from the document image; and

outputting a segmentation result generated based on the adjusted segmented image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2006
From: NISHIDA, HIROBUMI
To: RICOH COMPANY, LTD.
Reel/Frame 017865/0102 →
Priority Claims (1)
JP 2005-064513 · Mar 8, 2005 · national
Continuity (1)
Related Publication 20060204095A1 · Sep 14, 2006