IP Library Granted Patent US 7,643,682
Granted Patent B2
US 7,643,682 · App. 11/405,771 · Granted Jan 5, 2010

Method of identifying redundant text in an electronic document

Assignee: PDFlib GmbH
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,643,682
App. No.
11/405,771
Granted
Jan 5, 2010
Kind
B2
Abstract

A method of identifying redundant text fragments, which create artificial artifacts only, in an electronic page description language document includes a) providing a page having a plurality of text fragments, each text fragment comprising at least one glyph, the document including Unicode values for all glyphs and geometric information of all text fragments on the page and page description language parameters of all glyphs, b) identifying two text fragments as redundant candidates, if the Unicode sequence of the text fragments have identical corresponding Unicode sequences, c) defining a bounding box of quadrangular shape for each of the two redundant candidates according to their font characteristics, d) calculating the overlapping area of the two bounding boxes, and e) determining whether the two candidates form redundant text fragments by comparing the ratio of the overlapping area to the area of the smaller bounding box of both text fragments with a predetermined threshold.

Claims (17)

1. A method of identifying redundant text fragments, which create artificial artifacts only, in an electronic document, comprising: operating a computer to carry out the following steps

a) providing an electronic document being described in a page description language, the document comprising at least one page having a plurality of text fragments, each text fragment comprising at least one glyph, the document further comprising Unicode values for all glyphs as well as geometric information including position and width of all text fragments on the page and page description language parameters of all glyphs that include at least one of font size, character spacing, and text distortion;

b) identifying two text fragments as redundant candidates, if corresponding Unicode sequences of the two text fragments are identical;

c) defining a bounding box of quadrangular shape for each of the two redundant candidates according to their font characteristics wherein the height of the bounding box is essentially equal to the font size of the first glyph in a text fragment, and wherein the width of the bounding box is essentially equal to the accumulated widths of all glyphs in the text fragment;

d) calculating the overlapping area of the two bounding boxes; and

e) determining whether the two candidates form redundant text fragments wherein a ratio of the overlapping area to the area of the smaller bounding box of both text fragments is calculated and the ratio is compared with a predetermined threshold.

2. The method of identifying redundant text fragments according to claim 1 , further comprising sorting the text fragments on the page according to their x/y position.

3. The method of identifying redundant text fragments according to claim 1 , wherein the predetermined threshold is between 0.5 and 0.7.

4. The method of identifying redundant text fragments according to claim 3 , wherein the predetermined threshold is between 0.55 and 0.65.

5. The method of identifying redundant text fragments according to claim 1 , further comprising discarding one of the two redundant text fragments for further text processing steps.

6. The method of identifying redundant text fragments according to claim 5 wherein the discarded redundant text fragment has a lower page index according to an original page description.

7. A program storage device readable by a computer, tangibly embodying a program of instructions executable by the computer to perform an operation of identifying redundant text fragments in an electronic document, the operation comprising the steps of:

a) providing an electronic document being described in a page description language, the document comprising at least one page having a plurality of text fragments, each text fragment comprising at least one glyph, the document further comprising Unicode values for all glyphs as well as geometric information including position and width of all text fragments on the page and page description language parameters of all glyphs that include at least one of font size, character spacing, and text distortion;

b) identifying two text fragments as redundant candidates, if corresponding Unicode sequences of the two text fragments are identical;

c) defining a bounding box of quadrangular shape for each of the two redundant candidates according to their font characteristics wherein a height of the bounding box is essentially equal to the font size of the first glyph in the text fragment, and wherein a width of the bounding box is essentially equal to accumulated widths of all glyphs in the text fragment;

d) calculating the overlapping area of the two bounding boxes; and

e) determining whether the two candidates form redundant text fragments wherein a ratio of the overlapping area to the area of the smaller bounding box of both text fragments is calculated and this ratio is compared with a predetermined threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2009
From: BRONSTEIN, SERGE
To: PDFLIB GMBH
Reel/Frame 023527/0152 →
Priority Claims (1)
EP 05012452 · Jun 9, 2005 · regional
Continuity (1)
Related Publication 20060282769A1 · Dec 14, 2006