IP Library Granted Patent US 9,098,581
Granted Patent B2
US 9,098,581 · App. 11/995,650 · Granted Aug 4, 2015

Method for finding text reading order in a document

Inventors: Sherif Yacoub (Barcelona, ES); Daniel Ortega (Barcelona, ES); Paolo Faraboschi (Barcelona, ES); Jose Abad Peiro (Barcelona, ES)
Assignee: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
G06F17/30864G06K9/00463G06F17/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,098,581
App. No.
11/995,650
Granted
Aug 4, 2015
Kind
B2
Abstract

Example methods for finding a text reading order in a document are described in which text zones are determined, the text zones are clustered using semantic measure and correlation and a reading order is found within each of the clusters.

Claims (39)

1. A method of finding a text reading order in a document, the method comprising, with a processor executing instructions from a memory:

determining a plurality of text zones within the document;

for each text zone, generating a semantic measure of text within the text zone; and

calculating a plurality of pairwise correlations between respective zone pairs;

clustering the plurality of text zones into a plurality of clusters in dependence upon the plurality of pairwise correlations; and

determining a separate reading order for the plurality of text zones within each respective cluster.

2. The method of claim 1 , wherein the semantic measure is multi-valued, and a first pairwise correlation of the plurality of pairwise correlations is determined in dependence upon a combined similarity of respective values between a first text zone of each pair and a second text zone.

3. The method of claim 1 , wherein the plurality of clusters are determined by:

(a) selecting an initial text zone of the plurality of text zones as the start of a first cluster of the plurality of clusters;

(b) adding to the first cluster any of the plurality of text zones which has a first pairwise correlation of the plurality of pairwise correlations of greater than a cut-off value with a second text zone already in the first cluster;

(c) repeating (b) until no further text zones are added to the first cluster; and

(d) repeating (a) to (c) until there are no more clusters to create.

4. The method of claim 3 , wherein the cut-off value is user selectable.

5. The method of claim 1 , further comprising, after the clustering, assigning to an existing cluster any text zone which has not already been assigned to a cluster by the clustering.

6. The method of claim 5 , wherein a non-clustered zone is assigned to an existing cluster based on layout proximity.

7. The method of claim 1 , further comprising determining whether any two clusters have interleaving zones, and if so merging the two clusters.

8. The method of claim 1 , further comprising determining whether the text zones of a given cluster, considered in order of layout flow, are interrupted by one or more text zones of another cluster, and if so splitting the given cluster at a point of interruption.

9. The method of claim 1 , further comprising, after the generating, excluding unwanted text zones from consideration in dependence upon their respective semantic measures.

10. A computer program comprising a series of computer-operable instructions recorded on a non-transitory computer-readable medium that, when executed by a processor, cause the processor to perform:

determining a plurality of text zones within a document;

for each text zone, generating a semantic measure of text within the text zone; and

calculating a plurality of pairwise correlations between respective zone pairs;

clustering the plurality of text zones into a plurality of clusters in dependence upon the plurality of pairwise correlations; and

determining a separate reading order for the plurality of text zones within each respective cluster.

11. A method of finding a text reading order in a document that has been rendered into electronic form, the method comprising, with a computing system:

for each of a plurality of text zones within the document, generating a semantic measure of text within that text zone;

calculating a correlation factor for each of a number of pairs of zones from within the plurality of text zones, the correlation factor being based on the semantic measure for each text zone in a zone pair;

clustering the plurality of text zones into a plurality of clusters by comparing the correlation factor generated for the number of pairs of zones to a minimum threshold; and

determining a separate reading order for the plurality of text zones within each of the respective clusters.

12. The method of claim 11 , wherein the semantic measure is based on any of a number of characters, a number of strings, a number of spaces, an average word length, a presence of particular words, an absence of particular words, classes or words, a type of punctuation and a density of punctuation.

13. The method of claim 11 , further comprising discarding a number of the text zones prior to the clustering based on the semantic measure.

14. The method of claim 11 , further comprising, for a first text zone of the plurality of text zones not assigned to any plurality of clusters by the clustering, assigning the first text zone to a first cluster of the plurality of clusters based on layout proximity.

15. The method of claim 11 , further comprising, after the clustering, merging two clusters into a single cluster if the two clusters have interleaving zones.

16. The method of claim 11 , further comprising, after the clustering, splitting a first cluster of the plurality of clusters into two separate clusters if a first text zone from a second cluster of the plurality of clusters interrupts the layout flow of the first cluster.

17. A non-transitory computer-readable memory having a computer program stored thereon to determine a text reading order within an electronically rendered document, that when executed by a processor of a computer causes the computer to:

for each of a plurality of text zones in the document, generate a semantic measure of text within the text zone;

calculate a correlation factor for pairs of zones from within the plurality of text zones, the correlation factor being based on the semantic measure for each text zone in the zone pair;

cluster the plurality of text zones into a plurality of clusters by comparing the correlation factors generated for the pairs of zones to a minimum threshold; and

determine a separate reading order for the plurality of text zones within each of the respective clusters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2008
From: HEWLETT-PACKARD ESPANOLA, S.L.
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 020687/0907 →
Continuity (1)
Related Publication 20100198827A1 · Aug 5, 2010