IP Library › Granted Patent US 9,223,756
Granted Patent B2
US 9,223,756 · App. 13/800,242 · Granted Dec 29, 2015

Method and apparatus for identifying logical blocks of text in a document

Inventor: Ram Bhushan Agrawal (Noida, IN)
Assignee: ADOBE SYSTEMS INCORPORATED
G06F17/21
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,223,756
App. No.
13/800,242
Granted
Dec 29, 2015
Kind
B2
Abstract

A computer implemented method and apparatus for identifying logical blocks of text in a document where document structure information is absent. The method comprises accessing a document, wherein the document comprises a plurality of words; identifying word information for each word in the plurality of words; creating a plurality of text lines based on the word information, wherein each text line in the plurality of text lines comprises one or more words in the plurality of words; and creating a plurality of text blocks derived from the plurality of text lines.

Claims (50)

1. A computer implemented method comprising:

accessing a document, wherein the document comprises a plurality of words;

calculating a threshold horizontal gap between words based only on the average vertical height of the plurality of words, wherein horizontal and vertical are orthogonal directions within the document;

creating a plurality of text lines from the plurality of words based on the threshold horizontal gap between words based on the average vertical height of the plurality of words, wherein each text line in the plurality of text lines comprises one or more words in the plurality of words; and

creating a plurality of text blocks derived from the plurality of text lines.

2. The method of claim 1 , further comprising:

repairing interfering text blocks in the plurality of text blocks by merging text blocks when a first text block comprises a second text block; and

identifying justification alignment of each text block in the plurality of text blocks.

3. The method of claim 1 , wherein the document comprises a position-based text layout format and lacks document structure information.

4. The method of claim 1 , further comprising identifying a position of word extents in an x,y coordinate plane and one or more Unicode characters representing a text of the word.

5. The method of claim 1 , wherein creating a plurality of text lines comprises comparing information about a current word with information about an immediately previous word, wherein the information comprises at least one of word information of the current word and the immediately previous word, a position of each word in the document, or a difference in relative heights of the words.

6. The method of claim 1 , wherein creating a plurality of text blocks comprises comparing information about a current text line with information about an immediately previous text line, wherein the information comprises at least one of a difference in relative heights of the current text line and the immediately previous text line, a size of a vertical gap between the current text line and the immediately previous text line, word information of the words in the current text line, or a horizontal vector line between the current text line and the immediately previous text line.

7. The method of claim 1 , wherein creating the plurality of text lines from the plurality of words comprises:

searching the document for one or more drop cap characters, one or more list labels, one or more list items, one or more superscripts, or one or more vertical lines between words; and

identifying words to include in one or more of the plurality of text lines based on the results of the search for the one or more drop cap characters, the one or more list labels, the one or more list items, the one or more superscripts, or the one or more vertical lines between words.

8. The method of claim 1 , wherein creating a plurality of text blocks from the plurality of text lines further comprises:

identifying a number of numerical digits in one or more text lines;

searching for one or more horizontal lines between text lines, one or more list labels, one or more list items, or one or more drop cap characters;

comparing the width of at least one text block to a pre-determined width threshold; and

identifying lines to include in one or more of the plurality of text blocks based on:

the identified number of numerical digits in a text line;

the results of the search for the one or more horizontal line between two text lines, the one or more list labels, the one or more list items, or the one or more drop cap characters; and

the comparison between the width of the at least one text block and the pre-determined width threshold.

9. A system comprising:

at least one processor; and

at least one non-transitory computer readable storage medium storing instructions thereon that, when executed by the at least one processor, cause the system to:

identify a plurality of text blocks in a document, wherein the document comprises a plurality of words;

calculate the average vertical height of the plurality of words, wherein horizontal and vertical are orthogonal directions within the document;

set a threshold horizontal gap equal to the calculated average vertical height of the plurality of words;

create a plurality of text lines from the plurality of words based on the threshold horizontal gap between words based on the average vertical height of the plurality of words, wherein each text line in the plurality of text lines comprises one or more words in the plurality of words; and

create a plurality of text blocks based on the plurality of text lines.

10. The system of claim 9 , wherein the document comprises a position-based text layout format and lacks document structure information.

11. The system of claim 9 , further comprising instructions stored thereon that, when executed by the at least one processor, cause the system to identify a position of word extents in an x,y coordinate plane and one or more Unicode characters representing a text of the word.

12. The system of claim 9 , wherein the instructions stored thereon that, when executed by the at least one processor, cause the system to:

create text lines by comparing information about a current word with information about an immediately previous word, wherein the information comprises at least one of word information of the current word and the immediately previous word, a position of each word in the document, or a difference in relative heights of the words, and

create text blocks by comparing information about a current text line with information about an immediately previous text line, wherein the information comprises at least one of a difference in relative heights of the current text line and the immediately previous text line, a size of a vertical gap between the current text line and the immediately previous text line, word information of the words in the current text line, or a horizontal vector line between the current text line and the immediately previous text line.

13. A non-transitory computer readable medium for storing computer instructions that, when executed by at least one processor cause the at least one processor to:

access a document, wherein the document comprises a plurality of words;

calculate a threshold horizontal gap between words based only on the average vertical height of the plurality of words, wherein horizontal and vertical are orthogonal directions within the document;

create a plurality of text lines from the plurality of words based on the threshold horizontal gap between words based on the average vertical height of the plurality of words, wherein each text line in the plurality of text lines comprises one or more words in the plurality of words; and

create a plurality of text blocks derived from the plurality of text lines.

14. The non-transitory computer readable medium of claim 13 , further comprising computer instructions that, when executed by at least one processor, cause the at least one processor to:

correct interfering text blocks in the plurality of text blocks; and

identify justification alignment of each text block in the plurality of text blocks.

15. The non-transitory computer readable medium of claim 13 , wherein the document comprises a position-based text layout format and lacks document structure information.

16. The non-transitory computer readable medium of claim 13 , further comprising instructions thereon that, when executed by the at least one processor, cause the at least one processor to identify a position of word extents in an x,y coordinate plane and one or more Unicode characters representing a text of the word.

17. The non-transitory computer readable medium of claim 13 , wherein the computer instructions, when executed by at least one processor, cause the at least one processor to create text lines by comparing information about a current word with information about an immediately previous word, wherein the information comprises at least one of word information of the current word and the immediately previous word, a position of each word in the document, or a difference in relative heights of the words.

18. The non-transitory computer readable medium of claim 13 , wherein the computer instructions, when executed by at least one processor, cause the at least one processor to create text blocks by comparing information about a current text line with information about an immediately previous text line, wherein the information comprises at least one of a difference in relative heights of the current text line and the immediately previous text line, a size of a vertical gap between the current text line and the immediately previous text line, word information of the words in the current text line, or a horizontal vector line between the current text line and the immediately previous text line.

19. The non-transitory computer readable medium of claim 14 , wherein the computer instructions, when executed by at least one processor, cause the at least one processor to correct interfering text blocks by merging text blocks when a first text block comprises a second text block and splitting a text block into two or more smaller text blocks to remove the interference.

20. The non-transitory computer readable medium of claim 14 , wherein the computer instructions, when executed by at least one processor, cause the at least one processor to justify alignment by classifying a text block as right-justified, left-justified, fully-justified, or centered, or splitting a text block into two or more smaller text blocks when a steep change a left or right margin of the text block or a plurality of paragraphs is identified.

Assignments (2)
CHANGE OF NAME Recorded Apr 8, 2019
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 048867/0882 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2013
From: AGRAWAL, RAM BHUSHAN
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 029991/0485 →
Continuity (1)
Related Publication 20140281939A1 · Sep 18, 2014