IP Library Granted Patent US 10,242,257
Granted Patent B2
US 10,242,257 · App. 15/639,399 · Granted Mar 26, 2019

Methods and devices for extracting text from documents

Inventors: Raghavendra Hosabettu (Bangalore, IN); Sendil Kumar Jaya Kumar (Bangalore, IN); Raghottam Mannopantar (Bangalore, IN)
Assignee: Wipro Limited
G06K9/00449G06K9/00463G06K9/325
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,242,257
App. No.
15/639,399
Granted
Mar 26, 2019
Kind
B2
Abstract

Methods, devices, and non-transitory computer readable storage media for extracting text from documents are disclosed. The method includes performing layout analysis on the document to identify a plurality of regions within a plurality of pages in the document. The method further includes identifying a table region from within the plurality of regions based on homogeneity between a plurality of textual lines in a page from the plurality of pages. The method includes identifying at least two rows and at least two columns within the table region. The method further includes identifying a plurality of cells within the table region based on the at least two rows and the at least two columns. The method includes extracting text from each of the plurality of cells.

Claims (59)

1. A method for extracting text from a document, the method comprising:

performing, by a text extraction device, a layout analysis on the document to identify a plurality of regions within a plurality of pages in the document;

identifying, by the text extraction device, a table region from within the plurality of regions based on homogeneity between a plurality of textual lines in a page from the plurality of pages, wherein the homogeneity is computed based on a plurality of preselected textual parameters associated with the plurality of textual lines, wherein identifying the table region further comprises:

determining, for each textual line in a page column within the page, values for the plurality of preselected textual parameters comprising at least one of pixel length, number of words, total pixel space between adjacent words, or number of characters;

computing, for each textual line in the page column, a variance of value of at least one of the plurality of preselected textual parameter from an associated average parameter value determined for all textual lines within the page column;

determining, for each textual line in the page column, an average variance based on the variance computed for each of the at least one of the plurality of preselected textual parameters; and

computing, for each textual line in the page column, a covariance based on a difference between the average variance of each textual line and an associated contiguous textual line within the page column;

identifying, by the text extraction device, at least two rows and at least two columns within the table region;

identifying, by the text extraction device, a plurality of cells within the table region based on the at least two rows and the at least two columns; and

extracting, by the text extraction device, text from each of the plurality of cells.

2. The method of claim 1 , wherein the plurality of regions comprises at least one of at least one header, at least one footer, at least one page column, or at least one image.

3. The method of claim 2 , wherein at least one page column is identified based on a threshold number of characters and a threshold number of words associated with a page column width.

4. The method of claim 1 , wherein identifying the table region further comprises:

computing a homogeneity index for each textual line in a page column within the page, wherein a homogeneity index for a textual line is computed based on a number of characters in the textual line and the plurality of preselected textual parameters; and

identifying a set of contiguous textual lines having a same homogeneity index, wherein the set of contiguous textual lines form the table region.

5. The method of claim 1 , wherein identifying the at least two rows and the at least two columns within the table region further comprises:

identifying a plurality of sets of contiguous pixels comprising a predefined color within the table region;

comparing each of the plurality of sets of contiguous pixels along the horizontal direction of the document with a minimum row pixel threshold, to identify the at least two rows; and

comparing each of the plurality of sets of contiguous pixels along the vertical direction of the document with a minimum column pixel threshold, to identify the at least two columns.

6. A text extraction device for extracting text from a document, the text extraction device comprises:

a processor; and

a memory communicatively coupled to the processor, wherein the memory stores instructions, which, on execution by the processor, causes the processor to:

perform a layout analysis on the document to identify a plurality of regions within a plurality of pages in the document;

identify a table region from within the plurality of regions based on homogeneity between a plurality of textual lines in a page from the plurality of pages, wherein the homogeneity is computed based on a plurality of preselected textual parameters associated with the plurality of textual lines, wherein identifying the table region further comprises:

determine for each textual line in a page column within the page, values for the plurality of preselected textual parameters comprising at least one of pixel length, number of words, total pixel space between adjacent words, or number of characters;

compute, for each textual line in the page column, a variance of value of at least one of the plurality of preselected textual parameter from an associated average parameter value determined for all textual lines within the page column;

determine, for each textual line in the page column, an average variance based on the variance computed for each of the at least one of the plurality of preselected textual parameters; and

compute, for each textual line in the page column, a covariance based on a difference between the average variance of each textual line and an associated contiguous textual line within the page column;

identify at least two rows and at least two columns within the table region;

identify a plurality of cells within the table region based on the at least two rows and the at least two columns; and

extract text from each of the plurality of cells.

7. The text extraction device of claim 6 , wherein the plurality of regions comprises at least one of at least one header, at least one footer, at least one page column, or at least one image.

8. The text extraction device of claim 7 , wherein at least one page column is identified based on a threshold number of characters and a threshold number of words associated with a page column width.

9. The text extraction device of claim 6 , wherein the instructions, on execution by the processor, further cause the processor to:

compute a homogeneity index for each textual line in a page column within the page, wherein a homogeneity index for a textual line is computed based on a number of characters in the textual line and the plurality of preselected textual parameters; and

identify a set of contiguous textual lines having a same homogeneity index, wherein the set of contiguous textual lines form the table region.

10. The text extraction device of claim 6 , wherein the instructions, on execution by the processor, further cause the processor to:

identify a plurality of sets of contiguous pixels comprising a predefined color within the table region;

compare each of the plurality of sets of contiguous pixels along the horizontal direction of the document with a minimum row pixel threshold, to identify the at least two rows; and

compare each of the plurality of sets of contiguous pixels along the vertical direction of the document with a minimum column pixel threshold, to identify the at least two columns.

11. A non-transitory computer-readable storage medium comprising a set of executable instructions stored thereon that, when executed by one or more processors, cause the processors to:

perform a layout analysis on the document to identify a plurality of regions within a plurality of pages in the document;

identify a table region from within the plurality of regions based on homogeneity between a plurality of textual lines in a page from the plurality of pages, wherein the homogeneity is computed based on a plurality of preselected textual parameters associated with the plurality of textual lines, wherein identifying the table region further comprises:

determine for each textual line in a page column within the page, values for the plurality of preselected textual parameters comprising at least one of pixel length, number of words, total pixel space between adjacent words, or number of characters;

compute, for each textual line in the page column, a variance of value of at least one of the plurality of preselected textual parameter from an associated average parameter value determined for all textual lines within the page column;

determine, for each textual line in the page column, an average variance based on the variance computed for each of the at least one of the plurality of preselected textual parameters; and

compute, for each textual line in the page column, a covariance based on a difference between the average variance of each textual line and an associated contiguous textual line within the page column;

identify at least two rows and at least two columns within the table region;

identify a plurality of cells within the table region based on the at least two rows and the at least two columns; and

extract text from each of the plurality of cells.

12. The non-transitory computer-readable storage medium of claim 11 , wherein the plurality of regions comprises at least one of at least one header, at least one footer, at least one page column, or at least one image.

13. The non-transitory computer-readable storage medium of claim 12 , wherein at least one page column is identified based on a threshold number of characters and a threshold number of words associated with a page column width.

14. The non-transitory computer-readable storage medium of claim 11 , wherein the instructions, when executed by the processors, further causes the processor to:

compute a homogeneity index for each textual line in a page column within the page, wherein a homogeneity index for a textual line is computed based on a number of characters in the textual line and the plurality of preselected textual parameters; and

identify a set of contiguous textual lines having a same homogeneity index, wherein the set of contiguous textual lines form the table region.

15. The non-transitory computer-readable storage medium of claim 11 , wherein the instructions, when executed by the processors, further causes the processor to:

identify a plurality of sets of contiguous pixels comprising a predefined color within the table region;

compare each of the plurality of sets of contiguous pixels along the horizontal direction of the document with a minimum row pixel threshold, to identify the at least two rows; and

compare each of the plurality of sets of contiguous pixels along the vertical direction of the document with a minimum column pixel threshold, to identify the at least two columns.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2017
From: HOSABETTU, RAGHAVENDRA; KUMAR, SENDIL KUMAR JAYA; MANNOPANTAR, RAGHOTTAM
To: WIPRO LIMITED
Reel/Frame 043136/0368 →
Priority Claims (1)
IN 201741017499 · May 18, 2017 · national
Continuity (1)
Related Publication 20180336404A1 · Nov 22, 2018
Cited By (1)
US 12,260,662