IP Library Granted Patent US 7,310,773
Granted Patent B2
US 7,310,773 · App. 10/346,795 · Granted Dec 18, 2007

Removal of extraneous text from electronic documents

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,310,773
App. No.
10/346,795
Granted
Dec 18, 2007
Kind
B2
Abstract

Method and apparatus for removing lines of extraneous text from a document. Similarities are identified between lines of text on each page and corresponding lines on a selected subset of pages. Different weight values are associated with different line numbers of text on a page, each weight value indicating a degree of likelihood that a line of text contains extraneous text. One or more lines of text are selectively removed from a page as a function of the similarities and associated weight values of line numbers of the lines of text.

Claims (51)

1. A method for removing lines of extraneous text from a plurality of pages of text, comprising:

identifying similarities between lines of text on each page and corresponding lines on a selected subset of pages;

associating different weight values with different line numbers of a page, each weight value indicating a degree of likelihood that a line of text at the line number contains extraneous text;

selectively removing one or more lines of text at corresponding line numbers from each of the plurality of the pages as a function of the similarities and associated weight values of line numbers of the lines of text; identifying non-body pages of text;

logically grouping body pages of text in a first set of pages and non-body pages in a second set of pages;

wherein the selected subset of pages used in identifying similarities between corresponding lines of text includes pages of the first set and the second set; and

identifying similarities between lines of text of each page in the second set and corresponding lines on a selected subset of the second set of pages.

2. The method of claim 1 , further comprising:

receiving input data in which each word has associated attributes that indicate a page-relative bounding box; and

for each word for which a selected portion of a height of the page-relative bounding box overlaps a line, assigning the word to the line on a page.

3. The method of claim 1 , wherein identifying similarities comprises comparing for equality of characters at corresponding positions of corresponding lines of text.

4. The method of claim 3 , wherein identifying similarities comprises quantifying similarities as a function of a number of equal characters.

5. The method of claim 4 , wherein identifying similarities comprises quantifying similarities as a function of bounding boxes of the corresponding lines.

6. The method of claim 3 , wherein identifying similarities comprises processing all digit characters as being equal to a selected character.

7. The method of claim 1 , wherein each page has a top and a bottom, the method further comprising selecting a number of lines of text at the top and bottom of each page for identifying similarities.

8. An apparatus for removing lines of extraneous text from a plurality of pages of text, comprising:

means for identifying similarities between lines of text on each page and corresponding lines on a selected subset of pages;

means for associating different weight values with different line numbers of a page, each weight value indicating a degree of likelihood that a line of text at the line number contains extraneous text;

means for selectively removing one or more lines of text at corresponding line numbers from each of the plurality of the pages as a function of the similarities and associated weight values of line numbers of the lines of text; means for identifying non-body pages of text;

means for logically grouping body pages of text in a first set of pages and non-body pages in a second set of pages;

wherein the selected subset of pages used in identifying similarities between corresponding lines of text includes pages of the first set and the second set; and

means for identifying similarities between lines of text of each page in the second set and corresponding lines on a selected subset of the second set of pages.

9. A document processing system, comprising:

at least one data processing system, wherein the at least one data processing system is configured with: a memory;

a document retrieval component arranged to obtain an electronic document including a plurality of pages of text;

a text extraction component coupled to the document retrieval component, the text extraction component configured to identify similarities between lines of text on each page and corresponding lines on a selected subset of pages, associate different weight values with different line numbers of text on a page, each weight value indicating a degree of likelihood that a line of text at the line number contains extraneous text, and selectively remove one or more lines of text at corresponding line numbers from each of the plurality of the pages as a function of the similarities and associated weight values of line numbers of the lines of text; and

an application coupled to the text extraction component, the application configured to process the document without the extraneous text; wherein the text extraction component is configured to identify non-body pages of text, logically group body pages of text in a first set of pages and non-body pages in a second set of pages, wherein the selected subset of pages used in identifying similarities between corresponding lines of text includes pages of the first set and the second set, and identify similarities between lines of text of each page in the second set and corresponding lines on a selected subset of the second set of pages.

10. The system of claim 9 , wherein:

the document retrieval component is configured to generate words of text, each word having an associated page-relative bounding box; and

the text extraction is configured to, for each word for which a selected portion of a height of the page-relative bounding box overlaps a line, assign the word to the line on a page.

11. The system of claim 9 , wherein the text extraction component is configured to compare for equality of characters at corresponding positions of corresponding lines of text.

12. The system of claim 11 , wherein the text extraction component is configured to quantify similarities as a function of a number of equal characters.

13. The system of claim 12 , wherein the text extraction component is configured to quantify similarities as a function of bounding boxes of the corresponding lines.

14. The system of claim 11 , wherein the text extraction component is configured to process all digit characters as being equal to a selected character.

15. The system of claim 9 , wherein each page has a top and a bottom, the system further comprising selecting a number of lines of text at the top and bottom of each page for identifying similarities.

16. An article of manufacture, comprising:

a computer-readable medium configured with instructions for causing a computer to remove lines of extraneous text from a plurality of pages of text by performing the steps of,

identifying similarities between lines of text on each page and corresponding lines on a selected subset of pages;

associating different weight values with different line numbers of text on a page, each weight value indicating a degree of likelihood that a line of text at the line number contains extraneous text;

selectively removing one or more lines of text at corresponding line numbers from each of the plurality of the pages as a function of the similarities and associated weight values of line numbers of the lines of text; identifying non-body pages of text;

logically grouping body pages of text in a first set of pages and non-body pages in a second set of pages;

wherein the selected subset of pages used in identifying similarities between corresponding lines of text includes pages of the first set and the second set; and

identifying similarities between lines of text of each page in the second set and corresponding lines on a selected subset of the second set of pages.

17. The article of manufacture of claim 16 , wherein the computer-readable medium is further configured to cause a computer to perform the steps, further comprising:

receiving input data in which each word has associated attributes that indicate a page-relative bounding box; and

for each word for which a selected portion of a height of the page-relative bounding box overlaps a line, assigning the word to the line on a page.

18. The article of manufacture of claim 16 , wherein identifying similarities comprises comparing for equality of characters at corresponding positions of corresponding lines of text.

19. The article of manufacture of claim 18 , wherein identifying similarities comprises quantifying similarities as a function of a number of equal characters.

20. The article of manufacture of claim 19 , wherein identifying similarities comprises quantifying similarities as a function of bounding boxes of the corresponding lines.

21. The article of manufacture of claim 18 , wherein identifying similarities comprises processing all digit characters as being equal to a selected character.

22. The article of manufacture of claim 16 , wherein each page has a top and a bottom, and the steps further comprise selecting a number of lines of text at the top and bottom of each page for identifying similarities.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2021
From: OT PATENT ESCROW, LLC
To: VALTRUS INNOVATIONS LIMITED
Reel/Frame 056157/0492 →
PATENT ASSIGNMENT, SECURITY INTEREST, AND LIEN AGREEMENT Recorded Jan 26, 2021
From: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP; HEWLETT PACKARD ENTERPRISE COMPANY
To: OT PATENT ESCROW, LLC
Reel/Frame 055269/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2015
From: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 037079/0001 →