IP Library Granted Patent US 9,569,413
Granted Patent B2
US 9,569,413 · App. 13/465,833 · Granted Feb 14, 2017

Document text processing using edge detection

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,569,413
App. No.
13/465,833
Granted
Feb 14, 2017
Kind
B2
Abstract

A document is received that has a plurality of lines with text. This document includes text associated with at least one topic of interest and text not associated with the at least one topic of interest. Thereafter, it is determined, for each line in the document, a length of the line and a number of off-topic indicators with the off-topic indicators characterizing portions of the document as likely being not being associated with the at least one topic of interest. Thereafter, a density for each line can be determined based on the determined line length and the determined number of off-topic indicators. The determined densities for each line are used to identify portions of the documents likely associated with the at least one topic of interest so that data characterizing the identified portions of the document can be provided. Related apparatus, systems, techniques and articles are also described.

Claims (37)

1. A method comprising:

receiving a document having a plurality of lines with text, the document comprising text associated with at least one topic of interest and text not associated with the at least one topic of interest;

determining, for each line in the document, a length of the line and a number of off-topic indicators, the off-topic indicators characterizing portions of the document as likely being not being associated with the at least one topic of interest;

determining a density for each line based on the determined line length and the determined number of off-topic indicators, the density for each line being proportional to the determined line length and inversely proportional to an exponential base greater than one of the number of off-topic indicators;

identifying, using the determined densities for each line, portions of the document likely associated with the at least one topic of interest; and

providing data characterizing the identified portions of the document.

2. A method as in claim 1 , wherein, the determined number of off-topic indicators is based on a number of hyperlinks in the line.

3. A method as in claim 1 , wherein, the determined number of off-topic indicators is based on one or more of: a number of formatting portions, spaces, font tags, font information, HTML Tags, spacing, non-printed text formatting and spacing elements, or characters found within the line.

4. A method as in claim 2 , wherein the number of off-topic indicators for each line is proportional to a weighted sum of the number of hyperlinks and a number of two consecutive space characters.

5. A method as in claim 1 , further comprising: applying a density smoothing filter to each line to result in smoothed densities, and wherein the portions of the documents likely associated with the at least one topic of interest are identified using the smoothed densities.

6. A method as in claim 1 , wherein the portions of the documents likely associated with the at least one topic of interest are identified by: for every line in file, if a pre-defined density related condition is met, (i) growing an existing text island if an immediately prior line is part of the existing text island, and (ii) creating a new text island if the immediately prior line is not part of an existing text island, and if a pre-defined density condition is not met, skipping the line.

7. A method as in claim 6 , wherein the pre-defined density relation condition comprises whether the density for the line is above a pre-defined value.

8. A method as in claim 6 , wherein the pre-defined density condition is experimentally determined using a plurality of historical documents.

9. A method as in claim 6 , wherein the pre-defined density condition is determined using a model trained using a plurality of historical documents.

10. A method as in claim 9 , wherein the model comprises: a supervised machine learning algorithm or a regression model.

11. A method as in claim 7 , wherein after the text islands have all been generated, for every island, dropping each text island that does not meet a pre-defined retention condition, and writing each remaining text island to a cleaned document file, the remaining text islands corresponding to the identified portions of the documents likely associated with the at least one topic of interest.

12. A method as in claim 11 , wherein the pre-defined retention condition comprises a minimum number of lines for each text island.

13. A method as in claim 11 , wherein the pre-defined retention condition comprises a spatial location of the text island in comparison to other text islands.

14. A method as in claim 1 , wherein providing data comprises one or more of storing the data, transmitting the data, and displaying the data.

15. A method as in claim 1 , wherein providing data comprises generating a cleaned document file removing the portions of the document not likely associated with the at least one topic of interest.

16. A non-transitory computer-readable medium encoding instructions that, when executed by at least one data processor, cause the at least one data processor to perform operations comprising:

receiving a document having a plurality of lines with text, the document comprising text associated with at least one topic of interest and text not associated with the at least one topic of interest;

determining, for each line in the document, a length of the line and a number of off-topic indicators, the off-topic indicators characterizing portions of the document as likely being not being associated with the at least one topic of interest;

determining a density for each line based on the determined line length and the determined number of off-topic indicators, the density for each line being proportional to the determined line length and inversely proportional to an exponential base greater than one of the number of off-topic indicators;

identifying, using the determined densities for each line, portions of the documents likely associated with the at least one topic of interest; and

providing data characterizing the identified portions of the document.

17. A computer-readable medium as in claim 16 , wherein the operations further comprise: applying a density smoothing filter to each line to result in smoothed densities, and wherein the portions of the documents likely associated with the at least one topic of interest are identified using the smoothed densities.

18. A computer-readable medium as in claim 16 , wherein the portions of the documents likely associated with the at least one topic of interest are identified by: for every line in file, if a pre-defined density related condition is met, (i) growing an existing text island if an immediately prior line is part of the existing text island, or (ii) creating a new text island if the immediately prior line is not part of an existing text island, and if a pre-defined density condition is not met, skipping the line.

19. A system comprising:

at least one data processor;

memory storing instructions, which when executed by the at least one data processor result in operations comprising:

receiving a document having a plurality of lines with text, the document comprising text associated with at least one topic of interest and text not associated with the at least one topic of interest;

determining, for each line in the document, a length of the line and a number of off-topic indicators, the off-topic indicators includes a number of hyperlinks in the line, the off-topic indicators characterizing portions of the document as likely being not being associated with the at least one topic of interest;

determining a density for each line based on the determined line length and the determined number of off-topic indicators, the density for each line being proportional to the determined line length and inversely proportional to an exponential base greater than one of the number of off-topic indicators;

identifying, using the determined densities for each line, portions of the documents likely associated with the at least one topic of interest; and

providing data characterizing the identified portions of the document;

wherein the portions of the documents likely associated with the at least one topic of interest are identified by: for every line in file, if a pre-defined density related condition is met, (i) growing an existing text island if an immediately prior line is part of the existing text island, and (ii) creating a new text island if the immediately prior line is not part of an existing text island, and if a pre-defined density condition is not met, skipping the line.

Assignments (2)
CHANGE OF NAME Recorded Aug 26, 2014
From: SAP AG
To: SAP SE
Reel/Frame 033625/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2012
From: SHAMI, MOHAMMAD; HERMAN, DAVID; BOTROS, SHERIF
To: SAP AG
Reel/Frame 028173/0434 →