IP Library Granted Patent US 9,336,202
Granted Patent B2
US 9,336,202 · App. 13/753,645 · Granted May 10, 2016

Method and system relating to salient content extraction for electronic content

Inventor: Shahzad Khan (Ottawa, CA)
Assignee: Whyz Technologies Limited
G06F17/2785G06F17/30684G06F17/30699G06F17/30707G06F17/274G06F17/289
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,336,202
App. No.
13/753,645
Granted
May 10, 2016
Kind
B2
Abstract

Individuals receive overwhelming barrage of information which must be filtered, processed, analyzed, reviewed, consolidated and distributed or acted upon. Automatic approaches to “scraping” salient content from sources of content are provided allowing the salient content to be provided to the user or subjected to further processing such as clustering or sentiment analysis for example. Embodiments of the invention provide for: automated scraper induction based on document and/or contextual semantic cues and document structure analysis. identifying salient text, removing boiler-plate text, off-topic content and other non-salient content; deriving reusable descriptive extraction patterns for subsequent documents; applying descriptive extraction patterns for extraction from subsequent documents from the same source; intelligent identification of extraction success confidence score, using historical success scores; and employing confidence scores to automatically trigger new extraction pattern identification if extracted confidence is below an acceptable confidence threshold.

Claims (33)

1. A method comprising:

a) receiving an item of content;

b) identifying within the item of content using a microprocessor a set of lexical pattern cues for core content of the item of content and selecting a segment of the item of content having a highest likelihood as being the core content based upon a structural analysis of the item of content in dependence upon at least the set of lexical pattern cues;

c) parsing the item of content to generate a hierarchy of content within the item of content;

d) ranking the hierarchy of content in dependence upon at least the lexical pattern cues and sorting the resulting ranking;

e) identifying a gap when searching down the ranking meeting a predetermined threshold and removing those portions of the hierarchy of content below the gap to generate truncated content;

f) finding all occurrences for portions of the hierarchy of content with closest match to the lexical pattern cues closest to the start of the item of content;

g) determining whether multiple matches to the lexical pattern cues exist and establishing an action in dependence upon at least whether multiple matches exist or not;

h) performing the action, wherein the action is at least one of:

establishing the occurrence for the portion of the hierarchy of content as the core content of the item of content when the determination of multiple matches is negative; and

establishing the occurrence for the portion of the hierarchy of content that at least one of contains the largest portion of the item of content and is the first occurrence as the core content of the item of content when the determination of multiple matches is positive.

2. The method according to claim 1 further comprising:

i) establishing a truncation point within the remaining portion of the hierarchy of content, the truncation point being the start of trailing extraneous content established by semantic analysis of the truncated content; and

j) removing that portion of the hierarchy of content after the truncation point from the truncated content.

3. The method according to claim 1 , wherein the item of content is a web page and the hierarchy of content is a document object model tree.

4. The method according to claim 1 further comprising:

i) establishing a core topic relating to the core content;

j) assessing the next portions of the truncated content for cohesion with the core topic and discarding those that are not cohesive;

k) evaluating a retained portion of the truncated content to determine whether each portion stays related to the core content and truncating those portions that go off topic;

l) repeating steps (j) and (k) until all portions of the truncated content have been analysed;

m) storing remaining truncated content as final content.

5. The method according to claim 4 further comprising:

n) comparing the final content to any other occurrences of portions of the hierarchy of content matching the lexical pattern cues for a closer match than the current selection and selecting said if a closer match; and

o) storing the resulting active portion of the hierarchy of content in a database together with an association to the item of content.

6. The method according to claim 4 further comprising:

determining in step ( 1 ) whether a threshold is reached in terms of a rate of discarding and truncating portions of truncated content compared to assessing and evaluating them; and

removing all subsequent portions of truncated content when the threshold is reached.

7. The method according to claim 1 further comprising:

i) employing the location of final text within the hierarchy of content to describe a descriptive extraction pattern that can be employed to identify the final text in the hierarchy of content; and

j) storing this descriptive extraction pattern in association with a label that can identify the portion of the selected content in which the final text is found as it is located within the hierarchy of content.

8. The method according to claim 1 further comprising:

i) determining a confidence metric in dependence upon at least a comparison of the truncated text against the hierarchy of content; and

j) storing the confidence metric together with at least one of the item of content, a reference to the item of content, and the truncated content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2013
From: KHAN, SHAHZAD
To: WHYZ TECHNOLOGIES LIMITED
Reel/Frame 029737/0538 →
Continuity (2)
Provisional Application 61647183 · May 15, 2012
Related Publication 20130311169A1 · Nov 21, 2013