IP Library Granted Patent US 9,449,114
Granted Patent B2
US 9,449,114 · App. 12/761,272 · Granted Sep 20, 2016

Removing non-substantive content from a web page by removing its text-sparse nodes and removing high-frequency sentences of its text-dense nodes using sentence hash value frequency across a web page collection

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,449,114
App. No.
12/761,272
Granted
Sep 20, 2016
Kind
B2
Abstract

A method and system for removing chrome from a web page is provided. An example system includes a parsing module, a text density analyzer, a content node selector 206 , and a text extractor. The parsing module may be configured to parse a web page into a tree structure. The text density analyzer may be configured to determine a text density score value for each node from the tree structure. The content node selector may be configured to identify one or more nodes from the tree structure as content nodes based on their respective text density score values. The text extractor may be configured to extract text from the content nodes only.

Claims (45)

1. A method comprising:

accessing a web page, the web page represented by a hierarchical mark-up language;

parsing the web page into a tree structure based on a hierarchy associated with the web page;

for each node from the tree structure, determining a text density score value, the text density score value calculated as a sum of a word count associated with a node and weighted word counts associated with adjacent nodes from the tree structure, an adjacent node from the adjacent nodes comprising a sibling of the node;

identifying those nodes from the tree structure for which a text density value is above the threshold value as content nodes;

extracting text from the content nodes associated with the web page and ignoring nodes that were not identified as content nodes;

breaking up the extracted text and further text into a plurality of sentences;

hashing each sentence from the plurality of sentences;

calculating frequency of each sentence in the extracted text and the further text; and

identifying sentences from one or more sentences having a frequency value above a frequency threshold value as indicative of boilerplate language.

2. The method of claim 1 , comprising:

providing the extracted text associated with the web page to a sentence analyzer; and

removing from the extracted text the one or more sentences indicative of boilerplate language to produce subject text.

3. The method of claim 2 , further comprising providing the subject text to a consuming application.

4. The method of claim 1 , wherein the further text is associated with one or more further web pages, the web page is associated with a first web page and a further web page from the one or more further web pages is associated with a second web page.

5. The method of claim 1 , comprising removing structured chrome from the web page prior to the parsing of the web page into the tree structure.

6. The method of claim 5 , wherein structured chrome comprises embedded Cascading Style Sheets (CSS).

7. The method of claim 1 , wherein the web page is represented by HyperText Markup Language (HTML).

8. The method of claim 7 , wherein a node from the nodes that were not identified as content nodes is associated with navigation, advertising, or HTML comments.

9. A computer system comprising:

a parsing module, implemented using one or more processors, to parse, using the at least one processor, a web page into a tree structure based on hierarchy of a hierarchical mark-up language associated with the web page;

a text density analyzer, implemented using one or more processors, to determine, using the at least one processor, a text density score value for each node from the tree structure, the text density score value calculated as a sum of a word count associated with a node and weighted word counts associated with adjacent nodes from the tree structure, an adjacent node from the adjacent nodes comprising a sibling of the node, a sibling of the node or a child of the node;

a content node selector, implemented using one or more processors, to identify, using the at least one processor, those nodes from the tree structure for which a text density value is above the threshold value as content nodes;

a text extractor, implemented using one or more processors, to extract, using the at least one processor, text from the content nodes associated with the web page and ignoring nodes that were not identified as content nodes; and

a sentence analyzer, implemented using one or more processors, to, using the at least one processor:

break up the extracted text and further text into a plurality of sentences,

hash each sentence from the plurality of sentences,

calculate frequency of each sentence from the plurality of sentences in the extracted text and the further text, and

identify a sentence from the plurality of sentences having a frequency value above a frequency threshold value as indicative of boilerplate language.

10. The system of claim 9 , comprising:

a boilerplate filter to remove from the extracted text sentences indicative of boilerplate language to produce subject text.

11. The system of claim 10 , further comprising a communications module to provide the subject text to a consuming application.

12. The system of claim 9 , wherein the further text is associated with one or more further web pages, the web page is associated with a first web page and a further web page from the one or more further web pages is associated with a second web page.

13. The system of claim 9 , comprising a structured chrome filter to remove structured chrome from the web page.

14. The system of claim 13 , wherein structured chrome comprises embedded CSS.

15. The system of claim 9 , wherein the web page is represented by HyperText Markup Language (HTML).

16. A machine-readable non-transitory storage medium having instruction data to cause a machine to perform operations comprising:

parsing a web page into a tree structure based on a hierarchy of a mark-up language associated with the web page;

determining a text density score value for each node from the tree structure, the text density score value calculated as a sum of a word count associated with a node and weighted word counts associated with adjacent nodes from the tree structure, an adjacent node from the adjacent nodes comprising a sibling of the node;

identifying those nodes from the tree structure for which a text density value is above the threshold value as content nodes; and

extracting text from the content nodes associated with the web page and ignoring nodes that were not identified as content nodes, the extracted text associated with the web page;

breaking up the extracted text and further text into a plurality of sentences;

hashing each sentence from the plurality of sentences;

calculating frequency of each sentence from the plurality of sentences in the extracted text and the further text; and

identifying a sentence from the plurality of sentences having a frequency value above a frequency threshold value as indicative of boilerplate language.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2015
From: EBAY INC.
To: PAYPAL, INC.
Reel/Frame 036169/0707 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 20, 2010
From: ROPER, JOHN; GLASGOW, DANE
To: EBAY INC.
Reel/Frame 024419/0591 →